X:马东锡 NLP
@dongxi_nlp · X
切换来源
@dongxi_nlp@dongxi_nlpAI 评分4040 @dongxi_nlp@dongxi_nlpAI 评分6464 引用@OpenAI@OpenAIThis is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast. https://t.co/gDd0IsewJw
@dongxi_nlp@dongxi_nlpAI 评分2727
引用@dongxi_nlp@dongxi_nlp人类常常先学会一个个单词,随后在漫长的经历中,把它们组合起来,表达越来越复杂的意义。 比如,从“流”“星”“雨”,慢慢理解“流星”,又理解“流星雨”。 人通过组合,观看、想象和使用,逐渐理解语言的意义。 反观 AI,起点非常粗暴。它从庞大的语料中建立 token 与编号之间的对应关系: 1,流;2,星;3,雨;4,流星;5,流星雨。 在最初的这一刻,那里只有 1、2、3、4、5,还没有夜空,没有坠落的光,也没有人在流星雨下许愿。 随后,庞大的神经网络开始学习这些编号在无数文本中的联系,由此逐渐建立 token 与 token 之间的关系,并从关系之中逼近语言的意义。 但无论如何,AI 的这种苍白感,在一开始就注定了。
@dongxi_nlp@dongxi_nlpAI 评分2020 非常厉害! Software World——一个由智能体运营的“GitHub”。https://t.co/Gpj0u1aQIR
@dongxi_nlp@dongxi_nlpAI 评分3030 非常厉害! Video Delta Net (VDN): 用于实时文本生成视频的混合注意力,画质近乎无损。https://t.co/TKIEz5PRYt
@dongxi_nlp@dongxi_nlpAI 评分2424 @dongxi_nlp@dongxi_nlpAI 评分3333 Gemini 3.8 发布了,表现非常亮眼。 Gemini 似乎在复制去年 2.5 pro - 3 的节奏,夏天后开始发力,到年底来一波 ultra 年票打折,绑走用户。 然后拉半年。
@dongxi_nlp@dongxi_nlpAI 评分4949
引用@natolambert@natolambertOver the weekend I had Codex parse 500K arXiv AI/ML papers since ChatGPT to understand which open models are used for research. In 2024, ~30% of papers mentioned an American open model and only 10% a Chinese model. Today, ~40% of papers mention a Chinese (open) LLM, and only 25-30% an American one. Chinese models are the default for research. Chinese mentions are still growing while American open models are stagnating. When looking at this data it's important to remember that papers substantially lag model releases, as research takes a long time. Qwen's steady growth is reflective of this, but so is Llama's lasting power. Some more observations: 1. Qwen has been steadily growing, and today 1/3 of papers which mention any LLM mention qwen. OpenAI's closed models are the highest overall, at ~37%. 2. Llama peaked around April of 2025 at 30% of papers which mention any LLM (including ChatGPT etc). Llama 4 was released at about the same time, and Llama has been declining since. 3. Gemini and Claude are less common than the leading open models, mentioned in 10-15% of papers puts them behind all of Qwen, Llama, and DeepSeek. Open models should be and are the foundations of open research. The % of papers mentioning any LLM have been steadily climbing since 2023. | Year | January | April | July | October | | 2023 | 10.43% | 15.39% | 18.69% | 32.18% | | 2024 | 29.70% | 33.93% | 35.70% | 44.25% | | 2025 | 39.23% | 45.28% | 44.94% | 53.52% | | 2026 | 55.49% | 57.26% | 53.14% | TBD Now over 50% of AI papers, from 10% in 2023. Other notes: - Gemma and Mistral hover around 5-10%. - Our beloved fully-open Olmo models have been ~1% since the first release in Jan. 2024. - DeepSeek has a clear jump after R1 in Jan. 2025 - Data derived from the most popular ML arXiv categories: cs. AI, cs. CL, cs. CV, cs. LG, stat. ML Just like our downloads and derivative model data, this is updated daily on the Interconnects Open Model Dashboard.
@dongxi_nlp@dongxi_nlpAI 评分1616 @dongxi_nlp@dongxi_nlpAI 评分2222 检验结果,审计过程。 面向发现式智能的 benchmark:TRACES https://t.co/SrputhGyqB
引用@dongxi_nlp@dongxi_nlphttps://t.co/su2424d9MB
@dongxi_nlp@dongxi_nlpAI 评分66 @dongxi_nlp@dongxi_nlpAI 评分1616 TRACES 那个不信任正确答案的基准测试。https://t.co/EDstfG6Gzm
引用@dongxi_nlp@dongxi_nlphttps://t.co/c2lM28Bc2e
@dongxi_nlp@dongxi_nlpAI 评分66
@dongxi_nlp@dongxi_nlpAI 评分4949
引用@thsottiaux@thsottiauxWe've investigated a few messages about codex usage limits being different. That's not something we change without engaging the community and being transparent. What we did see is that when talking to affected users many were using sub2api. Converting a subscription into api traffic to then re-serve or share across many users is not something we support and this type of usage gets flagged by our fraud-prevention systems. You are completely fine if you use your subscription through Sign in With ChatGPT, either through the official clients or through one of the many OSS clients (Pi, OpenCode, ...) that support signing in with your account and using your included usage.
@dongxi_nlp@dongxi_nlpAI 评分2323
@dongxi_nlp@dongxi_nlpAI 评分5656 引用@jietang@jietangThoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
@dongxi_nlp@dongxi_nlpAI 评分2828
@dongxi_nlp@dongxi_nlpAI 评分2323 @dongxi_nlp@dongxi_nlpAI 评分2626