X:Rohan Paul
@rohanpaul_ai · X
切换来源
@rohanpaul_ai@rohanpaul_aiAI 评分2626 
@rohanpaul_ai@rohanpaul_aiAI 评分4343 
@rohanpaul_ai@rohanpaul_aiAI 评分4444 
@rohanpaul_ai@rohanpaul_ai精选AI 评分6565 
推荐理由:论文给出技能库持久化风险的量化证据,并附检测基准与修复方案,可迁移到智能体安全评估。
@rohanpaul_ai@rohanpaul_aiAI 评分4747
引用@jietang@jietangThoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
@rohanpaul_ai@rohanpaul_aiAI 评分3535 完全没有后段曲线刹车的余地 好奇感知回路怎么会漏掉那么近的障碍物 在北京国家速滑馆举行的第二届世界人形机器人运动会热身/测试期间 https://t.co/HXUqgIDWXK

@rohanpaul_ai@rohanpaul_aiAI 评分1818 @rohanpaul_ai@rohanpaul_aiAI 评分2525 
@rohanpaul_ai@rohanpaul_aiAI 评分4040 
@rohanpaul_ai@rohanpaul_ai精选AI 评分6767 
推荐理由:泄露的投资人信给出 OpenRouter 的 token 消耗增速,可看到支付公司如何把智能体当作经济参与者来定价。
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,您提供的主推文内容仅包含一个链接(https://t.co/sPX7Rc2Jbr),没有可翻译的正文文字。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分4646 
@rohanpaul_ai@rohanpaul_aiAI 评分4848 
@rohanpaul_ai@rohanpaul_aiAI 评分66 @rohanpaul_ai@rohanpaul_aiAI 评分4545 OpenAI 预览 Private Safety Processing,让符合条件的零数据保留(ZDR)客户在请求处理后仍可使用前沿模型,且 OpenAI 不保留其提示词与回复。
引用@OpenAI@OpenAIWe will continue to offer Zero Data Retention for frontier models. As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions. To help address those risks, we're previewing Private Safety Processing, which is designed to improve safety without giving OpenAI personnel access to the underlying content.
@rohanpaul_ai@rohanpaul_aiAI 评分99 @rohanpaul_ai@rohanpaul_aiAI 评分3030 
@rohanpaul_ai@rohanpaul_aiAI 评分6363 
@rohanpaul_ai@rohanpaul_aiAI 评分6060 
@rohanpaul_ai@rohanpaul_aiAI 评分4545 比赛开始前,一个营的 Booster T2 机器人同步行进。 好奇它们在那样的速度下如何保持横向间距如此均匀 所有机器人的步频看起来都锁定在同一个时钟上 https://t.co/ObxmX6q2xI

@rohanpaul_ai@rohanpaul_aiAI 评分3131 
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,主推文内容仅包含一个链接(https://t.co/Zz84x5Q4mI),没有可翻译的正文文本。请提供推文的实际文字内容,以便我进行翻译和标题拟定。
@rohanpaul_ai@rohanpaul_aiAI 评分2121 
@rohanpaul_ai@rohanpaul_aiAI 评分2727 
@rohanpaul_ai@rohanpaul_aiAI 评分2424 中国机器人实时平衡控制的精彩展示,保持重心稳定。 这个人形机器人把弯道当成了直道。 在弯道处横向力控制得如此干净利落,速度丝毫不减。 https://t.co/ZAGN9e1kil

@rohanpaul_ai@rohanpaul_aiAI 评分6060 开源模型家族 Ornith-1.5 发布,包含 9B Dense、35B MoE 和 397B MoE 三个版本。

@rohanpaul_ai@rohanpaul_aiAI 评分2727 📘 Ornith-1.5 技术报告:https://t.co/JUd4KU6i87 🤗Huggingface:https://t.co/6evFvzrNkC
@rohanpaul_ai@rohanpaul_aiAI 评分1010 下一个基准测试:检测出人类,然后随便选任何其他地方。 https://t.co/YayhBbnuTv

@rohanpaul_ai@rohanpaul_aiAI 评分3737 
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,您提供的主推文内容仅为一个链接(https://t.co/wLPHyW2OjV),没有可翻译的正文文字。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分3434 

@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分4949 斯坦福新论文发现,在 3 个企业级 Agent benchmark 上,不到 3% 的分数差异来自 Agent 本身,7–23% 来自 Agent 与具体任务的交互,排行榜排名高未必适合你的工作流。

@rohanpaul_ai@rohanpaul_aiAI 评分6464
引用@rohanpaul_ai@rohanpaul_aiFor the first time, a majority of Americans under 30 are more concerned than excited about AI, reaching 55%. Fresh Pew Research data. - The employment fear is even broader, with 73% of under-30s expecting AI to reduce U.S. jobs over 20 years, up from 61% in 2024. - Only 5% of U.S. adults now think AI will create more jobs. - Most under-30s still use chatbots, so widespread use is coexisting with rising concern and mixed views about AI's effect on creativity. - Most under-30s also use chatbots, yet Pew's earlier work finds they are equally likely to say those tools hurt or help their creativity.
@rohanpaul_ai@rohanpaul_aiAI 评分77 抱歉,主推文内容仅包含一个链接(https://t.co/1toQgEXoQJ),没有可翻译的文字内容。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_ai精选AI 评分7373 Anthropic 公布 Claude 自主设计 de novo 蛋白质结合体的实验结果:在 1320 个可测设计中 354 个结合目标,命中率 26.8%,15 个靶点中 14 个找到结合体。
引用@AnthropicAI@AnthropicAIMany drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work. We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets. We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
推荐理由:论文给出 Claude 自主完成蛋白质设计全流程的可复现流程和实测命中率,便于评估智能体在科研实验中的实际边界。
@rohanpaul_ai@rohanpaul_aiAI 评分2020 来看看 @thehypedotnews 推出的 24x7 电台形式的 AI News。 听当天的 AI 新闻,相当舒缓、好听。 https://t.co/nSakICqBOE
@rohanpaul_ai@rohanpaul_aiAI 评分4545 3D 编程测试显示,DeepSeek-V4-Pro-0813 的 token 消耗量是 Muse Spark 1.2 的 48 倍。

@rohanpaul_ai@rohanpaul_aiAI 评分66 抱歉,您提供的主推文内容只有一个链接(https://t.co/Tucoekm2DR),没有实际的推文文字。我无法翻译不存在的文本。 请提供推文的实际文字内容,我会立即为您翻译并拟定标题。
@rohanpaul_ai@rohanpaul_aiAI 评分5757 