唉:GPT-Astra 似乎距离发布还有“数周”,而 Anthropic 在此之前大概也不会发布 Fable 5.1,尽管他们确实需要更好的公关。https://t.co/VSJJcmxjkQ
X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(537)
@kimmonismus@kimmonismusAI 评分2424 @cohere@cohereAI 评分55 
@OpenBMB@OpenBMBAI 评分6161 


@rohanpaul_ai@rohanpaul_aiAI 评分4343 
@ClementDelangue@ClementDelangueAI 评分2727 同意 @gdb 的观点,我们需要用 API 和开放模型来武装网络防御者,力度要远超现在。AI 真正的网络安全风险在于攻击者与防御者之间在权力、能力和资源上的不对称!
引用@gdb@gdbdefenders can see the future, and have a narrow window to uplevel their cybersecurity practices now. key is to uplevel fundamentals and apply the best AI tools. what we’re doing at OpenAI, and where other organizations can start: https://t.co/P3IMZkV234
@thexpin@thexpinAI 评分6464 
@lifesinger@lifesingerAI 评分1616 @rohanpaul_ai@rohanpaul_aiAI 评分4444 
@rohanpaul_ai@rohanpaul_ai精选AI 评分6565 
推荐理由:论文给出技能库持久化风险的量化证据,并附检测基准与修复方案,可迁移到智能体安全评估。
@OpenBMB@OpenBMBAI 评分5050 

@frxiaobei@frxiaobeiAI 评分22 看看喝了多少糖! https://t.co/VBphQkMJ1H

@kimmonismus@kimmonismusAI 评分3535 
@rohanpaul_ai@rohanpaul_aiAI 评分4747
引用@jietang@jietangThoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
@Kling_ai@Kling_aiAI 评分66 掌控每一个角度。抢占聚光灯。✨ https://t.co/YF4ONSXZkj

@PixVerse_@PixVerse_AI 评分22 @PixVerse_@PixVerse_AI 评分2525 
@elonmusk@elonmuskAI 评分11 @elonmusk@elonmuskAI 评分44 @elonmusk@elonmuskAI 评分1010 @elonmusk@elonmuskAI 评分88 @foxshuo@foxshuoAI 评分4242 智谱 GLM-5.3 在 AA 竞技场开榜,与 Kimi K3 并列 60 分,成为唯二挤进 T1 阵营的国产模型,紧咬 GPT-5.6 和 Claude Opus 5。

@elonmusk@elonmuskAI 评分3838 Grok @Bot 拥有自己的远程计算机 @SpaceXAI,所以即使你合上笔记本或重启台式机,它也能继续工作。 大不同。 https://t.co/7dUQeLFrtA
@elonmusk@elonmuskAI 评分2525 Grok Build https://t.co/ppqui4HYMm https://t.co/KfRSQ1FfiT
引用@cb_doge@cb_dogeThis is why @Grok Build is a game changer. I had never used Blender in my life, yet on Day 1, with zero experience, I created this iPhone render from scratch with the help of Grok Build. https://t.co/fUaKmAkjCS
@elonmusk@elonmuskAI 评分44 @elonmusk@elonmuskAI 评分66 @cb_doge@cb_dogeAI 评分1616 
@kimmonismus@kimmonismus精选AI 评分7171 
推荐理由:原文给出 Asana 大型前端迁移的工时与成本对比,可作为判断智能体承接遗留代码改造的参考。
@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分1919 AI 极大满足了人大多数时候只想说、只想听自己想说想听的内容,而非真正想讨论的需求。这让"与不同的人做开放、差异化讨论"的需求被显化,也让人与人交流中真正不可替代的部分变得清晰。

@testingcatalog@testingcatalog精选AI 评分7272 
推荐理由:Cursor 让子智能体在独立虚拟机中运行并支持定时任务,读者可据此了解云端编码智能体的隔离与自动化方向。
@kimmonismus@kimmonismusAI 评分2222 我问这个的原因是:Sonnet 5 跟比如 5.6 Luna 相比,又慢又贵。我看不出它有什么使用场景,尤其是我用 Claude 越来越少了。
@kimmonismus@kimmonismusAI 评分1010 @PixVerse_@PixVerse_AI 评分11 @alibaba_cloud@alibaba_cloudAI 评分2525 @kimmonismus@kimmonismusAI 评分3030 Slack 的发展很有意思。最近宣布了 Claude 深度集成到 Slack(连 Andrej Karpathy 都写了相关内容),现在又在实现编程功能。https://t.co/l1cIyI8HRm
@cb_doge@cb_dogeAI 评分1313 “我是一名技术专家。看到人们享受我公司制造的产品,我从中获得快乐。” —— Elon Musk https://t.co/wcb7OEREVb

@dongxi_nlp@dongxi_nlpAI 评分1212 模型们飞速的进化,benchmark 数字,dgx spark 玩家千篇一律的多少 token/s 太过苍白。 准备开个新坑,系统地写写后训练。
@gabriel1@gabriel1AI 评分66 就像某件你被迫全年无休、持续多年去做的事,其中大多数最佳实践早已被摸索清楚,却还承受着巨大的社会压力,要求你把每个动作都做到完美。
@kimmonismus@kimmonismusAI 评分5555 引用@ClaudeDevs@ClaudeDevsYou can now set Claude Code's output style to Concise. Claude leads with the result, keeps responses short, and still gives full detail when you ask. Turn it on in /config → Output style, or set "outputStyle": "Concise" in settings.json. https://t.co/XYg7bHeVT2
@Baidu_Inc@Baidu_IncAI 评分55 @Baidu_Inc@Baidu_IncAI 评分4747 