X:Kim
@kimmonismus · X
切换来源
@kimmonismus@kimmonismusAI 评分55 @kimmonismus@kimmonismusAI 评分3333 
@kimmonismus@kimmonismusAI 评分55 @kimmonismus@kimmonismusAI 评分5454 Cerebras 发布 CS-4 机架级 AI 推理系统,在沿用 5nm 晶圆、4 万亿晶体管和 90 万个 AI 核心的情况下,通过重新设计供电与散热让晶圆以两倍时钟频率运行,推理性能提升近一倍。


@kimmonismus@kimmonismusAI 评分2525 @kimmonismus@kimmonismusAI 评分2424 唉:GPT-Astra 似乎距离发布还有“数周”,而 Anthropic 在此之前大概也不会发布 Fable 5.1,尽管他们确实需要更好的公关。https://t.co/VSJJcmxjkQ
@kimmonismus@kimmonismusAI 评分3535 
@kimmonismus@kimmonismus精选AI 评分7171 
推荐理由:原文给出 Asana 大型前端迁移的工时与成本对比,可作为判断智能体承接遗留代码改造的参考。
@kimmonismus@kimmonismusAI 评分2222 我问这个的原因是:Sonnet 5 跟比如 5.6 Luna 相比,又慢又贵。我看不出它有什么使用场景,尤其是我用 Claude 越来越少了。
@kimmonismus@kimmonismusAI 评分1010 @kimmonismus@kimmonismusAI 评分3030 Slack 的发展很有意思。最近宣布了 Claude 深度集成到 Slack(连 Andrej Karpathy 都写了相关内容),现在又在实现编程功能。https://t.co/l1cIyI8HRm
@kimmonismus@kimmonismusAI 评分5555 引用@ClaudeDevs@ClaudeDevsYou can now set Claude Code's output style to Concise. Claude leads with the result, keeps responses short, and still gives full detail when you ask. Turn it on in /config → Output style, or set "outputStyle": "Concise" in settings.json. https://t.co/XYg7bHeVT2
@kimmonismus@kimmonismus精选AI 评分6565 
推荐理由:CNBC 披露的这组用户与营收数字,可用来观察 OpenAI 编程产品的采用规模与商业化节奏。
@kimmonismus@kimmonismusAI 评分2525 Stripe 告诉投资者“奇点已经开始了”,2026 年 1 月 1 日。 确实,我同意这一点。2026 年是转折点。一切都在加速,现在大家挂在嘴边的词是“递归自我改进”。

@kimmonismus@kimmonismusAI 评分77 7/ 仓库:https://t.co/KFAJ170gBL 如果它对你有用,点个 star 能帮其他开发者发现它。MIT 许可,免费,可自托管。
@kimmonismus@kimmonismusAI 评分5656 @kimmonismus@kimmonismusAI 评分3333 
@kimmonismus@kimmonismusAI 评分2222 4/ 发布时即模型无关:OpenAI、Anthropic、Google,以及 Kimi、GLM 和 DeepSeek 等开放权重模型。将每个任务路由到符合你成本和延迟预算的模型。
@kimmonismus@kimmonismusAI 评分3535 @kimmonismus@kimmonismusAI 评分3131 



@kimmonismus@kimmonismus精选AI 评分6666 
推荐理由:开源发布的智能体运行框架给出本地与托管两种部署路径,并列出同一模型与开放模型下的单次运行成本对比。
@kimmonismus@kimmonismusAI 评分00 @kimmonismus@kimmonismusAI 评分44 @kimmonismus@kimmonismusAI 评分00 哇,谢谢 @axios。真的很荣幸能被报道 <3 https://t.co/GSYvHDTedv https://t.co/2UgAfkfiE7

@kimmonismus@kimmonismusAI 评分88 7/ 它弥合了“设计概念”与“我能在此基础上构建的东西”之间的差距。对于独立开发者或小团队来说,这能节省大量时间和成本。 你可以在这里查看:https://t.co/NllA1PJTGl
@kimmonismus@kimmonismusAI 评分2626 
@kimmonismus@kimmonismusAI 评分1616 
@kimmonismus@kimmonismusAI 评分2626 

@kimmonismus@kimmonismusAI 评分3232 


@kimmonismus@kimmonismusAI 评分2828 
@kimmonismus@kimmonismusAI 评分2222 
@kimmonismus@kimmonismusAI 评分6363 
引用@Replit@ReplitReplit Free Mode, powered by @OpenAI GPT-5.6 Luna. Let’s make intelligence accessible to everyone. https://t.co/UDcrYl5HZL
@kimmonismus@kimmonismusAI 评分44 @kimmonismus@kimmonismusAI 评分44 @kimmonismus@kimmonismus精选AI 评分6767 Moderna 与 Merck 的 AI 辅助个体化 mRNA 癌症疗法在涉及 1,137 名患者的 Phase 3 试验中取得成功,作者称这是该疗法首次在 Phase 3 试验中成功。


推荐理由:内容给出 AI 筛选肿瘤新抗原并定制 mRNA 疗法的流程,以及 Phase 2 与 Phase 3 的对照数据,便于了解 AI 在个体化癌症治疗中的位置。
@kimmonismus@kimmonismusAI 评分2929 @kimmonismus@kimmonismusAI 评分99 @kimmonismus@kimmonismus精选AI 评分7575 
推荐理由:推文并列了 OpenAI 与 Anthropic 最新季度的营收和运营亏损,可用于比较两家公司在 IPO 前的财务位置。
@kimmonismus@kimmonismusAI 评分5656 引用@jietang@jietangThoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
@kimmonismus@kimmonismusAI 评分11