X:Kim
@kimmonismus · X
切换来源
@kimmonismus@kimmonismusAI 评分1919 @kimmonismus@kimmonismusAI 评分3737
引用sui ☄️ (@birdabo)@birdabo‼️it seems Anthropic is ready to publicly launch a new version of Mythos, something better than Mythos Preview. a codenamed model “Oceanus” was given access to some red teamers yesterday according to @synthwavedd. it’s apparently been paused already, due to someone reselling access through a Chinese API proxy lmao 💀 Mythos pricing might also end up at with $16 Input, $80 Output according to @scaling01
@kimmonismus@kimmonismus精选AI 评分6868 引用Chubby♨️ (@kimmonismus)@kimmonismusHoly moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about
推荐理由:汇总了 Anthropic 博客关于递归自我改进的关键数据与三种未来路径,可据此判断自动化编码的推进速度。
@kimmonismus@kimmonismus精选AI 评分7373 


引用Anthropic (@AnthropicAI)@AnthropicAIOur internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. anthropic.com/institute/recu…
推荐理由:转述 Anthropic 内部数据,读者可据此了解递归自我改进讨论背后的具体加速指标。
@kimmonismus@kimmonismusAI 评分3737 @kimmonismus@kimmonismusAI 评分4949 @kimmonismus@kimmonismus精选AI 评分6969 引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlysNVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest US open weights models, Gemma 4 31B (39.2), Nemotron 3 Super (36.0) and gpt-oss-120b (33.3), but behind the Chinese-led open weights frontier (Kimi K2.6 at 53.9). We partnered with @NVIDIA to evaluate this model for intelligence and speed ahead of its public release. These figures use the final NVFP4 weights that NVIDIA recommends for inference, but our tests show minimal intelligence impact compared to BF16 testing, with higher precision resulting in an Artificial Analysis Intelligence Index score of 48.2 vs. the NVFP4 score of 47.7. Key Takeaways: ➤ Nemotron 3 Ultra leads in speed for its intelligence: through BlackBox AI ahead of release, Nemotron 3 Ultra is served at over 400 output tokens per second - this is slightly faster than the typical serving speed of gpt-oss-120b despite being >4X larger, and comes with significantly greater intelligence ➤ Largest Nemotron 3 model so far: with approximately 550 billion total parameters and 55 billion active, Nemotron 3 Ultra is significantly larger than its siblings and is the largest and most intelligent US open weights model release ever ➤ Nemotron 3 Ultra is the leading US open weights model on the Artificial Analysis Intelligence and Agentic Indexes by far, but Gemma 4 31B scores ~1 point higher on the Coding Index (comprised of Terminal-Bench Hard and SciCode)
推荐理由:Artificial Analysis 的评测让读者能横向比较 Nemotron 3 Ultra 与美国及中国开源权重模型的智能与速度表现。
@kimmonismus@kimmonismus精选AI 评分8181 

引用NVIDIA AI (@NVIDIAAI)@NVIDIAAIToday we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models. Video
推荐理由:原文给出 550B 开源模型的权重、训练数据与完整配方,并说明其在长任务智能体上的吞吐表现,读者可据此判断开源前沿模型的可复现程度。
@kimmonismus@kimmonismus精选AI 评分7272 

推荐理由:原文引用 OpenAI 关于递归自我改进早期迹象的表述,可看到其对竞争压力与治理难题的判断。
@kimmonismus@kimmonismusAI 评分55 @kimmonismus@kimmonismusAI 评分6060 

@kimmonismus@kimmonismusAI 评分5252 
@kimmonismus@kimmonismusAI 评分6464 
@kimmonismus@kimmonismusAI 评分2828 



@kimmonismus@kimmonismusAI 评分55 
@kimmonismus@kimmonismusAI 评分2727 引用eric (@ericlim)@ericlimchadeks chad gippity subahapp 🔜
@kimmonismus@kimmonismusAI 评分1212 我有点困惑,同时又很兴奋。我感觉 OpenAI 正在为一些重大发布做准备。 超级应用?5.6?来吧! 视频

@kimmonismus@kimmonismus精选AI 评分8181
引用Google (@Google)@GoogleToday we’re introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run locally with just 16GB of VRAM. It’s open and accessible for everyone to use under a permissive Apache 2.0 license. This is all made possible by our new, unified architecture that removes separate multimodal encoders. Here’s how we did it 🧵
推荐理由:它把视觉与音频编码器并入主干,让 12B 模型能在 16GB 显存本地运行,读者可据此判断端侧多模态的门槛变化。
@kimmonismus@kimmonismusAI 评分5151 引用Chubby♨️ (@kimmonismus)@kimmonismusFirst hands-on with Microsoft’s new Surface Laptop Ultra. Microsoft is clearly positioning this as a new class of creator and AI laptop, powered by new NVIDIA silicon with an RTX GPU built for local AI, creative workflows, and gaming. A few standout specs: -New NVIDIA chip with RTX GPU -Up to 1 petaflop of AI compute -Up to 128GB unified memory -15-inch mini-LED PixelSense Ultra touchscreen -3:2 aspect ratio -262 PPI -Up to 2,000 nits peak HDR brightness -Less than 18mm thick Video Video Video
@kimmonismus@kimmonismusAI 评分1515 这大概就是 GPT-5.6。我猜要么明天,要么下周。 准备好,朋友们。我们要迎来一场狂野之旅了!
引用leo 🐾 (@synthwavedd)@synthwaveddmercury-alpha
@kimmonismus@kimmonismusAI 评分5252 引用Aoden Teo (@AodenTeoMT)@AodenTeoMTToday, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of latency. We’ve open-sourced the model weights, with API access coming soon. Hear how Miso One sounds in the thread below. Video
@kimmonismus@kimmonismusAI 评分55 
@kimmonismus@kimmonismusAI 评分1414 更多视频和用例,比如本地运行 Open Claw、blender 等等 Video Video Video



@kimmonismus@kimmonismusAI 评分5454 作者首次上手微软新款 Surface Laptop Ultra,微软将其定位为面向创作者与本地 AI 的新品类笔记本,搭载新 NVIDIA 芯片与 RTX GPU,用于本地 AI、创作流程和游戏。




@kimmonismus@kimmonismusAI 评分55 @kimmonismus@kimmonismusAI 评分5959 
@kimmonismus@kimmonismusAI 评分5656 引用elie (@eliebakouch)@eliebakouchmicrosoft MAI tech report is a gold mine, one of the most transparent for a model at this scale. this model uses zero synthetic data or distillation from previous models. this means reasoning, agentic behavior, tool use are all learned fully during post-training with no cold start. bold choice that makes it harder and requires more iterations to reach sota, but you get FULL control over your model series and it proves they are serious about being a frontier lab. the tech report is insanely detailed and precise about numbers. to give an example, they give the exact MFU across all the iterations of the model, with the exact changes etc. they also share the full scaling ladder recipe, to my knowledge this is the first time i've seen this in a tech report at this scale let's look at all of this in this likely very long thread 🧵
@kimmonismus@kimmonismusAI 评分55 @kimmonismus@kimmonismusAI 评分22 周环比 50% @kimmonismus:太疯狂了,不是吗?
引用Chubby♨️ (@kimmonismus)@kimmonismusInsane, isn’t it?
@kimmonismus@kimmonismusAI 评分5555 @kimmonismus@kimmonismusAI 评分4848 
@kimmonismus@kimmonismusAI 评分44 非常感谢你,@satyanadella,给我与你交流的机会!


@kimmonismus@kimmonismusAI 评分1919 关于 Opus/sonnet:据我了解,他们将其与 sonnet 4.6 对比。只有在 SWE pro 上与 opus 相当。如果我没理解错的话,原话是“与 sonnet 4.6 并驾齐驱”
@kimmonismus@kimmonismusAI 评分3131 刚发现 „Mai“-1 thinking 代表的是: Microsoft AI-thinking。 🤯
引用Chubby♨️ (@kimmonismus)@kimmonismusMai-1 thinking: Mid size model, 45b active parameter, MoE, side by side with sonnet 4.6 0 distillation „Microsoft’s first reasoning model“
@kimmonismus@kimmonismusAI 评分1919 等等什么?和 Gemini 3.1 pro 相同的训练 FLOPs?
引用swyx (@swyx)@swyxuhhh did Mustafa just leak the Mythos FLOP count?? was this public knowledge before, even if its an estimate i dont get what you gain out of this
@kimmonismus@kimmonismusAI 评分1010 „大家都讨厌 AI slop“ „我们要来决定:这是 vibe,还是 slop?“ 这活动听起来挺有意思 :D

@kimmonismus@kimmonismusAI 评分4141 在"no prior"直播播客中,也花了大量时间讨论社区和数据中心扩张。似乎存在真正巨大的阻力。这个话题占据了讨论的很大一部分。有人反复强调,数据中心扩张带来繁荣,并不会导致社区成本增加。
引用Chubby♨️ (@kimmonismus)@kimmonismusIt is interesting how much focus is being placed on data centers and the community. Recently, there were numerous reports regarding resistance to data center expansion; now comes the promise from Microsoft: no increase in electricity costs due to data centers, along with resource conservation.
@kimmonismus@kimmonismusAI 评分3030 非常期待这期“no prior”节目! 很想多了解他们的 Solaris 项目,他们的智能体掌机
引用Chubby♨️ (@kimmonismus)@kimmonismusThis came as a surprise: Microsoft has unveiled handheld and desktop devices designed to control one's agents. It reminds me of what I had expected from OpenAI’s hardware-standalone devices for controlling agents. Video
@kimmonismus@kimmonismusAI 评分5959
引用Chubby♨️ (@kimmonismus)@kimmonismusMustafa Suleyman, Microsoft AI: 7 new Microsoft Models, no end in sight when it comes to development, orders of magnitude in the next few years Video
@kimmonismus@kimmonismusAI 评分3535 Mustafa Suleyman,Microsoft AI:7 款新的 Microsoft 模型,开发看不到尽头,未来几年将有数量级的提升 Video

引用Chubby♨️ (@kimmonismus)@kimmonismusOpen claw windows companion app