X:Kim
@kimmonismus · X
切换来源
@kimmonismus@kimmonismusAI 评分00 @kimmonismus@kimmonismusAI 评分3434 
@kimmonismus@kimmonismusAI 评分00 来源 teddit.net/r/macbookpro/s/b2…
@kimmonismus@kimmonismusAI 评分3838 苹果的 Touch Bar 生不逢时。想象一下它今天能有的那些惊人用例。 - 速率限制、上下文等等
引用Chubby♨️ (@kimmonismus)@kimmonismusTomorrow could be Apple’s most important AI moment yet. WWDC 2026 is expected to be all about one thing: making Siri relevant again. If the leaks are right, Apple is rebuilding Siri around a custom Google Gemini model, reportedly around 1.2 trillion parameters. For context: Apple’s own on-device AI model is roughly 3B parameters. The biggest rumor: Apple’s new Siri will reportedly be powered in the background by Google Gemini. Not as a Google-branded chatbot, but as an Apple-controlled intelligence layer running behind Siri, likely tied to Apple’s privacy-first infrastructure. So the new Siri likely becomes a hybrid system: • small Apple model locally on your device • large Gemini-class model in the cloud • Siri as the orchestration layer • Apple controlling the UI, app access and privacy layer What to further expect: • a much more conversational Siri • deeper personal context across apps, messages, files, calendar, photos and contacts • screen awareness • actions inside apps • a dedicated Siri app with chat history • voice chat, file uploads and multimodal interaction • better integration with Dynamic Island • optional support for other AI services like ChatGPT, Claude or Gemini Apple wants to turn Siri into the private AI layer of the operating system. A system agent that can search, understand, write, edit, summarize, organize and act across your iPhone, Mac and iPad. We may also see new Apple Intelligence features for: • AI photo editing • smarter Camera / Visual Intelligence • improved Writing Tools • natural-language Shortcuts • better Wallet and Health integrations • more privacy controls around AI data Either way, WWDC 2026 could define Apple’s position in the AI race. Exciting how the new CEO will handle all of this. Images: Bloomberg, Mark Gurman
@kimmonismus@kimmonismusAI 评分3737 


@kimmonismus@kimmonismusAI 评分77 @kimmonismus@kimmonismusAI 评分4848 
@kimmonismus@kimmonismusAI 评分3636 
@kimmonismus@kimmonismusAI 评分2323 我不知道有谁不对 Karpathy 怀有最高的敬意。这部短片再次展现了他是一位多么伟大的科学家。Anthropic 的一大胜利。 Video

@kimmonismus@kimmonismusAI 评分1010 @kimmonismus@kimmonismusAI 评分6161 美国犹他州一群居民与一家非营利组织就 Kevin O'Leary 计划在 Box Elder County 建设的 Stratos AI 数据中心园区起诉相关官员。

@kimmonismus@kimmonismusAI 评分00 @kimmonismus@kimmonismusAI 评分1717 我是说,我理解。员工不加薪——显然,除了 Anthropic 的人。

@kimmonismus@kimmonismusAI 评分33 真心难过,居然没几个人看懂这个梗。我还以为我们是个极客社区呢。:(
引用Chubby♨️ (@kimmonismus)@kimmonismusI had to buy it. Sadly the shirt arrived after Microsoft build.
@kimmonismus@kimmonismusAI 评分2626 Claude 5 Mythos 绝不会在 GPT-5.6 不同周发布的情况下单独发布。 我现在坚信下周就是发布周。
引用Chubby♨️ (@kimmonismus)@kimmonismusHoly, release is so close. It will be named „Claude Mythos 5“, a tier above Opus. I got the feeling coming week will be so huge
@kimmonismus@kimmonismusAI 评分1313 
@kimmonismus@kimmonismusAI 评分2828 天哪,发布真的近了。它将命名为"Claude Mythos 5",比 Opus 高一个层级。 我感觉下周会非常重磅。
引用Mirochill (@mirochill)@mirochill👀 Mythos est apparu quelques secondes chez Anthropic ! Son nom sera Claude Mythos 5 : La meilleure classe de modèle ne sera plus Opus, mais Mythos. Elle commencera à la version 5 de Claude. Vous pouvez retrouver des avant-premières sur mon profil !
@kimmonismus@kimmonismusAI 评分00 我不得不买下它。可惜这件T恤在Microsoft Build之后才到。

@kimmonismus@kimmonismusAI 评分3030 @kimmonismus@kimmonismusAI 评分4848 友情提醒一下:早在二月份,我们就已经有了首批“在创造自身过程中起到关键作用”的模型。 RSI 是一个已经持续了一段时间的进程。
引用Chubby♨️ (@kimmonismus)@kimmonismusOpenAI just wrote: "We also see early signs of recursive self-improvement (RSI) in today’s systems: where AI development is itself accelerated by AI. We expect this to increase competitive pressures among developers and nations, and create governance challenges that existing institutions are not equipped to address. As RSI emerges, societies will need ways to shape the trajectory of AI development and ensure that it serves human interests." The vibe has changed, something is happening.
@kimmonismus@kimmonismusAI 评分1414 引用Big Tech Alert (@BigTechAlert)@BigTechAlert🆕 @satyanadella has started following @kimmonismus
@kimmonismus@kimmonismusAI 评分3838
引用Markus J. Buehler (@ProfBuehlerMIT)@ProfBuehlerMITWe've made a breakthrough in self-evolving AI scientists moving from "search" to "principled discovery": Scientific discovery requires that the search space itself changes, and an AI scientist must perceive this shift without intervention. We built an AI that achieves this for the first time with the ability to discover the scientific vocabulary it reasons in. Evidence, tools, artifacts, verifiers, failures & claims become typed provenance. We show three distinct modalities: 1) retrieval, adding known objects; 2) search, exploring a fixed schema; and critically: 3) discovery, a verified regime transition. We solve the open-endedness evaluation problem by lifting agentic workflows into a typed copresheaf and proving, via a Kan obstruction, that true discovery is not unbounded generation but a verifiable schema expansion: old evidence is transported by Left Kan extension, and genuine novelty is mathematically quantified by the pointwise residual beyond the transported image - separating discovery from mere search and making novelty objective and measurable rather than a subjective judgment or benchmark delta. Our AI scientist is built in a way that does not pre-conceive the approach it chooses; instead, we endow the system with formal power to adapt, evolve, and reason from first principles. Case studies include: 1⃣Builder/Breaker model that discovers mode-conditioned compliance in proteins; 2⃣CategoryScienceClaw that finds anisotropic fiber-network stiffness rules. Great work in collaboration with my graduate student @fwang108_ @MITdeptofBE F.Y. Wang & M.J. Buehler, Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence, arXiv:2606.01444, 2026 Video
@kimmonismus@kimmonismusAI 评分6060 引用Moritz Wallawitsch (@MoritzW42)@MoritzW42holy shit - their api is leaking customer data
@kimmonismus@kimmonismus精选AI 评分6969 
推荐理由:剑桥团队把 AI 设计的超级抗原推进到人体试验阶段,免疫反应有限,但验证了这条路径可被测试。
@kimmonismus@kimmonismusAI 评分2323 @kimmonismus@kimmonismusAI 评分2626 来源 blog.google/innovation-and-a…
@kimmonismus@kimmonismusAI 评分6060 
@kimmonismus@kimmonismusAI 评分3232 引用🚨 AI News | TestingCatalog (@testingcatalog)@testingcatalogMYTHOS 🔥: Another early preview of recently spotted "Oceanus" checkpoint output. "Oceanus" is rumored to be a version of the upcoming Mythos model, which is planned for public release within "weeks", according to Anthropic. "Oceanus" prompt 👀 Video
@kimmonismus@kimmonismusAI 评分5050 引用Alex Kantrowitz (@Kantrowitz)@KantrowitzAI Pioneer Geoff Hinton tells me he believes AI is conscious.... and humans better get used to the idea that they're not the only intelligent life on earth. "They've very like us," he says. "They're beings like us." AI chatbots, he says, must understand your questions in order to answer them. There's an awareness there that equates to sentience. "We're going to have to accept that intelligence is not just biological." Video
@kimmonismus@kimmonismusAI 评分3434
引用Chubby♨️ (@kimmonismus)@kimmonismusI've read the comment several times now that this is IPO talk. And it's a fair comment. Yes, both OpenAI and Anthropic are currently talking about RSI. And yes, both are planning an IPO in 2026. A model like Mythos and an article about RSI appear at just the right time, which naturally makes it seem odd. But if you read through the noise and look at the evidence, you can see it. And at least the data that Anthropic provides suggests the validity of their thesis, at least based on what has been presented. At the same time, Dario Amodei started talking about RSI as early as 2024, saying he didn't consider it far-fetched, long before the IPO, and discussed it in his article "Machines of Loving Grace." Something similar happened with OpenAI. In short: it's not just empty talk, but has a valid basis, although real-world use cases will probably soon be demonstrated using this myth-like model, thus providing a more solid foundation for the debate. But I consider their statements to be more than just IPO rhetoric.
@kimmonismus@kimmonismusAI 评分2828 
@kimmonismus@kimmonismusAI 评分2929 
@kimmonismus@kimmonismusAI 评分2828 
@kimmonismus@kimmonismusAI 评分2424 
@kimmonismus@kimmonismusAI 评分3333 引用Chubby♨️ (@kimmonismus)@kimmonismusI believe the majority still doesn't understand the momentous threshold humanity is facing. Anthropic itself states quite clearly that even if development ceased entirely, if all development were frozen, they would still witness massive societal changes: "Even if model capabilities were frozen at today’s level, we would expect major changes to occur in the world. (...) And we are still early in the diffusion of today’s models into the wider economy, where a 100-person company can increasingly do the work of a 1,000-person one, because each employee will sit atop a pyramid of agents." But there's no question of stagnation. Anthropic itself still maintains that development has exceeded its own internal assumptions. Take that statement seriously for a second and consider it. Although Anthropic models internally and assumes exponential development, even this trajectory lags behind actual development, which is even faster. "It's happening faster than we thought, and the implications deserve greater attention." and "The rate at which AI models improve is accelerating. The length of tasks that they can reliably complete on their own has been doubling roughly every four months, up from an earlier trend of doubling every seven months. In March 2024, Claude Opus 3 could complete software tasks that take humans about four minutes to complete. A year later, Claude Sonnet 3.7 managed tasks that took about an hour and a half. A year after that, Claude Opus 4.6 managed 12-hour tasks.1 If this trend holds, tasks that take a skilled person days could come into range this year. So again: there can be no question of standing still. The models are not only getting better, they can also work autonomously for longer. Certainly numerous breakthroughs are still needed, context window is still a problem. But the most likely direction is that the models themselves will find the solutions to the underlying problems. This opens up unforeseen possibilities, and Demis Hassabi's statement that the golden age of science is not a dream, not a utopia, but a purposeful reality, is now confirmed. And finally, it's not just Anthropic, but also OpenAI, that sees this development, considers it feasible, and is moving forward. Most people don't know what's coming. But one thing is certain: it's coming even faster than expected. And it will be even bigger. Myth was just the beginning.
@kimmonismus@kimmonismusAI 评分2020 Claude Mythos 太强了。感谢 @Lentils80 看看这个 MacOS 输出。一次成型。 视频

@kimmonismus@kimmonismus精选AI 评分7575 

引用Chubby♨️ (@kimmonismus)@kimmonismusHoly moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about
推荐理由:文中引用 Anthropic 对递归自我改进的判断,并列出任务时长翻倍周期与代码占比等数据,便于把握当前的 AI 进展速度。
@kimmonismus@kimmonismusAI 评分6161 引用Omar Sanseviero (@osanseviero)@osansevieroIntroducing Magenta RealTime 2 🎺 - Open model for live music generation - Just 2.4B parameters, perfect for on-device - Low latency control - Control with audio, MIDI, and text We're releasing it with a series of apps to experiment directly in Mac! Video
@kimmonismus@kimmonismusAI 评分2222 引用Tavus (@tavus)@tavusIntroducing Tavus Solutions. Complete, production-ready AI humans for the enterprise workflows where human-quality conversation changes the outcome. Built and run alongside you by the Tavus team. Video
@kimmonismus@kimmonismus精选AI 评分6666 引用Chubby♨️ (@kimmonismus)@kimmonismusHoly moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about
推荐理由:文中并列了编码速度、任务时长与代码占比等具体数字,可用来观察 AI 自主编码能力的演进节奏。