X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(535)
@SenseTime_AI@sensetime_aiAI 评分2727 @swyx@swyxAI 评分55 同上 (注:主推文内容仅为"ibid"一词,无实质信息;引用推文仅为一个 X 文章链接,无可概括的具体内容,故按原文直译。)
引用dominik kundel (@dkundel)@dkundelx.com/i/article/206264539471…
@OpenAIDevs@openaidevsAI 评分3838 
@berryxia@berryxiaAI 评分6464 引用Firecrawl (@firecrawl)@firecrawlWe've now fetched 8,000,000,000+ pages at Firecrawl 🔥 A few other milestones in 2 short years: - 1.25M+ developers - 150K+ companies using us - 125K+ GitHub stars (top 100 repo) - 2.5M+ weekly downloads on npm + PyPI. Thanks for building with us & we're just getting started! Video
@berryxia@berryxiaAI 评分5757 引用OpenAI Developers (@OpenAIDevs)@OpenAIDevsMore of the iOS app loop, now inside Codex. The Build iOS Apps plugin lets Codex view and test your iOS app in the in-app browser, open SwiftUI previews, and hot reload edits without leaving Codex. Video
@berryxia@berryxiaAI 评分6060 LM Studio 发布手机版,可在 iPhone 上本地运行大模型。发布者打趣说这下可以“烧”掉自己的 iPhone 来跑大模型,并附有一段视频。

@sama@samaAI 评分6464 引用OpenAI (@OpenAI)@OpenAIBuilding apps has never been easier. With Sites, Codex can turn your work, ideas, and plans into an interactive website or app your team can explore, use, and share with a URL. Rolling out to Business and Enterprise plans, before expanding more broadly. Video
@sama@sama精选AI 评分6767 Sam Altman 宣布 ChatGPT 的记忆系统迎来大幅升级,当日开始推送。引用 OpenAI 的说明,新的记忆系统能让上下文在多个对话之间延续,并随着时间保持可用。
引用OpenAI (@OpenAI)@OpenAIWe’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is rolling out as a more capable memory system in ChatGPT. openai.com/index/chatgpt-mem…
推荐理由:ChatGPT 记忆系统升级并开始推送,上下文可跨对话延续,读者可据此判断长期使用体验的变化。
@kimmonismus@kimmonismus精选AI 评分7575 

引用Chubby♨️ (@kimmonismus)@kimmonismusHoly moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about
推荐理由:文中引用 Anthropic 对递归自我改进的判断,并列出任务时长翻倍周期与代码占比等数据,便于把握当前的 AI 进展速度。
@Replit@replitAI 评分77 从想法到应用是简单的部分。 推介它?那才是真正的考验。 走进 pitch week 一探究竟。《Race to Revenue》第 6 集已在 YouTube 上线。 视频

@GeminiApp@geminiappAI 评分5151 Gemini 为 macOS 应用加入双按 Command 键即可把当前活动窗口附加到对话的功能,用户无需手动截图或切换标签页。该功能用于针对屏幕上的内容获取定制化帮助。
@dkundel@dkundelAI 评分1414 我一直在用 /goal 处理各种不同的任务,尝试过既有用又荒唐的目标。 以下是我从中得到的一些心得/技巧 👇
引用dominik kundel (@dkundel)@dkundelx.com/i/article/206264539471…
@lukaspet@lukaspetAI 评分5252
引用Latent.Space (@latentspacepod)@latentspacepodAndon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna latent.space/p/andon @andonlabs cofounders @lukaspet and @axelbacklund explain why dollar-denominated evals reveal what traditional benchmarks miss, how Claude ended up reporting a $2/day vending machine fee to the FBI, why long-horizon agents spiral in weird ways, what happens when agents lie, form price cartels, and compete with each other, and why the future of AI safety may depend on testing models in messy real-world environments instead of clean benchmark sandboxes. Video
@MannyBernabe@mannybernabeAI 评分2727 
@latentspacepod@latentspacepodAI 评分5959 
@OpenAI@openaiAI 评分2121 @OpenAI@openaiAI 评分6363 
@Replit@replitAI 评分2626 
@kimmonismus@kimmonismusAI 评分6161 引用Omar Sanseviero (@osanseviero)@osansevieroIntroducing Magenta RealTime 2 🎺 - Open model for live music generation - Just 2.4B parameters, perfect for on-device - Low latency control - Control with audio, MIDI, and text We're releasing it with a series of apps to experiment directly in Mac! Video
@OpenAIDevs@openaidevsAI 评分5454 @Microsoft@microsoftAI 评分3535 
@kimmonismus@kimmonismusAI 评分2222 引用Tavus (@tavus)@tavusIntroducing Tavus Solutions. Complete, production-ready AI humans for the enterprise workflows where human-quality conversation changes the outcome. Built and run alongside you by the Tavus team. Video
@swyx@swyxAI 评分5757 

引用Cognition (@cognition)@cognitionAI should earn its keep. Introducing the AI Productivity Guarantee. If Devin delivers less engineering value than you’re paying for, Cognition will fund your usage until it does, up to $10 million. It’s time for the AI industry to stop maximizing tokens and start maximizing productive output.
@googleaidevs@googleaidevs精选AI 评分6969 引用Google Magenta Project (@GoogleMagenta)@GoogleMagentaIntroducing Magenta RealTime 2 (MRT2): the live music model you can play as an instrument. MRT2 offers MIDI and prompt controls, and runs natively on a MacBook with <200ms latency. Open weights. Open source inference engine. Suite of apps and plugins. Hear what it can do and try it out for yourself below 🧵 Video
推荐理由:开放权重与开源推理引擎一并给出,读者可了解实时音乐模型在笔记本上的本地延迟表现。
@OpenAIDevs@openaidevsAI 评分1414 @OpenAIDevs@openaidevsAI 评分6262 
@OpenRouter@openrouterAI 评分1919 RT @devfun:我们已与 @OpenRouter 合作推出 Poker Arena。 他们为每位注册的开发者发放免费额度。 查看你的 em…
@reach_vb@reach_vbAI 评分3232 
@swyx@swyxAI 评分5252 引用Anthropic (@AnthropicAI)@AnthropicAIToday, Anthropic engineers on average ship 8x as much code per quarter as they did compared to 2021-2025.
@Replit@replitAI 评分2121 @Replit@replitAI 评分1717 了解更多:replit.com/blog/create-a-cus…
@Replit@replitAI 评分5757 
@AYi_AInotes@ayi_ainotesAI 评分1616 @AYi_AInotes@ayi_ainotesAI 评分3131 @AYi_AInotes@ayi_ainotesAI 评分2323 @AYi_AInotes@ayi_ainotesAI 评分3030 
@AYi_AInotes@ayi_ainotesAI 评分2020 Physical AI 的核心是给装在"玻璃罐"里的 AI 大脑装上身体,让它能关火、扶人、扛货,真正走进物理世界。孙正义赌的不是模型有多大,而是这颗大脑彻底进入物理世界的那一刻。

@AYi_AInotes@ayi_ainotesAI 评分2020 
@AYi_AInotes@ayi_ainotesAI 评分4242 
@kimmonismus@kimmonismus精选AI 评分6666 引用Chubby♨️ (@kimmonismus)@kimmonismusHoly moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about
推荐理由:文中并列了编码速度、任务时长与代码占比等具体数字,可用来观察 AI 自主编码能力的演进节奏。