跳到正文

X

关注 AI 研究者、开发者与机构的动态

当前显示全部 AI 相关新闻
按账号或来源筛选(535)
全部X新闻X:Rohan Paul1450 条X:Kim1216 条X:阿易 AI Notes917 条X:阿里云 / Alibaba Cloud465 条X:Testing Catalog410 条X:cb_doge356 条X:Elvis Saravia355 条X:Elon Musk338 条X:OpenRouter311 条X:Ethan Mollick289 条X:Artificial Analysis285 条X:PixVerse276 条X:Alexandr Wang(Scale AI 创始人/Meta 首席 AI 官)274 条X:ZHO236 条X:Replit226 条X:SemiAnalysis226 条X:OpenAI Developers212 条X:小北184 条X:swyx162 条X:MiniMax157 条X:Gemini156 条X:Dex Horthy(HumanLayer)146 条X:Emad Mostaque135 条X:Tibo130 条X:OpenAI125 条X:Epoch AI119 条X:AI Safety Memes115 条X:X.PIN114 条X:Google AI for Developers103 条X:Thomas Wolf(Hugging Face 联创/CSO)102 条X:蚂蚁百灵101 条X:Runway101 条X:面壁智能 OpenBMB92 条X:Luma AI91 条X:Claude Devs90 条X:Aravind Srinivas(Perplexity CEO)89 条X:马东锡 NLP87 条X:Jason Liu86 条X:Frank Wang 玉伯85 条X:洪明83 条X:Sam Altman81 条X:fofr78 条X:Gabriel77 条X:Perplexity72 条X:Yuchen Jin71 条X:阑夕70 条X:赵纯想70 条X:Nathan Lambert69 条X:Francois Chollet67 条X:Cohere65 条X:Tencent WorkBuddy63 条X:opencode62 条X:Greg Brockman61 条X:Eric Zakariasson60 条X:Claude56 条X:Krea AI56 条X:Microsoft Research54 条X:Peter Steinberger52 条X:Suno52 条X:腾讯混元51 条X:Anthropic49 条X:karminski48 条X:Clément Delangue(Hugging Face CEO)47 条X:Google DeepMind45 条X:Thariq45 条X:OpenClaw44 条X:商汤 SenseTime (@SenseTime_AI)42 条X:Charlie Holtz41 条X:通义千问 / Qwen39 条X:Google AI38 条X:AK36 条X:Boris Cherny34 条X:可灵 Kling AI32 条X:Peter McCrory(Anthropic 首席经济学家)31 条X:百度 Baidu30 条X:Deedy Das30 条X:AI at Meta28 条X:Aidan Gomez(Cohere CEO)27 条X:Josh Woodward26 条X:Mustafa Suleyman(Microsoft AI CEO)26 条X:Viggle AI26 条X:Logan Kilpatrick25 条X:Noam Brown21 条X:李继刚20 条X:Odyssey20 条X:Andrew Milich19 条X:硅基流动 SiliconFlow18 条X:华为云18 条X:SpaceXAI17 条X:Tianyi Cui17 条X:DeepSeek16 条X:Sundar Pichai14 条X:PixVerse (@PixVerse)13 条X:智谱 Z.ai12 条X:Demis Hassabis12 条X:Mark Zuckerberg12 条X:Mistral AI12 条X:ARC Prize (@arcprize)11 条X:fal (@fal)11 条X:唐杰10 条X:Ammaar Reshi10 条X:Andrew Ng(DeepLearning.AI 创始人)10 条X:Arena (@arena)10 条X:Barret 李靖10 条X:DAIR.AI (@dair_ai)10 条X:ElevenLabs (@elevenlabs)10 条X:Eric Mitchell10 条X:Karina Nguyen10 条X:Lee Robinson10 条X:Noah Zweben10 条X:Arthur Mensch(Mistral CEO) (@arthurmensch)9 条X:Fei-Fei Li9 条X:Jensen Huang9 条X:Kimi.ai9 条X:Philipp Schmid(Google DeepMind 开发者体验) (@_philschmid)9 条Tripo(官方 X)8 条X:卡兹克 (@Khazix0918)8 条X:Cognition (@cognition)8 条X:Cursor (@cursor_ai)8 条X:Figure AI (@Figure_robot)8 条X:Gemini Notebook (@Gemini_Notebook)8 条X:Georgi Gerganov(llama.cpp) (@ggerganov)8 条X:Hao AI Lab8 条X:Higgsfield AI (@higgsfield)8 条X:Jeff Dean8 条X:Manus (@ManusAI)8 条X:Meshy (@MeshyAI)8 条X:MiniMax Design (H3) (@Hailuo_AI)8 条X:谢赛宁7 条X:Jim Fan7 条X:Michael Truell7 条X:張小珺 Xiaojùn6 条X:Lisa Su(AMD CEO) (@LisaSu)6 条X:NotebookLM0 条@berryxia · 历史来源430 条@vista8 · 历史来源262 条@op7418 · 历史来源231 条@google · 历史来源30 条@getsuperintel · 历史来源9 条@latentspacepod · 历史来源9 条@android · 历史来源8 条@dreamlabla · 历史来源8 条@mannybernabe · 历史来源8 条@karpathy · 历史来源7 条@alexatallah · 历史来源6 条@ryanleeminimax · 历史来源5 条@theo · 历史来源5 条@aidotengineer · 历史来源4 条@dkundel · 历史来源4 条@reach_vb · 历史来源4 条@dotey · 历史来源3 条@eliebakouch · 历史来源3 条@googlechrome · 历史来源3 条@kilocode · 历史来源3 条@maxforai · 历史来源3 条@newsfromgoogle · 历史来源3 条@richardssutton · 历史来源3 条@skylermiao7 · 历史来源3 条@victorsuortiz · 历史来源3 条@ajambrosino · 历史来源2 条@akashi203 · 历史来源2 条@anatolikopadze · 历史来源2 条@andrewcurran_ · 历史来源2 条@antirez · 历史来源2 条@barrnanas · 历史来源2 条@coreyching · 历史来源2 条@deanwball · 历史来源2 条@designarena · 历史来源2 条@fellmentke · 历史来源2 条@gergelyorosz · 历史来源2 条@gmi_cloud · 历史来源2 条@gravicle · 历史来源2 条@hxiao · 历史来源2 条@id_aa_carmack · 历史来源2 条@jackminong · 历史来源2 条@lennysan · 历史来源2 条@mada299 · 历史来源2 条@microsoft · 历史来源2 条@mikastars39 · 历史来源2 条@mitchellh · 历史来源2 条@nabeelqu · 历史来源2 条@rudrank · 历史来源2 条@sebastienbubeck · 历史来源2 条@zan2434 · 历史来源2 条@___harald___ · 历史来源1 条@_boraturan · 历史来源1 条@0xjaniak · 历史来源1 条@0xkato · 历史来源1 条@47fucb4r8c69323 · 历史来源1 条@559hkdt · 历史来源1 条@aaliya_va · 历史来源1 条@abhikatte42 · 历史来源1 条@abhishekpatiil · 历史来源1 条@aboutberlin · 历史来源1 条@addyosmani · 历史来源1 条@agi2asi · 历史来源1 条@aiaicreate · 历史来源1 条@aimlapi · 历史来源1 条@aisaonehq · 历史来源1 条@aisystemprompt · 历史来源1 条@alemtuzlak · 历史来源1 条@alexxubyte · 历史来源1 条@alupsasca · 历史来源1 条@amasad · 历史来源1 条@ampcode · 历史来源1 条@anas_build_ · 历史来源1 条@aniketmaurya · 历史来源1 条@anitakirkovska · 历史来源1 条@anneliesgamble · 历史来源1 条@antigravity · 历史来源1 条@arafatkatze · 历史来源1 条@argofowl · 历史来源1 条@ashiknewazaj · 历史来源1 条@atabarrok · 历史来源1 条@atomic_chat_hq · 历史来源1 条@awe_automation · 历史来源1 条@awesomekling · 历史来源1 条@awscloud · 历史来源1 条@ayushagarwal · 历史来源1 条@baaadas · 历史来源1 条@bai_agi · 历史来源1 条@bbuddha_xyz · 历史来源1 条@bclavie · 历史来源1 条@beccalytics · 历史来源1 条@benfleming__ · 历史来源1 条@benhylak · 历史来源1 条@benjamineyliu · 历史来源1 条@bfl_ml · 历史来源1 条@bleysg · 历史来源1 条@bolna_dev · 历史来源1 条@bosmeny · 历史来源1 条@boxmining · 历史来源1 条@bozhou_ai · 历史来源1 条@brexhq · 历史来源1 条@brian_armstrong · 历史来源1 条@brianchew · 历史来源1 条@bridgemindai · 历史来源1 条@budgetpixel · 历史来源1 条@cahidarda · 历史来源1 条@calmpromptshq · 历史来源1 条@ce_zhang · 历史来源1 条@cedric_chee · 历史来源1 条@chaitralikakde · 历史来源1 条@chatgpt · 历史来源1 条@chatgptapp · 历史来源1 条@christinetyip · 历史来源1 条@christofsalis · 历史来源1 条@clark__labs · 历史来源1 条@cloudflaredev · 历史来源1 条@cnorth_13 · 历史来源1 条@cnzoecomeback · 历史来源1 条@cocohearts · 历史来源1 条@code_star · 历史来源1 条@codebyaurelia · 历史来源1 条@commandcodeai · 历史来源1 条@consensusnlp · 历史来源1 条@contralabs_ai · 历史来源1 条@cozyblaze265065 · 历史来源1 条@crimedecoder · 历史来源1 条@crtr0 · 历史来源1 条@damnventures · 历史来源1 条@daniellockyer · 历史来源1 条@darioamodei · 历史来源1 条@davidmaliglowka · 历史来源1 条@davidondrej1 · 历史来源1 条@davidsacks · 历史来源1 条@dbirker78883 · 历史来源1 条@deryatr_ · 历史来源1 条@devfun · 历史来源1 条@diegocabezas01 · 历史来源1 条@digitalocean · 历史来源1 条@dimillian · 历史来源1 条@dimitrispapail · 历史来源1 条@discussingfilm · 历史来源1 条@dkthomp · 历史来源1 条@dmitryrybin1 · 历史来源1 条@dmsobol · 历史来源1 条@douglasyaody · 历史来源1 条@duckduckgo · 历史来源1 条@easyrouterio · 历史来源1 条@edgardobriban · 历史来源1 条@eisokant · 历史来源1 条@elliotarledge · 历史来源1 条@encrypted · 历史来源1 条@endpointarena · 历史来源1 条@envato · 历史来源1 条@escanorreloaded · 历史来源1 条@esrtweet · 历史来源1 条@ethanhe_42 · 历史来源1 条@eu_commission · 历史来源1 条@fba · 历史来源1 条@fdavidsont · 历史来源1 条@figmaweave · 历史来源1 条@finn_meeks · 历史来源1 条@first_tree_ai · 历史来源1 条@flavioad · 历史来源1 条@flowith · 历史来源1 条@fminzhou · 历史来源1 条@freddie_spirit · 历史来源1 条@frydwia · 历史来源1 条@futurestacked · 历史来源1 条@garrettlord · 历史来源1 条@garrytan · 历史来源1 条@gavinsbaker · 历史来源1 条@GayaniFigma · 历史来源1 条@genspark_ai · 历史来源1 条@gitlawb · 历史来源1 条@gneubig · 历史来源1 条@gokulr · 历史来源1 条@goodfireai · 历史来源1 条@goodnesmbakara · 历史来源1 条@googleaistudio · 历史来源1 条@gordic_aleksa · 历史来源1 条@gro_tsen · 历史来源1 条@hangsiin · 历史来源1 条@happycapyai · 历史来源1 条@haydenbleasel · 历史来源1 条@helloiamleonie · 历史来源1 条@hey_asiif · 历史来源1 条@hilbertspaess · 历史来源1 条@howtoprompt__ · 历史来源1 条@hq4ai · 历史来源1 条@hypersoren · 历史来源1 条@ianbremmer · 历史来源1 条@interaction · 历史来源1 条@intology · 历史来源1 条@iron_redux · 历史来源1 条@ithilgore · 历史来源1 条@itsreallyvivek · 历史来源1 条@jamesjyu · 历史来源1 条@jameszmsun · 历史来源1 条@jason_young1231 · 历史来源1 条@jawad_rahman_ · 历史来源1 条@jaydendavisnc · 历史来源1 条@jeffbarg · 历史来源1 条@jenzhuscott · 历史来源1 条@jiayuan_jy · 历史来源1 条@jilles · 历史来源1 条@jimcramer · 历史来源1 条@jimsyoung_ · 历史来源1 条@jinjingliang · 历史来源1 条@jjacky · 历史来源1 条@jjackyliang · 历史来源1 条@joefioti · 历史来源1 条@joi___ai · 历史来源1 条@joinhandshake · 历史来源1 条@joinpursuit · 历史来源1 条@joulee · 历史来源1 条@jsconfasia · 历史来源1 条@jsrailton · 历史来源1 条@juminoz · 历史来源1 条@kaizero_ainta · 历史来源1 条@karanganesan · 历史来源1 条@kdaigle · 历史来源1 条@kentherogers · 历史来源1 条@kevinsays · 历史来源1 条@khudonogov · 历史来源1 条@koraykv · 历史来源1 条@kotekjedi_ml · 历史来源1 条@kuberwastaken · 历史来源1 条@kurz_gesagt · 历史来源1 条@kwindla · 历史来源1 条@lafalcemateo · 历史来源1 条@lakshyaaagrawal · 历史来源1 条@larrylv · 历史来源1 条@layoffai · 历史来源1 条@levinstanley · 历史来源1 条@lifeofjer · 历史来源1 条@livekit · 历史来源1 条@lostinlatencyx · 历史来源1 条@lotte_verheyden · 历史来源1 条@lqiao · 历史来源1 条@luciushq · 历史来源1 条@luckeyfaraday · 历史来源1 条@lukaspet · 历史来源1 条@madhavsinghal_ · 历史来源1 条@manassharmahere · 历史来源1 条@markiewagner · 历史来源1 条@marksaroufim · 历史来源1 条@marsxiang_ · 历史来源1 条@maseehg_ · 历史来源1 条@mattshumer_ · 历史来源1 条@mem0ai · 历史来源1 条@mengto · 历史来源1 条@merettm · 历史来源1 条@micahcarroll · 历史来源1 条@michael_chomsky · 历史来源1 条@michaelarnaldi · 历史来源1 条@microsoftai · 历史来源1 条@mike_acton · 历史来源1 条@mikeyyyzhao · 历史来源1 条@minchoi · 历史来源1 条@minimaxagent · 历史来源1 条@minu_who · 历史来源1 条@mkbhd · 历史来源1 条@modal · 历史来源1 条@moritzthuening · 历史来源1 条@moxie · 历史来源1 条@mstockton · 历史来源1 条@mtslive · 历史来源1 条@multimodalart · 历史来源1 条@neelnanda5 · 历史来源1 条@neilrahilly · 历史来源1 条@nickbaumann_ · 历史来源1 条@nirantk · 历史来源1 条@noemititarenco · 历史来源1 条@notjazii · 历史来源1 条@nousresearch · 历史来源1 条@oblomovius · 历史来源1 条@ollama · 历史来源1 条@onlyterp · 历史来源1 条@onlyzhynx · 历史来源1 条@organicgpt · 历史来源1 条@orgrem · 历史来源1 条@p0 · 历史来源1 条@palantirtech · 历史来源1 条@palmerluckey · 历史来源1 条@pandatalk8 · 历史来源1 条@parishilton · 历史来源1 条@patrickcarlyle · 历史来源1 条@patricktoulme · 历史来源1 条@paulg · 历史来源1 条@paulsolt · 历史来源1 条@pbdtokenrouter · 历史来源1 条@pererabinoy · 历史来源1 条@philhchen · 历史来源1 条@pirroh · 历史来源1 条@pjaccetturo · 历史来源1 条@postlive · 历史来源1 条@pranaveight · 历史来源1 条@prathamdby · 历史来源1 条@prince_canuma · 历史来源1 条@pumpkherm · 历史来源1 条@pvncher · 历史来源1 条@qiaoqiao2001 · 历史来源1 条@rajveerbach · 历史来源1 条@randyhaddad6 · 历史来源1 条@rauchg · 历史来源1 条@raveeshbhalla · 历史来源1 条@rayanpal_ · 历史来源1 条@rayfernando1337 · 历史来源1 条@redpoint · 历史来源1 条@ric_rtp · 历史来源1 条@richardsocher · 历史来源1 条@rileybrown · 历史来源1 条@robertvaradan · 历史来源1 条@ronshepherd · 历史来源1 条@rosmine · 历史来源1 条@rthiago · 历史来源1 条@ruben_kostard · 历史来源1 条@runware · 历史来源1 条@rvivek · 历史来源1 条@ryanjunejo · 历史来源1 条@safaricheung · 历史来源1 条@samuelstroschei · 历史来源1 条@sanmking · 历史来源1 条@saranormous · 历史来源1 条@savinovnikolay · 历史来源1 条@scale_ai · 历史来源1 条@scaling01 · 历史来源1 条@sdaily_ai · 历史来源1 条@secscottbessent · 历史来源1 条@seltaa_ · 历史来源1 条@sergiopaniego · 历史来源1 条@servasyy_ai · 历史来源1 条@sethltx · 历史来源1 条@shashankgoyal95 · 历史来源1 条@sherryyanjiang · 历史来源1 条@sherylhsu02 · 历史来源1 条@shl · 历史来源1 条@sighjith · 历史来源1 条@simistern · 历史来源1 条@southpkcommons · 历史来源1 条@sriramkri · 历史来源1 条@sshoaibali · 历史来源1 条@stalkermustang · 历史来源1 条@status_effects · 历史来源1 条@stevencheng · 历史来源1 条@stockanalystpro · 历史来源1 条@suekhim · 历史来源1 条@sultanalfardan · 历史来源1 条@suraj_sharma14 · 历史来源1 条@swisscheese4299 · 历史来源1 条@swmansion · 历史来源1 条@systematicls · 历史来源1 条@teksedge · 历史来源1 条@tftc21 · 历史来源1 条@theahmadosman · 历史来源1 条@themidasproj · 历史来源1 条@theonejvo · 历史来源1 条@therealadamg · 历史来源1 条@timsoulo · 历史来源1 条@tmuxvim · 历史来源1 条@tobi · 历史来源1 条@togethercompute · 历史来源1 条@trackernetwork · 历史来源1 条@trustkerneltech · 历史来源1 条@ttunguz · 历史来源1 条@tuhinchakr · 历史来源1 条@twistartups · 历史来源1 条@ubermenscchh · 历史来源1 条@udayan_w · 历史来源1 条@usefastlane · 历史来源1 条@uzyn · 历史来源1 条@valeriocapraro · 历史来源1 条@vasuman · 历史来源1 条@vdbergrianne · 历史来源1 条@vibeguessing · 历史来源1 条@victoriakimse · 历史来源1 条@victoriawu77 · 历史来源1 条@victortaelin · 历史来源1 条@vikaskansalhq · 历史来源1 条@volchika · 历史来源1 条@walden_yan · 历史来源1 条@warpdotdev · 历史来源1 条@waynesutton · 历史来源1 条@wesroth · 历史来源1 条@whosamberella · 历史来源1 条@xdinodeer · 历史来源1 条@xicilion · 历史来源1 条@xucian_ · 历史来源1 条@yacinemtb · 历史来源1 条@yaojingang · 历史来源1 条@yevr19 · 历史来源1 条@yoheinakajima · 历史来源1 条@yongquanyq · 历史来源1 条@youtubejocoding · 历史来源1 条@yusufg · 历史来源1 条@zachbussey · 历史来源1 条@zeddotdev · 历史来源1 条@zeroxkyle · 历史来源1 条@zhenthebuilder · 历史来源1 条@zicohacks · 历史来源1 条@zixuanli_ · 历史来源1 条@zymazza · 历史来源1 条
14,584 条AI 相关新闻 · 最新在前
6月5日周五
  1. @berryxia64

    Firecrawl公布两年内抓取超80亿网页,开发者超125万、企业用户超15万,GitHub星标超12.5万,npm和PyPI周下载量超250万次。作者据此认为,智能体的上限取决于能否稳定、持续、低成本获取真实世界的最新网页数据,AI竞争的焦点已从模型能力转向web上下文层。

    引用Firecrawl (@firecrawl)@firecrawl

    We've now fetched 8,000,000,000+ pages at Firecrawl 🔥 A few other milestones in 2 short years: - 1.25M+ developers - 150K+ companies using us - 125K+ GitHub stars (top 100 repo) - 2.5M+ weekly downloads on npm + PyPI. Thanks for building with us & we're just getting started! Video

  2. @berryxia57

    OpenAI Developers 上线 Build iOS Apps 插件,让 Codex 可在 in-app browser 中查看和测试 iOS 应用、打开 SwiftUI 预览,并在不离开 Codex 的情况下热重载编辑。

    引用OpenAI Developers (@OpenAIDevs)@OpenAIDevs

    More of the iOS app loop, now inside Codex. The Build iOS Apps plugin lets Codex view and test your iOS app in the in-app browser, open SwiftUI previews, and hot reload edits without leaving Codex. Video

  3. @sama64

    OpenAI 在 ChatGPT 中推出 Sites,Codex 可以把工作、想法和计划转成可交互的网站或应用,并用 URL 分享给团队使用。该功能先面向 Business 和 Enterprise 计划推出,之后会扩大范围。Sam Altman 转发时表示希望自己小时候就有这个工具,同时怀念 Hypercard。

    引用OpenAI (@OpenAI)@OpenAI

    Building apps has never been easier. With Sites, Codex can turn your work, ideas, and plans into an interactive website or app your team can explore, use, and share with a URL. Rolling out to Business and Enterprise plans, before expanding more broadly. Video

  4. @sama67

    Sam Altman 宣布 ChatGPT 的记忆系统迎来大幅升级,当日开始推送。引用 OpenAI 的说明,新的记忆系统能让上下文在多个对话之间延续,并随着时间保持可用。

    引用OpenAI (@OpenAI)@OpenAI

    We’ve been researching new ways for ChatGPT memory to carry context across conversations and keep it useful over time. Today, that work is rolling out as a more capable memory system in ChatGPT. openai.com/index/chatgpt-mem…

    推荐理由:ChatGPT 记忆系统升级并开始推送,上下文可跨对话延续,读者可据此判断长期使用体验的变化。

  5. @kimmonismus75

    Anthropic 在一篇博客中称 AI 进展快于其内部预期,模型能可靠独立完成的任务时长约每四个月翻倍,此前趋势为每七个月。博客提到 Claude 已编写 Anthropic 代码库中 80% 以上的合并代码,并描述 AI 自主设计后继模型的递归自我改进前景。X 用户 @kimmonismus 转述该文并认为,即便模型能力冻结在当前水平,社会仍会因现有模型扩散而出现重大变化。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:文中引用 Anthropic 对递归自我改进的判断,并列出任务时长翻倍周期与代码占比等数据,便于把握当前的 AI 进展速度。

  6. @lukaspet52

    Andon Labs 联合创始人 @lukaspet 转发并赞同“Reality is humanity's real last exam”的说法,表示每天为 AI 模型设计无法打败的测试都越来越难。他引用的 latentspacepod 节目介绍了 Andon Labs 的实景 AI 评测,包括以美元计价的评测为何能暴露传统基准遗漏之处、Claude 把 2 美元/天的自动售货机费用上报 FBI、长时程智能体如何以奇怪方式失控,以及智能体说谎、形成价格联盟并相互竞争等现象。

    引用Latent.Space (@latentspacepod)@latentspacepod

    Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna latent.space/p/andon @andonlabs cofounders @lukaspet and @axelbacklund explain why dollar-denominated evals reveal what traditional benchmarks miss, how Claude ended up reporting a $2/day vending machine fee to the FBI, why long-horizon agents spiral in weird ways, what happens when agents lie, form price cartels, and compete with each other, and why the future of AI safety may depend on testing models in messy real-world environments instead of clean benchmark sandboxes. Video

  7. @kimmonismus61

    Magenta RealTime 2 发布,这是一个面向实时音乐生成的开源模型,仅 2.4B 参数,适合端侧运行,支持低延迟控制,可通过音频、MIDI 和文本进行控制,并随附一系列可在 Mac 上直接体验的应用。转发该消息的作者认为这个模型很有创意,并提到在长途航班上也可以用它现场创作音乐。

    引用Omar Sanseviero (@osanseviero)@osanseviero

    Introducing Magenta RealTime 2 🎺 - Open model for live music generation - Just 2.4B parameters, perfect for on-device - Low latency control - Control with audio, MIDI, and text We're releasing it with a series of apps to experiment directly in Mac! Video

  8. @kimmonismus22

    这是企业 AI 的下一个重大突破:不是聊天机器人,而是嵌入真实工作流中的 AI 人类。 Tavus Solutions 将困难的部分抽象化:人设、对话设计、集成、调优和部署。 企业带来工作流。 Tavus 带来 AI 人类层。 从“构建 AI 基础设施”到“部署人类级 AI 界面”的转变。 感觉科幻正在变成现实。

    引用Tavus (@tavus)@tavus

    Introducing Tavus Solutions. Complete, production-ready AI humans for the enterprise workflows where human-quality conversation changes the outcome. Built and run alongside you by the Tavus team. Video

  9. @swyx57

    swyx 评论 Cognition 发布首个评测,其私有企业评测时长上限达 100 小时,而 METR 约 16 小时。METR 数据集为 7 名技术人员 34 场会话、rlog 0.83;Cognition 数据集为 126 名用户 258 场会话、留出集 rlog 0.74。

    引用Cognition (@cognition)@cognition

    AI should earn its keep. Introducing the AI Productivity Guarantee. If Devin delivers less engineering value than you’re paying for, Cognition will fund your usage until it does, up to $10 million. It’s time for the AI industry to stop maximizing tokens and start maximizing productive output.

  10. @googleaidevs69

    Google 发布开放权重的实时音乐模型 Magenta RealTime 2(MRT2),支持 MIDI 与提示词控制,可在 MacBook 上本地运行且延迟低于 200ms。官方同时提供开放权重、开源推理引擎及一系列应用和插件,演示中还展示了用 MIDI 键盘、实时文本提示乃至手势来演奏。

    引用Google Magenta Project (@GoogleMagenta)@GoogleMagenta

    Introducing Magenta RealTime 2 (MRT2): the live music model you can play as an instrument. MRT2 offers MIDI and prompt controls, and runs natively on a MacBook with <200ms latency. Open weights. Open source inference engine. Suite of apps and plugins. Hear what it can do and try it out for yourself below 🧵 Video

    推荐理由:开放权重与开源推理引擎一并给出,读者可了解实时音乐模型在笔记本上的本地延迟表现。

  11. @kimmonismus66

    Anthropic 发布博客探讨递归自我改进,称距离能完全自主设计并构建后继模型的 AI 已不远,但强调这尚未到来、也非必然,只是可能比多数机构预想的更早。文中引用数据称 Anthropic 工程师如今每季度交付代码量约为 2021–2025 年的 8 倍,AI 能可靠完成的任务时长约每 4 个月翻一倍,截至 2026 年 5 月 Claude 撰写了并入其代码库 80%+ 的代码。博客还给出三种未来路径,认为人类设定方向、效率持续复利提升是可能路径,而完全递归自我改进的对齐结果最不确定。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:文中并列了编码速度、任务时长与代码占比等具体数字,可用来观察 AI 自主编码能力的演进节奏。