跳到正文

X:Kim

@kimmonismus · X

当前显示全部 AI 相关新闻
切换来源
全部X新闻X:Rohan Paul1321 条X:Kim1179 条X:阿易 AI Notes872 条X:阿里云 / Alibaba Cloud441 条X:Testing Catalog372 条X:Elvis Saravia332 条X:cb_doge324 条X:Elon Musk322 条X:OpenRouter291 条X:Artificial Analysis275 条X:Ethan Mollick269 条X:PixVerse252 条X:Alexandr Wang(Scale AI 创始人/Meta 首席 AI 官)228 条X:Replit217 条X:ZHO214 条X:SemiAnalysis208 条X:OpenAI Developers204 条X:小北178 条X:swyx162 条X:MiniMax157 条X:Gemini156 条X:Dex Horthy(HumanLayer)136 条X:Emad Mostaque127 条X:Tibo124 条X:OpenAI119 条X:X.PIN105 条X:AI Safety Memes104 条X:Google AI for Developers103 条X:蚂蚁百灵101 条X:Thomas Wolf(Hugging Face 联创/CSO)96 条X:Epoch AI93 条X:Luma AI91 条X:Runway90 条X:Claude Devs86 条X:Jason Liu85 条X:马东锡 NLP84 条X:面壁智能 OpenBMB83 条X:Aravind Srinivas(Perplexity CEO)83 条X:Frank Wang 玉伯81 条X:Sam Altman80 条X:洪明75 条X:fofr74 条X:Gabriel74 条X:Perplexity69 条X:阑夕68 条X:Francois Chollet66 条X:Yuchen Jin65 条X:赵纯想64 条X:Nathan Lambert64 条X:Tencent WorkBuddy61 条X:Greg Brockman60 条X:opencode60 条X:Eric Zakariasson59 条X:Cohere56 条X:Microsoft Research54 条X:Peter Steinberger52 条X:腾讯混元51 条X:Suno50 条X:Claude47 条X:Clément Delangue(Hugging Face CEO)45 条X:Anthropic44 条X:Krea AI44 条X:OpenClaw43 条X:Charlie Holtz41 条X:Google DeepMind40 条X:通义千问 / Qwen39 条X:Google AI38 条X:Thariq38 条X:karminski35 条X:AK32 条X:Peter McCrory(Anthropic 首席经济学家)31 条X:可灵 Kling AI30 条X:Boris Cherny30 条X:百度 Baidu29 条X:Deedy Das29 条X:AI at Meta28 条X:Aidan Gomez(Cohere CEO)27 条X:Josh Woodward26 条X:Viggle AI26 条X:Logan Kilpatrick25 条X:Mustafa Suleyman(Microsoft AI CEO)25 条X:李继刚20 条X:Noam Brown20 条X:Odyssey20 条X:硅基流动 SiliconFlow18 条X:Andrew Milich18 条X:SpaceXAI17 条X:Tianyi Cui17 条X:华为云16 条X:DeepSeek16 条X:Sundar Pichai14 条X:智谱 Z.ai12 条X:Mark Zuckerberg12 条X:Demis Hassabis11 条X:唐杰10 条X:Ammaar Reshi10 条X:Andrew Ng(DeepLearning.AI 创始人)10 条X:Eric Mitchell10 条X:Karina Nguyen10 条X:Lee Robinson10 条X:Mistral AI10 条X:Noah Zweben10 条X:Barret 李靖9 条X:Fei-Fei Li9 条X:Jensen Huang9 条X:Kimi.ai9 条X:Hao AI Lab8 条X:Jeff Dean8 条X:谢赛宁7 条X:Jim Fan7 条X:Michael Truell7 条X:張小珺 Xiaojùn6 条X:NotebookLM0 条@berryxia · 历史来源430 条@vista8 · 历史来源262 条@op7418 · 历史来源231 条@sensetime_ai · 历史来源34 条@google · 历史来源30 条@getsuperintel · 历史来源9 条@latentspacepod · 历史来源9 条@android · 历史来源8 条@dreamlabla · 历史来源8 条@mannybernabe · 历史来源8 条@karpathy · 历史来源7 条@alexatallah · 历史来源6 条@ryanleeminimax · 历史来源5 条@theo · 历史来源5 条PixVerse (@PixVerse) · 历史来源5 条@aidotengineer · 历史来源4 条@dkundel · 历史来源4 条@reach_vb · 历史来源4 条@dotey · 历史来源3 条@eliebakouch · 历史来源3 条@googlechrome · 历史来源3 条@kilocode · 历史来源3 条@maxforai · 历史来源3 条@newsfromgoogle · 历史来源3 条@richardssutton · 历史来源3 条@skylermiao7 · 历史来源3 条@victorsuortiz · 历史来源3 条@ajambrosino · 历史来源2 条@akashi203 · 历史来源2 条@anatolikopadze · 历史来源2 条@andrewcurran_ · 历史来源2 条@antirez · 历史来源2 条@barrnanas · 历史来源2 条@coreyching · 历史来源2 条@deanwball · 历史来源2 条@designarena · 历史来源2 条@fal · 历史来源2 条@fellmentke · 历史来源2 条@gergelyorosz · 历史来源2 条@gmi_cloud · 历史来源2 条@gravicle · 历史来源2 条@hxiao · 历史来源2 条@id_aa_carmack · 历史来源2 条@jackminong · 历史来源2 条@lennysan · 历史来源2 条@mada299 · 历史来源2 条@microsoft · 历史来源2 条@mikastars39 · 历史来源2 条@mitchellh · 历史来源2 条@nabeelqu · 历史来源2 条@rudrank · 历史来源2 条@sebastienbubeck · 历史来源2 条@zan2434 · 历史来源2 条@___harald___ · 历史来源1 条@_boraturan · 历史来源1 条@0xjaniak · 历史来源1 条@0xkato · 历史来源1 条@47fucb4r8c69323 · 历史来源1 条@559hkdt · 历史来源1 条@aaliya_va · 历史来源1 条@abhikatte42 · 历史来源1 条@abhishekpatiil · 历史来源1 条@aboutberlin · 历史来源1 条@addyosmani · 历史来源1 条@agi2asi · 历史来源1 条@aiaicreate · 历史来源1 条@aimlapi · 历史来源1 条@aisaonehq · 历史来源1 条@aisystemprompt · 历史来源1 条@alemtuzlak · 历史来源1 条@alexxubyte · 历史来源1 条@alupsasca · 历史来源1 条@amasad · 历史来源1 条@ampcode · 历史来源1 条@anas_build_ · 历史来源1 条@aniketmaurya · 历史来源1 条@anitakirkovska · 历史来源1 条@anneliesgamble · 历史来源1 条@antigravity · 历史来源1 条@arafatkatze · 历史来源1 条@arena · 历史来源1 条@argofowl · 历史来源1 条@arthurmensch · 历史来源1 条@ashiknewazaj · 历史来源1 条@atabarrok · 历史来源1 条@atomic_chat_hq · 历史来源1 条@awe_automation · 历史来源1 条@awesomekling · 历史来源1 条@awscloud · 历史来源1 条@ayushagarwal · 历史来源1 条@baaadas · 历史来源1 条@bai_agi · 历史来源1 条@bbuddha_xyz · 历史来源1 条@bclavie · 历史来源1 条@beccalytics · 历史来源1 条@benfleming__ · 历史来源1 条@benhylak · 历史来源1 条@benjamineyliu · 历史来源1 条@bfl_ml · 历史来源1 条@bleysg · 历史来源1 条@bolna_dev · 历史来源1 条@bosmeny · 历史来源1 条@boxmining · 历史来源1 条@bozhou_ai · 历史来源1 条@brexhq · 历史来源1 条@brian_armstrong · 历史来源1 条@brianchew · 历史来源1 条@bridgemindai · 历史来源1 条@budgetpixel · 历史来源1 条@cahidarda · 历史来源1 条@calmpromptshq · 历史来源1 条@ce_zhang · 历史来源1 条@cedric_chee · 历史来源1 条@chaitralikakde · 历史来源1 条@chatgpt · 历史来源1 条@chatgptapp · 历史来源1 条@christinetyip · 历史来源1 条@christofsalis · 历史来源1 条@clark__labs · 历史来源1 条@cloudflaredev · 历史来源1 条@cnorth_13 · 历史来源1 条@cnzoecomeback · 历史来源1 条@cocohearts · 历史来源1 条@code_star · 历史来源1 条@codebyaurelia · 历史来源1 条@cognition · 历史来源1 条@commandcodeai · 历史来源1 条@consensusnlp · 历史来源1 条@contralabs_ai · 历史来源1 条@cozyblaze265065 · 历史来源1 条@crimedecoder · 历史来源1 条@crtr0 · 历史来源1 条@damnventures · 历史来源1 条@daniellockyer · 历史来源1 条@darioamodei · 历史来源1 条@davidmaliglowka · 历史来源1 条@davidondrej1 · 历史来源1 条@davidsacks · 历史来源1 条@dbirker78883 · 历史来源1 条@deryatr_ · 历史来源1 条@devfun · 历史来源1 条@diegocabezas01 · 历史来源1 条@digitalocean · 历史来源1 条@dimillian · 历史来源1 条@dimitrispapail · 历史来源1 条@discussingfilm · 历史来源1 条@dkthomp · 历史来源1 条@dmitryrybin1 · 历史来源1 条@dmsobol · 历史来源1 条@douglasyaody · 历史来源1 条@duckduckgo · 历史来源1 条@easyrouterio · 历史来源1 条@edgardobriban · 历史来源1 条@eisokant · 历史来源1 条@elliotarledge · 历史来源1 条@encrypted · 历史来源1 条@endpointarena · 历史来源1 条@envato · 历史来源1 条@escanorreloaded · 历史来源1 条@esrtweet · 历史来源1 条@ethanhe_42 · 历史来源1 条@eu_commission · 历史来源1 条@fba · 历史来源1 条@fdavidsont · 历史来源1 条@figmaweave · 历史来源1 条@finn_meeks · 历史来源1 条@first_tree_ai · 历史来源1 条@flavioad · 历史来源1 条@flowith · 历史来源1 条@fminzhou · 历史来源1 条@freddie_spirit · 历史来源1 条@frydwia · 历史来源1 条@futurestacked · 历史来源1 条@garrettlord · 历史来源1 条@garrytan · 历史来源1 条@gavinsbaker · 历史来源1 条@GayaniFigma · 历史来源1 条@genspark_ai · 历史来源1 条@gitlawb · 历史来源1 条@gneubig · 历史来源1 条@gokulr · 历史来源1 条@goodfireai · 历史来源1 条@goodnesmbakara · 历史来源1 条@googleaistudio · 历史来源1 条@gordic_aleksa · 历史来源1 条@gro_tsen · 历史来源1 条@hangsiin · 历史来源1 条@happycapyai · 历史来源1 条@haydenbleasel · 历史来源1 条@helloiamleonie · 历史来源1 条@hey_asiif · 历史来源1 条@hilbertspaess · 历史来源1 条@howtoprompt__ · 历史来源1 条@hq4ai · 历史来源1 条@hypersoren · 历史来源1 条@ianbremmer · 历史来源1 条@interaction · 历史来源1 条@intology · 历史来源1 条@iron_redux · 历史来源1 条@ithilgore · 历史来源1 条@itsreallyvivek · 历史来源1 条@jamesjyu · 历史来源1 条@jameszmsun · 历史来源1 条@jason_young1231 · 历史来源1 条@jawad_rahman_ · 历史来源1 条@jaydendavisnc · 历史来源1 条@jeffbarg · 历史来源1 条@jenzhuscott · 历史来源1 条@jiayuan_jy · 历史来源1 条@jilles · 历史来源1 条@jimcramer · 历史来源1 条@jimsyoung_ · 历史来源1 条@jinjingliang · 历史来源1 条@jjacky · 历史来源1 条@jjackyliang · 历史来源1 条@joefioti · 历史来源1 条@joi___ai · 历史来源1 条@joinhandshake · 历史来源1 条@joinpursuit · 历史来源1 条@joulee · 历史来源1 条@jsconfasia · 历史来源1 条@jsrailton · 历史来源1 条@juminoz · 历史来源1 条@kaizero_ainta · 历史来源1 条@karanganesan · 历史来源1 条@kdaigle · 历史来源1 条@kentherogers · 历史来源1 条@kevinsays · 历史来源1 条@khudonogov · 历史来源1 条@koraykv · 历史来源1 条@kotekjedi_ml · 历史来源1 条@kuberwastaken · 历史来源1 条@kurz_gesagt · 历史来源1 条@kwindla · 历史来源1 条@lafalcemateo · 历史来源1 条@lakshyaaagrawal · 历史来源1 条@larrylv · 历史来源1 条@layoffai · 历史来源1 条@levinstanley · 历史来源1 条@lifeofjer · 历史来源1 条@livekit · 历史来源1 条@lostinlatencyx · 历史来源1 条@lotte_verheyden · 历史来源1 条@lqiao · 历史来源1 条@luciushq · 历史来源1 条@luckeyfaraday · 历史来源1 条@lukaspet · 历史来源1 条@madhavsinghal_ · 历史来源1 条@manassharmahere · 历史来源1 条@markiewagner · 历史来源1 条@marksaroufim · 历史来源1 条@marsxiang_ · 历史来源1 条@maseehg_ · 历史来源1 条@mattshumer_ · 历史来源1 条@mem0ai · 历史来源1 条@mengto · 历史来源1 条@merettm · 历史来源1 条@micahcarroll · 历史来源1 条@michael_chomsky · 历史来源1 条@michaelarnaldi · 历史来源1 条@microsoftai · 历史来源1 条@mike_acton · 历史来源1 条@mikeyyyzhao · 历史来源1 条@minchoi · 历史来源1 条@minimaxagent · 历史来源1 条@minu_who · 历史来源1 条@mkbhd · 历史来源1 条@modal · 历史来源1 条@moritzthuening · 历史来源1 条@moxie · 历史来源1 条@mstockton · 历史来源1 条@mtslive · 历史来源1 条@multimodalart · 历史来源1 条@neelnanda5 · 历史来源1 条@neilrahilly · 历史来源1 条@nickbaumann_ · 历史来源1 条@nirantk · 历史来源1 条@noemititarenco · 历史来源1 条@notjazii · 历史来源1 条@nousresearch · 历史来源1 条@oblomovius · 历史来源1 条@ollama · 历史来源1 条@onlyterp · 历史来源1 条@onlyzhynx · 历史来源1 条@organicgpt · 历史来源1 条@orgrem · 历史来源1 条@p0 · 历史来源1 条@palantirtech · 历史来源1 条@palmerluckey · 历史来源1 条@pandatalk8 · 历史来源1 条@parishilton · 历史来源1 条@patrickcarlyle · 历史来源1 条@patricktoulme · 历史来源1 条@paulg · 历史来源1 条@paulsolt · 历史来源1 条@pbdtokenrouter · 历史来源1 条@pererabinoy · 历史来源1 条@philhchen · 历史来源1 条@pirroh · 历史来源1 条@pjaccetturo · 历史来源1 条@postlive · 历史来源1 条@pranaveight · 历史来源1 条@prathamdby · 历史来源1 条@prince_canuma · 历史来源1 条@pumpkherm · 历史来源1 条@pvncher · 历史来源1 条@qiaoqiao2001 · 历史来源1 条@rajveerbach · 历史来源1 条@randyhaddad6 · 历史来源1 条@rauchg · 历史来源1 条@raveeshbhalla · 历史来源1 条@rayanpal_ · 历史来源1 条@rayfernando1337 · 历史来源1 条@redpoint · 历史来源1 条@ric_rtp · 历史来源1 条@richardsocher · 历史来源1 条@rileybrown · 历史来源1 条@robertvaradan · 历史来源1 条@ronshepherd · 历史来源1 条@rosmine · 历史来源1 条@rthiago · 历史来源1 条@ruben_kostard · 历史来源1 条@runware · 历史来源1 条@rvivek · 历史来源1 条@ryanjunejo · 历史来源1 条@safaricheung · 历史来源1 条@samuelstroschei · 历史来源1 条@sanmking · 历史来源1 条@saranormous · 历史来源1 条@savinovnikolay · 历史来源1 条@scale_ai · 历史来源1 条@scaling01 · 历史来源1 条@sdaily_ai · 历史来源1 条@secscottbessent · 历史来源1 条@seltaa_ · 历史来源1 条@sergiopaniego · 历史来源1 条@servasyy_ai · 历史来源1 条@sethltx · 历史来源1 条@shashankgoyal95 · 历史来源1 条@sherryyanjiang · 历史来源1 条@sherylhsu02 · 历史来源1 条@shl · 历史来源1 条@sighjith · 历史来源1 条@simistern · 历史来源1 条@southpkcommons · 历史来源1 条@sriramkri · 历史来源1 条@sshoaibali · 历史来源1 条@stalkermustang · 历史来源1 条@status_effects · 历史来源1 条@stevencheng · 历史来源1 条@stockanalystpro · 历史来源1 条@suekhim · 历史来源1 条@sultanalfardan · 历史来源1 条@suraj_sharma14 · 历史来源1 条@swisscheese4299 · 历史来源1 条@swmansion · 历史来源1 条@systematicls · 历史来源1 条@teksedge · 历史来源1 条@tftc21 · 历史来源1 条@theahmadosman · 历史来源1 条@themidasproj · 历史来源1 条@theonejvo · 历史来源1 条@therealadamg · 历史来源1 条@timsoulo · 历史来源1 条@tmuxvim · 历史来源1 条@tobi · 历史来源1 条@togethercompute · 历史来源1 条@trackernetwork · 历史来源1 条@trustkerneltech · 历史来源1 条@ttunguz · 历史来源1 条@tuhinchakr · 历史来源1 条@twistartups · 历史来源1 条@ubermenscchh · 历史来源1 条@udayan_w · 历史来源1 条@usefastlane · 历史来源1 条@uzyn · 历史来源1 条@valeriocapraro · 历史来源1 条@vasuman · 历史来源1 条@vdbergrianne · 历史来源1 条@vibeguessing · 历史来源1 条@victoriakimse · 历史来源1 条@victoriawu77 · 历史来源1 条@victortaelin · 历史来源1 条@vikaskansalhq · 历史来源1 条@volchika · 历史来源1 条@walden_yan · 历史来源1 条@warpdotdev · 历史来源1 条@waynesutton · 历史来源1 条@wesroth · 历史来源1 条@whosamberella · 历史来源1 条@xdinodeer · 历史来源1 条@xicilion · 历史来源1 条@xucian_ · 历史来源1 条@yacinemtb · 历史来源1 条@yaojingang · 历史来源1 条@yevr19 · 历史来源1 条@yoheinakajima · 历史来源1 条@yongquanyq · 历史来源1 条@youtubejocoding · 历史来源1 条@yusufg · 历史来源1 条@zachbussey · 历史来源1 条@zeddotdev · 历史来源1 条@zeroxkyle · 历史来源1 条@zhenthebuilder · 历史来源1 条@zicohacks · 历史来源1 条@zixuanli_ · 历史来源1 条@zymazza · 历史来源1 条
1,179 条AI 相关新闻 · 最新在前
6月5日周五
  1. @kimmonismus37

    朋友们,准备好了。Anthropic 似乎正在准备发布其 Mythos 级模型。 定价:每 1M 输入 token $16 / 每 1M 输出 token $80。 发布可能非常临近,甚至可能与 GPT-5.6 同一周。竞争再次升温。 Gemini 3.5 Pro 即将面临严峻压力。最好是个狠货。

    引用sui ☄️ (@birdabo)@birdabo

    ‼️it seems Anthropic is ready to publicly launch a new version of Mythos, something better than Mythos Preview. a codenamed model “Oceanus” was given access to some red teamers yesterday according to @synthwavedd. it’s apparently been paused already, due to someone reselling access through a Chinese API proxy lmao 💀 Mythos pricing might also end up at with $16 Input, $80 Output according to @scaling01

  2. @kimmonismus68

    Anthropic 发布博客文章讨论递归自我改进(RSI),称距离能完全自主设计和构建后继模型的 AI 已不远,但强调这尚未实现也并非必然。文中数据包括 Anthropic 工程师如今每季度交付的代码量约为 2021 至 2025 年的 8 倍,截至 2026 年 5 月 Claude 编写了并入其代码库 80% 以上的代码,AI 能可靠完成的任务时长约每 4 个月翻一倍。文章提出三种未来路径,其中人类仍掌握方向的复合效率提升被视为最可能的走向。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Holy moly, Anthropic is getting very serious about recursive self-improvement! One word: acceleration. Insane blog article. Tl;dr: •We are close to an AI capable of fully autonomously designing and building its own successor •They stress this isn’t here yet and isn’t inevitable, but could arrive sooner than most institutions are ready for •Anthropic engineers now ship on average 8x as much code per quarter as they did in 2021–2025 •Task length AI can reliably complete is doubling roughly every 4 months (up from every 7 months) •Opus 3 (Mar 2024) handled ~4-minute tasks; Sonnet 3.7 (a year later) ~90-minute tasks; Opus 4.6 (a year after that) 12-hour tasks •SWE-bench went from low single digits to saturated in two years; CORE-bench (research reproduction) went ~20% to saturated in 15 months •METR found Claude Mythos Preview could work “at least” 16 hours, at the top of what they can currently measure •As of May 2026, Claude authored 80%+ of code merged into Anthropic’s codebase (low single digits before Claude Code launched in Feb 2025) •A March 2026 poll of 130 research staff: median respondent estimated ~4x output with Mythos Preview •One April 2026 example: Claude shipped 800+ fixes cutting a class of API errors 1,000x, work an engineer estimated would have taken a human four years •Claude-written code quality: worse than human in late 2025, roughly at parity now, expected to be strictly better within the year •On the hardest open-ended tasks, Claude’s success rate hit 76% in May 2026, up 50 points in six months •Code-speedup test: Opus 4 averaged ~3x speedup (May 2025), Mythos Preview ~52x (April 2026); a skilled human needs 4–8 hours to hit 4x •In an AI-safety research project, Claude agents recovered 97% of a performance gap (vs ~23% for two human researchers in a week), over 800 compute-hours and ~$18K •On picking the better “next step” in research sessions, the best model beat the human choice 51% (Nov 2025, Opus 4.5) rising to 64% (April 2026, Mythos Preview) •Human comparative advantage, for now: research taste and judgment, i.e. choosing which problems matter and when an approach is a dead end Three possible futures •The trend stalls (S-curve), but today’s capabilities still diffuse widely; they consider this least likely •Compounding efficiency gains, with humans still setting direction; 100-person firms doing the work of 10,000+; they think this is the likely path •Full recursive self-improvement, where AI builds its successors and pace is set by compute; the alignment outcome here is what they’re least certain about

    推荐理由:汇总了 Anthropic 博客关于递归自我改进的关键数据与三种未来路径,可据此判断自动化编码的推进速度。

  3. @kimmonismus73

    Anthropic 发布博客称其内部数据显示 Claude 正在加速 AI 研发,存在走向递归自我改进的可能,并强调这尚未到来、也并非必然。博客列举的指标包括:Anthropic 工程师每季度交付代码量约为 2021–2025 年平均水平的 8 倍,AI 能可靠完成的任务时长约每 4 个月翻倍(此前为每 7 个月)。

    引用Anthropic (@AnthropicAI)@AnthropicAI

    Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. anthropic.com/institute/recu…

    推荐理由:转述 Anthropic 内部数据,读者可据此了解递归自我改进讨论背后的具体加速指标。

6月4日周四
  1. @kimmonismus69

    Artificial Analysis 发布了对 NVIDIA Nemotron 3 Ultra 的评测,其智能指数为 47.7,领先 Gemma 4 31B 的 39.2、Nemotron 3 Super 的 36.0 和 gpt-oss-120b 的 33.3,但低于 Kimi K2.6 的 53.9。该模型约 5500 亿总参数、550 亿激活参数,评测使用 NVIDIA 推荐的 NVFP4 权重,比 BF16 测试小幅下降(48.2 对 47.7)。它在 BlackBox AI 上以每秒超 400 output tokens 的速度提供,略快于 gpt-oss-120b,但体量超过后者 4 倍以上。

    引用Artificial Analysis (@ArtificialAnlys)@ArtificialAnlys

    NVIDIA has just released Nemotron 3 Ultra, the new most intelligent US open weights model, with leading speed for its intelligence Nemotron 3 Ultra scores 47.7 on the Artificial Analysis Intelligence Index, well ahead of the next strongest US open weights models, Gemma 4 31B (39.2), Nemotron 3 Super (36.0) and gpt-oss-120b (33.3), but behind the Chinese-led open weights frontier (Kimi K2.6 at 53.9). We partnered with @NVIDIA to evaluate this model for intelligence and speed ahead of its public release. These figures use the final NVFP4 weights that NVIDIA recommends for inference, but our tests show minimal intelligence impact compared to BF16 testing, with higher precision resulting in an Artificial Analysis Intelligence Index score of 48.2 vs. the NVFP4 score of 47.7. Key Takeaways: ➤ Nemotron 3 Ultra leads in speed for its intelligence: through BlackBox AI ahead of release, Nemotron 3 Ultra is served at over 400 output tokens per second - this is slightly faster than the typical serving speed of gpt-oss-120b despite being >4X larger, and comes with significantly greater intelligence ➤ Largest Nemotron 3 model so far: with approximately 550 billion total parameters and 55 billion active, Nemotron 3 Ultra is significantly larger than its siblings and is the largest and most intelligent US open weights model release ever ➤ Nemotron 3 Ultra is the leading US open weights model on the Artificial Analysis Intelligence and Agentic Indexes by far, but Gemma 4 31B scores ~1 point higher on the Coding Index (comprised of Terminal-Bench Hard and SciCode)

    推荐理由:Artificial Analysis 的评测让读者能横向比较 Nemotron 3 Ultra 与美国及中国开源权重模型的智能与速度表现。

  2. @kimmonismus81

    NVIDIA 发布 Nemotron 3 Ultra,一款完全开源的 550B MoE 模型,激活参数 55B,权重、训练数据与完整配方全部公开。该模型采用混合 Mamba-Attention MoE 架构,NVIDIA 称其在长输出智能体任务上的吞吐量约为同类开源模型的 6 倍,同时保持相同准确率。

    引用NVIDIA AI (@NVIDIAAI)@NVIDIAAI

    Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models. Video

    推荐理由:原文给出 550B 开源模型的权重、训练数据与完整配方,并说明其在长任务智能体上的吞吐表现,读者可据此判断开源前沿模型的可复现程度。

  3. @kimmonismus72

    OpenAI 在文中表示,当前系统已出现递归自我改进(RSI)的早期迹象,即 AI 发展本身被 AI 加速。其预计这会加大开发者与国家之间的竞争压力,并带来现有机构难以应对的治理挑战。引用这段话的 @kimmonismus 评论称,氛围已经改变。

    推荐理由:原文引用 OpenAI 关于递归自我改进早期迹象的表述,可看到其对竞争压力与治理难题的判断。

  4. @kimmonismus52

    一项斯坦福牵头的盲评研究对近 3000 组匿名对比显示,16 所法学院的教授在合同法学生问答中有 75% 的情况下更偏好 AI 生成的答案,而非同行教授的答案。研究还显示,教授判断 AI 答案在教学上有害的可能性远低于人类答案,分别为 3.5% 与 12%。作者引用研究说明团队测试了包括商业辅导工具和 Google NotebookLM 在内的多种系统,并称可以想象模型在 6-12 个月后的表现。

  5. @kimmonismus81

    Google 发布 Gemma 4 12B 开源模型,采用 Apache 2.0 许可,可在 16GB 显存笔记本上本地运行,支持智能体推理、视觉与音频,作者称其质量接近 Google 的 26B 模型。

    引用Google (@Google)@Google

    Today we’re introducing Gemma 4 12B — our latest open model that brings advanced agentic reasoning, vision and audio directly to your laptop. It delivers performance nearing our larger Gemma models with a much smaller total memory footprint, while being small enough to run locally with just 16GB of VRAM. It’s open and accessible for everyone to use under a permissive Apache 2.0 license. This is all made possible by our new, unified architecture that removes separate multimodal encoders. Here’s how we did it 🧵

    推荐理由:它把视觉与音频编码器并入主干,让 12B 模型能在 16GB 显存本地运行,读者可据此判断端侧多模态的门槛变化。

  6. @kimmonismus51

    作者 @kimmonismus 在微软 behind-the-scenes 活动中现场查看了新款 Surface Laptop Ultra 并录制了视频。他认为微软意在直接对标 Apple、挑战 MacBook Pro,做工、散热、屏幕和 NVIDIA 芯片令人印象深刻,但他未做真实场景测试。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    First hands-on with Microsoft’s new Surface Laptop Ultra. Microsoft is clearly positioning this as a new class of creator and AI laptop, powered by new NVIDIA silicon with an RTX GPU built for local AI, creative workflows, and gaming. A few standout specs: -New NVIDIA chip with RTX GPU -Up to 1 petaflop of AI compute -Up to 128GB unified memory -15-inch mini-LED PixelSense Ultra touchscreen -3:2 aspect ratio -262 PPI -Up to 2,000 nits peak HDR brightness -Less than 18mm thick Video Video Video

  7. @kimmonismus52

    Miso One 发布,这是一个 8B 参数的开源权重文本转语音模型,支持从短样本一键克隆声音,延迟 110ms。模型权重已在 GitHub 免费开放,可自行托管,音频数据不必离开本机,也无需 API;官方在发布中表示 API 访问即将推出,并提供了可直接试听的 demo。

    引用Aoden Teo (@AodenTeoMT)@AodenTeoMT

    Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a human and responds faster than a human, with just 110 milliseconds of latency. We’ve open-sourced the model weights, with API access coming soon. Hear how Miso One sounds in the thread below. Video

6月3日周三
  1. @kimmonismus56

    微软 MAI 技术报告被解读为一款 1T 参数、35B 激活、在 33.5T token 上训练的模型,训练中未使用合成数据,也未从此前模型蒸馏。解读指出,该模型的推理、智能体行为与工具使用能力全部在后训练阶段习得,没有冷启动,这种做法难度更高、需要更多迭代才能达到 SOTA。报告还给出了各轮迭代的精确 MFU 以及完整的缩放配方(scaling ladder recipe)。

    引用elie (@eliebakouch)@eliebakouch

    microsoft MAI tech report is a gold mine, one of the most transparent for a model at this scale. this model uses zero synthetic data or distillation from previous models. this means reasoning, agentic behavior, tool use are all learned fully during post-training with no cold start. bold choice that makes it harder and requires more iterations to reach sota, but you get FULL control over your model series and it proves they are serious about being a frontier lab. the tech report is insanely detailed and precise about numbers. to give an example, they give the exact MFU across all the iterations of the model, with the exact changes etc. they also share the full scaling ladder recipe, to my knowledge this is the first time i've seen this in a tech report at this scale let's look at all of this in this likely very long thread 🧵

  2. @kimmonismus41

    在"no prior"直播播客中,也花了大量时间讨论社区和数据中心扩张。似乎存在真正巨大的阻力。这个话题占据了讨论的很大一部分。有人反复强调,数据中心扩张带来繁荣,并不会导致社区成本增加。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    It is interesting how much focus is being placed on data centers and the community. Recently, there were numerous reports regarding resistance to data center expansion; now comes the promise from Microsoft: no increase in electricity costs due to data centers, along with resource conservation.

  3. @kimmonismus59

    微软推出 MAI-1 thinking,作者称其为微软首个推理模型,中等规模、45B 激活参数、MoE 架构,且没有进行蒸馏,可与 Claude Sonnet 4.6 并列。作者引用的 Mustafa Suleyman 内容提到微软此次共发布 7 个新模型,并称未来几年仍会成倍推进开发。

    引用Chubby♨️ (@kimmonismus)@kimmonismus

    Mustafa Suleyman, Microsoft AI: 7 new Microsoft Models, no end in sight when it comes to development, orders of magnitude in the next few years Video