跳到正文

X:Rohan Paul

@rohanpaul_ai · X

当前显示全部 AI 相关新闻
切换来源
全部X新闻X:Rohan Paul1273 条X:Kim1142 条X:阿易 AI Notes846 条X:阿里云 / Alibaba Cloud433 条X:Testing Catalog347 条X:Elvis Saravia325 条X:cb_doge308 条X:Elon Musk307 条X:OpenRouter288 条X:Artificial Analysis261 条X:Ethan Mollick261 条X:PixVerse237 条X:Replit215 条X:Alexandr Wang(Scale AI 创始人/Meta 首席 AI 官)205 条X:SemiAnalysis204 条X:ZHO203 条X:OpenAI Developers202 条X:小北171 条X:swyx161 条X:MiniMax157 条X:Gemini156 条X:Dex Horthy(HumanLayer)131 条X:Emad Mostaque122 条X:Tibo122 条X:OpenAI119 条X:X.PIN104 条X:Google AI for Developers103 条X:蚂蚁百灵101 条X:AI Safety Memes98 条X:Thomas Wolf(Hugging Face 联创/CSO)95 条X:Epoch AI93 条X:Luma AI90 条X:Runway88 条X:Claude Devs84 条X:Aravind Srinivas(Perplexity CEO)82 条X:Jason Liu82 条X:马东锡 NLP80 条X:面壁智能 OpenBMB79 条X:Frank Wang 玉伯79 条X:Sam Altman79 条X:洪明73 条X:fofr72 条X:Gabriel71 条X:阑夕67 条X:Francois Chollet66 条X:Nathan Lambert62 条X:Yuchen Jin62 条X:Perplexity61 条X:opencode60 条X:赵纯想59 条X:Eric Zakariasson59 条X:Greg Brockman58 条X:Cohere55 条X:Tencent WorkBuddy55 条X:Microsoft Research54 条X:Peter Steinberger51 条X:腾讯混元50 条X:Suno50 条X:Claude45 条X:Anthropic44 条X:Clément Delangue(Hugging Face CEO)44 条X:OpenClaw42 条X:Charlie Holtz41 条X:Krea AI41 条X:通义千问 / Qwen39 条X:Google DeepMind39 条X:Google AI37 条X:Thariq36 条X:karminski32 条X:AK31 条X:Peter McCrory(Anthropic 首席经济学家)31 条X:Boris Cherny30 条X:百度 Baidu29 条X:AI at Meta28 条X:Deedy Das28 条X:可灵 Kling AI27 条X:Aidan Gomez(Cohere CEO)27 条X:Viggle AI26 条X:Mustafa Suleyman(Microsoft AI CEO)25 条X:Logan Kilpatrick24 条X:Josh Woodward23 条X:Noam Brown20 条X:李继刚19 条X:硅基流动 SiliconFlow18 条X:Andrew Milich18 条X:Odyssey18 条X:SpaceXAI17 条X:Tianyi Cui17 条X:DeepSeek16 条X:华为云15 条X:Sundar Pichai14 条X:智谱 Z.ai12 条X:Mark Zuckerberg12 条X:Demis Hassabis11 条X:唐杰10 条X:Ammaar Reshi10 条X:Andrew Ng(DeepLearning.AI 创始人)10 条X:Eric Mitchell10 条X:Karina Nguyen10 条X:Lee Robinson10 条X:Mistral AI10 条X:Noah Zweben10 条X:Barret 李靖9 条X:Fei-Fei Li9 条X:Jensen Huang9 条X:Kimi.ai9 条X:Hao AI Lab8 条X:Jeff Dean8 条X:谢赛宁7 条X:Jim Fan7 条X:Michael Truell7 条X:張小珺 Xiaojùn6 条X:NotebookLM0 条@berryxia · 历史来源430 条@vista8 · 历史来源262 条@op7418 · 历史来源231 条@sensetime_ai · 历史来源34 条@google · 历史来源30 条@getsuperintel · 历史来源9 条@latentspacepod · 历史来源9 条@android · 历史来源8 条@dreamlabla · 历史来源8 条@mannybernabe · 历史来源8 条@karpathy · 历史来源7 条@alexatallah · 历史来源6 条@ryanleeminimax · 历史来源5 条@theo · 历史来源5 条PixVerse (@PixVerse) · 历史来源5 条@aidotengineer · 历史来源4 条@dkundel · 历史来源4 条@reach_vb · 历史来源4 条@dotey · 历史来源3 条@eliebakouch · 历史来源3 条@googlechrome · 历史来源3 条@kilocode · 历史来源3 条@maxforai · 历史来源3 条@newsfromgoogle · 历史来源3 条@richardssutton · 历史来源3 条@skylermiao7 · 历史来源3 条@victorsuortiz · 历史来源3 条@ajambrosino · 历史来源2 条@akashi203 · 历史来源2 条@anatolikopadze · 历史来源2 条@andrewcurran_ · 历史来源2 条@antirez · 历史来源2 条@barrnanas · 历史来源2 条@coreyching · 历史来源2 条@deanwball · 历史来源2 条@designarena · 历史来源2 条@fal · 历史来源2 条@fellmentke · 历史来源2 条@gergelyorosz · 历史来源2 条@gmi_cloud · 历史来源2 条@gravicle · 历史来源2 条@hxiao · 历史来源2 条@id_aa_carmack · 历史来源2 条@jackminong · 历史来源2 条@lennysan · 历史来源2 条@mada299 · 历史来源2 条@microsoft · 历史来源2 条@mikastars39 · 历史来源2 条@mitchellh · 历史来源2 条@nabeelqu · 历史来源2 条@rudrank · 历史来源2 条@sebastienbubeck · 历史来源2 条@zan2434 · 历史来源2 条@___harald___ · 历史来源1 条@_boraturan · 历史来源1 条@0xjaniak · 历史来源1 条@0xkato · 历史来源1 条@47fucb4r8c69323 · 历史来源1 条@559hkdt · 历史来源1 条@aaliya_va · 历史来源1 条@abhikatte42 · 历史来源1 条@abhishekpatiil · 历史来源1 条@aboutberlin · 历史来源1 条@addyosmani · 历史来源1 条@agi2asi · 历史来源1 条@aiaicreate · 历史来源1 条@aimlapi · 历史来源1 条@aisaonehq · 历史来源1 条@aisystemprompt · 历史来源1 条@alemtuzlak · 历史来源1 条@alexxubyte · 历史来源1 条@alupsasca · 历史来源1 条@amasad · 历史来源1 条@ampcode · 历史来源1 条@anas_build_ · 历史来源1 条@aniketmaurya · 历史来源1 条@anitakirkovska · 历史来源1 条@anneliesgamble · 历史来源1 条@antigravity · 历史来源1 条@arafatkatze · 历史来源1 条@arena · 历史来源1 条@argofowl · 历史来源1 条@arthurmensch · 历史来源1 条@ashiknewazaj · 历史来源1 条@atabarrok · 历史来源1 条@atomic_chat_hq · 历史来源1 条@awe_automation · 历史来源1 条@awesomekling · 历史来源1 条@awscloud · 历史来源1 条@ayushagarwal · 历史来源1 条@baaadas · 历史来源1 条@bai_agi · 历史来源1 条@bbuddha_xyz · 历史来源1 条@bclavie · 历史来源1 条@beccalytics · 历史来源1 条@benfleming__ · 历史来源1 条@benhylak · 历史来源1 条@benjamineyliu · 历史来源1 条@bfl_ml · 历史来源1 条@bleysg · 历史来源1 条@bolna_dev · 历史来源1 条@bosmeny · 历史来源1 条@boxmining · 历史来源1 条@bozhou_ai · 历史来源1 条@brexhq · 历史来源1 条@brian_armstrong · 历史来源1 条@brianchew · 历史来源1 条@bridgemindai · 历史来源1 条@budgetpixel · 历史来源1 条@cahidarda · 历史来源1 条@calmpromptshq · 历史来源1 条@ce_zhang · 历史来源1 条@cedric_chee · 历史来源1 条@chaitralikakde · 历史来源1 条@chatgpt · 历史来源1 条@chatgptapp · 历史来源1 条@christinetyip · 历史来源1 条@christofsalis · 历史来源1 条@clark__labs · 历史来源1 条@cloudflaredev · 历史来源1 条@cnorth_13 · 历史来源1 条@cnzoecomeback · 历史来源1 条@cocohearts · 历史来源1 条@code_star · 历史来源1 条@codebyaurelia · 历史来源1 条@cognition · 历史来源1 条@commandcodeai · 历史来源1 条@consensusnlp · 历史来源1 条@contralabs_ai · 历史来源1 条@cozyblaze265065 · 历史来源1 条@crimedecoder · 历史来源1 条@crtr0 · 历史来源1 条@damnventures · 历史来源1 条@daniellockyer · 历史来源1 条@darioamodei · 历史来源1 条@davidmaliglowka · 历史来源1 条@davidondrej1 · 历史来源1 条@davidsacks · 历史来源1 条@dbirker78883 · 历史来源1 条@deryatr_ · 历史来源1 条@devfun · 历史来源1 条@diegocabezas01 · 历史来源1 条@digitalocean · 历史来源1 条@dimillian · 历史来源1 条@dimitrispapail · 历史来源1 条@discussingfilm · 历史来源1 条@dkthomp · 历史来源1 条@dmitryrybin1 · 历史来源1 条@dmsobol · 历史来源1 条@douglasyaody · 历史来源1 条@duckduckgo · 历史来源1 条@easyrouterio · 历史来源1 条@edgardobriban · 历史来源1 条@eisokant · 历史来源1 条@elliotarledge · 历史来源1 条@encrypted · 历史来源1 条@endpointarena · 历史来源1 条@envato · 历史来源1 条@escanorreloaded · 历史来源1 条@esrtweet · 历史来源1 条@ethanhe_42 · 历史来源1 条@eu_commission · 历史来源1 条@fba · 历史来源1 条@fdavidsont · 历史来源1 条@figmaweave · 历史来源1 条@finn_meeks · 历史来源1 条@first_tree_ai · 历史来源1 条@flavioad · 历史来源1 条@flowith · 历史来源1 条@fminzhou · 历史来源1 条@freddie_spirit · 历史来源1 条@frydwia · 历史来源1 条@futurestacked · 历史来源1 条@garrettlord · 历史来源1 条@garrytan · 历史来源1 条@gavinsbaker · 历史来源1 条@GayaniFigma · 历史来源1 条@genspark_ai · 历史来源1 条@gitlawb · 历史来源1 条@gneubig · 历史来源1 条@gokulr · 历史来源1 条@goodfireai · 历史来源1 条@goodnesmbakara · 历史来源1 条@googleaistudio · 历史来源1 条@gordic_aleksa · 历史来源1 条@gro_tsen · 历史来源1 条@hangsiin · 历史来源1 条@happycapyai · 历史来源1 条@haydenbleasel · 历史来源1 条@helloiamleonie · 历史来源1 条@hey_asiif · 历史来源1 条@hilbertspaess · 历史来源1 条@howtoprompt__ · 历史来源1 条@hq4ai · 历史来源1 条@hypersoren · 历史来源1 条@ianbremmer · 历史来源1 条@interaction · 历史来源1 条@intology · 历史来源1 条@iron_redux · 历史来源1 条@ithilgore · 历史来源1 条@itsreallyvivek · 历史来源1 条@jamesjyu · 历史来源1 条@jameszmsun · 历史来源1 条@jason_young1231 · 历史来源1 条@jawad_rahman_ · 历史来源1 条@jaydendavisnc · 历史来源1 条@jeffbarg · 历史来源1 条@jenzhuscott · 历史来源1 条@jiayuan_jy · 历史来源1 条@jilles · 历史来源1 条@jimcramer · 历史来源1 条@jimsyoung_ · 历史来源1 条@jinjingliang · 历史来源1 条@jjacky · 历史来源1 条@jjackyliang · 历史来源1 条@joefioti · 历史来源1 条@joi___ai · 历史来源1 条@joinhandshake · 历史来源1 条@joinpursuit · 历史来源1 条@joulee · 历史来源1 条@jsconfasia · 历史来源1 条@jsrailton · 历史来源1 条@juminoz · 历史来源1 条@kaizero_ainta · 历史来源1 条@karanganesan · 历史来源1 条@kdaigle · 历史来源1 条@kentherogers · 历史来源1 条@kevinsays · 历史来源1 条@khudonogov · 历史来源1 条@koraykv · 历史来源1 条@kotekjedi_ml · 历史来源1 条@kuberwastaken · 历史来源1 条@kurz_gesagt · 历史来源1 条@kwindla · 历史来源1 条@lafalcemateo · 历史来源1 条@lakshyaaagrawal · 历史来源1 条@larrylv · 历史来源1 条@layoffai · 历史来源1 条@levinstanley · 历史来源1 条@lifeofjer · 历史来源1 条@livekit · 历史来源1 条@lostinlatencyx · 历史来源1 条@lotte_verheyden · 历史来源1 条@lqiao · 历史来源1 条@luciushq · 历史来源1 条@luckeyfaraday · 历史来源1 条@lukaspet · 历史来源1 条@madhavsinghal_ · 历史来源1 条@manassharmahere · 历史来源1 条@markiewagner · 历史来源1 条@marksaroufim · 历史来源1 条@marsxiang_ · 历史来源1 条@maseehg_ · 历史来源1 条@mattshumer_ · 历史来源1 条@mem0ai · 历史来源1 条@mengto · 历史来源1 条@merettm · 历史来源1 条@micahcarroll · 历史来源1 条@michael_chomsky · 历史来源1 条@michaelarnaldi · 历史来源1 条@microsoftai · 历史来源1 条@mike_acton · 历史来源1 条@mikeyyyzhao · 历史来源1 条@minchoi · 历史来源1 条@minimaxagent · 历史来源1 条@minu_who · 历史来源1 条@mkbhd · 历史来源1 条@modal · 历史来源1 条@moritzthuening · 历史来源1 条@moxie · 历史来源1 条@mstockton · 历史来源1 条@mtslive · 历史来源1 条@multimodalart · 历史来源1 条@neelnanda5 · 历史来源1 条@neilrahilly · 历史来源1 条@nickbaumann_ · 历史来源1 条@nirantk · 历史来源1 条@noemititarenco · 历史来源1 条@notjazii · 历史来源1 条@nousresearch · 历史来源1 条@oblomovius · 历史来源1 条@ollama · 历史来源1 条@onlyterp · 历史来源1 条@onlyzhynx · 历史来源1 条@organicgpt · 历史来源1 条@orgrem · 历史来源1 条@p0 · 历史来源1 条@palantirtech · 历史来源1 条@palmerluckey · 历史来源1 条@pandatalk8 · 历史来源1 条@parishilton · 历史来源1 条@patrickcarlyle · 历史来源1 条@patricktoulme · 历史来源1 条@paulg · 历史来源1 条@paulsolt · 历史来源1 条@pbdtokenrouter · 历史来源1 条@pererabinoy · 历史来源1 条@philhchen · 历史来源1 条@pirroh · 历史来源1 条@pjaccetturo · 历史来源1 条@postlive · 历史来源1 条@pranaveight · 历史来源1 条@prathamdby · 历史来源1 条@prince_canuma · 历史来源1 条@pumpkherm · 历史来源1 条@pvncher · 历史来源1 条@qiaoqiao2001 · 历史来源1 条@rajveerbach · 历史来源1 条@randyhaddad6 · 历史来源1 条@rauchg · 历史来源1 条@raveeshbhalla · 历史来源1 条@rayanpal_ · 历史来源1 条@rayfernando1337 · 历史来源1 条@redpoint · 历史来源1 条@ric_rtp · 历史来源1 条@richardsocher · 历史来源1 条@rileybrown · 历史来源1 条@robertvaradan · 历史来源1 条@ronshepherd · 历史来源1 条@rosmine · 历史来源1 条@rthiago · 历史来源1 条@ruben_kostard · 历史来源1 条@runware · 历史来源1 条@rvivek · 历史来源1 条@ryanjunejo · 历史来源1 条@safaricheung · 历史来源1 条@samuelstroschei · 历史来源1 条@sanmking · 历史来源1 条@saranormous · 历史来源1 条@savinovnikolay · 历史来源1 条@scale_ai · 历史来源1 条@scaling01 · 历史来源1 条@sdaily_ai · 历史来源1 条@secscottbessent · 历史来源1 条@seltaa_ · 历史来源1 条@sergiopaniego · 历史来源1 条@servasyy_ai · 历史来源1 条@sethltx · 历史来源1 条@shashankgoyal95 · 历史来源1 条@sherryyanjiang · 历史来源1 条@sherylhsu02 · 历史来源1 条@shl · 历史来源1 条@sighjith · 历史来源1 条@simistern · 历史来源1 条@southpkcommons · 历史来源1 条@sriramkri · 历史来源1 条@sshoaibali · 历史来源1 条@stalkermustang · 历史来源1 条@status_effects · 历史来源1 条@stevencheng · 历史来源1 条@stockanalystpro · 历史来源1 条@suekhim · 历史来源1 条@sultanalfardan · 历史来源1 条@suraj_sharma14 · 历史来源1 条@swisscheese4299 · 历史来源1 条@swmansion · 历史来源1 条@systematicls · 历史来源1 条@teksedge · 历史来源1 条@tftc21 · 历史来源1 条@theahmadosman · 历史来源1 条@themidasproj · 历史来源1 条@theonejvo · 历史来源1 条@therealadamg · 历史来源1 条@timsoulo · 历史来源1 条@tmuxvim · 历史来源1 条@tobi · 历史来源1 条@togethercompute · 历史来源1 条@trackernetwork · 历史来源1 条@trustkerneltech · 历史来源1 条@ttunguz · 历史来源1 条@tuhinchakr · 历史来源1 条@twistartups · 历史来源1 条@ubermenscchh · 历史来源1 条@udayan_w · 历史来源1 条@usefastlane · 历史来源1 条@uzyn · 历史来源1 条@valeriocapraro · 历史来源1 条@vasuman · 历史来源1 条@vdbergrianne · 历史来源1 条@vibeguessing · 历史来源1 条@victoriakimse · 历史来源1 条@victoriawu77 · 历史来源1 条@victortaelin · 历史来源1 条@vikaskansalhq · 历史来源1 条@volchika · 历史来源1 条@walden_yan · 历史来源1 条@warpdotdev · 历史来源1 条@waynesutton · 历史来源1 条@wesroth · 历史来源1 条@whosamberella · 历史来源1 条@xdinodeer · 历史来源1 条@xicilion · 历史来源1 条@xucian_ · 历史来源1 条@yacinemtb · 历史来源1 条@yaojingang · 历史来源1 条@yoheinakajima · 历史来源1 条@yongquanyq · 历史来源1 条@youtubejocoding · 历史来源1 条@yusufg · 历史来源1 条@zachbussey · 历史来源1 条@zeddotdev · 历史来源1 条@zeroxkyle · 历史来源1 条@zhenthebuilder · 历史来源1 条@zicohacks · 历史来源1 条@zixuanli_ · 历史来源1 条@zymazza · 历史来源1 条
1,273 条AI 相关新闻 · 最新在前
9月8日周二
  1. @rohanpaul_ai61

    微软 AI 的新论文研究了 LLM 智能体在长程多步任务中的退化,跨 9 个模型的测试显示成功率通常随依赖步数增加而下降。在 ToolQA 上,短程近乎完美的模型到 16 步时成功率降到 0-33%;上下文长度并非主因,缩短上下文反而让退化更严重。论文建议按实际工作流长度测试智能体、衡量单步可靠性,并在坏步骤影响后续之前加入检查或检查点。

  2. @rohanpaul_ai58

    芝加哥大学法学院与伯克利法学院正在限制 AI 在教学中的使用。芝加哥大学法学院将于今年秋季试点在一年级核心课程中禁用电子设备,把 AI 排除在基础教学之外;伯克利法学院 2026 年夏季生效的默认政策禁止 AI 用于构思、列提纲、起草、修改、翻译或编辑计入学分的作业。FT 指出,这一收紧发生在法律实务界加快采用 AI 的背景下。

  3. @rohanpaul_ai41

    研究者给智能体提供相同历史轨迹的两种形式——保留更多执行细节的 Workflow Memory 与蒸馏后的 SKILL.md,技能版本表现高出 6.06 个百分点。轨迹分析显示 65.7% 的技能案例通过程序性锚定生效,仅 4.5% 靠补充缺失知识。技能主要用于执行层面(先做什么、用哪些工具、验证什么、避免哪些错误),但用错场景或过于死板执行时仍会拖累表现。

9月7日周一
  1. @rohanpaul_ai45

    公开的前沿模型永远是一个延迟很久的 checkpoint。 顶级实验室已经生活在未来,领先 2/3 个 checkpoint。 OpenAI/Anthropic 可能已经在使用我们直到 2027 年中才能接触到的模型。

    引用@thsottiaux@thsottiaux

    Astra was probably our biggest competitive advantage while it wasn’t generally available. Since we’ve had it our productivity jumped so much that we shifted some of our plans 6 months ahead and will ship them at DevDay instead of mid next year.

  2. @rohanpaul_ai27

    个人 AI 的难点不在给模型更多上下文,而在判断哪些上下文该真正改变答案。Rohan Paul 指出,记忆排序不能只靠相似度,还需考虑时效性、当前任务,以及新记忆是否与旧记忆矛盾,有时检索到某条记忆也应决定不让它影响回答。他认为这正是 TodayAI 做个人 AI 的思路。

    引用@kimmonismus@kimmonismus

    An AI can complete a task perfectly and still spend your time on the wrong thing. Imagine an assistant preparing notes for tomorrow’s meeting. A message arrives that changes the deadline, and suddenly another project needs your attention first. The notes may be excellent. The assistant still needs to recognize that your priorities have changed. As AI takes on more work, these decisions become increasingly important. Which task should come first? When is an interruption justified? When should the system ask you? That’s the question behind @TodayAIofficial’s approach to personal AI: how can an agent learn what matters to a particular person and use that context throughout the day? Memory is part of the answer. Remembering a deadline helps. Connecting it to a promise you made last week, noticing that the plan has changed, and bringing it up while you can still act takes more. It also requires a way to correct the system. People change their minds. A preference from three months ago may no longer apply. An assistant should make its assumptions visible and let you update them. I think this is a useful direction for personal AI. There’s considerable value in software that can connect scattered information and help you decide where to focus. The test will be how well it handles an ordinary, messy day: catching the commitment you might miss, explaining why it needs attention, and leaving the decision with you.

  3. @rohanpaul_ai48

    我认为“时间跨度”正在成为讨论智能体能力更有用的方式之一。 在 OpenAI,随着任务变长,无需人工干预的成功率急剧下降,从 15 分钟以内任务的 86% 降至最长的 64-128 小时区间的仅约 16%,而需要干预的成功运行变得更加常见。 一个模型可以在局部极其强大,但在长执行轨迹上仍然不可靠。 不是 token。不是基准分数。系统能在人类必须介入之前对任务保持有效控制多久? 也许我们正在走向一种基准,类似 每个完成任务所消耗的人类分钟数。

    引用@rohanpaul_ai@rohanpaul_ai

    I think "time horizon" is becoming one of the more useful ways to talk about agent capability. At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common. A model can be extremely capable locally and still be unreliable over a long execution trajectory. Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene? maybe we are moving towards a benchmark, something like human minutes consumed per completed task.

  4. @rohanpaul_ai56

    Rohan Paul 认为「时间跨度」正成为衡量智能体能力更实用的指标,并引用 OpenAI 数据指出任务越长、无需人工干预的成功率越低,从 15 分钟以内的 86% 降至 64-128 小时区间的约 16%。他指出模型可以在局部表现很强,但在长执行轨迹上仍可能不可靠。他由此提出,或许可以转向类似「每完成一个任务消耗多少人工分钟」这样的基准。

    引用@rohanpaul_ai@rohanpaul_ai

    OpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that would take a skilled researcher quite few days. inside OpenAI research, agent runtime has already crossed human labor by a wide margin. 3.1-to-1 agent-to-human ratio “In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.” That ratio measures agent runtime rather than equivalent productivity, but it captures how deeply parallel agent work has entered OpenAI research.

  5. @rohanpaul_ai41

    研究发现编码模型互相编辑代码时倾向过度修改,而严格提示词无法稳定解决这一问题。研究者用"保持编辑最小但通过构建与测试"两个信号后训练 Olmo3 7B,CROCODIL 将各模型实现的编辑距离大致减半,同时提升构建与全测试通过率。该研究仅限 Rust 函数编辑,但提示混用编码模型的团队应基准测试跨模型编辑并衡量多余 diff 大小。

  6. @rohanpaul_ai64

    OpenAI 研究人员的实验速度已升至 2025 年基线的 1.6 倍,每位在职研究员能测试的想法明显增多。作者引用的内容称 OpenAI 已达到其“自动化研究实习生”里程碑,agent 与人类的工时比为 3.1 比 1,即按标准 8 小时工作日计,截至 8 月中旬研究组织每 1 个人类工作日对应 3.1 个 agent 工作日;该比例衡量的是 agent 运行时长而非等效产出。

    引用@rohanpaul_ai@rohanpaul_ai

    OpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that would take a skilled researcher quite few days. inside OpenAI research, agent runtime has already crossed human labor by a wide margin. 3.1-to-1 agent-to-human ratio “In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.” That ratio measures agent runtime rather than equivalent productivity, but it captures how deeply parallel agent work has entered OpenAI research.

  7. @rohanpaul_ai56

    OpenAI 公布的数据显示,公司工程师人均代码变更量约为 2025 年前基准的 7 倍,图表中该指标在 2026 年初明显抬升。引用内容称 OpenAI 已达成“自动化研究实习生”里程碑,即人类监督下可完成需熟练研究员数天完成的明确任务;截至 8 月中旬,研究组织每 1 个人工工作日对应 3.1 个 agent 工作日。该比例衡量的是 agent 运行时长,而非等效产出。

    引用@rohanpaul_ai@rohanpaul_ai

    OpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that would take a skilled researcher quite few days. inside OpenAI research, agent runtime has already crossed human labor by a wide margin. 3.1-to-1 agent-to-human ratio “In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.” That ratio measures agent runtime rather than equivalent productivity, but it captures how deeply parallel agent work has entered OpenAI research.