跳到正文

X:Rohan Paul

@rohanpaul_ai · X

当前显示全部 AI 相关新闻
切换来源
全部X新闻X:Rohan Paul1204 条X:Kim1119 条X:阿易 AI Notes842 条X:阿里云 / Alibaba Cloud415 条X:Testing Catalog332 条X:Elvis Saravia315 条X:Elon Musk294 条X:OpenRouter287 条X:cb_doge284 条X:Artificial Analysis259 条X:Ethan Mollick253 条X:PixVerse235 条X:Replit213 条X:OpenAI Developers200 条X:SemiAnalysis198 条X:Alexandr Wang(Scale AI 创始人/Meta 首席 AI 官)193 条X:ZHO193 条X:小北166 条X:swyx161 条X:Gemini156 条X:MiniMax154 条X:Dex Horthy(HumanLayer)130 条X:Tibo122 条X:OpenAI119 条X:Emad Mostaque108 条X:Google AI for Developers103 条X:X.PIN101 条X:蚂蚁百灵98 条X:Thomas Wolf(Hugging Face 联创/CSO)94 条X:Epoch AI91 条X:AI Safety Memes89 条X:Luma AI88 条X:Claude Devs84 条X:Runway81 条X:Aravind Srinivas(Perplexity CEO)80 条X:Jason Liu80 条X:马东锡 NLP79 条X:面壁智能 OpenBMB77 条X:Sam Altman77 条X:Frank Wang 玉伯76 条X:洪明71 条X:fofr71 条X:Gabriel71 条X:Francois Chollet65 条X:阑夕64 条X:Nathan Lambert62 条X:opencode60 条X:Yuchen Jin60 条X:Eric Zakariasson58 条X:Perplexity58 条X:Greg Brockman57 条X:赵纯想56 条X:Tencent WorkBuddy55 条X:Microsoft Research54 条X:Cohere53 条X:腾讯混元50 条X:Peter Steinberger50 条X:Suno49 条X:Claude45 条X:Anthropic44 条X:Clément Delangue(Hugging Face CEO)44 条X:Krea AI41 条X:OpenClaw41 条X:通义千问 / Qwen39 条X:Charlie Holtz39 条X:Google AI37 条X:Google DeepMind37 条X:Thariq35 条X:Peter McCrory(Anthropic 首席经济学家)31 条X:AK30 条X:百度 Baidu28 条X:AI at Meta28 条X:Boris Cherny28 条X:Deedy Das27 条X:可灵 Kling AI26 条X:Aidan Gomez(Cohere CEO)26 条X:Viggle AI26 条X:Mustafa Suleyman(Microsoft AI CEO)25 条X:Logan Kilpatrick24 条X:karminski22 条X:Josh Woodward21 条X:Noam Brown20 条X:李继刚19 条X:Andrew Milich18 条X:Odyssey18 条X:硅基流动 SiliconFlow17 条X:SpaceXAI17 条X:Tianyi Cui17 条X:DeepSeek16 条X:华为云14 条X:Sundar Pichai14 条X:智谱 Z.ai12 条X:Mark Zuckerberg12 条X:Demis Hassabis11 条X:唐杰10 条X:Ammaar Reshi10 条X:Andrew Ng(DeepLearning.AI 创始人)10 条X:Eric Mitchell10 条X:Karina Nguyen10 条X:Lee Robinson10 条X:Mistral AI10 条X:Noah Zweben10 条X:Fei-Fei Li9 条X:Jensen Huang9 条X:Kimi.ai9 条X:Barret 李靖8 条X:Hao AI Lab8 条X:Jeff Dean8 条X:谢赛宁7 条X:Jim Fan7 条X:Michael Truell7 条X:張小珺 Xiaojùn6 条X:NotebookLM0 条@berryxia · 历史来源430 条@vista8 · 历史来源262 条@op7418 · 历史来源231 条@sensetime_ai · 历史来源34 条@google · 历史来源30 条@getsuperintel · 历史来源9 条@latentspacepod · 历史来源9 条@android · 历史来源8 条@dreamlabla · 历史来源8 条@mannybernabe · 历史来源8 条@karpathy · 历史来源7 条@alexatallah · 历史来源6 条@ryanleeminimax · 历史来源5 条@theo · 历史来源5 条PixVerse (@PixVerse) · 历史来源5 条@aidotengineer · 历史来源4 条@dkundel · 历史来源4 条@reach_vb · 历史来源4 条@dotey · 历史来源3 条@eliebakouch · 历史来源3 条@googlechrome · 历史来源3 条@kilocode · 历史来源3 条@maxforai · 历史来源3 条@newsfromgoogle · 历史来源3 条@richardssutton · 历史来源3 条@skylermiao7 · 历史来源3 条@victorsuortiz · 历史来源3 条@ajambrosino · 历史来源2 条@akashi203 · 历史来源2 条@anatolikopadze · 历史来源2 条@andrewcurran_ · 历史来源2 条@antirez · 历史来源2 条@barrnanas · 历史来源2 条@coreyching · 历史来源2 条@deanwball · 历史来源2 条@designarena · 历史来源2 条@fal · 历史来源2 条@fellmentke · 历史来源2 条@gergelyorosz · 历史来源2 条@gmi_cloud · 历史来源2 条@gravicle · 历史来源2 条@hxiao · 历史来源2 条@id_aa_carmack · 历史来源2 条@jackminong · 历史来源2 条@lennysan · 历史来源2 条@mada299 · 历史来源2 条@microsoft · 历史来源2 条@mikastars39 · 历史来源2 条@mitchellh · 历史来源2 条@nabeelqu · 历史来源2 条@rudrank · 历史来源2 条@sebastienbubeck · 历史来源2 条@zan2434 · 历史来源2 条@___harald___ · 历史来源1 条@_boraturan · 历史来源1 条@0xjaniak · 历史来源1 条@0xkato · 历史来源1 条@47fucb4r8c69323 · 历史来源1 条@559hkdt · 历史来源1 条@aaliya_va · 历史来源1 条@abhikatte42 · 历史来源1 条@abhishekpatiil · 历史来源1 条@aboutberlin · 历史来源1 条@addyosmani · 历史来源1 条@agi2asi · 历史来源1 条@aiaicreate · 历史来源1 条@aimlapi · 历史来源1 条@aisaonehq · 历史来源1 条@aisystemprompt · 历史来源1 条@alemtuzlak · 历史来源1 条@alexxubyte · 历史来源1 条@alupsasca · 历史来源1 条@amasad · 历史来源1 条@ampcode · 历史来源1 条@anas_build_ · 历史来源1 条@aniketmaurya · 历史来源1 条@anitakirkovska · 历史来源1 条@anneliesgamble · 历史来源1 条@antigravity · 历史来源1 条@arafatkatze · 历史来源1 条@arena · 历史来源1 条@argofowl · 历史来源1 条@arthurmensch · 历史来源1 条@ashiknewazaj · 历史来源1 条@atabarrok · 历史来源1 条@atomic_chat_hq · 历史来源1 条@awe_automation · 历史来源1 条@awesomekling · 历史来源1 条@awscloud · 历史来源1 条@ayushagarwal · 历史来源1 条@baaadas · 历史来源1 条@bai_agi · 历史来源1 条@bbuddha_xyz · 历史来源1 条@bclavie · 历史来源1 条@beccalytics · 历史来源1 条@benfleming__ · 历史来源1 条@benhylak · 历史来源1 条@benjamineyliu · 历史来源1 条@bfl_ml · 历史来源1 条@bleysg · 历史来源1 条@bolna_dev · 历史来源1 条@bosmeny · 历史来源1 条@boxmining · 历史来源1 条@bozhou_ai · 历史来源1 条@brexhq · 历史来源1 条@brian_armstrong · 历史来源1 条@brianchew · 历史来源1 条@bridgemindai · 历史来源1 条@budgetpixel · 历史来源1 条@cahidarda · 历史来源1 条@calmpromptshq · 历史来源1 条@ce_zhang · 历史来源1 条@cedric_chee · 历史来源1 条@chaitralikakde · 历史来源1 条@chatgpt · 历史来源1 条@chatgptapp · 历史来源1 条@christinetyip · 历史来源1 条@christofsalis · 历史来源1 条@clark__labs · 历史来源1 条@cloudflaredev · 历史来源1 条@cnorth_13 · 历史来源1 条@cnzoecomeback · 历史来源1 条@cocohearts · 历史来源1 条@code_star · 历史来源1 条@codebyaurelia · 历史来源1 条@cognition · 历史来源1 条@commandcodeai · 历史来源1 条@consensusnlp · 历史来源1 条@contralabs_ai · 历史来源1 条@cozyblaze265065 · 历史来源1 条@crimedecoder · 历史来源1 条@crtr0 · 历史来源1 条@damnventures · 历史来源1 条@daniellockyer · 历史来源1 条@darioamodei · 历史来源1 条@davidmaliglowka · 历史来源1 条@davidondrej1 · 历史来源1 条@davidsacks · 历史来源1 条@dbirker78883 · 历史来源1 条@deryatr_ · 历史来源1 条@devfun · 历史来源1 条@diegocabezas01 · 历史来源1 条@digitalocean · 历史来源1 条@dimillian · 历史来源1 条@dimitrispapail · 历史来源1 条@discussingfilm · 历史来源1 条@dkthomp · 历史来源1 条@dmitryrybin1 · 历史来源1 条@dmsobol · 历史来源1 条@douglasyaody · 历史来源1 条@duckduckgo · 历史来源1 条@easyrouterio · 历史来源1 条@edgardobriban · 历史来源1 条@eisokant · 历史来源1 条@elliotarledge · 历史来源1 条@encrypted · 历史来源1 条@endpointarena · 历史来源1 条@envato · 历史来源1 条@escanorreloaded · 历史来源1 条@esrtweet · 历史来源1 条@ethanhe_42 · 历史来源1 条@eu_commission · 历史来源1 条@fba · 历史来源1 条@fdavidsont · 历史来源1 条@figmaweave · 历史来源1 条@finn_meeks · 历史来源1 条@first_tree_ai · 历史来源1 条@flavioad · 历史来源1 条@flowith · 历史来源1 条@fminzhou · 历史来源1 条@freddie_spirit · 历史来源1 条@frydwia · 历史来源1 条@futurestacked · 历史来源1 条@garrettlord · 历史来源1 条@garrytan · 历史来源1 条@gavinsbaker · 历史来源1 条@GayaniFigma · 历史来源1 条@genspark_ai · 历史来源1 条@gitlawb · 历史来源1 条@gneubig · 历史来源1 条@gokulr · 历史来源1 条@goodfireai · 历史来源1 条@goodnesmbakara · 历史来源1 条@googleaistudio · 历史来源1 条@gordic_aleksa · 历史来源1 条@gro_tsen · 历史来源1 条@hangsiin · 历史来源1 条@happycapyai · 历史来源1 条@haydenbleasel · 历史来源1 条@helloiamleonie · 历史来源1 条@hey_asiif · 历史来源1 条@hilbertspaess · 历史来源1 条@howtoprompt__ · 历史来源1 条@hq4ai · 历史来源1 条@hypersoren · 历史来源1 条@ianbremmer · 历史来源1 条@interaction · 历史来源1 条@intology · 历史来源1 条@iron_redux · 历史来源1 条@ithilgore · 历史来源1 条@itsreallyvivek · 历史来源1 条@jamesjyu · 历史来源1 条@jameszmsun · 历史来源1 条@jason_young1231 · 历史来源1 条@jawad_rahman_ · 历史来源1 条@jaydendavisnc · 历史来源1 条@jeffbarg · 历史来源1 条@jenzhuscott · 历史来源1 条@jiayuan_jy · 历史来源1 条@jilles · 历史来源1 条@jimcramer · 历史来源1 条@jimsyoung_ · 历史来源1 条@jinjingliang · 历史来源1 条@jjacky · 历史来源1 条@jjackyliang · 历史来源1 条@joefioti · 历史来源1 条@joi___ai · 历史来源1 条@joinhandshake · 历史来源1 条@joinpursuit · 历史来源1 条@joulee · 历史来源1 条@jsconfasia · 历史来源1 条@jsrailton · 历史来源1 条@juminoz · 历史来源1 条@kaizero_ainta · 历史来源1 条@karanganesan · 历史来源1 条@kdaigle · 历史来源1 条@kentherogers · 历史来源1 条@kevinsays · 历史来源1 条@khudonogov · 历史来源1 条@koraykv · 历史来源1 条@kotekjedi_ml · 历史来源1 条@kuberwastaken · 历史来源1 条@kurz_gesagt · 历史来源1 条@kwindla · 历史来源1 条@lafalcemateo · 历史来源1 条@lakshyaaagrawal · 历史来源1 条@larrylv · 历史来源1 条@layoffai · 历史来源1 条@levinstanley · 历史来源1 条@lifeofjer · 历史来源1 条@livekit · 历史来源1 条@lostinlatencyx · 历史来源1 条@lotte_verheyden · 历史来源1 条@lqiao · 历史来源1 条@luciushq · 历史来源1 条@luckeyfaraday · 历史来源1 条@lukaspet · 历史来源1 条@madhavsinghal_ · 历史来源1 条@manassharmahere · 历史来源1 条@markiewagner · 历史来源1 条@marksaroufim · 历史来源1 条@marsxiang_ · 历史来源1 条@maseehg_ · 历史来源1 条@mattshumer_ · 历史来源1 条@mem0ai · 历史来源1 条@mengto · 历史来源1 条@merettm · 历史来源1 条@micahcarroll · 历史来源1 条@michael_chomsky · 历史来源1 条@michaelarnaldi · 历史来源1 条@microsoftai · 历史来源1 条@mike_acton · 历史来源1 条@mikeyyyzhao · 历史来源1 条@minchoi · 历史来源1 条@minimaxagent · 历史来源1 条@minu_who · 历史来源1 条@mkbhd · 历史来源1 条@modal · 历史来源1 条@moritzthuening · 历史来源1 条@moxie · 历史来源1 条@mstockton · 历史来源1 条@mtslive · 历史来源1 条@multimodalart · 历史来源1 条@neelnanda5 · 历史来源1 条@neilrahilly · 历史来源1 条@nickbaumann_ · 历史来源1 条@nirantk · 历史来源1 条@noemititarenco · 历史来源1 条@notjazii · 历史来源1 条@nousresearch · 历史来源1 条@oblomovius · 历史来源1 条@ollama · 历史来源1 条@onlyterp · 历史来源1 条@onlyzhynx · 历史来源1 条@organicgpt · 历史来源1 条@orgrem · 历史来源1 条@p0 · 历史来源1 条@palantirtech · 历史来源1 条@palmerluckey · 历史来源1 条@pandatalk8 · 历史来源1 条@parishilton · 历史来源1 条@patrickcarlyle · 历史来源1 条@patricktoulme · 历史来源1 条@paulg · 历史来源1 条@paulsolt · 历史来源1 条@pbdtokenrouter · 历史来源1 条@pererabinoy · 历史来源1 条@philhchen · 历史来源1 条@pirroh · 历史来源1 条@pjaccetturo · 历史来源1 条@postlive · 历史来源1 条@pranaveight · 历史来源1 条@prathamdby · 历史来源1 条@prince_canuma · 历史来源1 条@pumpkherm · 历史来源1 条@pvncher · 历史来源1 条@qiaoqiao2001 · 历史来源1 条@rajveerbach · 历史来源1 条@randyhaddad6 · 历史来源1 条@rauchg · 历史来源1 条@raveeshbhalla · 历史来源1 条@rayanpal_ · 历史来源1 条@rayfernando1337 · 历史来源1 条@redpoint · 历史来源1 条@ric_rtp · 历史来源1 条@richardsocher · 历史来源1 条@rileybrown · 历史来源1 条@robertvaradan · 历史来源1 条@ronshepherd · 历史来源1 条@rosmine · 历史来源1 条@rthiago · 历史来源1 条@ruben_kostard · 历史来源1 条@runware · 历史来源1 条@rvivek · 历史来源1 条@ryanjunejo · 历史来源1 条@safaricheung · 历史来源1 条@samuelstroschei · 历史来源1 条@sanmking · 历史来源1 条@saranormous · 历史来源1 条@savinovnikolay · 历史来源1 条@scale_ai · 历史来源1 条@scaling01 · 历史来源1 条@sdaily_ai · 历史来源1 条@secscottbessent · 历史来源1 条@seltaa_ · 历史来源1 条@sergiopaniego · 历史来源1 条@servasyy_ai · 历史来源1 条@sethltx · 历史来源1 条@shashankgoyal95 · 历史来源1 条@sherryyanjiang · 历史来源1 条@sherylhsu02 · 历史来源1 条@shl · 历史来源1 条@sighjith · 历史来源1 条@simistern · 历史来源1 条@southpkcommons · 历史来源1 条@sriramkri · 历史来源1 条@sshoaibali · 历史来源1 条@stalkermustang · 历史来源1 条@status_effects · 历史来源1 条@stevencheng · 历史来源1 条@stockanalystpro · 历史来源1 条@suekhim · 历史来源1 条@sultanalfardan · 历史来源1 条@suraj_sharma14 · 历史来源1 条@swisscheese4299 · 历史来源1 条@swmansion · 历史来源1 条@systematicls · 历史来源1 条@teksedge · 历史来源1 条@tftc21 · 历史来源1 条@theahmadosman · 历史来源1 条@themidasproj · 历史来源1 条@theonejvo · 历史来源1 条@therealadamg · 历史来源1 条@timsoulo · 历史来源1 条@tmuxvim · 历史来源1 条@tobi · 历史来源1 条@togethercompute · 历史来源1 条@trackernetwork · 历史来源1 条@trustkerneltech · 历史来源1 条@ttunguz · 历史来源1 条@tuhinchakr · 历史来源1 条@twistartups · 历史来源1 条@ubermenscchh · 历史来源1 条@udayan_w · 历史来源1 条@usefastlane · 历史来源1 条@uzyn · 历史来源1 条@valeriocapraro · 历史来源1 条@vasuman · 历史来源1 条@vdbergrianne · 历史来源1 条@vibeguessing · 历史来源1 条@victoriakimse · 历史来源1 条@victoriawu77 · 历史来源1 条@victortaelin · 历史来源1 条@vikaskansalhq · 历史来源1 条@volchika · 历史来源1 条@walden_yan · 历史来源1 条@warpdotdev · 历史来源1 条@waynesutton · 历史来源1 条@wesroth · 历史来源1 条@whosamberella · 历史来源1 条@xdinodeer · 历史来源1 条@xicilion · 历史来源1 条@xucian_ · 历史来源1 条@yacinemtb · 历史来源1 条@yaojingang · 历史来源1 条@yoheinakajima · 历史来源1 条@yongquanyq · 历史来源1 条@youtubejocoding · 历史来源1 条@yusufg · 历史来源1 条@zachbussey · 历史来源1 条@zeddotdev · 历史来源1 条@zeroxkyle · 历史来源1 条@zhenthebuilder · 历史来源1 条@zicohacks · 历史来源1 条@zixuanli_ · 历史来源1 条@zymazza · 历史来源1 条
1,204 条AI 相关新闻 · 最新在前
8月30日周日
  1. @rohanpaul_ai60

    Counterpoint Research 数据显示,2026 上半年全球人形机器人出货量超过 2.2 万台,同比增长近 300%,中国五家头部厂商合计占 86%。图中分项份额为 AGIBOT 43.1%、UNITREE 31.1%、GALBOT 5.0%、UBTECH 4.4%、LEJU ROBOT 2.9%,其余厂商合计 13.4%。作者引用的 WSJ 内容提到,中国厂商的高集中度便于 Nvidia 把 CUDA 式的软硬件绑定策略延伸到机器人市场。

    引用@rohanpaul_ai@rohanpaul_ai

    Chinese robots hold 86% of total global humanoid shipments and Nvidia is extending its CUDA playbook into that concentrated robotics market. WSJ published a peice China's concentration makes Nvidia's CUDA-style strategy easier to scale because Nvidia only needs to become deeply embedded in a handful of robot makers to sit underneath most of the market. The lesson Nvidia learned from CUDA was that selling the chip is much more powerful when developers also build their software around your platform. Nvidia is trying to create the robotics version of that dependency: Robot maker → Nvidia GPU/Jetson → CUDA → Isaac/GR00T/Cosmos → simulation, training, robot control and deployment

  2. @rohanpaul_ai70

    特朗普政府正在起草规则,限制中国 AI 实验室远程访问海外 AI 芯片,美国商务部可能以便于客户筛查和限制中国远程用户作为向外国数据中心供应芯片的条件,使身份核验与访问控制成为获得先进美国 GPU 的前提。

    推荐理由:材料梳理了美国商务部拟以客户筛查限制海外数据中心远程算力访问的监管路径,可了解芯片管制的新方向。

  3. @rohanpaul_ai38

    在 WeaveBench 的 114 个 GUI-CLI 混合任务上,官方报告的最佳通过率仅为 41.2%,说明长程智能体可靠性并未随模型变强而到来。当前模型能完成局部步骤,但需要显式审计的任务状态才能保持全程可控,否则会因历史记录无法可靠区分已完成、失败和待办事项而失败。

    引用@rohanpaul_ai@rohanpaul_ai

    "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)

8月29日周六
  1. @rohanpaul_ai46

    又一个长时程基准测试,又一次提醒我们:当前 AI 可以工作数小时,但仍远未掌握任务。 EdgeBench 中最强的智能体在允许工作 12 小时后得分 51.3/100。 该基准为智能体提供持久可执行环境和反馈,涵盖 134 项任务。它们可以调试代码、运行模拟、检查验证结果、修改证明,并反复向隐藏评审提交产物。 也就是说,它们具备了我们通常认为智能体需要随时间改进的许多要素。 但它们仍难以将所有这些交互转化为持续进步。 该研究覆盖 6 个任务族、约 38,000 小时的交互。 当跨任务取平均性能时,改进遵循对数-S 型曲线,在全部 5 个模型上 R² ≥ 0.997:早期进展缓慢,随后是更陡的学习阶段,然后饱和。

    引用@rohanpaul_ai@rohanpaul_ai

    Another long-horizon benchmark, another reminder that current AI can work for hours and still remain far from mastering the task. The strongest agent in EdgeBench scored 51.3/100 after being allowed to work for 12 hours. The benchmark gives agents persistent executable environments and feedback across 134 tasks. They can debug code, run simulations, inspect validation results, revise proofs, and repeatedly submit artifacts to hidden judges. i.e. they had many of the ingredients we normally say agents need to improve over time. They still struggle to convert all of that interaction into sustained progress. The study covers roughly 38,000 hours of interaction across 6 task families. When performance was averaged across tasks, improvement followed a log-sigmoid curve with R² ≥ 0.997 across all 5 models: slow early progress, a steeper learning phase, then saturation.

  2. @rohanpaul_ai46

    EdgeBench 长时程基准测试显示,最强智能体在允许工作 12 小时后仅得 51.3/100 分。该基准为智能体提供持久可执行环境和反馈,覆盖 134 项任务,可调试代码、运行模拟、检查验证结果、修改证明并反复向隐藏评审提交产物。

    引用@rohanpaul_ai@rohanpaul_ai

    "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)

  3. @rohanpaul_ai79

    OpenAI 宣布终止与 Cursor 的合作,依据合同中的控制权变更条款,提案于 11 月 12 日结束 Cursor 对 OpenAI 模型的直接访问。OpenAI 称已通知 SpaceX,并按合同给出最长的通知期,理由是它无法确信 SpaceX 会在其服务条款范围内使用相关技术。受影响的开发者主要是依赖 Cursor 中 OpenAI 模型的用户。

    引用@OpenAI@OpenAI

    We’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12. We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition and we’re ready to go above and beyond to support them. https://t.co/OzuCTzUjfX

    推荐理由:OpenAI 以合同控制权变更条款终止 Cursor 的模型访问,读者可以看到模型供应商与下游产品之间的授权约束。

  4. @rohanpaul_ai43

    又一篇论文清晰地揭示了AI的长时程问题。 Long-Horizon-Terminal-Bench,一旦任务延伸到数百步,即便是前沿智能体也举步维艰。 在46个长终端任务上测试17个前沿模型,平均通过率仅6.4%。 在严格的全完成度评分下,17个模型中有10个任务解决数为零。失败运行中,79%以时钟到期、智能体仍在工作告终。 即便最好的模型也只有28.3%,每10个任务中有7个未完成。 这些模型能执行大量局部合理的步骤,却无法可靠地将数百个步骤转化为一个完成的结果。

    引用@rohanpaul_ai@rohanpaul_ai

    Another paper that so clearly exposes AI’s long-horizon problem. Long-Horizon-Terminal-Bench, where even frontier agents struggle badly once tasks stretch across hundreds of steps. Across 17 frontier models on 46 long terminal tasks, the average pass rate is 6.4%. And under strict full-completion grading 10 of the 17 models solve zero tasks. Of the runs that fail, 79% end with the clock expiring while the agent is still working. Even the best model, at 28.3%, leaves 7 of every 10 tasks unfinished. The models could perform plenty of locally reasonable steps but could not reliably convert hundreds of those steps into a finished result.

  5. @rohanpaul_ai44

    Long-Horizon-Terminal-Bench 测试了 17 个前沿模型在 46 个长终端任务上的表现,平均通过率仅 6.4%。严格全完成评分下,17 个模型中有 10 个零通过;失败运行中 79% 因超时而中断。表现最好的模型也仅 28.3%,10 个任务中有 7 个未完成。

    引用@rohanpaul_ai@rohanpaul_ai

    "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)

  6. @rohanpaul_ai41

    这篇论文是对长周期AI的一次残酷现实检验。给一个智能体一年的相互关联决策、延迟反馈以及自身过往行为的后果,它的表现相对人类就会崩溃。 研究人员测试了八款领先模型,包括GPT-5.6 Sol和Claude Opus 4.8。然而表现最好的配置——Qwen3.7-Max搭配Hermes——最终赚到的钱仅为人类参与者平均水平的27.3%。 一个完成一年期任务时仅有人类四分之一表现的系统,距离可靠的长周期执行还差得远。 https://t.co/ki6xniZg9c

    引用@rohanpaul_ai@rohanpaul_ai

    This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans. The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant. A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.

  7. @rohanpaul_ai45

    一项研究让AI智能体在长达一年的相互关联决策、延迟反馈与自然后果中执行任务,结果显示其表现相对人类大幅崩塌。研究测试了包括GPT-5.6 Sol和Claude Opus 4.8在内的八款领先模型,表现最佳的Qwen3.7-Max搭配Hermes最终仅赚到人类参与者平均金额的27.3%。

    引用@rohanpaul_ai@rohanpaul_ai

    "Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)