X:Rohan Paul
@rohanpaul_ai · X
切换来源
@rohanpaul_ai@rohanpaul_aiAI 评分4242 
@rohanpaul_ai@rohanpaul_aiAI 评分4545 
@rohanpaul_ai@rohanpaul_aiAI 评分44 @rohanpaul_ai@rohanpaul_aiAI 评分3737
@rohanpaul_ai@rohanpaul_aiAI 评分3636 Antioch Robotics 完成 3200 万美元 A 轮融资,用于构建让场景本身成为软件对象的仿真基础设施,使物理 AI 团队能像软件团队跑测试套件一样测试新版本。

@rohanpaul_ai@rohanpaul_aiAI 评分3838 
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,主推文内容仅包含一个链接(https://t.co/5vDr0XAjg8),没有可翻译的文字正文。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分2626 
@rohanpaul_ai@rohanpaul_aiAI 评分1919 – https://t.co/1bQgSysiOd 标题:"Skill Following:评估检索增强型 LLM 智能体中的实际技能使用"
@rohanpaul_ai@rohanpaul_aiAI 评分4444 
@rohanpaul_ai@rohanpaul_aiAI 评分2424 
@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分2222 – https://t.co/p6oMCckUtR 标题:《智能体腐化有多快?面向生产决策的 LLM 智能体长时程退化实证研究》
@rohanpaul_ai@rohanpaul_aiAI 评分6161 
@rohanpaul_ai@rohanpaul_aiAI 评分3535 特朗普总统:“谁赢得AI,谁就赢得一切。这就是力量。这比互联网还重大。这是一场革命。过去200年里有过几次革命,但这是一场革命。” https://t.co/Tu2t74o28e

@rohanpaul_ai@rohanpaul_aiAI 评分5858 
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,该推文内容仅包含一个链接(https://t.co/MA2vGd40BK),没有可翻译的正文文本。请提供推文的实际文字内容,以便我进行翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分5656 
@rohanpaul_ai@rohanpaul_aiAI 评分4646 GPT-6 Astra 在 MazeBench 无 Python 赛道得分 14%,是 Claude Fable 5.1 的 2% 的 7 倍,也超过 GPT-5.6 Sol 开启代码辅助的 13%。

@rohanpaul_ai@rohanpaul_aiAI 评分4444 
@rohanpaul_ai@rohanpaul_aiAI 评分2020 – https://t.co/AhKiNT01D7 标题:"ArcticSwarm:在长时程多智能体研究中推迟早期共识"
@rohanpaul_ai@rohanpaul_aiAI 评分2222 – https://t.co/gVhUMbl85t 标题:"揭秘 Agent Skills:为什么它们有效——直到不再有效"
@rohanpaul_ai@rohanpaul_aiAI 评分4141 
@rohanpaul_ai@rohanpaul_aiAI 评分4040 
@rohanpaul_ai@rohanpaul_aiAI 评分4040 
@rohanpaul_ai@rohanpaul_aiAI 评分3030 GPT 6 Astra 通关了《I'm Not a Robot》全部 48 关。 网站将不再问你是不是人类,而是开始问你付不付费。 https://t.co/IGtHnY8zda

@rohanpaul_ai@rohanpaul_aiAI 评分5555 
@rohanpaul_ai@rohanpaul_aiAI 评分2222 – https://t.co/5Z8zAEiVsQ 标题:“为什么 CLAUDE.md 不断膨胀?智能体编程中的灾难性记忆”
@rohanpaul_ai@rohanpaul_aiAI 评分4949 
@rohanpaul_ai@rohanpaul_aiAI 评分4545
引用@thsottiaux@thsottiauxAstra was probably our biggest competitive advantage while it wasn’t generally available. Since we’ve had it our productivity jumped so much that we shifted some of our plans 6 months ahead and will ship them at DevDay instead of mid next year.
@rohanpaul_ai@rohanpaul_aiAI 评分3434 David Sacks:AI 末日论叙事背后已经有大量资金支持。 而 Anthropic IPO 可能会让这一资金基础变得更加庞大。 https://t.co/1AerZ63OMp

@rohanpaul_ai@rohanpaul_aiAI 评分5959 
@rohanpaul_ai@rohanpaul_aiAI 评分44 @rohanpaul_ai@rohanpaul_aiAI 评分3131 
@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分4141 
@rohanpaul_ai@rohanpaul_aiAI 评分2727 引用@kimmonismus@kimmonismusAn AI can complete a task perfectly and still spend your time on the wrong thing. Imagine an assistant preparing notes for tomorrow’s meeting. A message arrives that changes the deadline, and suddenly another project needs your attention first. The notes may be excellent. The assistant still needs to recognize that your priorities have changed. As AI takes on more work, these decisions become increasingly important. Which task should come first? When is an interruption justified? When should the system ask you? That’s the question behind @TodayAIofficial’s approach to personal AI: how can an agent learn what matters to a particular person and use that context throughout the day? Memory is part of the answer. Remembering a deadline helps. Connecting it to a promise you made last week, noticing that the plan has changed, and bringing it up while you can still act takes more. It also requires a way to correct the system. People change their minds. A preference from three months ago may no longer apply. An assistant should make its assumptions visible and let you update them. I think this is a useful direction for personal AI. There’s considerable value in software that can connect scattered information and help you decide where to focus. The test will be how well it handles an ordinary, messy day: catching the commitment you might miss, explaining why it needs attention, and leaving the decision with you.
@rohanpaul_ai@rohanpaul_aiAI 评分4848 引用@rohanpaul_ai@rohanpaul_aiI think "time horizon" is becoming one of the more useful ways to talk about agent capability. At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common. A model can be extremely capable locally and still be unreliable over a long execution trajectory. Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene? maybe we are moving towards a benchmark, something like human minutes consumed per completed task.