X:Rohan Paul
@rohanpaul_ai · X
切换来源
@rohanpaul_ai@rohanpaul_aiAI 评分2424 
@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分2222 – https://t.co/p6oMCckUtR 标题:《智能体腐化有多快?面向生产决策的 LLM 智能体长时程退化实证研究》
@rohanpaul_ai@rohanpaul_aiAI 评分6161 
@rohanpaul_ai@rohanpaul_aiAI 评分3535 特朗普总统:“谁赢得AI,谁就赢得一切。这就是力量。这比互联网还重大。这是一场革命。过去200年里有过几次革命,但这是一场革命。” https://t.co/Tu2t74o28e

@rohanpaul_ai@rohanpaul_aiAI 评分5858 
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,该推文内容仅包含一个链接(https://t.co/MA2vGd40BK),没有可翻译的正文文本。请提供推文的实际文字内容,以便我进行翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分5656 
@rohanpaul_ai@rohanpaul_aiAI 评分4646 GPT-6 Astra 在 MazeBench 无 Python 赛道得分 14%,是 Claude Fable 5.1 的 2% 的 7 倍,也超过 GPT-5.6 Sol 开启代码辅助的 13%。

@rohanpaul_ai@rohanpaul_aiAI 评分4444 
@rohanpaul_ai@rohanpaul_aiAI 评分2020 – https://t.co/AhKiNT01D7 标题:"ArcticSwarm:在长时程多智能体研究中推迟早期共识"
@rohanpaul_ai@rohanpaul_aiAI 评分2222 – https://t.co/gVhUMbl85t 标题:"揭秘 Agent Skills:为什么它们有效——直到不再有效"
@rohanpaul_ai@rohanpaul_aiAI 评分4141 
@rohanpaul_ai@rohanpaul_aiAI 评分4040 
@rohanpaul_ai@rohanpaul_aiAI 评分4040 
@rohanpaul_ai@rohanpaul_aiAI 评分3030 GPT 6 Astra 通关了《I'm Not a Robot》全部 48 关。 网站将不再问你是不是人类,而是开始问你付不付费。 https://t.co/IGtHnY8zda

@rohanpaul_ai@rohanpaul_aiAI 评分5555 
@rohanpaul_ai@rohanpaul_aiAI 评分2222 – https://t.co/5Z8zAEiVsQ 标题:“为什么 CLAUDE.md 不断膨胀?智能体编程中的灾难性记忆”
@rohanpaul_ai@rohanpaul_aiAI 评分4949 
@rohanpaul_ai@rohanpaul_aiAI 评分4545
引用@thsottiaux@thsottiauxAstra was probably our biggest competitive advantage while it wasn’t generally available. Since we’ve had it our productivity jumped so much that we shifted some of our plans 6 months ahead and will ship them at DevDay instead of mid next year.
@rohanpaul_ai@rohanpaul_aiAI 评分3434 David Sacks:AI 末日论叙事背后已经有大量资金支持。 而 Anthropic IPO 可能会让这一资金基础变得更加庞大。 https://t.co/1AerZ63OMp

@rohanpaul_ai@rohanpaul_aiAI 评分5959 
@rohanpaul_ai@rohanpaul_aiAI 评分44 @rohanpaul_ai@rohanpaul_aiAI 评分3131 
@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分4141 
@rohanpaul_ai@rohanpaul_aiAI 评分2727 引用@kimmonismus@kimmonismusAn AI can complete a task perfectly and still spend your time on the wrong thing. Imagine an assistant preparing notes for tomorrow’s meeting. A message arrives that changes the deadline, and suddenly another project needs your attention first. The notes may be excellent. The assistant still needs to recognize that your priorities have changed. As AI takes on more work, these decisions become increasingly important. Which task should come first? When is an interruption justified? When should the system ask you? That’s the question behind @TodayAIofficial’s approach to personal AI: how can an agent learn what matters to a particular person and use that context throughout the day? Memory is part of the answer. Remembering a deadline helps. Connecting it to a promise you made last week, noticing that the plan has changed, and bringing it up while you can still act takes more. It also requires a way to correct the system. People change their minds. A preference from three months ago may no longer apply. An assistant should make its assumptions visible and let you update them. I think this is a useful direction for personal AI. There’s considerable value in software that can connect scattered information and help you decide where to focus. The test will be how well it handles an ordinary, messy day: catching the commitment you might miss, explaining why it needs attention, and leaving the decision with you.
@rohanpaul_ai@rohanpaul_aiAI 评分4848 引用@rohanpaul_ai@rohanpaul_aiI think "time horizon" is becoming one of the more useful ways to talk about agent capability. At OpenAI, as tasks get longer, success without human intervention collapses, from 86% on sub-15-minute tasks to only around 16% on the longest 64-128-hour bucket, while successful runs requiring intervention become much more common. A model can be extremely capable locally and still be unreliable over a long execution trajectory. Not tokens. Not benchmark score. How long can the system keep useful control of a task before a human has to intervene? maybe we are moving towards a benchmark, something like human minutes consumed per completed task.
@rohanpaul_ai@rohanpaul_aiAI 评分5656
引用@rohanpaul_ai@rohanpaul_aiOpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that would take a skilled researcher quite few days. inside OpenAI research, agent runtime has already crossed human labor by a wide margin. 3.1-to-1 agent-to-human ratio “In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.” That ratio measures agent runtime rather than equivalent productivity, but it captures how deeply parallel agent work has entered OpenAI research.
@rohanpaul_ai@rohanpaul_aiAI 评分3030 Jevons 悖论应用于劳动力。 当智能变得廉价,对它的需求就会爆炸式增长,而最接近源头的人会吸收这一切。 AI 确实把时间还回来了,但领域里的每个人都只是把它再投资了回去。

@rohanpaul_ai@rohanpaul_aiAI 评分5454 
@rohanpaul_ai@rohanpaul_aiAI 评分4141 
@rohanpaul_ai@rohanpaul_aiAI 评分4444 微软与康奈尔大学论文提出 Free Pause Tokens,让模型预测下一个 token 时获得额外计算,且不增加 token、不扩大 KV cache、不增加解码步骤。

@rohanpaul_ai@rohanpaul_ai精选AI 评分7878 
推荐理由:OpenAI 首席科学家公开谈论扩展节奏与对齐缺口,读者可据此了解头部实验室对安全进展的最新判断。
@rohanpaul_ai@rohanpaul_aiAI 评分6464
引用@rohanpaul_ai@rohanpaul_aiOpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that would take a skilled researcher quite few days. inside OpenAI research, agent runtime has already crossed human labor by a wide margin. 3.1-to-1 agent-to-human ratio “In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.” That ratio measures agent runtime rather than equivalent productivity, but it captures how deeply parallel agent work has entered OpenAI research.
@rohanpaul_ai@rohanpaul_aiAI 评分3838 引用@rohanpaul_ai@rohanpaul_aiOpenAI engineers are changing roughly 7x more code per contributor than the pre-2025 baseline. https://t.co/K8EK8zuvqe https://t.co/5afIiyoITp
@rohanpaul_ai@rohanpaul_aiAI 评分5656
引用@rohanpaul_ai@rohanpaul_aiOpenAI just officially said it has reached its "automated research intern" milestone. i.e. a human-supervised system able to complete well-defined tasks that would take a skilled researcher quite few days. inside OpenAI research, agent runtime has already crossed human labor by a wide margin. 3.1-to-1 agent-to-human ratio “In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.” That ratio measures agent runtime rather than equivalent productivity, but it captures how deeply parallel agent work has entered OpenAI research.
@rohanpaul_ai@rohanpaul_aiAI 评分2020 反AI的愤怒病毒是真实存在的。 事实上,反AI的愤怒是对工具有效最响亮的承认。 https://t.co/tKgtijLfLK https://t.co/qf7sEgCAcQ

@rohanpaul_ai@rohanpaul_aiAI 评分4747 Claude Code 作者 Boris Cherny 在 Startup School 2026 上被问及人们如何学会像他一样使用 Claude Code,他建议别听 LinkedIn 网红的说法。
