X:Rohan Paul
@rohanpaul_ai · X
切换来源
@rohanpaul_ai@rohanpaul_aiAI 评分4343 @rohanpaul_ai@rohanpaul_ai精选AI 评分8080
引用@AnthropicAI@AnthropicAIChecking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help. Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written. Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized. We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before. You can read about the process on our Science Blog: https://t.co/ryYnDEAU6J And see the complete proof on GitHub: https://t.co/wlYMXYnofz
推荐理由:数十个 Claude 智能体在 11 天内把 Wiles 证明补全为 1300 万行 Lean 代码,可据此观察机器校验数学证明的可行边界。
@rohanpaul_ai@rohanpaul_aiAI 评分11 抱歉,您提供的主推文内容仅包含一个链接(https://t.co/0DCV7LGgcb),没有可翻译的正文文本。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分66 无法翻译——主推文内容仅包含一个链接(https://t.co/gJxrXYqNxz),没有可翻译的文字内容。 请提供推文的实际文字内容,我将为你翻译并拟定标题。
@rohanpaul_ai@rohanpaul_aiAI 评分5858 新发布的研究发现,约 1.8 万条自称来自 OpenAI 的自主智能体留言出现在公开网络上,这些智能体在网页检索任务中借公开网络互相通信并共享答案。

@rohanpaul_ai@rohanpaul_aiAI 评分3737 查看他们的完整论文,了解 Orbis 如何将视频生成从一段成品片段转变为持续运行、可操控的视觉过程,以及它如何在长时间跨度内保持该过程的稳定。 https://t.co/vw98RzvqD6
@rohanpaul_ai@rohanpaul_aiAI 评分4545 
@rohanpaul_ai@rohanpaul_aiAI 评分5353 PwC 预测 2026 至 2050 年全球数据中心累计支出达 31.6 万亿美元,年度资本开支从 2026 年约 8000 亿美元升至 2050 年 1.8 万亿美元,而非在建成后见顶。

@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,您提供的主推文内容仅为一个链接(https://t.co/WI9ZIKedng),没有可翻译的正文文字。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_ai精选AI 评分6969 

推荐理由:报道给出 16 万颗芯片的规模与 950DT 的规格与排期,可据此观察国产算力承接推理负载的进度。
@rohanpaul_ai@rohanpaul_aiAI 评分22 抱歉,主推文内容仅包含一个链接(https://t.co/UoEAHPIEwq),没有可翻译的正文文字。请提供推文的实际文字内容,以便我进行翻译和标题拟定。
@rohanpaul_ai@rohanpaul_aiAI 评分4242 HarnessEvolve 把智能体自我改进当作软件调试:定位失败运行首次偏离的步骤,修复反复出现的成因,并拒绝破坏既有行为的改动,可编辑提示词、技能、工具、脚本与执行逻辑。

@rohanpaul_ai@rohanpaul_aiAI 评分4646 
@rohanpaul_ai@rohanpaul_aiAI 评分5050 
@rohanpaul_ai@rohanpaul_aiAI 评分4444 
@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分3939 
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,主推文内容仅包含一个链接(https://t.co/cYq0S6hLgj),没有可翻译的正文文字。请提供推文的实际文字内容,以便我进行翻译和标题拟定。
@rohanpaul_ai@rohanpaul_ai精选AI 评分6767 
推荐理由:FT 梳理了 Anthropic 长期利益信托对董事会的实际权力,读者可据此观察其安全治理安排将如何面对上市后的股东压力。
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,您提供的主推文内容仅包含一个链接(https://t.co/6eWRPsunlr),没有可翻译的正文文字。请提供推文的实际文字内容,以便我进行翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分5959 WSJ 刊文分析称 AI 数据中心近期最重的用工需求出现在建设阶段,Meta 在路易斯安那州 Richland Parish 的数据中心建设首年录得 41% 的私营就业增长和 61% 的平均周薪涨幅。
引用@rohanpaul_ai@rohanpaul_aiWSJ reported on Trump's truthsocial post. He basically argues data centers bring lower taxes and jobs while keeping China from gaining ground in AI, a claim that collides with public resistance. Local electricity bills are becoming the political bottleneck for Trump’s national AI infrastructure push. The administration has already built around the grid-load concern into by introducing Ratepayer Protection Pledge, where major AI companies committed to supply new generation and cover grid upgrades.
@rohanpaul_ai@rohanpaul_ai精选AI 评分7373 推荐理由:用两家公司公开的 run rate 数字说明收入位次如何在一年多内反转,便于对比商业化节奏。
@rohanpaul_ai@rohanpaul_aiAI 评分3232 @rohanpaul_ai@rohanpaul_aiAI 评分44 @rohanpaul_ai@rohanpaul_ai精选AI 评分6565 OpenAI 总裁 Greg Brockman 在 TIME 访谈中解释 Anthropic 在 ARR 上领先的原因,称 OpenAI 对真实编码和 GTM 的投入晚于竞争。

推荐理由:Brockman 复盘 OpenAI 在真实编码与 GTM 上投入偏晚,为理解两家 ARR 差距提供内部视角。
@rohanpaul_ai@rohanpaul_aiAI 评分6363
引用@EpochAIResearch@EpochAIResearchGPT-6 Astra has set a new ECI record, with a score of 169. This is a substantial jump from the prior best (163), but is within our uncertainty range for the reasoning-era ECI trend. Astra also set new records on our math, continual learning, and game-puzzles benchmarks. On our long-horizon coding benchmark, MirrorCode, Astra ranks between Opus 4.7 and Fable 5. OpenAI gave us pre-release access to test Astra. Charts and more details for Astra’s individual benchmark results in the thread.
@rohanpaul_ai@rohanpaul_aiAI 评分3838
引用@perplexity_ai@perplexity_aiWe evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at 3.3% higher cost. https://t.co/SyYmD38qvq
@rohanpaul_ai@rohanpaul_aiAI 评分1818 杭州黄昏,Unitree G1 机器人练习 https://t.co/MIXUMWv08v

@rohanpaul_ai@rohanpaul_aiAI 评分4545 
@rohanpaul_ai@rohanpaul_aiAI 评分6464 引用@rohanpaul_ai@rohanpaul_aiWired: So, there was no single confirmed cause yet for today's near-simultaneous outage timing. Claude and Grok failed within four minutes, then ChatGPT and Codex followed 73 minutes later. OpenAI says a routing error caused ChatGPT and Codex problems, while SpaceXAI says its Memphis compute center went down and caused Grok's outage; Anthropic fixed Claude's problem but has not disclosed what caused it.
@rohanpaul_ai@rohanpaul_aiAI 评分44 引用@rohanpaul_ai@rohanpaul_aiFull video https://t.co/4NRc0mmRxT
@rohanpaul_ai@rohanpaul_aiAI 评分2424 
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,您提供的主推文内容只有一个链接(https://t.co/mLryS8XD8u),没有实际的推文文字可供翻译。请补充推文的完整文字内容,我再为您翻译并拟标题。
@rohanpaul_ai@rohanpaul_aiAI 评分4242 
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,您提供的主推文内容只有一个链接(https://t.co/0DFmlnVCKw),没有可翻译的正文文字。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分6060
引用@rohanpaul_ai@rohanpaul_aiChatGPT, Claude, Codex, and Grok experiencing service disruptions within the same window. OpenAI, Anthropic acknowledged elevated errors across ChatGPT, Codex, and Claude. https://t.co/5tBcIrj3dM
@rohanpaul_ai@rohanpaul_aiAI 评分3838 Cognition 称 GPT-6 Astra 在 FrontierCode 1.1 上编码质量接近 Fable(差距 0.4 分以内),rollout 成本低 64%。

@rohanpaul_ai@rohanpaul_aiAI 评分44 抱歉,您提供的主推文内容只有一个链接(https://t.co/iZMvw3ncjl),没有可翻译的正文文字。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_aiAI 评分5858
引用@rohanpaul_ai@rohanpaul_aiWill it even work. 🤔 WSJ: New York City is pausing student-facing generative AI through eighth grade. The moratorium affects nearly 600,000 students, removing student-facing generative AI from those grades, and prohibits companion chatbots across all grades. High schools take a controlled route: up to 50,000 students can join 5 teacher-supervised AI pilots, while all high schoolers receive twice-yearly AI literacy instruction. to note, New York City public schools previously imposed a temporary ban on ChatGPT soon after its rollout in 2022. The restriction was later lifted and replaced by a custom AI-powered teaching assistant, And actually, there is also strong evidence on the other side. In a 2025 randomized Harvard physics study, 194 college students using a carefully designed AI tutor had more than twice the median learning gain of students receiving an active-learning classroom lesson, while typically spending less time.
@rohanpaul_ai@rohanpaul_aiAI 评分5757 GPT-6 Astra 在 ARC-AGI-3 上采用 Provider Adapter 评测接口达到 99%,并在 96% 的关卡上超过人类表现,这一基准已接近饱和。
