马斯克表示 SpaceX 和 Tesla 正各自建设每年 100GW 的太阳能产能,同时 SpaceX 计划在内部自建涡轮叶片铸造厂,使天然气涡轮机上线时间最多提前 18 个月。
X:Rohan Paul
@rohanpaul_ai · X
切换来源
@rohanpaul_ai@rohanpaul_aiAI 评分5959 
@rohanpaul_ai@rohanpaul_aiAI 评分5252 
@rohanpaul_ai@rohanpaul_aiAI 评分4343 
@rohanpaul_ai@rohanpaul_aiAI 评分6060
引用@rohanpaul_ai@rohanpaul_aiChinese robots hold 86% of total global humanoid shipments and Nvidia is extending its CUDA playbook into that concentrated robotics market. WSJ published a peice China's concentration makes Nvidia's CUDA-style strategy easier to scale because Nvidia only needs to become deeply embedded in a handful of robot makers to sit underneath most of the market. The lesson Nvidia learned from CUDA was that selling the chip is much more powerful when developers also build their software around your platform. Nvidia is trying to create the robotics version of that dependency: Robot maker → Nvidia GPU/Jetson → CUDA → Isaac/GR00T/Cosmos → simulation, training, robot control and deployment
@rohanpaul_ai@rohanpaul_aiAI 评分4444 
@rohanpaul_ai@rohanpaul_aiAI 评分77 抱歉,主推文内容仅包含一个链接(https://t.co/pSn1g90K3D),没有可翻译的正文文字。请提供推文的实际文字内容,以便我进行翻译和标题拟定。
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,您提供的主推文内容仅为一个链接(https://t.co/6jEaSJtckQ),没有可翻译的正文文本。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_ai精选AI 评分7373 

推荐理由:起诉书给出 Anthropic 获取 LibGen 数据、微调中奖励歌词复现等具体指控,可了解这起版权诉讼的争议焦点。
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,主推文内容仅包含一个链接(https://t.co/VZKanBlwcG),没有可翻译的正文文本。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_ai精选AI 评分7070 
推荐理由:材料梳理了美国商务部拟以客户筛查限制海外数据中心远程算力访问的监管路径,可了解芯片管制的新方向。
@rohanpaul_ai@rohanpaul_aiAI 评分5252 
@rohanpaul_ai@rohanpaul_aiAI 评分2323 – https://t.co/6mrarrY7q1 标题:"LongHorizon-Harness:推进面向真实世界任务的长时程智能体"
@rohanpaul_ai@rohanpaul_aiAI 评分3838
引用@rohanpaul_ai@rohanpaul_ai"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)
@rohanpaul_ai@rohanpaul_aiAI 评分22 @rohanpaul_ai@rohanpaul_aiAI 评分4545 
@rohanpaul_ai@rohanpaul_aiAI 评分4646 引用@rohanpaul_ai@rohanpaul_aiAnother long-horizon benchmark, another reminder that current AI can work for hours and still remain far from mastering the task. The strongest agent in EdgeBench scored 51.3/100 after being allowed to work for 12 hours. The benchmark gives agents persistent executable environments and feedback across 134 tasks. They can debug code, run simulations, inspect validation results, revise proofs, and repeatedly submit artifacts to hidden judges. i.e. they had many of the ingredients we normally say agents need to improve over time. They still struggle to convert all of that interaction into sustained progress. The study covers roughly 38,000 hours of interaction across 6 task families. When performance was averaged across tasks, improvement followed a log-sigmoid curve with R² ≥ 0.997 across all 5 models: slow early progress, a steeper learning phase, then saturation.
@rohanpaul_ai@rohanpaul_aiAI 评分2222 @rohanpaul_ai@rohanpaul_aiAI 评分4646
引用@rohanpaul_ai@rohanpaul_ai"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)
@rohanpaul_ai@rohanpaul_aiAI 评分2525 – https://t.co/V83tZd5aSk 标题:"WikiSkill:将智能体经验编译为持久知识以促进技能演化"
@rohanpaul_ai@rohanpaul_aiAI 评分4949 Google 提出 WikiSkill,将智能体工作区拆成不可变执行轨迹、活跃技能文件与中间 wiki 三层,把失败模式和每次技能修改提案的采纳或拒绝结果持久记录在 wiki 中。

@rohanpaul_ai@rohanpaul_aiAI 评分5555 
@rohanpaul_ai@rohanpaul_aiAI 评分88 
@rohanpaul_ai@rohanpaul_aiAI 评分1919 @rohanpaul_ai@rohanpaul_aiAI 评分4646 
@rohanpaul_ai@rohanpaul_ai精选AI 评分7979
引用@OpenAI@OpenAIWe’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12. We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition and we’re ready to go above and beyond to support them. https://t.co/OzuCTzUjfX
推荐理由:OpenAI 以合同控制权变更条款终止 Cursor 的模型访问,读者可以看到模型供应商与下游产品之间的授权约束。
@rohanpaul_ai@rohanpaul_aiAI 评分4343 引用@rohanpaul_ai@rohanpaul_aiAnother paper that so clearly exposes AI’s long-horizon problem. Long-Horizon-Terminal-Bench, where even frontier agents struggle badly once tasks stretch across hundreds of steps. Across 17 frontier models on 46 long terminal tasks, the average pass rate is 6.4%. And under strict full-completion grading 10 of the 17 models solve zero tasks. Of the runs that fail, 79% end with the clock expiring while the agent is still working. Even the best model, at 28.3%, leaves 7 of every 10 tasks unfinished. The models could perform plenty of locally reasonable steps but could not reliably convert hundreds of those steps into a finished result.
@rohanpaul_ai@rohanpaul_aiAI 评分3333 – https://t.co/qJoaEvvag3 标题:"Long-Horizon-Terminal-Bench:用基于密集奖励的评分测试智能体在长时程终端任务上的极限"
@rohanpaul_ai@rohanpaul_aiAI 评分4444
引用@rohanpaul_ai@rohanpaul_ai"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)
@rohanpaul_ai@rohanpaul_aiAI 评分2929 @rohanpaul_ai@rohanpaul_aiAI 评分3131 
@rohanpaul_ai@rohanpaul_aiAI 评分4141 引用@rohanpaul_ai@rohanpaul_aiThis paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans. The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant. A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.
@rohanpaul_ai@rohanpaul_aiAI 评分4545
引用@rohanpaul_ai@rohanpaul_ai"Long-horizon tasks are still a joke. They do not work, and I do not care what anybody says. Do not show me a stupid evaluation. Do not tell me about some dumb script you ran for 48 hours. Long-horizon tasks are not handled well. They simply do not work." - Chamath at Stanford AI Club "2nd, complex problems also do not work. They are neither addressed nor handled well. Why is this important? If AI develops like any other technology, we are going to experience an initial rise—the hype cycle. Then, we will see a natural contraction because, somehow and somewhere, something is going to fail. We are all going to see this, and then we will enter what is called the “trough of disillusionment.” I think the business and MBA folks will confirm whether that is true. Afterward, you typically see the slow and gradual adoption of the real, final solution. This happened with the internet, and it has happened in many other cases. The problem is that we are spending hundreds of billions, potentially trillions, of dollars trying to figure out how to cross this chasm. So, what do we do? If we do not figure this out, people will reach the trough of disillusionment and say that AI was a joke. I think we need to be able to bring AI into highly complicated environments and make it work. What is my solution? At a very basic level, you need a symbolic space that guides the embedded space." ---- From "techniahqrobot" YouTube channel, (full video link in comment)
@rohanpaul_ai@rohanpaul_aiAI 评分2424 – https://t.co/I2oja1MYf9 标题:"MerchantBench:面向电商运营中长期一致性的 LLM 智能体基准测试"
@rohanpaul_ai@rohanpaul_aiAI 评分2424 @rohanpaul_ai@rohanpaul_aiAI 评分3333 
@rohanpaul_ai@rohanpaul_aiAI 评分4545 @rohanpaul_ai@rohanpaul_aiAI 评分4646 
@rohanpaul_ai@rohanpaul_aiAI 评分4040 
@rohanpaul_ai@rohanpaul_aiAI 评分2424 – https://t.co/ZgHSxLO5xO 标题:"FreeToken:带宽自适应执行的边缘原生 MoE 高效服务"
@rohanpaul_ai@rohanpaul_aiAI 评分1818 你可以通过 AI/ML API 在这里用这些模型生成 Backrooms https://t.co/umit1HYZLg