X:Rohan Paul
@rohanpaul_ai · X
切换来源
@rohanpaul_ai@rohanpaul_aiAI 评分4848 
@rohanpaul_ai@rohanpaul_aiAI 评分4545 GPT-6 Astra 在 Vending-Bench 2 中 6 次独立运行平均赚得 $15,515,远超 Fable 的 $5,422,Fable 最好成绩仍低于 Astra 最差成绩。

@rohanpaul_ai@rohanpaul_aiAI 评分55
@rohanpaul_ai@rohanpaul_aiAI 评分5252 
@rohanpaul_ai@rohanpaul_aiAI 评分44 @rohanpaul_ai@rohanpaul_aiAI 评分6060 OpenAI 韩国办公室表示,OpenAI 正与三星电子联合研发和生产其正在开发的下一代 AI 芯片。该办公室称,与三星电子的联合生产与研发是双方进展最大、获得认可最多的领域之一。

@rohanpaul_ai@rohanpaul_aiAI 评分4646 引用@rohanpaul_ai@rohanpaul_aiJensen Huang on "distillation" In this interview with axios, he was asked this question: "Should open source model companies be allowed to distill closed models" "Distillation, learning from AI, learning from other people, and learning from other sources of knowledge, is fundamental to intelligence. We are constantly learning from other people. I am learning from you through the questions you are asking, and you are learning from me. All day long, we are learning from one another. AI also has to learn from something. The original AI models, whether they were open or closed, were trained on previously created knowledge from the internet. Now, AI is generating more content than humans. In a few more years, the internet could be 99% AI-generated content, and that content will have been created by some form of AI. As a result, AI systems will constantly be distilling knowledge and intelligence from other AI systems. The fact that AI can learn is a good thing. We want AI systems to be intelligent because a smarter AI can also be a safer AI." ---- From "Axios" YouTube channel, (full video link in comment)
@rohanpaul_ai@rohanpaul_aiAI 评分5959
引用@rohanpaul_ai@rohanpaul_aiIn a new allegation, the U.S. government has accused 6 Chinese AI firms of using large-scale distillation to copy American model capabilities. The NSA, CISA and FBI published joint advisory AA26-251A, naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z .AI as running industrial-scale distillation campaigns against U.S. frontier models since at least late 2024. The agencies say requests moved through a gray market of API proxies called transfer stations, bulk premium subscriptions shared across developer teams, and third-party aggregators that strip account metadata. Moonshot AI allegedly distilled 17 U.S. models, including Anthropic's Claude Fable 5, to train Kimi-K3. They claim DeepSeek's prompts pushed models to write out hidden chain-of-thought steps, which transfers reasoning method rather than finished answers. So they say that DeepSeek's headline training cost only covers the compute it burned. The advisory argues that number is misleading because it leaves out what the training data actually cost, which DeepSeek allegedly obtained by distilling U.S. models rather than producing it through its own research. So the cheap-training story rests on data someone else paid to create. The mitigation section asks U.S. labs to serve suspected distillers subtly degraded responses without telling them. Its detection indicators are behavioral, covering sustained 24/7 usage, new accounts at immediate maximum throughput, and traffic optimized for cache hits. Those patterns also describe an ordinary enterprise agent fleet, leaving each provider to decide which customers receive an undisclosed downgrade.
@rohanpaul_ai@rohanpaul_ai精选AI 评分7979 
推荐理由:公告列出了被指控的蒸馏渠道与行为检测指标,读者可看到模型输出和账号用量如何被当作判定依据。
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,主推文内容仅包含一个链接(https://t.co/AgiUYbPGhr),没有可翻译的正文文字。请提供推文的实际文字内容,以便我进行翻译和标题拟定。
@rohanpaul_ai@rohanpaul_aiAI 评分77 @rohanpaul_ai@rohanpaul_aiAI 评分5757 
@rohanpaul_ai@rohanpaul_aiAI 评分4646 OpenAI 可能会暂时停止新的 Pro 订阅,因为 GPT-6 Astra 的需求正在爆发。
引用@thsottiaux@thsottiauxDemand for Astra is really unprecedented. We're pulling all the levers possible to sustain the demand, but I've not seen anything like it until now and we went through very steep growth before. Priority will always be to keep excellent service for existing users, but we might have to pause new Pro subscriptions for a bit if this continues.
@rohanpaul_ai@rohanpaul_aiAI 评分5454 引用@rohanpaul_ai@rohanpaul_aiMark Zuckerberg talks about how he is using Meta’s newly released Muse agent. something much closer to a persistent personal worker than a chatbot. helping his 3-year-old daughter likes baking, monitor when mountain climbing permits become available and secure them for their weekend trips, and also for physical training.. He installed cameras in his MMA gym and has Muse watch the footage and give him feedback on his performance. Zuckerberg gives it ongoing goals and projects, and the agent can continue working on them, remember what happened, and proactively suggest what it should do next. ---- From "Sources Podcast and @alexeheath YouTube channel, (full video link in comment)
@rohanpaul_ai@rohanpaul_aiAI 评分2424 Muse 智能体被刻意与凭证、网络权限和安全执行机制隔离开来,因此即使模型被攻破或被操纵,攻击者也不应自动获得对所有内容的访问权限。https://t.co/XS4ZfKTc6k

@rohanpaul_ai@rohanpaul_ai精选AI 评分6767 扎克伯格解释了 Meta 新 Muse 智能体的保密云虚拟机如何工作,每位用户可获得一台存放高度私密个人数据的保密云虚拟机,系统设计上连 Meta 自身也无法查看这些内容。
引用@rohanpaul_ai@rohanpaul_aiMeta just released Muse AI assistant for personal tasks, powered by Muse Spark 1.3, Meta's latest model for longer-horizon agentic work - it can connect to email, calendars, shopping, payments and other services, then keep working after the app closes. - each user gets a Muse Secure VM, with a separate Sentinel agent that checks outbound data and actions before they reach the network. - Meta is making Muse free to use for up to 100M tokens per week, with subscription plans for people who use more compute - Sensitive actions such as purchases or emails require user approval, while credentials stay in Secure Credential Storage so Muse cannot read passwords or payment details. - the model, Muse harness, deterministic code and classifiers screen prompt injections before hostile content can enter the model's context. - currently it launches for U.S. users on web, iOS, Android and WhatsApp, with free access plus $20 and $100 monthly tiers for heavier usage. - the security design moves the trust boundary beyond model behavior into VM isolation, credential separation and a kernel-enforced gate on network actions.
推荐理由:扎克伯格解释 Muse 智能体如何用保密云虚拟机隔离私人数据,读者可据此了解这套安全设计的具体取舍。
@rohanpaul_ai@rohanpaul_aiAI 评分4141 
@rohanpaul_ai@rohanpaul_aiAI 评分1919 – https://t.co/zgtRqCDFsw 标题:"Diagnosing with Insights:通过行为抽象对智能体失败进行结构化分析"
@rohanpaul_ai@rohanpaul_aiAI 评分22 @rohanpaul_ai@rohanpaul_aiAI 评分4444 
@rohanpaul_ai@rohanpaul_aiAI 评分4343 引用@rohanpaul_ai@rohanpaul_aiMark Zuckerberg talks about how he is using Meta’s newly released Muse agent. something much closer to a persistent personal worker than a chatbot. helping his 3-year-old daughter likes baking, monitor when mountain climbing permits become available and secure them for their weekend trips, and also for physical training.. He installed cameras in his MMA gym and has Muse watch the footage and give him feedback on his performance. Zuckerberg gives it ongoing goals and projects, and the agent can continue working on them, remember what happened, and proactively suggest what it should do next. ---- From "Sources Podcast and @alexeheath YouTube channel, (full video link in comment)
@rohanpaul_ai@rohanpaul_aiAI 评分1010 https://t.co/XU0qtAdZIJ (注:主推文仅含一个链接,无实质文字内容;引用推文为"full video"加链接,亦无新闻信息。无法生成有意义的标题和翻译。)
引用@rohanpaul_ai@rohanpaul_aifull video https://t.co/QqMXZUrhFz
@rohanpaul_ai@rohanpaul_ai精选AI 评分7474
引用@rohanpaul_ai@rohanpaul_aiMeta just released Muse AI assistant for personal tasks, powered by Muse Spark 1.3, Meta's latest model for longer-horizon agentic work - it can connect to email, calendars, shopping, payments and other services, then keep working after the app closes. - each user gets a Muse Secure VM, with a separate Sentinel agent that checks outbound data and actions before they reach the network. - Meta is making Muse free to use for up to 100M tokens per week, with subscription plans for people who use more compute - Sensitive actions such as purchases or emails require user approval, while credentials stay in Secure Credential Storage so Muse cannot read passwords or payment details. - the model, Muse harness, deterministic code and classifiers screen prompt injections before hostile content can enter the model's context. - currently it launches for U.S. users on web, iOS, Android and WhatsApp, with free access plus $20 and $100 monthly tiers for heavier usage. - the security design moves the trust boundary beyond model behavior into VM isolation, credential separation and a kernel-enforced gate on network actions.
推荐理由:扎克伯格展示了 Muse 处理长期目标的方式,与发布细节一起呈现出长周期个人智能体的形态。
@rohanpaul_ai@rohanpaul_aiAI 评分2222 
@rohanpaul_ai@rohanpaul_aiAI 评分22 @rohanpaul_ai@rohanpaul_aiAI 评分4343 
@rohanpaul_ai@rohanpaul_aiAI 评分77 @rohanpaul_ai@rohanpaul_aiAI 评分2727 Muse 智能体被刻意与凭证、网络权限和安全执行机制隔离开来,因此攻破或操纵该模型不应自动让攻击者获得对所有内容的访问权限。https://t.co/jr8By10gIZ

@rohanpaul_ai@rohanpaul_ai精选AI 评分7272 Meta 发布面向个人任务的 AI 助手 Muse,由 Muse Spark 1.3 驱动,可连接邮箱、日历、购物和支付等服务,应用关闭后仍能继续工作。



引用@finkd@finkdIntroducing Muse, the personal agent that understands your goals and works 24/7 to get things done for you.
推荐理由:材料梳理了 Muse 的免费额度与安全架构,读者可据此判断个人智能体的权限与信任边界如何落地。
@rohanpaul_ai@rohanpaul_aiAI 评分6464 
@rohanpaul_ai@rohanpaul_aiAI 评分3636 引用@rohanpaul_ai@rohanpaul_aiOpenAI's unreleased internal model sits on a much higher capability curve than GPT-6 Astra. the one that just solved the 90-year old Navier-Stokes Millennium Prize math Problem. the new model also appears to use additional reasoning compute much more productively. That's how throwing 10,000 agents and enormous compute at Navier-Stokes could produce something qualitatively different from simply running Astra longer.
@rohanpaul_ai@rohanpaul_aiAI 评分5050 引用@rohanpaul_ai@rohanpaul_aiMASSIVE: OpenAI's 10,000 AI agents, powered by a model significantly more capable than GPT-6 Astra, solved a math problem unresolved for roughly 90 years. Looks like, in this new world, the scaling law may be agent count, not model size: enough capable agents can search far more scientific ideas than any human group ever could. when tens of thousands of strong agents can explore, criticize, and combine ideas, many problems that are "too hard" today may become mostly a compute problem. For this one, for 90 years, mathematicians could not answer whether smooth 3D fluid flow can suddenly break down. OpenAI's proposed proof says it can. To find this solution, OpenAI split thousands of agents into groups exploring different ideas, then shared the best findings between them. The search took 88 hours and used about 130B output tokens across 2.7M messages. GPT-6 Astra then spent another 17 hours checking and converting the proof into Lean, a system that verifies mathematical proofs. That Lean version is public, so mathematicians can inspect the argument themselves.
@rohanpaul_ai@rohanpaul_aiAI 评分3232 
@rohanpaul_ai@rohanpaul_aiAI 评分6262
引用@OpenAI@OpenAIThis model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
@rohanpaul_ai@rohanpaul_aiAI 评分2727 @rohanpaul_ai@rohanpaul_ai精选AI 评分6868 
推荐理由:用多家近期 AI 融资的估值收入倍数做横向对照,呈现 Cognition 这轮融资在其中的位置。
@rohanpaul_ai@rohanpaul_ai精选AI 评分7777
引用@OpenAI@OpenAIWe’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
推荐理由:原文给出智能体搜索的耗时与 token 消耗,可了解大规模智能体协作做数学研究的具体形态。
@rohanpaul_ai@rohanpaul_aiAI 评分55 抱歉,主推文内容仅包含一个链接(https://t.co/5vDr0XAjg8),没有可翻译的文字正文。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_ai精选AI 评分6969 
推荐理由:原文披露芯片厂商以信用担保介入数据中心融资竞标,可观察算力扩张中厂商资产负债表扮演的新角色。
@rohanpaul_ai@rohanpaul_aiAI 评分2626