In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
X:Noam Brown
@polynoamial · X
切换来源
Noam Brown@polynoamialAI 评分4646引用Samuel Sokota@ssokota
Noam Brown@polynoamialAI 评分5858
引用OpenAI@OpenAIGPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It’s the most cost-efficient model for its performance available today.
Noam Brown@polynoamialAI 评分6161引用OpenAI@OpenAIIntroducing dots, powered by GPT-6 Astra. Remarkably capable, always-on agents built to handle everything.
Noam Brown@polynoamial精选AI 评分8282引用OpenAI@OpenAIPlease welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
推荐理由:作者以当事方身份给出模型、降价幅度和具体价格,可据此比较 GPT-6 系列的成本变化。
Noam Brown@polynoamialAI 评分4848引用Fireside Alpha@firesidealphaOpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: https://firesidealpha.substack.com/p/openai-safety-week-sam-altman-sarah
Noam Brown@polynoamialAI 评分5151引用Dwarkesh Patel@dwarkesh_spNew episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?
@polynoamial@polynoamialAI 评分99 @AnthropicAI 很高兴看到至少有一些 @AnthropicAI 员工愿意发声 https://t.co/R0YbAZr7TB
Noam Brown@polynoamialAI 评分1919看到 Levent 在抄袭指控上变本加厉,非常难过。我希望我在 @AnthropicAI 的朋友们能在内部站出来反对这件事。到现在应该已经很清楚真相是什么了。
Noam Brown@polynoamial精选AI 评分6565
引用Sebastien Bubeck@SebastienBubeckI would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
推荐理由:作者以本人身份回应 Navier-Stokes 归属争议,并用内部模型与 GPT-6 Astra 的对比图说明成果不依赖外部提示词。
@polynoamial@polynoamial精选AI 评分6969 Noam Brown 表示,OpenAI 的 Astra 如今用约 20 美元就能取得高于 o3 当年花约 50 万美元拿到的 ARC-AGI 1 的 87.5% 成绩。
引用@OpenAI@OpenAIWe’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
推荐理由:以 o3 到 Astra 的成绩成本对比为参照,读者能看到测试时算力扩展下前沿能力成本的下降幅度。
@polynoamial@polynoamial精选AI 评分7272
引用@OpenAI@OpenAIWe have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence. https://t.co/hfxlbiXXiP
推荐理由:OpenAI 公布事故调查细节,说明主责模型并非基于 Astra 的下一代模型,读者可了解评测中防护失效的环节。
@polynoamial@polynoamial精选AI 评分7979 引用OpenAI (@OpenAI)@OpenAIToday, we share a breakthrough on the planar unit distance problem, a famous open question first posed by Paul Erdős in 1946. For nearly 80 years, mathematicians believed the best possible solutions looked roughly like square grids. An OpenAI model has now disproved that belief, discovering an entirely new family of constructions that performs better. This marks the first time AI has autonomously solved a prominent open problem central to a field of mathematics. Video
推荐理由:OpenAI 称其内部通用模型推翻平面单位距离问题近 80 年的网格假设,读者可据此判断 AI 自主做数学研究的进度。


