GPT-6 Astra 已上线 YouMind 确实非常强大 非常值得体验 https://t.co/tbWZ3ydg69
X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(541)
@lifesinger@lifesingerAI 评分2828 
@rohanpaul_ai@rohanpaul_aiAI 评分44 @rohanpaul_ai@rohanpaul_aiAI 评分3030 
@elonmusk@elonmuskAI 评分2626 @ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分2020 用户已开始对 GPT-6 Astra 进行测试,先基于传统工作流程和软件验证其表现,若达标再进入创新流程与建构测试。推文未披露具体测试项目、benchmark 分数或结果,仅表示"坐等结果"。

@opencode@opencodeAI 评分3131 最近加入 Zen: - GLM 5.3 - GLM 5.3 Flash - Muse Spark 1.3 - DeepSeek V4 Flash Vision Exp
@dexhorthy@dexhorthyAI 评分55 @dexhorthy@dexhorthyAI 评分1111 SlopCodeBench 更新 - Astra 已加入 https://t.co/bClNb0CDg2

@Replit@ReplitAI 评分2525 @Replit@ReplitAI 评分5454 @Replit@ReplitAI 评分5555 @Replit@ReplitAI 评分4444 
@cb_doge@cb_dogeAI 评分3232 人们涌向奥斯汀去乘坐 Cybercab:https://t.co/QKhHF7EQph

@emollick@emollickAI 评分3333 
@cb_doge@cb_dogeAI 评分3030 
@hongming731@hongming731AI 评分3232 引用@hongming731@hongming731https://t.co/wcJuMxEHT5
@hongming731@hongming731AI 评分55 @emollick@emollickAI 评分3434 
@jxnlco@jxnlcoAI 评分55 
@omarsar0@omarsar0AI 评分1616 @cb_doge@cb_dogeAI 评分3434 
@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分1919 卧槽!GPT-6 Astra 有了!!! 赶紧冲啊朋友们!!!!!!!! https://t.co/dfjZl1jWZP

@SemiAnalysis_@SemiAnalysis_AI 评分55 https://t.co/SRSZy2ixlL https://t.co/sElpE0L5lb

@thsottiaux@thsottiauxAI 评分44 @thsottiaux@thsottiaux精选AI 评分6767 推荐理由:官方公布 Astra 提前上线,并同步重置 Plus、Pro 与 Business 的额度,读者可了解本轮开放的适用档位与时间节点。
@cb_doge@cb_dogeAI 评分5757 
@rohanpaul_ai@rohanpaul_aiAI 评分55 https://t.co/S4WgxVTCA2 (注:主推文仅含一个链接,无实质文字内容;引用推文仅含"Full video"及链接,同样缺乏可概括的新闻信息,无法生成符合规则的标题与正文翻译。)
引用@rohanpaul_ai@rohanpaul_aiFull video https://t.co/Ci7NIPQl70
@rohanpaul_ai@rohanpaul_aiAI 评分5151 
@cb_doge@cb_dogeAI 评分5151 
@elonmusk@elonmuskAI 评分66 @elonmusk@elonmuskAI 评分1919 @alexandr_wang@alexandr_wangAI 评分4444 引用@ArtificialAnlys@ArtificialAnlysAnnouncing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index v4.2 changelog: + AA-Briefcase, our agentic knowledge work evaluation with a private test set + @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! Intelligence Index v4.2 changes in detail: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5. ➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure. Key results: ➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google ➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier ➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
@kimmonismus@kimmonismusAI 评分2626 @omarsar0@omarsar0AI 评分4444 一项研究称 LLM 多智能体系统实际只需约六种通信拓扑:将码本容量从 8 扩到 64 后,通过奖励过滤存活下来的拓扑仍收敛到大致相同的六种。

@cb_doge@cb_dogeAI 评分6464 
@sama@sama精选AI 评分8989 引用@sama@samaGPT-6 Astra is now available to all Pro, Enterprise, and Business Premium users in Work/Codex, and is available in the API. We will start rollout to Plus and Business users next. Thank you for the patience.
推荐理由:GPT-6 Astra 的可用范围从 Pro 和 Enterprise 扩展到 Plus 与 Business 用户,读者可据此了解其开放进度。
@kimmonismus@kimmonismus精选AI 评分8787
引用@AnthropicAI@AnthropicAIChecking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help. Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written. Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized. We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before. You can read about the process on our Science Blog: https://t.co/ryYnDEAU6J And see the complete proof on GitHub: https://t.co/wlYMXYnofz
推荐理由:原文列出 Claude 形式化费马大定理所用代码规模、支撑定理数量与开源入口,读者可据此了解 AI 自动形式化的当前进展。
@testingcatalog@testingcatalogAI 评分6262 引用@thsottiaux@thsottiauxOK nevermind, the team and Astra did a good job and our systems are more scalable than we anticipated. Astra is now rolled out to all Plus and Business users too. Hope you have a blast and let us know how it goes!
@perplexity_ai@perplexity_ai精选AI 评分6666 引用@perplexity_ai@perplexity_aiWe evaluated GPT-6 Astra on WANDR. It scored 0.682 at $11.98 per task, the highest score of any model we tested. GPT-6-Astra scored 13.5% higher than Fable 5.1 at 6.1% lower cost, and 27.0% higher than Opus 5 at 3.3% higher cost. https://t.co/SyYmD38qvq
推荐理由:原文附有 WANDR 得分与单任务成本对比,读者可了解 GPT-6 Astra 在同类模型中的位置。
@OpenRouter@OpenRouterAI 评分1111