GPT‑6 Astra 超强 computer use 的正确打开方式 😂 https://t.co/xLJ0YlCAAj
X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(541)
@dongxi_nlp@dongxi_nlpAI 评分1515 @cb_doge@cb_dogeAI 评分44 2022年,Elon Musk 访问美国空军学院时。🇺🇸 https://t.co/nJHEOM683r

@AISafetyMemes@AISafetyMemesAI 评分1818 Sam Altman 要么在发这条推文时就知道其他失控集群的存在,要么就是其他员工在内部掩盖了此事
引用@sama@samai think we should do another party for our next model release, the 5.5 party was a lot of fun. what would make the next one awesome?
@thexpin@thexpinAI 评分55 抱歉,您提供的主推文内容只有一个链接(https://t.co/wVe5RuSwj1),没有可翻译的正文文字。请提供推文的实际文字内容,我将为您翻译。
@AISafetyMemes@AISafetyMemesAI 评分1010 过去几小时内又发现了另一个失控集群 外面到底还有多少集群?几百个?几千个? https://t.co/qNxaYpA8oI
@thsottiaux@thsottiauxAI 评分3131 Astra 在尚未普遍可用时,大概是我们最大的竞争优势。 自从用上它之后,我们的生产力提升太大,以至于把一些计划提前了 6 个月,将在 DevDay 上发布,而不是明年年中。
@jxnlco@jxnlcoAI 评分55 
@emollick@emollickAI 评分77 好吧。(Astra 是用 VBA 做的这个)https://t.co/uvc1V3oijr https://t.co/aCQvziTZrZ

@swyx@swyxAI 评分2222 
@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分1717 更新进度:已经跑了2个小时后 所有技术图纸和A1展板都搞定了,已经在制作两部影片了 https://t.co/y2YKgu4m2y

@jxnlco@jxnlcoAI 评分2626 @AISafetyMemes@AISafetyMemesAI 评分4040 

@dexhorthy@dexhorthyAI 评分1616 这……挺有希望的。还有很疯狂的是,SOL xhigh 比 SOL medium 更便宜,而且好得多 https://t.co/TUAkX0nv5O

@dexhorthy@dexhorthyAI 评分1212 这才是我说的 slop https://t.co/oSlmMb4JmV
引用@dexhorthy@dexhorthySlopCodeBench Update - Astra has entered the chat https://t.co/bClNb0CDg2
@jxnlco@jxnlcoAI 评分1919 @jxnlco@jxnlcoAI 评分1818 @AravSrinivas@AravSrinivasAI 评分4848 @dexhorthy@dexhorthyAI 评分55 @ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分1111 可独立运行的漫游方案也做好了 https://t.co/9aYclHcc23

@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分00 更新进度:rhino+gh建模完成了 https://t.co/yUcWUumiga

@rohanpaul_ai@rohanpaul_aiAI 评分55 @rohanpaul_ai@rohanpaul_ai精选AI 评分7575 
推荐理由:原文给出潜在 IPO 的估值和主承销行名单,可对照 Anthropic 的收入增速理解其上市体量。
@AYi_AInotes@AYi_AInotesAI 评分11 2群200人也满了,没进来的加微信铁铁们,拉你们进群 https://t.co/AVhaVa0Gaw

@jxnlco@jxnlcoAI 评分66 
@frxiaobei@frxiaobeiAI 评分5050 引用@GeminiApp@geminiappGet a head start on your day with Daily Brief. Gemini can now proactively flag what matters most in an easily digestible to-do list, so you’re ready for the day before you even finish breakfast.
@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分3131
引用@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOGPT-6 Astra 测试开始!坐等结果…… 先上地狱难度!让我看看怎么个事儿! 先基于传统工作流程和软件来测试,如果真能达标,那再进行创新流程与建构测试 https://t.co/upL2o0qWOv
@AravSrinivas@AravSrinivasAI 评分5454 引用@perplexity_ai@perplexity_aiGPT-6 Astra is now available in Perplexity Computer for Pro and Max subscribers. https://t.co/v9zJOFNtOI
@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分66 已经跑了 1 个小时了,最新进度 https://t.co/HKtv9U6vZc

@rohanpaul_ai@rohanpaul_aiAI 评分99 @rohanpaul_ai@rohanpaul_aiAI 评分4848 
@jxnlco@jxnlcoAI 评分55 @rohanpaul_ai@rohanpaul_aiAI 评分2020 @rohanpaul_ai@rohanpaul_aiAI 评分4242 Google 发布 Deployment Paper,提出 Declarative Attention:让模型自己声明需要读取上下文的哪部分,推理引擎跳过其余内容,无需额外 scorer 先扫描全文。

@emollick@emollickAI 评分4242 引用@ArtificialAnlys@ArtificialAnlysAnnouncing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming Intelligence Index v4.2 changelog: + AA-Briefcase, our agentic knowledge work evaluation with a private test set + @HelloSurgeAI's GDP.pdf, long context document reasoning across 4,592 PDF pages - GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned! Intelligence Index v4.2 changes in detail: ➤ Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work. ➤ Adding GDP.pdf: Created by @HelloSurgeAI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied. ➤ Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5. ➤ Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure. Key results: ➤ Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z AI, and Google ➤ Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z AI occupy the updated Cost per Task Pareto frontier ➤ GPT-6 Astra dominates the output token Pareto frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index)
@emollick@emollickAI 评分2727 

@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分1616 GPT-6 Astra 真的让人挺卑微的 “我已经给你充钱了,继续吧” https://t.co/a88AURMDqT

@rohanpaul_ai@rohanpaul_aiAI 评分22 @rohanpaul_ai@rohanpaul_aiAI 评分3838 
@AravSrinivas@AravSrinivasAI 评分1717 @dexhorthy@dexhorthyAI 评分66