X:Testing Catalog
@testingcatalog · X
切换来源
@testingcatalog@testingcatalogAI 评分22 @testingcatalog@testingcatalogAI 评分5151 
@testingcatalog@testingcatalogAI 评分2929 突发 🔥:ASTRA,也就是 GPT-6,将于今天发布! 今天有大量测试时间 👀 https://t.co/A3b4LoEV98 https://t.co/rGUDDA5f7a

引用@OpenAI@OpenAIhttps://t.co/U1TxCPDwwY
@testingcatalog@testingcatalogAI 评分2727 
@testingcatalog@testingcatalog精选AI 评分7070 推荐理由:原文给出模型权重之外还公开训练代码与两个架构组件,读者可据此判断这次开源发布的具体范围。
@testingcatalog@testingcatalogAI 评分6464 
@testingcatalog@testingcatalogAI 评分3333 引用@testingcatalog@testingcatalogDAILY AI BRIEF 🗞️ — Sept 2 OPENAI 🔥: > Official “Path to Astra” post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is “coming soon” — advanced cyber tools stay limited to testers / Daybreak Blue at first. > @M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha. GOOGLE 🔥: > Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway. > WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today. > Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later. ANTHROPIC 🔥: > Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads — about 25% cheaper typically, up to 45% on heavy agent runs. META 🔥: > Muse Voice Transcribe is live — MSL’s first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available. XAI 🔥: > Elon: “Grok 4.7 comes out in 10 days.” That’s ~Sept 12. Reply to Tobi on Grok 4.6. ALIBABA 🔥: > Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 — 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max. WORLD LABS 🔥: > Fei-Fei Li’s lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet. * Too much is happening, and I have some scoops planned for today, so I don't want to spam the algo. ** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.
@testingcatalog@testingcatalogAI 评分3535 Astralogy 🔮 https://t.co/3dTNqHQACa https://t.co/BNvKrmvmKZ

引用@testingcatalog@testingcatalogOPENAI 🔥: GPT-6-Astra model slug has been spotted on the APIs. If we will actually get it tomorrow, it would be a huge week. Routing first 👀 https://t.co/bFkV6Fw6Ki https://t.co/9CXsb6xdXa
@testingcatalog@testingcatalogAI 评分4848 @testingcatalog@testingcatalogAI 评分4646 
@testingcatalog@testingcatalogAI 评分6161 
引用@finkd@finkdMuse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API. Next up 🍉 and Muse Spark open weights releases coming soon. https://t.co/XQQEDEJGD7
@testingcatalog@testingcatalogAI 评分3737 

@testingcatalog@testingcatalogAI 评分3535 SPACEXAI 🔥:Grok Bot 现已登陆 Android 平台! Bot 测试时间 👀 https://t.co/8BepqL1CUW https://t.co/awbmsVOwMs

@testingcatalog@testingcatalogAI 评分1010 基准测试 👀 https://t.co/sZMZIdxPqv

@testingcatalog@testingcatalog精选AI 评分7272 
推荐理由:Flash 级模型与 Claude Opus 5 在 DeepSWE 1.1 上只差 3 个百分点,价格阶梯的变化也一并给出。
@testingcatalog@testingcatalogAI 评分22 @testingcatalog@testingcatalogAI 评分2020 @testingcatalog@testingcatalogAI 评分4747 
@testingcatalog@testingcatalogAI 评分3434 
@testingcatalog@testingcatalogAI 评分4848 @testingcatalog@testingcatalogAI 评分66 @testingcatalog@testingcatalogAI 评分4343 
@testingcatalog@testingcatalogAI 评分3636 

@testingcatalog@testingcatalogAI 评分22 @testingcatalog@testingcatalogAI 评分3838 
@testingcatalog@testingcatalogAI 评分3636 @testingcatalog@testingcatalogAI 评分4040 

@testingcatalog@testingcatalogAI 评分6464 
引用@OpenAI@OpenAIAs we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible. Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework. We're previewing how we evaluated the model, how its safeguards have advanced alongside its capabilities, and what we'll continue to learn and improve. https://t.co/OrrTgdU90K
@testingcatalog@testingcatalogAI 评分6060
引用@testingcatalog@testingcatalogBREAKING 🔥: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1! > It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. > On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. Rolling out on Claude now 👀 https://t.co/hxIt2yHzud https://t.co/SsGvoiWNvh
@testingcatalog@testingcatalog精选AI 评分7171 
引用@claudeai@claudeaiWe’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3
推荐理由:两款新模型在 Terminal-Bench 系列基准上的分数对比,可直观看出相对前代的提升幅度与推送状态。
@testingcatalog@testingcatalogAI 评分2929 

@testingcatalog@testingcatalogAI 评分5252
引用@finkd@finkdAlready powering dictation in the Meta desktop app and voice input in Muse Code. Live now on the Meta model API with a zero-data-retention tier. https://t.co/sppqZ0OSXz
@testingcatalog@testingcatalogAI 评分66 @testingcatalog@testingcatalogAI 评分2424 

@testingcatalog@testingcatalogAI 评分66 @testingcatalog@testingcatalogAI 评分5858 @testingcatalog@testingcatalogAI 评分5050 
@testingcatalog@testingcatalogAI 评分5050
引用@testingcatalog@testingcatalogPERPLEXITY 🔥: A new Hybrid mode is coming soon on Mac, letting local models handle some subtasks. > Gemma 4 (for 16 GB machines), Qwen3.6, and Perplexity's own model (For 32 GB machines) are currently supported. > "In Hybrid mode, Computer orchestrates in the cloud and delegates suitable subtasks to a model on your Mac." > "Local models run on your Mac's own memory and processor — the more capable your machine, the larger the models it can run." > "Privacy Gate is a small model that runs entirely on your Mac. It inspects anything headed to the cloud for personal or sensitive information — and if it finds any, you decide how to handle it before anything is sent." That would be huge 👀
@testingcatalog@testingcatalogAI 评分3434 Claude iOS 应用现在在 Max effort 选项旁新增了明确警告,提示其会多消耗 1.5 倍用量。 Max 1.5 ⚠️ https://t.co/VniimTrKe1


@testingcatalog@testingcatalogAI 评分2626 
