Google 上线 Gemini 3.8 Flash,六周内第三个 Flash 版本,主打编码与智能体,1M 上下文,12 月 31 日前优惠价 $0.75/$3.75;另有面向政府与关键基础设施的 Flash Cyber。
DAILY AI BRIEF 🗞 — Sept 3
GOOGLE 🔥:
> Gemini 3.8 Flash is live — third Flash in six weeks. Workhorse for coding and agents, 1M context. Intro price $0.75 / $3.75 through Dec 31.
> Flash Cyber is the defender twin (vuln find + auto-patch). Gated to Fairwind: governments, infra, trusted maintainers.
META 🔥:
> Muse Spark 1.3 is out in Muse Code and the Meta Model API. Zuck: biggest jump yet on coding and agents, “almost too cheap to meter.”
> Beat GPT-5.6 and Opus 5 on DeepSWE 1.1. Watermelon 🍉 and Spark open weights are next.
OPENAI 🔥:
> gpt-6-astra slug is on the APIs. Employees and testers are posting about a possible drop today. Still not public.
> Already tagged Critical for cyber under the Preparedness Framework — extra safeguards, limited tools first.
XAI 🔥:
> Grok Bot is live on Android. Elon: “Grok @Bot now on Android.” Play Store, same “give it real work” agent as iOS.
* Didn't plan to post it initially cuz all the news was covered already, but here you go :P
DAILY AI BRIEF 🗞️ — Sept 2 OPENAI 🔥: > Official “Path to Astra” post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is “coming soon” — advanced cyber tools stay limited to testers / Daybreak Blue at first. > @M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha. GOOGLE 🔥: > Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway. > WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today. > Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later. ANTHROPIC 🔥: > Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads — about 25% cheaper typically, up to 45% on heavy agent runs. META 🔥: > Muse Voice Transcribe is live — MSL’s first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available. XAI 🔥: > Elon: “Grok 4.7 comes out in 10 days.” That’s ~Sept 12. Reply to Tobi on Grok 4.6. ALIBABA 🔥: > Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 — 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max. WORLD LABS 🔥: > Fei-Fei Li’s lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet. * Too much is happening, and I have some scoops planned for today, so I don't want to spam the algo. ** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.在 X 查看被引用的帖子
来源:@testingcatalog · x.com