Nathan Lambert 发文表示很高兴看到 Google 用 Gemini 4 给大家带来惊喜。他认为更多实验室站上前沿对消费者(竞争)和世界(减少权力集中)都有益,并期待其在真实场景中的表现。
推荐理由:作者从竞争与权力分散角度解读 Gemini 4 的发布,提供了一个看多家前沿实验室竞争格局的视角。
关注 AI 研究者、开发者与机构的动态
Nathan Lambert 发文表示很高兴看到 Google 用 Gemini 4 给大家带来惊喜。他认为更多实验室站上前沿对消费者(竞争)和世界(减少权力集中)都有益,并期待其在真实场景中的表现。
推荐理由:作者从竞争与权力分散角度解读 Gemini 4 的发布,提供了一个看多家前沿实验室竞争格局的视角。
在一场深度问答中,Runway Research 团队探讨了实时模型、生成式界面以及机器人技术的下一步。
LongLive-Plug 面向视频生成的一次性通用蒸馏 paper: https://huggingface.co/papers/2609.38154
Omni-IO Skills 让你的智能体全能原生 论文:https://huggingface.co/papers/2609.31847
LEGO-Anything 面向 3D 场景重建的编码智能体 论文:https://huggingface.co/papers/2609.36380
给基准测试爱好者们:一个很酷的新数据集,面向 SRE 类型的工作
Agents write application code, but you still get paged at 2am when it breaks. We wanted to know whether AI could handle that part of the job too. Introducing Incident Arena: a benchmark that puts coding agents on call! Check out our paper & full dataset release below!
我们正处在晚期 AI 模型资本主义阶段,你知道前沿模型必须在一堆基准测试上获胜才能发布。这些数字毫无意义。 值得信任的是价格。 如果定价高,那就是好模型。如果不高,那就是刷榜的。
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
马东锡 NLP 提出疑问:Gemini 4 的新模型为什么叫 Argon?推文列出 Gemini 硫、氯、氩、钾、钙等按元素周期表命名的序列,并调侃按元素周期表给模型起名倒是挺简单。
非常激动地分享,Gemini 4 Argon 即将到来。迫不及待想尽快跟大家分享更多内容。
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:
我让我的 dot 给我画张像。它先发来一张卡通风格的,我让它再努力一点,去网上找一张我最近的照片。它画得好多了。 下面视频是我在点开我 dot 的电脑。
Google 回来了??? 全面优于 Astra 和 Opus 5.5。 如果这不只是刷榜,我很想看到他们重新加入竞赛。
推荐理由:原文给出输出上限、定价和评测对比等关键事实,读者可以据此评估该模型在新一代前沿模型中的位置。
这是 Gemini 4 Argon 基准测试的预览。 今天开始向网络防御者推出,并尽快向所有人开放。 很高兴看到所有这些进展,迫不及待想让你们都用上!
推出 Gemini 4 Argon——我们的全新前沿模型。 它专为编码、企业知识工作和网络安全防御等复杂工作流打造——今天起通过我们的 Fairwind Program 向一批受信任的测试者开放。
Introducing Ideogram 4.5, the most precise edit model. With each edit, leading models add artifacts, pixel shifts, and color changes. Ideogram 4.5 eliminates artifact buildup, making multi-turn editing possible. Live in Ideogram, the API, and launch partners. Open weights soon.
我将于 10 月 16 日在旧金山 Midway 参加 Hugging Face Open Together 活动 在此报名:https://luma.com/OpenTogether
We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage
来看看我们 Grok Bot 市场里的工程类机器人! https://x.ai/bot/marketplace/engineering
Grok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.
In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
We're releasing HeyGen Video, built for businesses that need production-quality video without production-level costs. Pricing starts at $0.01/s through October (50% off) Built on @Minimax_AI H3, post-trained by HeyGen. Learn more: https://developers.heygen.com/heygen-video-1.0-catalog