Omni-IO Skills 让你的智能体全能原生 论文:https://huggingface.co/papers/2609.31847
全部AI 动态
全部动态
今日 320 条
AK@_akhaliqAI 评分1818
AK@_akhaliqAI 评分2525LEGO-Anything 面向 3D 场景重建的编码智能体 论文:https://huggingface.co/papers/2609.36380

Dongxi 东锡 NLP@dongxi_nlpAI 评分3131
fofr@fofrAIAI 评分55
dex@dexhorthyAI 评分4444给基准测试爱好者们:一个很酷的新数据集,面向 SRE 类型的工作
引用andre --dangerously-skip-permissions@andrezfuAgents write application code, but you still get paged at 2am when it breaks. We wanted to know whether AI could handle that part of the job too. Introducing Incident Arena: a benchmark that puts coding agents on call! Check out our paper & full dataset release below!
Latent SpaceAI 评分7474 Latent Space DevDay 访谈:Ari Weinstein 谈 Computer Use 进展,Nikunj Handa 解析 OpenAI 新 API 栈
Latent Space 在 OpenAI DevDay 当天发布访谈,Computer Use 负责人 Ari Weinstein 表示 Computer Use 已与数月前 180 度不同。
Runway@runwaymlAI 评分4242
The Decoder精选AI 评分7979 Google 发布 Gemini 4 Argon,追赶 OpenAI 与 Anthropic 但未取得明确领先
Google 发布新旗舰模型 Gemini 4 Argon,是其七个多月来首款前沿模型,Artificial Analysis 测试中得 53 分,与 GPT-6 Astra (max)、Claude Fable 5.1 持平,但仍落后 Claude Opus 5.5 的 58 分。
推荐理由:原文汇总了第三方测试与定价细节,指出 Gemini 4 Argon 缩小差距但未领先,且单价优势来自低 token 价格而非效率。
Deedy@deedydasAI 评分3131我们正处在晚期 AI 模型资本主义阶段,你知道前沿模型必须在一堆基准测试上获胜才能发布。这些数字毫无意义。 值得信任的是价格。 如果定价高,那就是好模型。如果不高,那就是刷榜的。
Demis Hassabis@demishassabisAI 评分4646
Karina@karinanguyenAI 评分5252引用Google DeepMind@GoogleDeepMindIntroducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
Dongxi 东锡 NLP@dongxi_nlpAI 评分1818马东锡 NLP 提出疑问:Gemini 4 的新模型为什么叫 Argon?推文列出 Gemini 硫、氯、氩、钾、钙等按元素周期表命名的序列,并调侃按元素周期表给模型起名倒是挺简单。
Runway@runwaymlAI 评分3333
OpenRouter@OpenRouterAI 评分3939
fofr@fofrAIAI 评分3333非常激动地分享,Gemini 4 Argon 即将到来。迫不及待想尽快跟大家分享更多内容。
引用Sundar Pichai@sundarpichaiLots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:
Suno@sunoAI 评分2121
The Verge · AIAI 评分7373 Google 发布 Gemini 4 Argon,初期仅限受信任的网络防御者使用
Google 发布新一代前沿模型 Gemini 4 Argon,称其在软件工程、法律金融等企业知识工作和网络安全防御方面具有前沿性能。初期仅向一组受信任的网络防御者开放,Google 正参与美国政府预发布模型访问的自愿流程并逐步扩大访问。模型已用于 Google 内部工作流,如大规模代码库迁移;Google 将在更广泛发布前加强防范滥用和提示词注入攻击、监测错位等安全措施。
Josh Woodward@joshwoodwardAI 评分2222
Tibo@thsottiauxAI 评分2727我让我的 dot 给我画张像。它先发来一张卡通风格的,我让它再努力一点,去网上找一张我最近的照片。它画得好多了。 下面视频是我在点开我 dot 的电脑。

SemiAnalysis@SemiAnalysis_AI 评分4848
Runway@runwaymlAI 评分2626
Yuchen Jin@Yuchenj_UWAI 评分3030Google 回来了??? 全面优于 Astra 和 Opus 5.5。 如果这不只是刷榜,我很想看到他们重新加入竞赛。

ClaudeDevs@ClaudeDevsAI 评分3030
Ars Technica · AIAI 评分6565 Google 发布 Gemini 4 Argon 模型,暂未开放使用
Google 发布 Gemini 4 Argon,称其在编码、知识工作和网络安全方面性能领先,但模型仍处有限测试,普通用户暂无法使用。DeepSWE v1.1 达 77.9%,高于 GPT-6 Astra、Fable 5.1 和 Opus 5.5;API 定价为每百万输入 token $2、输出 $10,输出上限提升至 100 万 token(此前为 64,000)。
Ammaar Reshi@ammaarAI 评分4141这是 Gemini 4 Argon 基准测试的预览。 今天开始向网络防御者推出,并尽快向所有人开放。 很高兴看到所有这些进展,迫不及待想让你们都用上!

Google AI@GoogleAIAI 评分5353
Google DeepMind@GoogleDeepMindAI 评分3939推出 Gemini 4 Argon——我们的全新前沿模型。 它专为编码、企业知识工作和网络安全防御等复杂工作流打造——今天起通过我们的 Fairwind Program 向一批受信任的测试者开放。

Google DeepMind精选AI 评分7474 Google DeepMind 发布 Gemini 4 Argon,输出上限扩至 1M tokens
Google DeepMind 发布前沿模型 Gemini 4 Argon,先向 Fairwind Program 的可信网络防御者开放,后续将面向开发者、企业和消费者推出。
推荐理由:官方博客给出定价、输出上限和多个基准成绩,可帮读者评估该模型在编码与安全防御场景的实际定位。
Luma@LumaLabsAIAI 评分5454引用Ideogram@ideogram_aiIntroducing Ideogram 4.5, the most precise edit model. With each edit, leading models add artifacts, pixel shifts, and color changes. Ideogram 4.5 eliminates artifact buildup, making multi-turn editing possible. Live in Ideogram, the API, and launch partners. Open weights soon.
AK@_akhaliqAI 评分1414我将于 10 月 16 日在旧金山 Midway 参加 Hugging Face Open Together 活动 在此报名:https://luma.com/OpenTogether

Ars Technica · AIAI 评分6262 RFK Jr. 称 AI 支持其反疫苗观点,媒体实测 Gemini 与 ChatGPT 给出相反答案
美国卫生部长 Robert F. Kennedy 在 MAHA 活动上称 AI 比任何医生更有信息量,建议美国人用 AI 获取医疗第二意见,并称 AI 会推翻口罩、社交距离和疫苗防止传播等专家结论。Ars Technica 实测 Google 的 Gemini 与 ChatGPT,两者均明确回答口罩和社交距离能有效减少呼吸道传染病传播,与 Kennedy 的说法相反。
Aravind Srinivas@AravSrinivasAI 评分5353引用Perplexity@perplexity_aiWe built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage
eric zakariasson@ericzakariassonAI 评分4848来看看我们 Grok Bot 市场里的工程类机器人! https://x.ai/bot/marketplace/engineering
引用Grok Bot@botGrok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.
The DecoderAI 评分6161 OpenAI 与 Synopsys 合作开发芯片设计专用模型 GPT-Synopsys
OpenAI 与 Synopsys 签署多年战略合作,共同构建名为 GPT-Synopsys 的芯片设计专用 AI 模型,目标是让模型能对芯片设计和验证进行推理并直接操作 Synopsys 的 EDA 工具,工程师负责下达设计目标并审核输出。模型运行在 OpenAI 基础设施上,客户数据加密存储且不用于训练,半导体客户的早期测试已在进行,双方将联合销售并分享收入。
Noam Brown@polynoamialAI 评分4646引用Samuel Sokota@ssokotaIn our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
Claude Code GitHub ReleasesAI 评分3636 Claude Code v2.1.286 发布
Claude Code v2.1.286 发布,权限提示新增“2 of 5”计数,全屏列表的“N more”行支持鼠标点击跳转。本次修复涵盖多进程重复打开登录浏览器、claude --resume 与 --continue 丢失并行工具调用轮次、工具返回对象或数字导致的 API 400 错误,以及云会话因容器停止而无法唤醒等问题。
Perplexity@perplexity_aiAI 评分4646
OpenAI@OpenAIAI 评分2828
MiniMax (official)@MiniMax_AIAI 评分5555引用HeyGen@HeyGenWe're releasing HeyGen Video, built for businesses that need production-quality video without production-level costs. Pricing starts at $0.01/s through October (50% off) Built on @Minimax_AI H3, post-trained by HeyGen. Learn more: https://developers.heygen.com/heygen-video-1.0-catalog
Claude@claudeaiAI 评分1919