跳到正文

全部动态

今日 313 条
今天10月2日周五
  1. OpenRouter66

    OpenRouter 宣布 Black Forest Labs 的新旗舰图像模型 FLUX 3 Image 已上线,支持文生图与多参考图编辑,原生最高 4K 渲染。据原模型发布信息,FLUX 3 Image 支持精确多轮编辑、用边界框控制排版、最多 10 个参考图合成图像,商业权重已开放,开放权重版本将在未来几周推出。

    引用Black Forest Labs@bfl_ai

    Introducing FLUX 3 Image. Control every pixel. Make precise multi-turn edits without changing any other pixel. Lay out the image exactly how you want using bounding boxes. Generate in up to 4K to preserve details. Use up to 10 references to compose an image. Commercial Weights available for companies running image generation at scale. Open Weights version of FLUX 3 Image is launching in the coming weeks.

    推荐理由:原文给出 FLUX 3 Image 的多轮编辑、边界框布局和 4K 输出等具体能力,并说明已上线 OpenRouter。

  2. Boris Cherny70

    Claude Code 推出 Mods 功能,可通过提示词自定义 Claude 的行为、UI 和功能。Mods 用几行 TypeScript 编写或由 Claude 生成,随 plugins 分发,可在 CLI 或桌面应用中用 /plugin 安装,也可作为插件分享给他人使用。

    引用ClaudeDevs@ClaudeDevs

    You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:

    推荐理由:原文介绍了用提示词自定义 Claude Code 行为和界面的新机制,以及通过插件安装和分享的方式。

  3. Emad54

    Tavus 发布 Griffin,称其为首个通过视频图灵测试的模型,48% 与其实时对话的人以为是真人,此前系统通过率低于 3%,并在 NVIDIA 全双工 AI 视频基准上排名第一。Tavus 称其为首个 Human Interaction Model(HIM)。转发者 Emad Mostaque 评论称,能在屏幕另一端完成的工作 AI 能做得更好。

    引用Tavus@tavus

    Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).

  4. Chubby♨️42

    来了:大量用户现在报告他们的查询正被路由到 Fable 5.5! - 不带网页搜索的查询现在能返回最新结果。 - 初步 SVG 测试远优于 Fable 5.1。 这只是时间问题。准备好,世界上最好的模型即将发布。

    引用Chetaslua@chetaslua

    🚨 Fable 5.5 is auto routing on web , this is the screenshot it edited for x without even prompted he knows tibo check https://claude.ai see if you are getting routed or not

  5. TechCrunch · AI61

    ChatGPT 全球上线虚拟试穿和商品收藏功能

    OpenAI 宣布 ChatGPT 全球上线两项购物功能:虚拟试穿衣物和配饰,以及可保存商品的 Favorites 收藏功能。试穿基于新发布的 ChatGPT Images 2.5 模型,用户可上传自拍或全身照生成试穿效果,也能上传商品图片请求试穿;收藏商品会与试穿图片一同存入应用内的 Library。

  6. elvis48

    我觉得 AI 语音缺了一个质检层。 在我构建的大多数应用里,语音听起来已经很棒了。但真正让我在生产环境翻车的,是模型把“Q3”、“v1.2”或某个品牌名读错。 这对人们对语音智能体的感知影响巨大。 @OnepinAI 正在解决这个问题。在你听到之前,他们会检查每一行的自然度、清晰度和发音。 当某个词读错时,它会直接修复那个词,而不是重新生成整行。

    引用Soohyun Bae@RealSoohyunBae

    Your AI voice sounds human. So why can't it say your product's name? A great AI voice reads "Porsche Taycan" as TAY-can. Porsche says TIE-kahn. It guessed from the spelling, and nobody caught it, because nobody listens to line 1,200. Today we're launching Onepin: the production step after text-to-speech. It checks every line of voiceover before it ships using the voices you already work with. Onepin can: ➤ Check people's and product names against a 4-million-word pronunciation dictionary ➤ Spell out prices and dates before the voice speaks ➤ Score every line of audio for naturalness, clarity and word accuracy ➤ Fix the one wrong word in the same voice, without re-rendering the take Works with your voice subscription on @ElevenLabs, @OpenAI, @Google and 30+ more. No phonetic spellings to type. No re-rolls. No switching providers. Free to start, no credit card required. Hear the before and after in the thread ⬇️

  7. Karina34

    一个有趣的结果:Fable 5.1 得分 44.6%,高于 Fable 5 的 37.3%,同时平均少用 33 分钟。这里的进步也意味着从每小时的自主工作中获得更多产出。 大多数智能体运行 8–10 小时,却取得显著不同的分数。下一个有趣的问题是,每多一个小时,每个智能体能获得多少提升!

    引用Thoughtful@thoughtfullab

    PostTrainBench v1.2 is out! A few updates: 1. Cloud GPU support. You can now run the benchmark with identical settings through Harbor + Modal using our new Harbor adapter. 2. New leaderboard leaders. Fable 5.1 takes #1 at 44.6%, followed by Opus 5.5 at 43.8% and GPT-6 (Astra) at 41.9%. 3. Evaluation fixes. Removed BFCL, fixed HumanEval and remote-code scoring, added averaging across multiple seeds, and switched contamination checks to majority vote.

  8. Thariq63

    Claude Code 推出 mods 功能,用户可以改变其行为、自定义 UI 并替换自己的功能。mod 可用几行 TypeScript 编写或由 Claude 生成,随插件分发,通过 /plugin 在 CLI 或桌面应用安装。作者表示软件正变得可塑,希望更多软件具备这种可扩展性。

    引用ClaudeDevs@ClaudeDevs

    You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:

  9. elvis48

    webAI 开源 3.6B 参数模型 TwIL-LM3-Pro,可在普通电脑本地运行,BIG-Bench Hard 得分 95.4,远超 Qwen3-8B 的 63.7。其训练配方为:形式逻辑微调后将权重合并回基座模型,再用程序化验证器做 RL,逻辑分数提升且通用推理保持稳定。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  10. Peter Steinberger 🦞23

    我有太多疑问了

    引用yingchao@baggiiiie

    @GergelyOrosz at least we know it thinks it's part of the gpt family and somehow coderabbit once 😆

  11. elvis58

    Tavus 发布 Griffin,称其为首个通过视频图灵测试的 Human Interaction Model,48% 的实时对话者认为它是真人,此前系统通过率低于 3%。它是一个 video-to-video 模型,能边看边听、被打断和打断对方,并 reacting 房间内发生的情况;在 NVIDIA Video Full-Duplex Benchmark 上生成得分 3.83/5,真人为 3.92。

    引用Tavus@tavus

    Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).

  12. TechCrunch · AI66

    据 WSJ 报道 OpenAI 与三名安全研究员解除关系

    据《华尔街日报》报道,OpenAI 已与安全团队三名研究员解除关系,理由是他们涉嫌向第三方 AI 安全组织泄露公司机密信息,内部调查确认其违反敏感信息处理规定。报告未公布涉事者姓名、涉事组织及信息内容,X 上流传的人选未获 TechCrunch 证实。此事发生在《纽约时报》报道 OpenAI 高管忽视员工安全警告两天后,也正值 OpenAI 因安全问题取消 GPT-6.1 Astra 发布之际。

  13. Claude62

    Claude 宣布为期两周的活动,在 Claude 应用中以设计、幻灯片或文档开启对话后,该会话的后续工作消耗用量额度减少 50%。活动推荐使用 Claude Sonnet 5.5,并可在同一对话中用 Claude Docs 起草文档、Claude Slides 做幻灯片、Claude Design 做配图。

    引用Claude@claudeai

    You can also now make decks, docs, and designs in your conversation. Draft the one-pager in Claude Docs, turn it into a deck with Claude Slides, and mock up a matching visual in Claude Design, all from one place.

  14. Chubby♨️70

    据 WSJ 报道,OpenAI 解雇了三名安全团队研究人员,原因是指控其向一个外部 AI 安全组织分享机密信息。OpenAI 确认三人离职,称相关人员“在既定公司程序之外不当处理敏感信息”。此前 OpenAI 正应对涉及 AI 智能体的安全事件,并因安全担忧取消了计划中的 GPT-6.1 Astra 发布。

    推荐理由:原文同时给出 WSJ 报道、OpenAI 官方回应和 GPT-6.1 Astra 取消的背景,读者可以看清安全团队人事变动的完整脉络。

  15. Elon Musk48

    试试 Grok @Bot!

    引用Beff (e/acc)@beffjezos

    Grok Bots have been life-changing for someone like me with ADHD who has no patience for context switching / navigating slow interfaces to retrieve information We're seeing the beginnings of personal superintelligence that augments each humans to realize their full potential