跳到正文

全部动态

今日 24 条
今天10月2日周五
  1. IT Home62

    Black Forest Labs 发布生图模型 FLUX 3 Image:支持 4K 生成与元素精准排布

    Black Forest Labs 于 10 月 2 日发布图像生成模型 FLUX 3 Image,支持最高 4K 分辨率生成。模型基于 FLUX 3 基座,可在 0–1000 坐标网格上为元素指定 ID、描述和边界框实现精准排布,单次最多融入 10 张参考图像,并支持保持其他像素不变的指定区域编辑。目前为付费服务,开放模型版本将在数周内公开。

  2. Chubby♨️45

    webAI 发布 3.66B 参数形式逻辑模型 TwIL-LM3-Pro,可在笔记本本地运行。其综合逻辑评测与 Qwen3-8B 持平,参数量不足后者一半,并在全部六项形式逻辑任务上领先 VibeThinker-3B。该模型基于 IBM Granite 4.2 后训练,Q4 GGUF 权重仅 2.09 GiB,可通过 llama.cpp 本地推理。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  3. TechCrunch · AI63

    Google 发布 Gemini 4 Argon,称其为迄今最强模型

    Google(Alphabet)发布新模型 Gemini 4 Argon,主打防御性网络安全,称其可自主发现、验证并修复关键软件漏洞,目前仅通过 Fairwind 安全计划向部分网络安全合作伙伴开放。该模型也用于编码、调试和代码库迁移等日常工程工作,并称在多项基准上显著领先 GPT-6 Astra 与 Anthropic 的 Fable 和 Opus。

  4. OpenRouter66

    OpenRouter 宣布 Black Forest Labs 的新旗舰图像模型 FLUX 3 Image 已上线,支持文生图与多参考图编辑,原生最高 4K 渲染。据原模型发布信息,FLUX 3 Image 支持精确多轮编辑、用边界框控制排版、最多 10 个参考图合成图像,商业权重已开放,开放权重版本将在未来几周推出。

    引用Black Forest Labs@bfl_ai

    Introducing FLUX 3 Image. Control every pixel. Make precise multi-turn edits without changing any other pixel. Lay out the image exactly how you want using bounding boxes. Generate in up to 4K to preserve details. Use up to 10 references to compose an image. Commercial Weights available for companies running image generation at scale. Open Weights version of FLUX 3 Image is launching in the coming weeks.

    推荐理由:原文给出 FLUX 3 Image 的多轮编辑、边界框布局和 4K 输出等具体能力,并说明已上线 OpenRouter。

  5. Chubby♨️42

    来了:大量用户现在报告他们的查询正被路由到 Fable 5.5! - 不带网页搜索的查询现在能返回最新结果。 - 初步 SVG 测试远优于 Fable 5.1。 这只是时间问题。准备好,世界上最好的模型即将发布。

    引用Chetaslua@chetaslua

    🚨 Fable 5.5 is auto routing on web , this is the screenshot it edited for x without even prompted he knows tibo check https://claude.ai see if you are getting routed or not

  6. elvis48

    webAI 开源 3.6B 参数模型 TwIL-LM3-Pro,可在普通电脑本地运行,BIG-Bench Hard 得分 95.4,远超 Qwen3-8B 的 63.7。其训练配方为:形式逻辑微调后将权重合并回基座模型,再用程序化验证器做 RL,逻辑分数提升且通用推理保持稳定。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  7. elvis58

    Tavus 发布 Griffin,称其为首个通过视频图灵测试的 Human Interaction Model,48% 的实时对话者认为它是真人,此前系统通过率低于 3%。它是一个 video-to-video 模型,能边看边听、被打断和打断对方,并 reacting 房间内发生的情况;在 NVIDIA Video Full-Duplex Benchmark 上生成得分 3.83/5,真人为 3.92。

    引用Tavus@tavus

    Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).

  8. The Decoder60

    Ideogram 4.5 发布,主打局部编辑不动图像其余部分

    Ideogram 发布新模型 Ideogram 4.5,宣称编辑图像局部时保持其余部分不变,针对 GPT-Image 2.5 和 Nano Banana 仍易产生伪影的问题。模型提供四档质量,单张 0.8 到 22 美分,原生 2K 分辨率,已在 Ideogram 平台和 API 上线,合作方包括 Picsart、Runway、Pika 和 Leonardo AI,官方称开放权重版本即将发布。

  9. TechCrunch · AI56

    AWS 发布开源决策模型 Strands Decider 2B,基于 Qwen3.5-2B

    AWS 发布开源决策模型 Strands Decider 2B,灵感来自 TypeSafe 的 Jev,可在预设选项间高速低成本地做选择并给出置信度。模型完全开源、可本地运行,由 Amazon 杰出工程师 Marc Brooker 的内部项目改进而来,基于 Qwen3.5-2B 的架构但不生成文本,而是输出校准后的选择,同一周 OpenAI 也宣布了类似产品。

  10. Microsoft AI News61

    Microsoft AI 发布 MAI-Transcribe-2-Streaming 及 MAI-Voice-2.1 语音模型

    Microsoft AI 发布实时流式转写模型 MAI-Transcribe-2-Streaming,在 Artificial Analysis 的最终与部分转写准确率均排名第一,支持 60 种语言,接收到音频后约 100ms 内产出首个 partials,内部评测显示实时听写场景出词速度比最接近的竞品快 2x,年底前 introductory 价 $0.54 每小时音频。

    推荐理由:官方公布三款语音模型的具体定价、延迟数字和 Artificial Analysis 排名,可据此评估搭建语音智能体的成本与速度。

10月1日周四
  1. ViggleAI44

    ✨ Viggle Turbo v0.3 来了! 全新的 9-step 模式带来更干净的画面、更精细的细节,以及更好的小文字渲染。 现在在 ComfyUI 中更易使用——提供 LoRA 或单文件 int8/fp8/GGUF 模型。

    引用Yun Chen@t_mux

    Viggle Turbo v0.3 for Qwen-Image-2.1 is out! - Less grain than v0.2.1, a touch softer - New 9-step mode: finer detail, small text - ComfyUI: LoRA or single-file int8/fp8/GGUF Model: https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

  2. The Decoder79

    Google 发布 Gemini 4 Argon,追赶 OpenAI 与 Anthropic 但未取得明确领先

    Google 发布新旗舰模型 Gemini 4 Argon,是其七个多月来首款前沿模型,Artificial Analysis 测试中得 53 分,与 GPT-6 Astra (max)、Claude Fable 5.1 持平,但仍落后 Claude Opus 5.5 的 58 分。

    推荐理由:原文汇总了第三方测试与定价细节,指出 Gemini 4 Argon 缩小差距但未领先,且单价优势来自低 token 价格而非效率。

  3. Karina52

    Google DeepMind 发布新前沿模型 Gemini 4 Argon,面向编码、企业知识工作和网络安全防御等复杂工作流,即日起通过 Fairwind Program 向部分受信任测试者开放。作者引述其 PostTrainBench 得分 45.3%,高于 Gemini 3.1 Pro 的 21.99% 和 GPT-6 Astra 的 44.3%;评测表还显示其在自动化与智能体编码等多项基准领先,但在 FrontierSWE v2、Terminal-Bench 4.0 等项落后于对比模型。

    引用Google DeepMind@GoogleDeepMind

    Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.

  4. fofr33

    非常激动地分享,Gemini 4 Argon 即将到来。迫不及待想尽快跟大家分享更多内容。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  5. The Verge · AI73

    Google 发布 Gemini 4 Argon,初期仅限受信任的网络防御者使用

    Google 发布新一代前沿模型 Gemini 4 Argon,称其在软件工程、法律金融等企业知识工作和网络安全防御方面具有前沿性能。初期仅向一组受信任的网络防御者开放,Google 正参与美国政府预发布模型访问的自愿流程并逐步扩大访问。模型已用于 Google 内部工作流,如大规模代码库迁移;Google 将在更广泛发布前加强防范滥用和提示词注入攻击、监测错位等安全措施。