Black Forest Labs 发布 Flux 3 Image,支持多步骤局部编辑
Black Forest Labs 发布 Flux 3 模型家族的图像模型 Flux 3 Image,称可在多步骤编辑时不改动图像其他部分,覆盖文生图、图生图、文字渲染和照片级写实。
Black Forest Labs 发布 Flux 3 模型家族的图像模型 Flux 3 Image,称可在多步骤编辑时不改动图像其他部分,覆盖文生图、图生图、文字渲染和照片级写实。
Black Forest Labs 于 10 月 2 日发布图像生成模型 FLUX 3 Image,支持最高 4K 分辨率生成。模型基于 FLUX 3 基座,可在 0–1000 坐标网格上为元素指定 ID、描述和边界框实现精准排布,单次最多融入 10 张参考图像,并支持保持其他像素不变的指定区域编辑。目前为付费服务,开放模型版本将在数周内公开。
推荐理由:独立评测方自托管实测两个图像榜单排名,并对比上一代和同类开源模型,读者可了解其在开源阵营中的位置。
Introducing FLUX 3 Image. Control every pixel. Make precise multi-turn edits without changing any other pixel. Lay out the image exactly how you want using bounding boxes. Generate in up to 4K to preserve details. Use up to 10 references to compose an image. Commercial Weights available for companies running image generation at scale. Open Weights version of FLUX 3 Image is launching in the coming weeks.
推荐理由:原文给出 FLUX 3 Image 的多轮编辑、边界框布局和 4K 输出等具体能力,并说明已上线 OpenRouter。
Ideogram 发布新模型 Ideogram 4.5,宣称编辑图像局部时保持其余部分不变,针对 GPT-Image 2.5 和 Nano Banana 仍易产生伪影的问题。模型提供四档质量,单张 0.8 到 22 美分,原生 2K 分辨率,已在 Ideogram 平台和 API 上线,合作方包括 Picsart、Runway、Pika 和 Leonardo AI,官方称开放权重版本即将发布。
Viggle Turbo v0.3 for Qwen-Image-2.1 is out! - Less grain than v0.2.1, a touch softer - New 9-step mode: finer detail, small text - ComfyUI: LoRA or single-file int8/fp8/GGUF Model: https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo
viggle-turbo for Qwen-Image-2.1 isn't just faster — for most prompts, it's just as good as base.
inclusionAI 在 Hugging Face 开源 Ming-UniVision-16B-A3B,基于 MingTok 连续视觉 token,在单一自回归 NTP 框架内统一图像理解与生成,官方称联合训练收敛速度提升 3.5 倍。
inclusionAI 发布 Ming-Image-0.1-Design-Layer,可将一张平面设计图按指定层数拆解为多个 RGBA 图层,采用 MIT 许可证。模型推荐 1024 分辨率、12 步采样、CFG 2.0、BF16 精度,需单张 80 GiB 显存 CUDA GPU,可用 vLLM-Omni 部署,提示词增强支持 Ling-3.0-flash-VL 或 qwen3.8-27B。
inclusionAI 发布 Ming-Image-0.1-Design,一个面向 UI、信息图、海报等文字密集视觉设计的 6B 文生图模型,支持 RGBA 透明背景输出。
Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo
Hy Image3.5 preview is now available in ComfyUI. Professional-grade image generation, +30% win rate in human eval vs Hy Image3.0 → Text to image and Image to image in one model, up to 2K → Multilingual text, symbols, and small print that render correctly → Cinematic, comic, commercial photography, and illustration styles → Identity and product features that hold through scene, outfit, and style changes
Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!
推荐理由:官方确认 Qwen-Image-2.1 登顶两个图像 Arena 的开源榜首,榜单分数可用于同类模型的横向比较。
来自 @inteldevs 的 Day-0 OpenVINO 支持!🥳 Qwen-Image-2.1 已可在 Intel 硬件上优化运行。一个开放权重 checkpoint,同时支持生成和编辑。👇
We're excited to offer Day0 OpenVINO support for Qwen-Image-2.1 Read more about what you can accomplish here: https://ms.spr.ly/6019a5nm5
Generating high-quality images is cheaper and faster than ever. Muse Image, MAI-Image-2.6 and GPT Images 2.5 have substantially shifted the Text to Image Pareto frontiers for both price and speed in recent weeks.
Microsoft 于 9 月 4 日发布 MAI-Image-2.6,并面向 Microsoft Foundry 开发者推出低延迟版本 MAI-Image-2.6-Flash。
推荐理由:原文来自官方发布,给出排名、速度和价格效率数据,读者可据此评估两个版本在生产场景中的取舍。
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
Microsoft 发布 MAI-Image-2.6,在 Arena 文生图排行榜上排名第二,领先 Meta、Google 和 xAI 的模型。该模型总体 Elo 比 MAI-Image-2.5 提升 +79,文本渲染单项提升 +91 Elo,各评测类别均有进步。目前可在 Arena 上试用,本周晚些时候上线 MAI Playground,并将很快登陆 Microsoft Foundry 等产品。
Microsoft AI 发布 MAI-Image-2.5-Pro 和 MAI-Voice-2-Flash,两款模型均进入公开预览。
推荐理由:原文给出两款新模型的定价、性能数字和产品落地情况,可据此比较其质量速度成本取舍。
Google DeepMind 推出图像模型 Nano Banana 2 Lite(gemini-3.1-flash-lite-image),文生图延迟 4 秒。
推荐理由:官方公告给出了两款模型的定位、价格、延迟和已知限制,开发者可直接据此选型和接入。
Microsoft MAI 团队发布图像模型 MAI-Image-2.5 及更快更便宜的 MAI-Image-2.5-Flash,MAI-Image-2.5 在 Arena 图像编辑榜排名第 2(超过 Nano Banana 2),文生图榜排名第 3,较 MAI-Image-2 总分提升 75 分。
推荐理由:官方公布了 Arena 图像编辑第 2 的排名、与 Flash 双版本定价和 Foundry 入口,读者可据此评估生产图像工作流的选型。
Microsoft MAI 在 Build 2026 主题演讲中发布七款新模型,覆盖图像、语音、转录、推理和编码。
推荐理由:作者以当事方身份逐一公布七个新模型的定位、基准数字和开放渠道,读者可以对照竞品评估各自适用场景。
微软发布 MAI-Image-2-Efficient 图像模型,定位量产场景,成本比旗舰低 41%,适用于产品图、营销素材、UI 原型和批量管线,支持短文本和实时交互工作流。
Microsoft AI 宣布发布 MAI-Transcribe-1、MAI-Voice-1 和 MAI-Image-2 三款模型,现已在 Microsoft Foundry 和 MAI Playground 开放。
推荐理由:官方公布三款 MAI 模型的 Foundry 定价与基准对比,开发者可直接据此评估接入成本和可用性。
Microsoft AI 团队发布图像生成模型 MAI-Image-2,称其使 MAI 进入 Arena.ai 榜单前三的文生图实验室行列。
推荐理由:原文给出榜单排名、可用入口和面向创意工作的三大能力重点,读者可据此判断是否试用或接入。
Microsoft AI 于 2025 年 10 月 13 日宣布 MAI-Image 1,这是其首个完全在内部开发的图像生成模型,首日进入 LMArena 文生图模型前十。
推荐理由:微软首次完全自研的图像生成模型,官方说明了训练重点、擅长场景和后续落地产品,可了解其定位与可用入口。