跳到正文

#图像生成

今日 5 条
今天10月2日周五
  1. IT Home62

    Black Forest Labs 发布生图模型 FLUX 3 Image:支持 4K 生成与元素精准排布

    Black Forest Labs 于 10 月 2 日发布图像生成模型 FLUX 3 Image,支持最高 4K 分辨率生成。模型基于 FLUX 3 基座,可在 0–1000 坐标网格上为元素指定 ID、描述和边界框实现精准排布,单次最多融入 10 张参考图像,并支持保持其他像素不变的指定区域编辑。目前为付费服务,开放模型版本将在数周内公开。

  2. OpenRouter66

    OpenRouter 宣布 Black Forest Labs 的新旗舰图像模型 FLUX 3 Image 已上线,支持文生图与多参考图编辑,原生最高 4K 渲染。据原模型发布信息,FLUX 3 Image 支持精确多轮编辑、用边界框控制排版、最多 10 个参考图合成图像,商业权重已开放,开放权重版本将在未来几周推出。

    引用Black Forest Labs@bfl_ai

    Introducing FLUX 3 Image. Control every pixel. Make precise multi-turn edits without changing any other pixel. Lay out the image exactly how you want using bounding boxes. Generate in up to 4K to preserve details. Use up to 10 references to compose an image. Commercial Weights available for companies running image generation at scale. Open Weights version of FLUX 3 Image is launching in the coming weeks.

    推荐理由:原文给出 FLUX 3 Image 的多轮编辑、边界框布局和 4K 输出等具体能力,并说明已上线 OpenRouter。

  3. The Decoder60

    Ideogram 4.5 发布,主打局部编辑不动图像其余部分

    Ideogram 发布新模型 Ideogram 4.5,宣称编辑图像局部时保持其余部分不变,针对 GPT-Image 2.5 和 Nano Banana 仍易产生伪影的问题。模型提供四档质量,单张 0.8 到 22 美分,原生 2K 分辨率,已在 Ideogram 平台和 API 上线,合作方包括 Picsart、Runway、Pika 和 Leonardo AI,官方称开放权重版本即将发布。

10月1日周四
  1. ViggleAI44

    ✨ Viggle Turbo v0.3 来了! 全新的 9-step 模式带来更干净的画面、更精细的细节,以及更好的小文字渲染。 现在在 ComfyUI 中更易使用——提供 LoRA 或单文件 int8/fp8/GGUF 模型。

    引用Yun Chen@t_mux

    Viggle Turbo v0.3 for Qwen-Image-2.1 is out! - Less grain than v0.2.1, a touch softer - New 9-step mode: finer detail, small text - ComfyUI: LoRA or single-file int8/fp8/GGUF Model: https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

9月26日周六
9月24日周四
9月23日周三
  1. ViggleAI46

    我们为开源社区推出了首个 Qwen-Image-2.1 turbo。 试试 Viggle-Turbo,一个经 DMD 蒸馏的 Qwen-Image-2.1,仅需 4 个采样步即可生成和编辑,无需 classifier-free guidance。 权重:https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo

    引用Hugging Apps@HuggingApps

    Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo

  2. Tencent Hy60

    Hy Image3.5 preview 已可在 ComfyUI 中使用,人类评测胜率较 Hy Image3.0 提升 30%。单模型同时支持文生图和图生图,最高 2K 分辨率,可正确渲染多语言文字与符号,覆盖电影感、漫画、商业摄影和插画风格,身份与产品特征在场景、服装和风格切换中保持一致。

    引用ComfyUI@ComfyUI

    Hy Image3.5 preview is now available in ComfyUI. Professional-grade image generation, +30% win rate in human eval vs Hy Image3.0 → Text to image and Image to image in one model, up to 2K → Multilingual text, symbols, and small print that render correctly → Cinematic, comic, commercial photography, and illustration styles → Identity and product features that hold through scene, outfit, and style changes

  3. Qwen66

    千问(Qwen)宣布 Qwen-Image-2.1 在 Arena 的 Image Edit 和 Text-to-Image 两个榜单均排名第一的开源模型。@arena 引用称其 Image Edit Arena 得分 1367,总排名第 16,距第 15 名 GPT-Image-1.5-high-fidelity 仅差 3 分。

    引用Arena.ai@arena

    Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!

    推荐理由:官方确认 Qwen-Image-2.1 登顶两个图像 Arena 的开源榜首,榜单分数可用于同类模型的横向比较。

9月22日周二
9月19日周六
9月5日周六
9月2日周三
  1. Fei-Fei Li57

    Fei-Fei Li 宣布 World Labs 团队发布从零训练的多模态世界模型 Atlas。该模型支持像素级精准的相机控制生成帧、从单张输入图像重建大场景、通过重排视频帧模拟时空、从一或多张图像原生输出 3D 空间,以及将多张带位姿图像组合成一致的 3D 世界,潜在应用覆盖 VFX 到机器人。

    引用World Labs@theworldlabs

    Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.

8月11日周二
7月24日周五
7月1日周三
6月2日周二
  1. Microsoft AI News65

    Microsoft MAI-Image-2.5 发布,登 Arena 图像编辑榜第 2

    Microsoft MAI 团队发布图像模型 MAI-Image-2.5 及更快更便宜的 MAI-Image-2.5-Flash,MAI-Image-2.5 在 Arena 图像编辑榜排名第 2(超过 Nano Banana 2),文生图榜排名第 3,较 MAI-Image-2 总分提升 75 分。

    推荐理由:官方公布了 Arena 图像编辑第 2 的排名、与 Flash 双版本定价和 Foundry 入口,读者可据此评估生产图像工作流的选型。

4月14日周二
4月2日周四
3月20日周五
10月14日周二
  1. Microsoft AI News61

    Microsoft AI 发布首个完全自研图像生成模型 MAI-Image 1,登上 LMArena 文生图前十

    Microsoft AI 于 2025 年 10 月 13 日宣布 MAI-Image 1,这是其首个完全在内部开发的图像生成模型,首日进入 LMArena 文生图模型前十。

    推荐理由:微软首次完全自研的图像生成模型,官方说明了训练重点、擅长场景和后续落地产品,可了解其定位与可用入口。