跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-21AI 评分58
AI 导读

SenseNova 发布 8B 参数的开源轻量原生统一多模态模型 U1.5-Lite 完整版,覆盖视觉理解、生成与编辑。模型先训练文本渲染与信息图、美学质量、图像编辑等任务专项专家,再用 OPD 把能力迁移进单一轻量统一模型,推理时无需 router、专家切换或手动选模型。

正文

SenseNova's full U1.5-Lite release is basically a transition from "how many things can one model do?" to "can it do them together without falling apart?"

An 8B-param, open-source, lightweight native unified multimodal model for visual understanding, generation, and editing.

It did not chase a bigger model with U1.5-Lite; it chased a model that could reliably combine more visual skills at the same time.

SenseNova U1.5 Lite treats specialization as a training problem, then gives users 1 model for complex prompts, native 4K, text rendering, and local edits.

And because editing is native to the unified model, the source image, edit target, and generated result stay inside the same model workflow.

SenseNova first trains task-specialized experts for text rendering and infographics, aesthetic quality, and image editing. OPD then transfers those capabilities into one lightweight unified model, so inference does not require a router, expert switching, or manual model selection.

The full release also applies task-oriented RL around instruction adherence, visual preference, and edit fidelity.

That maps directly to the visible improvements: stronger complex-prompt handling, better composition and text layouts, stable native 2K/4K high-resolution generation, and local edits that preserve identity, geometry, and untouched regions.

来源:@rohanpaul_ai · x.com