AI 导读
蚂蚁百灵发布 Ling-3.0-flash-VL,在 Ling-3.0-flash 基础上增加图像和视频输入。该模型用 ViT 提取视觉特征,两层 MLP projector 将其与文本对齐,VideoRoPE 负责编码空间与时间信息。其 42 层混合骨干按 5:1 混合 KDA 与 Gated MLA,支持 1M token 上下文。
正文
Ling-3.0-flash-VL adds image and video inputs to Ling-3.0-flash. A ViT extracts visual features, a two layer MLP projector aligns them with text, and VideoRoPE encodes space and time. Its 42 layer hybrid backbone mixes KDA and Gated MLA at 5:1 and supports a 1M token context. https://t.co/T3ziRA0GPV
来源:@AntLingAGI · x.com