AI 导读
Meta 的 Muse Spark 1.2 支持从视觉生成可运行代码、将感知转化为物理动作等广泛多模态任务,并具备音视频理解能力以支撑企业级视频工作流。Meta 公布了新评测与演示,展示该模型视觉理解与推理能力的广度,包括解析多模态观测并调用工具引导机器人在非结构化环境中导航寻找橡皮鸭。
正文
Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use.
Today, we’re sharing new evals and demos that illustrate the breadth of the model’s visual understanding and reasoning capabilities.
Let’s start with a demo that shows how Muse Spark parses multimodal observations and calls tools to guide a robot to navigate in an unstructured environment to find a rubber duck.
🧵👇
来源:@AIatMeta · x.com