AI 导读
开发者 @speechjsp 基于 MiniCPM-o 4.5 微调出多模态双工交互智能体 Gander,其 Cerebellum 模块经额外训练与系统级优化,实现原生双工交互。Gander 将连续音视频感知、实时对话与异步智能体执行整合进一个持续交互循环,而非把语音交互与智能体执行当作分离步骤。
正文
🤖 What if an AI agent could keep seeing, listening, talking, and working — all at the same time?
Developer @speechjsp built Gander, a multimodal duplex interaction agent that combines continuous audio-visual perception, real-time conversation, and asynchronous agent execution.
Gander’s Cerebellum is fine-tuned from MiniCPM-o 4.5, with additional training and system-level improvements for native duplex interaction.
Instead of treating voice interaction and agent execution as separate steps, Gander brings them together in one continuous interaction loop.
来源:@OpenBMB · x.com