跳到正文
@berryxia· @berryxia · X·· 2026-06-11AI 评分62
AI 导读

mlx-vlm v0.6.3 上线,当天即支持 Google DeepMind 的 DiffusionGemma 与 Cohere North Mini Code 1.0 在 Mac 本地 MLX 运行。DiffusionGemma 采用新架构,以 256 token 整块并行生成,配合双向注意力与迭代自纠错,26B MoE 仅激活 3.8B,量化后 18GB 可跑;North Mini Code 为 30B MoE、3B active,BF16 下约 66 tok/s。用户可通过 uv pip install -U mlx-vlm 安装,并以 uv run mlx_vlm.server --model MODEL-REPO 启动服务体验。

正文

Prince Canuma直接把Google刚发布的DiffusionGemma和Cohere North Mini Code当天塞进Mac本地MLX,零等待直接把玩咯!

mlx-vlm v0.6.3刚上线,DiffusionGemma这个新架构直接生成256 token整块、双向注意力+迭代自纠错,26B MoE只激活3.8B,量化后18GB就能跑。

North Mini Code 30B MoE也只要3B active,BF16下66 tok/s起步。

全靠和Google DeepMind、Cohere的深度合作,Day-0支持拉满!

一键安装即可体验啊~

地址:huggingface.co/collections/m…

引用Prince Canuma (@Prince_Canuma)@Prince_Canuma
mlx-vlm v0.6.3 is here 🚀 Day-0 support for TWO new models from our partners we work closely with: 🔥 @GoogleDeepMind DiffusionGemma — a genuinely new architecture. Instead of token-by-token, it generates 256-token blocks in parallel with bi-directional attention and iteratively self-corrects the whole block, image-generator style. 26B MoE, only 3.8B active, fits in 18GB quantized. Day-0 MLX support via our Google DeepMind partnership, with long-context prefill tuned and ready. 🔥 @cohere's North Mini Code 1.0 — a 30B MoE with just 3B active, running ~66 tok/s in BF16 before any compression. Day-0 on MLX thanks to our close collaboration with the Cohere team. Get started today — install from source: > uv pip install -U mlx-vlm Then serve the model and point your favorite agent at it (pi, opencode, hermes, etc.): uv run mlx_vlm.server --model MODEL-REPO Model collection 👇🏽
在 X 查看被引用的帖子

来源:@berryxia · x.com