AI 导读
• 原生 256K 上下文,可通过 YaRN 扩展至 1M • 在编程和办公任务上强于 Qwen3.7-Plus,训练成本约为其 1/9 • 在 1M 上下文下,注意力 kernel 在 prefill 阶段最高快 7.6×,在 decoding 阶段最高快 4.9× super f*cking good
正文
• Native 256K context, extendable to 1M via YaRN
• Stronger than Qwen3.7-Plus on coding and office tasks at roughly 1/9 the training cost
• At 1M context, attention kernels are up to 7.6× faster during prefill and 4.9× faster during decoding
super f*cking good
来源:@kimmonismus · x.com