AI 导读
在 1M token 上下文长度下,QSA 的注意力 kernel 在 prefill 阶段最高快 7.6×,在 decode 阶段快 4.9×。在 90% 前缀缓存命中率下,Qwen3.8-Flash-Next 的 prefill 吞吐量是 Qwen3.7-Plus 的 8.6×。https://t.co/QZ0koCWvQU
正文
At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit rate, Qwen3.8-Flash-Next delivers 8.6× the prefill throughput of Qwen3.7-Plus. https://t.co/QZ0koCWvQU
来源:@Alibaba_Qwen · x.com