AI 导读
模型自行决定何时倾听。遇到难词时多等一会儿,遇到简单词则更快输出,利用自适应延迟预测每个 token 并提升准确率。借助自适应延迟,模型在速度与准确率的权衡上接近帕累托前沿。
正文
The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy. With adaptive delay, the model is near the pareto front on speed-accuracy tradeoff.
来源:@finkd · x.com