跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-30AI 评分44
AI 导读

腾讯论文发现,多模态模型切换到非思考模式虽能降低延迟,但用户可见的响应失败概率大幅上升。为此提出的 PatternEval 是一个 2415 条提示词的多模态基准,从思维链泄漏、重复、逻辑矛盾和表演式推理 4 个维度评估答案形态而非正确性,能捕捉正确性检查看不到的质量问题。

正文

Switching a multimodal model into non-thinking mode saves latency, but this Tencent paper finds it also makes user-visible response failures much more likely.

PatternEval scores the shape of an answer rather than its correctness, and finds that non-thinking inference breaks user-facing response quality far more often.

PatternEval is a 2,415-prompt multimodal benchmark that scores 4 of them: chain-of-thought leakage, repetition, logical contradiction, and performative reasoning.

That catches quality problems a correctness check never sees.

If you ship a fast mode, its answers need their own eval.

来源:@rohanpaul_ai · x.com