跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 2026-08-22AI 评分30
AI 导读

Artificial Analysis 推出 MLCR-AA,是其对 Wisedocs 基准的实现,仅运行私有的 Expert 与 Compound 层级,纯文本且使用干净上下文,衡量长有效上下文推理而非噪声容忍度。

正文

MLCR-AA is our implementation of the Wisedocs benchmark, with the following specifications:

➤ We run only the private Expert and Compound tiers, text-only and with clean context. We do not currently pad case files with filler documents, so this measures reasoning over long meaningful context rather than noise tolerance

➤ Grading uses a three-model judge panel (Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.5) and aggregates decisions via majority vote for each task and criterion. The top level score requires a response to be correct and complete by majority vote, and to pass the concision test

➤ Every question is run three times, and we report the mean across repeats

➤ Wisedocs has made various revisions to the dataset, and we use the latest version of the holdout set which was revised in August 2026

Scores on our leaderboard are therefore not directly comparable to Wisedocs' original announcement. MLCR-AA is not a component of the Artificial Analysis Intelligence Index.

来源:@ArtificialAnlys · x.com