AI 导读
Artificial Analysis 推出 MLCR-AA,从完整性、准确性、简洁性三个维度为模型回答打分。Claude Fable 5 达 73.8% 完整性与 90.1% 准确性,GPT-5.6 Terra (max) 准确性 93.7% 但完整性仅 33.9%。准确性较易达成,完整性才是模型分水岭;医疗记录审查中,准确但不完整的回答仍不可用。
正文
The latest models stay faithful to source information, but their findings and responses miss key details experts include. MLCR-AA scores each answer on completeness, accuracy, and concision. Accuracy is the easier of the two deciding dimensions, and completeness is where models separate: Claude Fable 5 reaches 73.8% completeness at 90.1% accuracy, while GPT-5.6 Terra (max) records 93.7% accuracy and 33.9% completeness. For medical record review, an answer that is accurate but incomplete can still be unsuitable.
来源:@ArtificialAnlys · x.com