AI 导读
腾讯HY实验室联合四家机构发布Chronicles-OCR基准,用2800张专家标注图像测试AI对3000年中国古文字的识别能力,28个前沿多模态模型几乎全线失败。
正文
鹅厂好的新基准测试,叫Chronicles-OCR。
腾讯HY实验室和四家机构一起做的,专门测AI对3000年中国古文字的识别能力。
2800张专家标注的图像,覆盖甲骨文、金文、篆书、隶书、楷书、行书、草书七大类。
结果28个前沿多模态模型全军覆没。
最强的VLLM在甲骨文上也只拿到14%的准确率。
端到端检测的H-mean最高才16.5%。
GPT-5和Gemini 2.5 Pro直接接近0。
更反直觉的是,开启reasoning模式反而让表现变差。
Chain-of-thought在感知失败的时候,反而放大了幻觉。
模型其实根本没在认字,它认的是载体。
古文字分类准确率能到96.7%,靠的是看到龟壳、青铜器这些容器,而不是看懂上面的字符。
到底非遗中的价值,AI的攻克只有九牛一毛。
The best VLLM scores only 14% on oracle bone script recognition. Chronicles-OCR, a new ancient Chinese character benchmark from Tencent HY and 4 institutions, just put 28 frontier models to the test across 3,000 years of Chinese writing. 🏺🗂️ modelscope.cn/datasets/Virtu… 2,800 expert-annotated images across 7 scripts: oracle bone, bronze, seal, clerical, regular, running, and cursive. Three findings worth knowing: 🔍 End-to-end detection collapses on ancient scripts. Best H-mean: 16.5. GPT-5 and Gemini 2.5 Pro near 0. 🧠 Reasoning makes it worse. Thinking mode hurts performance across almost all models — chain-of-thought amplifies hallucination when perception fails. 👁️ Models recognize the carrier, not the script. Ancient script classification hits 96.7% by spotting turtle shells and bronze vessels, not the characters themselves.在 X 查看被引用的帖子
来源:@berryxia · x.com