跳到正文
原文
The Decoder· Manuel Uth·· 3 小时前AI 评分53

Aleph Alpha 测试称中国 AI 模型在敏感话题上重复官方立场或拒答

Chinese AI models parrot state doctrine or refuse to answer on sensitive topics

AI 导读

Aleph Alpha 用自建基准测试 Qwen、DeepSeek、Kimi 等模型在 967 个敏感话题上的表现,AI 评分系统仅将 17% 至 41% 的回答评为平衡,其余重复官方立场、回避或拒答,Claude Sonnet 5 和 Mistral Small 平衡回答比例分别为 70% 和 92%。

来源:The Decoder · the-decoder.com