The Decoder· Manuel Uth·· 3 小时前AI 评分53
Aleph Alpha 测试称中国 AI 模型在敏感话题上重复官方立场或拒答
Chinese AI models parrot state doctrine or refuse to answer on sensitive topics
AI 导读
Aleph Alpha 用自建基准测试 Qwen、DeepSeek、Kimi 等模型在 967 个敏感话题上的表现,AI 评分系统仅将 17% 至 41% 的回答评为平衡,其余重复官方立场、回避或拒答,Claude Sonnet 5 和 Mistral Small 平衡回答比例分别为 70% 和 92%。
来源:The Decoder · the-decoder.com