跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 14 天前AI 评分51
AI 导读

OpenAI 发布心理健康对话基准 MentalHealthBench,GPT-6 Astra 得分 57.3,GPT-4o 为 32.1。该基准由 22 个国家的 80 多名执业心理学家和精神科医生共同参与创建,覆盖 19 种语言和近 20 个亚专科。

正文

OpenAI's new MentalHealthBench puts GPT-6 Astra at 57.3, versus GPT-4o's 32.1, on realistic mental health conversations.

Most mental health AI evaluations have centered on emergencies and broad safety criteria, leaving everyday and ambiguous conversations much less measured.

So OpenAI co-created this benchmark with more than 80 licensed psychologists and psychiatrists from 22 countries, spanning 19 languages and nearly 20 subspecialties.

Each synthetic conversation gets a custom expert rubric covering behaviors such as seeking context, preserving user agency, safety, and appropriate guidance.

At least 3 experts reviewed each case, and a criterion survived only when 2 agreed and a 3rd did not contradict it.

GPT-5.6 Sol then grades model answers against those human-written criteria, so score quality still partly depends on an LLM judge.

引用@OpenAI@OpenAI
We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was built with input from more than 80 mental health clinicians. We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work. https://t.co/VTm5ZgxJbl
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com