一篇 arXiv 论文对 LLM 作为合成问卷受访者做了心理测量审计,发现 LLM 能复现人类问卷结果的大致方向,但无法还原真实的响应分布。用 LLM 生成的问卷数据训练的模型在预测真人时平均 R² 为 -0.18,而用真人数据训练时为 0.28。作者认为 LLM 适合只需判断效应方向的低成本预调查,但在分布、效应量、相关性和下游分析上无法可靠替代真人受访者。
LLMs can reproduce the broad direction of human survey results, but not the actual human response distribution.
LLMs are useful for cheap pilot surveys where you mainly want the direction of an effect, but this paper finds they are not reliable replacements for actual human survey respondents when you care about realistic distributions, effect sizes, correlations, or downstream analysis.
Models trained on LLM-generated survey data averaged R² = -0.18 when predicting real humans, versus 0.28 when trained on human data.
– arxiv. org/abs/2608.14606
Title: "Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents"
来源:@rohanpaul_ai · x.com