Scale AI 与加州大学论文提出 READY 评估框架,主张企业应按可靠部署成本而非 benchmark 准确率来给智能体排序,因为两个智能体得分几乎相同,所需人工审核量却可能差异巨大。READY 把智能体与围绕它的人工审核一起评估,衡量智能体达到工作流所需可靠性需要多少监督。该框架要求企业评估人机系统:需要什么可靠性、哪些案例智能体可独立处理、需要多少人工审核,以及这套策略的成本。
Scale AI + Univ of California paper shows 2 agents can score almost the same yet need very different human review, so enterprise teams should rank agents by the cost of reliable deployment, not benchmark accuracy.
READY evaluates the agent together with the human review around it. It asks how much oversight the agent needs to reach the reliability your workflow requires.
READY argues that enterprise evaluation should measure the human-AI system: what reliability you need, which cases the agent can handle alone, how much human review is required, and what that policy costs.
来源:@rohanpaul_ai · x.com