Netflix 公开其"因为你看过"推荐理由背后的系统:一个模型生成解释文案,另一个模型在展示前逐条打分,人工每周审核。其新论文提出,生产环境中的 LLM-as-a-Judge 有生命周期,需像其他模型一样持续维护,流程分构建标注样本、调优评审模型、作为门控运行、监测漂移四阶段。
Netflix has explained the system behind those short "because you watched" lines: an AI writes them, another AI grades them, and humans audit weekly.
An AI judge is usually validated once and then trusted forever.
New Netflix paper argues a judge running in production has a lifecycle and needs maintaining like any other model.
At Netflix, one model writes the short lines telling members why a title was recommended, and another scores every one before it is shown.
Netflix splits that work into four stages: building labelled examples, tuning the judge, running it as a gate, and watching it for drift.
Checking whether the judge agrees with human labels is not enough.
It also has to reject a bad explanation for the same reason a person would, so tuning runs on written rationales rather than pass-fail marks.
A weekly human review then sets the bar by how far the raters disagree among themselves.
In a 5-week test against no explanation at all, members shifted slightly toward titles they had not watched and more often ended a browse by playing something.
– arxiv. org/abs/2608.18300
Title: "The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations"
来源:@rohanpaul_ai · x.com