Netflix 团队发表论文,把大规模推荐解释中的 LLM-as-a-Judge 描述为包含 Birth、Training、Deployment、Monitoring 四个阶段的完整生命周期,每周处理数十万条剧集级推荐解释并服务于数百万移动端会员。
This is one of the most useful writeups I have seen on keeping an LLM judge effective in production.
(bookmark it)
Netflix runs judges over hundreds of thousands of show-level recommendation explanations per week, served to millions of members on mobile.
They describe the judge as a lifecycle with four phases rather than an artifact that you only validate once.
> Birth defines multiple evaluation criteria and builds curated benchmarks with human labels and rationales.
> Training refines the judge's rubric through Reasoning-Aligned Rubric Tuning, using a meta-judge over reasoning output as the learning signal.
> Deployment puts one judge in two roles, quality gating and reflective generation.
> Monitoring runs continuous human-in-the-loop alignment that detects drift and triggers re-tuning behind a review gate.
A five-week A/B test over tens of millions of members shifted viewing toward previously unwatched content and increased successful browse-to-play sessions against a no-explanation control, with no quality-related takedowns.
Paper: https://t.co/H12znqteUi
Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX
来源:@omarsar0 · x.com