跳到正文
@omarsar0· @omarsar0 · X·· 2026-08-24AI 评分54
AI 导读

Netflix 团队发表论文,把大规模推荐解释中的 LLM-as-a-Judge 描述为包含 Birth、Training、Deployment、Monitoring 四个阶段的完整生命周期,每周处理数十万条剧集级推荐解释并服务于数百万移动端会员。

正文

This is one of the most useful writeups I have seen on keeping an LLM judge effective in production.

(bookmark it)

Netflix runs judges over hundreds of thousands of show-level recommendation explanations per week, served to millions of members on mobile.

They describe the judge as a lifecycle with four phases rather than an artifact that you only validate once.

> Birth defines multiple evaluation criteria and builds curated benchmarks with human labels and rationales.

> Training refines the judge's rubric through Reasoning-Aligned Rubric Tuning, using a meta-judge over reasoning output as the learning signal.

> Deployment puts one judge in two roles, quality gating and reflective generation.

> Monitoring runs continuous human-in-the-loop alignment that detects drift and triggers re-tuning behind a review gate.

A five-week A/B test over tens of millions of members shifted viewing toward previously unwatched content and increased successful browse-to-play sessions against a no-explanation control, with no quality-related takedowns.

Paper: https://t.co/H12znqteUi

Track more trending AI papers in our academy: https://t.co/1e8RZKs4uX

来源:@omarsar0 · x.com