跳到正文
@omarsar0· @omarsar0 · X·· 2026-09-03AI 评分63
AI 导读

Meta 的论文提出 CORAL,用 LLM 智能体框架在生产推荐系统上做持续优化,并通过 A/B 实验验证。智能体每轮观察运行信号,基于历史决策及其测量结果的记忆进行推理,并调用包括数值优化器在内的工具,把每次改动限制在固定预算内,策略只靠上下文从先前动作中改进而不更新参数。在两个大型社交平台上,同一框架在一个平台不增加服务成本提升互动,在另一个平台降低服务成本且未降低互动。

正文

Massive paper from Meta.

I like this one because it shows the use of agent harnesses for production-grade recommender systems.

Details below:

This is one of the more convincing agent deployments I've seen.

It runs against a live production recommender serving billions of people and reports A/B results.

Sustaining a recommender is continual optimization work. Content shifts, user behavior shifts, upstream models shift, and the choices governing retrieval, ranking and serving have to be revisited.

Human engineers test those changes through online experiments, which is slow enough that parts of the system go unrevised.

In CORAL, each cycle the agent observes operating signals, reasons over a memory of past decisions and their measured outcomes, and invokes tools including a numerical optimizer that keeps every change inside a fixed operating budget.

The policy improves in context from its own prior actions, with no parameter updates.

Across two large social platforms, the same harness improves engagement at no additional serving cost on one and reduces serving cost without degrading engagement on the other.

Performance improves as the loop iterates.

The guardrail design carries as much weight as the agent. A bounded change budget makes this safe to run against production.

Paper: https://t.co/G46EgVuPMR

Chat with Paper: https://t.co/KlYFT8dAFD

来源:@omarsar0 · x.com