Elvis Saravia 称用 Jev-as-a-Judge 做智能体评估是其目前见过最惊艳的 Jev 用例之一,但并非所有场景都该用 Jev,也不该在所有评估中都上前沿模型。他正在大量测试,早期结果显示兼顾准确率与成本的最优流程是:高置信度场景用 Jev,低置信度判定升级到前沿模型(GPT-6 或 Opus 5.5)。完整指南即将发布。
Don't sleep on using Jev-as-a-Judge for agent evaluation.
This is one of the most impressive Jev use cases I have found so far.
Jev is a natural fit as a Judge, but it doesn't mean you use it everywhere.
Similarly, you shouldn't use frontier models for evals everywhere.
I'm running lots of tests on this atm, but early results point to an optimized flow (balancing accuracy and cost) that combines Jev and frontier models.
Concretely, use Jev in high-confidence situations, and escalate to a frontier model (GPT-6 or Opus 5.5) in low-confidence verdicts.
Entire write-up coming soon. Let me know if you have questions as I build the full guide.
来源:@omarsar0 · x.com