Arena 在 Agent Arena 上用超 4700 个真实智能体会话评测了 typesafeai 的 Jev Router,结果显示其未推进当前帕累托前沿:达到 DeepSeek V4.1 Flash(Max)相近性能时成本高 38%,中位模型请求延迟高 1.7 倍。
We evaluated Jev Router by @typesafeai on Agent Arena.
Our tests spanning more than 4,700 real-world agentic sessions show:
1. Jev Router does not improve on the current Pareto frontier. For similar performance as DeepSeek V4.1 Flash (Max), the solution costs 38% more and its median model request latency is 1.7x higher.
2. However, it does mostly route to Pareto efficient models. The most LLM it picks is DeepSeek V4.1 Flash, with GPT-6.1 Sol and GPT-6 Luna also being frequent choices.
3. Jev Router’s key strength is steerability. Its score of +10% almost matches Claude Opus 5.5 (High) (+10.48). This highlights the benefit of Jev effectively routing to stronger LLM in response to user feedback.
More insights in the thread below.
来源:Arena.ai · x.com