TypeSafe 在 9 月 15 日发布 Jev 后,新基准 JevBench 随之推出,专门评测输出为有界软件决策而非开放式文本的模型。该基准将智能、校准、速度与成本合并为几何平均分,GPT-5.6 Luna 在困难案例准确率上明显高于 Jev 1.13.0,但 Jev 因延迟、校准和成本更优而综合领先。作者强调这只反映狭窄的类型化决策负载,并不代表 Jev 整体能力更强。
A new benchmark called JevBench just dropped.
for models whose output is a bounded software decision rather than open-ended prose and follows TypeSafe’s 15 Sept release of Jev, which takes application state plus fixed choices and returns a typed answer with probabilities instead of prose.
This benchmark's score deliberately combines Intelligence, Calibration, Speed and Cost because deployment can fail even when raw accuracy is high.
e.g. GPT-5.6 Luna records substantially higher hard-case accuracy than Jev 1.13.0, yet Jev leads the composite because the benchmark also prices latency, calibration and cost.
The geometric mean prevents exceptional performance on 1 axis from fully compensating for a weak one.
The result is evidence about a narrow typed-decision workload, not evidence that Jev is generally more capable than GPT-5.6 Luna.
来源:@rohanpaul_ai · x.com