跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 15 天前AI 评分60
AI 导读

Grok 4.7 在 Harvey 法律基准上得分 19.6%,高于 GPT-5.6 Sol Max 的 2.5% 和 Fable 5.1 Max 的 6.7%;在 Terminal-Bench 4.0 上达到 38.0%,比 Grok 4.6 的 20.3% 近乎翻倍。

正文

Legal work saw one of Grok 4.7’s biggest jumps.

On the Harvey Legal Agent Benchmark, Grok 4.7 scored 19.6%, while GPT-5.6 Sol Max scored 2.5% and Fable 5.1 Max scored 6.7%

Also, Long-running terminal work nearly doubled. Grok 4.7 reached 38.0% on Terminal-Bench 4.0, versus 20.3% for Grok 4.6, suggesting the larger improvement is in agents that must keep working through extended computer tasks.

And SpaceXAI made those gains without raising the headline token price. Grok 4.7 remains at $2/$6 per 1M input/output tokens.

引用@SpaceXAI@SpaceXAI
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed. https://t.co/H3OTBbXyvO
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com