AI 导读
突发:Grok 4.7 在法律 Agent 基准测试中占据主导地位。 它在真实、长期的法律任务中得分 19.6%,击败了所有其他受测模型。这几乎是 Fable 5.1 的 3 倍。 这是 Grok 4.7 和 SpaceXAI 的又一重大胜利。https://t.co/90ROx64SOI
正文
BREAKING: Grok 4.7 dominates the Legal Agent Benchmark.
It scored 19.6% on realistic, long-term legal tasks, beating every other model tested. That is nearly 3× Fable 5.1.
Another major win for Grok 4.7 and SpaceXAI. https://t.co/90ROx64SOI
来源:@cb_doge · x.com