跳到正文
@cb_doge· @cb_doge · X·· 15 天前AI 评分30
AI 导读

突发:Grok 4.7 在法律 Agent 基准测试中占据主导地位。 它在真实、长期的法律任务中得分 19.6%,击败了所有其他受测模型。这几乎是 Fable 5.1 的 3 倍。 这是 Grok 4.7 和 SpaceXAI 的又一重大胜利。https://t.co/90ROx64SOI

正文

BREAKING: Grok 4.7 dominates the Legal Agent Benchmark.

It scored 19.6% on realistic, long-term legal tasks, beating every other model tested. That is nearly 3× Fable 5.1.

Another major win for Grok 4.7 and SpaceXAI. https://t.co/90ROx64SOI

来源:@cb_doge · x.com