AI 导读
Grok 4.7 在 AA-Briefcase 上仅次于 Anthropic 模型,排名紧追 Opus 5,每任务成本约为其 50%。在 AA-Briefcase-Lite 尽调场景中,Grok 4.7 分析质量 Elo 从 1698 升至 1994,但展示 Elo 从 1531 微降至 1499。
正文
Grok 4.7 is behind only Anthropic models on AA-Briefcase, ranking just behind Opus 5 at ~50% of its Cost per Task
Grok 4.7’s improvements over Grok 4.6 are clear in AA-Briefcase-Lite, our public due diligence scenario where models are tasked with building market models and target assessment decks. Grok 4.7 gains significantly in Analytical Quality Elo (1698 → 1994) with a slight regression in Presentation Elo (1531 → 1499).
API cost to produce example decks: Grok 4.7 (xhigh) ~$8 vs. Grok 4.6 (xhigh) ~$4.40
来源:@ArtificialAnlys · x.com