跳到正文
@ArtificialAnlys· @ArtificialAnlys · X·· 15 天前AI 评分44
AI 导读

Grok Build 搭配 Grok 4.7(xhigh)在 Artificial Analysis 编码智能体指数上得分 56,高于 Grok 4.6(xhigh)的 47。 三个组成部分均有提升:DeepSWE v1.1 从 65% 升至 73%,Terminal-Bench 4.0 从 18% 升至 33%,SWE-Atlas-QnA 从 58% 升至 63%。 这些结果评估的是 Grok 搭配 Grok Build——其第一方编码智能体。它们与 Intelligence Index 的结果相互独立,后者统一了跨模型使用的评估框架。

正文

Grok Build with Grok 4.7 (xhigh) scores 56 on the Artificial Analysis Coding Agent Index, up from 47 with Grok 4.6 (xhigh).

It improves across all three components: DeepSWE v1.1 rises from 65% to 73%, Terminal-Bench 4.0 from 18% to 33%, and SWE-Atlas-QnA from 58% to 63%.

These results evaluate Grok with Grok Build, its first party coding agent. They are separate from the Intelligence Index results, which standardize the evaluation harness used across models.

来源:@ArtificialAnlys · x.com