跳到正文
@cb_doge· @cb_doge · X·· 15 天前精选AI 评分80
AI 导读

xAI 发布 Grok 4.7 模型卡,称其为 SpaceXAI 迄今最强的智能体模型,在长程编码、工程和办公知识任务上更强。模型现已可通过 SpaceXAI API、Grok Build、Cursor 以及 Word、PowerPoint、Excel 插件使用,消费端应用稍后上线。

推荐理由

模型卡列出了与 4.6 的编码基准对比及多处可用入口,读者可据此判断它在长程编码与智能体任务上的位置。

正文

Grok Bot Summary: Grok 4.7 Model Card

What it is
- SpaceXAI / xAI’s latest model
- Stronger at long, hard, agent-style work: coding, engineering, and office / knowledge tasks
- Gets more done with fewer steps and fewer output tokens than many peer frontier models
- Text in/out, plus image input
- Extra training on anonymized Cursor workflow data for coding agents
- Knowledge cutoff: June 2026 (some supplemental data through August 2026)

Where you can use it now
- SpaceXAI API
- Grok Build (default model in their terminal coding agent)
- Cursor (every plan)
- Office add-ins for Word, PowerPoint, Excel
- Model gateways (OpenRouter, Vercel, Cloudflare, Snowflake, Databricks, and others)
- Consumer apps / Grok-in-X: coming later

How it was trained
- Pretrain on public + licensed + internal data
- Longer supplemental training than 4.6
- SFT + RL on reasoning, agent harnesses, STEM, software, knowledge work, kernel work, web, and CAD

Coding (big jump vs 4.6)
- CursorBench 4.0: 46.3% (xhigh) / 43.9% (high) — long real Cursor sessions
- DeepSWE v1.1: 71.0% (high) — near GPT-5.6 Sol’s 72.7%
- Terminal-Bench 4.0: 38.0% (xhigh) vs 4.6’s 20.3%
- FrontierSWE V2: 29.0% (xhigh) vs 4.6’s 25.3%
- SWE-Marathon v1.1: 46.0% (high) vs 4.6’s 31.9%

Knowledge work
- Legal Agent Benchmark: 19.6% (xhigh) — leads the listed peers (4.6 was 15.8%)

Engineering / physical world
- EEBench (circuits / chip design): 66.0% (xhigh) vs 4.6’s 60.0%
- CADGenBench: 44.4% (high) vs 4.6’s 40.9% — tops the listed peers

Medical / biology (helpful, not autonomous doctor)
- HealthBench Professional: 56.7% (xhigh) vs 4.6’s 48.5%
- LatchBio Capabilities: 44.5% (xhigh) — real bio data analysis, not just quiz recall

Cyber (small gains; safer in production)
- Unrestricted capability probes are slightly up vs 4.6
- Best use case they emphasize: finding and fixing vulnerabilities, not end-to-end attacks
- Production safeguards measured separately; harmful/dual-use compliance is low (good)

Bio / chem safety
- Below their FAIF dual-use safety thresholds
- No dual-use bio capability increase vs 4.6; often lower on risky probes
- Dual-use refusal on severity-5 BioUseBench: 91.4%
- Bio refusal recall: 100%; chem: 99.9%

Jailbreaks, safety, behavior
- Standard jailbreak compliance: 0.01% (lower is better)
- Child-safety compliance: 0.0% (no failures in that eval)
- Dishonesty under pressure (MASK-Rectified): 0.00%
- Sycophancy: 0.03%
- Self-harm: refuses help for harm, still points people toward support

Important limits
- Not for autonomous high-stakes decisions in medicine, law, finance, or safety-critical systems without human experts
- Bound by SpaceXAI Acceptable Use Policy and terms
- They say they never silently downgrade the model or swap in a weaker one

One-line takeaway
Grok 4.7 is SpaceXAI’s strongest agent model yet — especially for long coding, terminal work, legal agent tasks, EE/CAD, and clinical communication — with tighter safety on dual-use bio and jailbreaks, available now in API, Cursor, Grok Build, and office add-ins.

Grok 4.7 Model Card:
https://t.co/VsJgjlilTL

来源:@cb_doge · x.com