SpaceXAI 发布 Grok 4.7,称其为面向编码与知识工作的最强模型,采用更大更新的基座模型,并在多小时复杂任务、长上下文和自我核查上做了强化。
原文列出 Grok 4.7 在编码与知识任务上的多项基准变化和 API 定价,读者可据此与同价位模型作横向比较。
BREAKING: SpaceXAI just released Grok 4.7!
Grok 4.7 is SpaceXAI’s most powerful model for coding and knowledge work. It is designed to work longer on difficult tasks, handle more context and carefully check its own work.
Major improvements:
• Uses a new and larger base model
• Trained longer on harder, multi-hour tasks
• Better at verifying its own answers
• Improved long-context performance
• Better at documents and presentations
• Natively understands the Grok Bot system
• Stronger safety and jailbreak protection
Benchmark results:
• CursorBench 4.0: 46.3%, up from 40.4%
• DeepSWE v1.1: 71.0%, up from 65.2%
• AA Briefcase v1.1: 1,657, up from 1,546
• Terminal-Bench 4.0: 38.0%, up from 20.3%
• Harvey Legal Benchmark: 19.6%, up from 15.8%
• HealthBench Professional: 56.7%, up from 48.5%
• EEBench: 64.0%, up from 53.0%
Grok 4.7 also scored higher than GPT-5.6 Sol and Fable 5.1 on the listed legal and electrical engineering benchmarks.
Safety improvements:
• 62.4% on LatchBio’s biosafety benchmark
• Only 3.3% of risky cyber prompts passed through on HackerBench
• SpaceXAI’s strongest model yet for refusing dangerous requests
• Rarely blocks legitimate cybersecurity work
• Select security partners are receiving invite-only red-team access
Pricing and availability:
• $2 per million input tokens
• $6 per million output tokens
• Same base price as Grok 4.6
• Fast version offers twice the output speed at twice the price
• Available now in Cursor, Grok Build and the Grok API
• Also available through coding tools, model routers and cloud platforms
SpaceXAI says Grok 4.7 is twice as fast and half the price of comparable models.
It beats Grok 4.6 across every listed benchmark while keeping the same base API price. Smarter, safer, stronger and still incredibly affordable.
SpaceXAI is moving at an unbelievable speed.
来源:@cb_doge · x.com