Synthetic Sciences 发布开源 OpenScience 智能体,在 Terminal-Bench-Science 的 70 个科研工作流中解决 53 项、得分 75.7%,高于 GPT-6 Astra 版 Codex 的 68.1%。
原文给出 OpenScience 与 Codex、Claude Code 在科研基准上的分数对比,可用于衡量科研智能体的当前水平。
Synthetic Sciences just launched its open-source OpenScience agent now outscores Codex and Claude Code on agentic science benchmarks.
A free Apache 2.0 workbench for any provider's model.
On Terminal-Bench-Science, 70 research workflows hosted by Stanford and the Laude Institute, OpenScience solved 53 tasks for 75.7%.
Codex with GPT-6 Astra scored 68.1%, the top entry on a September 23 mirror of the public leaderboard.
The gap widened on Terminal-Bench 4.0's 14 science tasks, where OpenScience scored 71.4% and Claude Code on Claude Fable 5.1 reached 60.0%.
GitHub link in comment
OpenScience is now the #1 scientific agent. Today it's out of beta and live on Product Hunt, with: • A new IDE: a faster, fully redesigned research workspace • OpenScience Ace: 30+ hand-picked models, including GPT-6 Astra, Claude Opus 5.5, Grok 4.7, Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash, in one pay-as-you-go wallet • Autoresearch: give it a metric and it hill-climbs through experiments on its own • One-click OAuth to bring your ChatGPT or Codex subscription • 300+ research skills, 50+ scientific tools and databases, and NVIDIA BioNeMo built in Already in use at 30+ universities and research labs. Free and fully open source.在 X 查看被引用的帖子
来源:@rohanpaul_ai · x.com