跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 10 天前精选AI 评分65
AI 导读

Synthetic Sciences 发布开源 OpenScience 智能体,在 Terminal-Bench-Science 的 70 个科研工作流中解决 53 项、得分 75.7%,高于 GPT-6 Astra 版 Codex 的 68.1%。

推荐理由

原文给出 OpenScience 与 Codex、Claude Code 在科研基准上的分数对比,可用于衡量科研智能体的当前水平。

正文

Synthetic Sciences just launched its open-source OpenScience agent now outscores Codex and Claude Code on agentic science benchmarks.

A free Apache 2.0 workbench for any provider's model.

On Terminal-Bench-Science, 70 research workflows hosted by Stanford and the Laude Institute, OpenScience solved 53 tasks for 75.7%.

Codex with GPT-6 Astra scored 68.1%, the top entry on a September 23 mirror of the public leaderboard.

The gap widened on Terminal-Bench 4.0's 14 science tasks, where OpenScience scored 71.4% and Claude Code on Claude Fable 5.1 reached 60.0%.

GitHub link in comment

引用@SynScience@SynScience
OpenScience is now the #1 scientific agent. Today it's out of beta and live on Product Hunt, with: • A new IDE: a faster, fully redesigned research workspace • OpenScience Ace: 30+ hand-picked models, including GPT-6 Astra, Claude Opus 5.5, Grok 4.7, Kimi K3, GLM-5.3, and DeepSeek V4.1 Flash, in one pay-as-you-go wallet • Autoresearch: give it a metric and it hill-climbs through experiments on its own • One-click OAuth to bring your ChatGPT or Codex subscription • 300+ research skills, 50+ scientific tools and databases, and NVIDIA BioNeMo built in Already in use at 30+ universities and research labs. Free and fully open source.
在 X 查看被引用的帖子

来源:@rohanpaul_ai · x.com