跳到正文
@berryxia· @berryxia · X·· 2026-05-25AI 评分58
AI 导读

开源深度研究系统 Onyx 在 DeepResearch Bench 上登顶,超越 Claude 和 ChatGPT,做法是把调度器的搜索权限彻底砍掉,只保留任务分解与评估。

正文

这个团队的研究也是有点反常识,对于LLM的研究调度也是有点不一样的研究。

一个开源团队发现把深度研究系统中最聪明的总指挥调节器直接剥夺搜索权限。

反而让整个系统在DeepResearch Bench上直接登顶吊打Claude和ChatGPT。

这个反直觉的设计让Onyx成为目前公开可用的最强深度研究员

它叫Onyx GitHub上已经完全开源你现在就能跑

故事听起来简单却戳穿了几乎所有大厂AI Agent的共同毛病。

传统深度研究系统包括OpenAI o1系列Anthropic和Google的方案都给调节器塞满了工具它能搜索网页打开链接读文档写报告一条龙到底。

结果呢?

调度器一拿到搜索权就忍不住自己动手它开始疯狂拉结果浅尝辄止根本不做高质量的任务分解最后产出的报告永远是表面级。

Onyx的团队观察到这个致命bug后做了件谁都没敢做的事把调度器的搜索工具彻底砍掉

它只能写任务brief只能分解query只能评估下级agent交回来的中间报告但它自己绝不能上网不能检索不能提前下结论

这一刀直接逼着调节器做真正的“高阶战略思考”

整个架构只保持两层上面一个纯策略的调节器下面最多6个独立的research agent。

三阶段流水线超级清晰

Phase 1 调节器无工具权限把一个复杂问题最多拆成6个聚焦的研究方向写出极度自洽的任务brief

Phase 2 把任务分发给3个隔离的研究agent每个agent最多跑8轮“搜索-阅读-思考”循环产出带引用来源的中间报告它们还能接入企业内部Confluence Slack等100+数据源并且严格做文档级权限控制

Phase 3 一个确定性步骤把所有报告去重重新编号生成统一引用地图输出最终高质量报告

因为调节器全程不碰原始数据它就不会被“看到第一个结果就想收工”的诱惑污染因为只两层传递信息不会在多层摘要里被扭曲

结果Onyx在DeepResearch Bench上拿下No.1全面超越闭源的Claude和ChatGPT

更狠的是它还能无缝接入企业内部知识库这点连很多付费方案都做不到。

你今晚就可以试

直接去Onyx GitHub仓库链接在下面star一下然后按照readme把整个系统跑起来用CrewAI做整体编排 Mistral的Voxtral做语音输入输出就能复刻一个完全开源的顶级深度研究员。

整个框架100%开源架构细节pipeline代码实验数据全在仓库里

Big Tech还在卷“给模型塞更多工具更多上下文”Onyx却用一个“故意阉割”的调节器告诉所有人最聪明的约束往往才是最强的能力。

nitter.net/i/status/2058837753954…

引用Avi Chawla (@_avichawla)@_avichawla
The No. 1 deep researcher beats Claude and ChatGPT with a trick neither uses. I studied the open-source architecture behind it. A counterintuitive thing I found is that the orchestrator agent that runs the entire research strategy has no search access. It can't query the web or open URLs. This looks wrong at first glance. Every other deep research system gives its coordinator far more capability. For instance: - OpenAI's approach trains a single model for many consecutive tool calls. It searches, reads, reasons, and writes the report in a long sequential chain. ↳ The researchers behind the No. 1 system (Onyx) observed that this causes the model to spend cycles on low-value searches instead of maintaining a high-level research strategy. - Anthropic and Google use an orchestrator-researcher pattern similar to Onyx's system. The key difference is how aggressively Onyx constrains the orchestrator. Most orchestrators have access to search and retrieval tools alongside dispatch capabilities. And the moment an orchestrator can search, it will. So instead of decomposing a query into focused research threads, it starts answering the question itself. It pulls a few results, skips proper task decomposition, and produces a surface-level report from whatever it found first. Stripping search from the orchestrator forces it to write self-contained and coherent task briefs for each research agent. The researchers also kept the architecture only two levels deep. When info passes through multiple agents, each one subtly distorts it through summarization/reinterpretation. Keeping it to two levels prevents this. These two constraints sit inside a larger three-phase pipeline (the visual below maps this): → Phase 1 decomposes the query into up to 6 research directions. No tool access prevents the model from prematurely answering. → Phase 2 dispatches 3 isolated research agents. Each runs up to 8 sub-cycles of search, read, and think, to produce an intermediate report with citations. The agents can also search internal enterprise docs (Confluence, Slack, 100+ connectors) with document-level permissions enforced, unlike proprietary solutions. → Phase 3 runs a deterministic step that renumbers and deduplicates to produce a report with a unified citation map. This pattern has been ranked No. 1 on DeepResearch Bench. The whole implementation is available on GitHub and you can try it yourself. Here's the Onyx Repo: github.com/onyx-dot-app/onyx (don't forget to star it ⭐ ) My co-founder wrote a detailed article on building a fully open-source deep researcher using Onyx as the deep research layer, CrewAI for orchestration, and Voxtral by Mistral for voice input/output. Read it below.
在 X 查看被引用的帖子

来源:@berryxia · x.com