跳到正文
原文
@omarsar0· @omarsar0 · X·· 2026-08-20精选AI 评分66
AI 导读

Elvis Saravia 介绍了 TrueFoundry 发布的开源 agent harness TrueForge,他获得早期访问并在本地运行了数天。TrueForge 负责工具调用循环、上下文管理、子智能体协调和沙箱代码执行,支持 OpenAI、Anthropic、Google 模型以及 Kimi、GLM、DeepSeek 等开源权重模型,模型路由是可配置项。

推荐理由

作者给出同一基准下的成本对比和多模型路由配置,便于判断自托管 agent harness 的实际取舍。

正文

New open-source agent harness just landed!

I got early access to TrueForge by TrueFoundry and have been running it locally for the past few days.

The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models.

TrueForge handles the runtime work that makes an agent reliable.

It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run.

A few things stood out from my testing and their published benchmarks.

Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it.

On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers).

Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12.

Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box.

It's time to own your agent harness.

Thanks to @truefoundry for partnering on this post.

来源:@omarsar0 · x.com