跳到正文
@testingcatalog· @testingcatalog · X·· 2026-08-19AI 评分61
AI 导读

Ornith-1.5-397B 公布评测结果,Terminal-Bench 2.1 为 85.1,SWE-bench Verified 为 86,DeepSWE 五次运行平均 56。35B 版本在这两项上分别为 68.5 和 79,每个 token 激活 3B 参数。此外还有量化后的 9B-Mobile 版本,面向 iPhone 和 Android,把编码智能体放到设备端。

正文

Ornith-1.5-397B reports 👀

> 85.1 on Terminal-Bench 2.1
> 86 on SWE-bench Verified.
> 56 on DeepSWE, averaged over five runs.

The 35B lands at 68.5 and 79 while activating 3B parameters per token.

There is also the quantized 9B-Mobile build that targets iPhone and Android, putting a coding agent on device.

Check out the weights and full tables 👀
https://t.co/Ijc9M6NXYO

来源:@testingcatalog · x.com