AI 导读
Ornith-1.5-397B 公布评测结果,Terminal-Bench 2.1 为 85.1,SWE-bench Verified 为 86,DeepSWE 五次运行平均 56。35B 版本在这两项上分别为 68.5 和 79,每个 token 激活 3B 参数。此外还有量化后的 9B-Mobile 版本,面向 iPhone 和 Android,把编码智能体放到设备端。
正文
Ornith-1.5-397B reports 👀
> 85.1 on Terminal-Bench 2.1
> 86 on SWE-bench Verified.
> 56 on DeepSWE, averaged over five runs.
The 35B lands at 68.5 and 79 while activating 3B parameters per token.
There is also the quantized 9B-Mobile build that targets iPhone and Android, putting a coding agent on device.
Check out the weights and full tables 👀
https://t.co/Ijc9M6NXYO
来源:@testingcatalog · x.com