Anthropic 表示期待通过正式发布提供 Mythos 级模型,这与此前称其"过于强大不宜发布"的表态相反。推文质疑这一转向:Mythos 据称已能发现其他模型从未发现的漏洞与利用方式,而 Anthropic 主要服务企业客户、算力紧张,未必需要此类 PR 操作。作者认为护栏到位并正式开放后,SWE 将获得显著提升,且从 benchmark 看目前没有模型能接近 Mythos。
"We look forward to making Mythos-class models available through general release"
I don't understand Anthropic's strategy regarding Mythos.
On the one hand, everyone is saying that Mythos has achieved the expected quality and is finding bugs and exploits that no other model has ever found.
On the other hand, precisely for this reason, Anthropic has repeatedly stated that it's "too powerful for release."
Why the sudden about-face? One explanation: PR. The preview, including a benchmark, combined with the statement that the model wouldn't be released due to its power, generated a lot of attention. But does Anthropic really need that?
Anthropic is so significant because they primarily serve enterprises. Their biggest problem: compute. Too many want Claude, too little compute to support it adequately. Therefore, this PR move wasn't necessary, and the IPO is still in the near future.
In short: it seems downright erratic to now do the exact opposite of what was stated.
Be that as it may, once the guardrails are in place and there is general availability, SWEs will receive a significant boost. Judging by the benchmarks, nothing even comes close to the myth so far.
Looks like they meant it.在 X 查看被引用的帖子
来源:@kimmonismus · x.com