跳到正文
🚨 AI News | TestingCatalog· @testingcatalog · X·· 2 小时前精选AI 评分66
AI 导读

据 VentureBeat,Mistral Large 4 在 DeepSWE 上得分 62%,超过 GLM-5.3。该模型还在 Finch 金融任务上得分 67%,在 Harvey Legal Agent Benchmark 法律任务上得分 15%,均为开源权重中的 SOTA。作者呼吁 Mistral 尽快发布技术报告。

推荐理由

转述了 Mistral Large 4 在 DeepSWE 等多项基准的具体分数,可与 GLM-5.3 等模型横向对比。

正文

Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.3, according to VentureBeat.

Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).

Le Chonk also scores 15% on Harvey’s Legal Agent Benchmark (Legal tasks, SOTA open-weight).

We need a tech report now 👀
h/t @AiBattle_

引用🚨 AI News | TestingCatalog@testingcatalog
BREAKING 🔥: Mistral announced Mistral Large 4 "Le Chonk", a new 1T-parameter open-weight model! > 49B active parameters, native multimodality. > Rolling out via APIs today; open-weight release is planned for the end of October. > SOTA on "critical workloads", including cyber defense. Le Chaton Fat "Le Chonk" is here 👀 https://x.com/MistralAI/status/2107457414387622310/video/1
在 X 查看被引用的帖子

来源:🚨 AI News | TestingCatalog · x.com