AI 这下真利好宠物经济了哈哈哈哈哈 等待 AI 运行的过程中撸猫实在太爽了 大家 AI 工作的时候快去猫咖或者干脆养一只宠物吧哈哈哈哈哈 https://t.co/M3QXDmiWNT
X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(541)
@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分1313 
@OpenBMB@OpenBMBAI 评分6262 引用@ArtificialAnlys@ArtificialAnlysOpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open weights model under 4B total parameters OpenBMB (@OpenBMB) is the open-source AI group behind the MiniCPM series of efficient small models. MiniCPM5-2B is a 2.6B parameter dense reasoning model with text input and output, released under Apache 2.0. Scoring 15 on the Intelligence Index, MiniCPM5-2B sits one point behind Ling 3.0 Tiny (16), which has ~3x the total parameters. Among open weights models under 4B total parameters, the next best score is Granite 4.2 3B (11). Key results: ➤ The highest Intelligence Index of any open weights model under 4B total parameters, setting a new Pareto-optimal point on Intelligence vs. Total Parameters: Its score of 15 is 4 points clear of Granite 4.2 3B (11). With 2.6B total parameters, it is 1 point ahead of Qwen3.5 4B (Reasoning, 14, estimated) with 44% fewer parameters, and level with Qwen3.5 9B (Reasoning, 15, estimated) at roughly 4x its size. As a dense model, its size advantage is in memory footprint rather than active-parameter compute. ➤ Strong agentic performance at this size: Its GDPval-AA v2 Elo of 831 leads <4B models, and on τ³-Banking it is joint-first with Ling 3.0 Tiny at 21%, compared to 8% for the next best model, Granite 4.2 8B. On AA-Briefcase, it placed second among the measured models in the comparison set with an Elo of 438, above Granite 4.2 8B (324) and just below Ling 3.0 Tiny (485). ➤ Knowledge, coding and long context are where it gives ground: MiniCPM5-2B places 7th in the set on Humanity's Last Exam (9%, behind Gemma 4 12B (Reasoning) at 16%), 8th on Terminal-Bench v2.1 (9%, behind Qwen3.5 9B (Reasoning) at 29%) and scores 0% on CritPt. On SciCode it is second of the five measured models at 26%, behind Granite 4.2 8B (31%). On AA-LCR v1.1 it scores 59%, 5th in the set, one point behind Ling 3.0 Tiny (60%). On GDP.pdf, our new professional document reasoning evaluation, it passes 1% of tasks outright, behind gpt-oss-20b (high) at 2%. ➤ Its AA-Omniscience score of -12 is earned by abstaining from answering rather than accuracy: MiniCPM5-2B attempts only 29% of AA-Omniscience questions, giving it a Non-Hallucination Rate of 78%. Its accuracy of 8% is a point below Ling 3.0 Tiny (9%) and half that of Qwen3.5 9B (Reasoning, 16%). Peers that attempt far more questions are penalized heavily, with Qwen3.5 9B (Reasoning) at -53 and gpt-oss-20b (high) at -63. ➤ It is token-efficient for a reasoning model: MiniCPM5-2B used 19k output tokens per Intelligence Index task, joint-lowest in the comparison model set with Granite 4.2 3B (19k). Ling 3.0 Tiny spends 56k, roughly 3x as many, for 1 more index point. Additional model details: ➤ Parameters: 2.6B (dense) ➤ Context window: 131k tokens ➤ Input modalities: Text only ➤ License: Apache 2.0
@thexpin@thexpinAI 评分11 @thexpin@thexpinAI 评分55 抱歉,您提供的主推文内容仅为一个链接(https://t.co/4SJ8O9m7mV),没有可翻译的正文文本。请提供推文的实际文字内容,我将为您翻译。
@omarsar0@omarsar0AI 评分6060 Google DeepMind、MIT 等机构的论文提出 SMART,一个主分支几乎不含代码的符号化机器学习性能建模库,其实现由编码子智能体在版本更新时从自然语言设计文档重新生成。

@ClementDelangue@ClementDelangueAI 评分2525 有时我会想,如果我们当初没有公开披露那起针对智能体的网络攻击,会发生什么。 这更加坚定了我的信念:AI 领域需要 100 倍的透明度。否则我们将陷入大麻烦!
@tianyi@tianyiAI 评分2121 @omarsar0@omarsar0AI 评分2121 这里的要点是,你不能害怕给 Astra 访问工具的权限。这个模型感觉没有极限。你之前做过但失败的一切,都应该用它再试一次。
@rohanpaul_ai@rohanpaul_aiAI 评分2222 – https://t.co/5Z8zAEiVsQ 标题:“为什么 CLAUDE.md 不断膨胀?智能体编程中的灾难性记忆”
@rohanpaul_ai@rohanpaul_aiAI 评分4949 
@tianyi@tianyiAI 评分4747 引用@tianyi@tianyiDeepSeek 社招资深后端/服务端工程师,新开放约 150 个 HC,好多好多新方向新系统新需求需要搞。简历建议直接扫码或邮箱投递,私信我简历的话会看但不一定能及时处理。
@PixVerse_@PixVerse_AI 评分1919 
@tianyi@tianyiAI 评分3535 DeepSeek 社招资深后端/服务端工程师,新开放约 150 个 HC,好多好多新方向新系统新需求需要搞。简历建议直接扫码或邮箱投递,私信我简历的话会看但不一定能及时处理。

@omarsar0@omarsar0AI 评分2222 
@omarsar0@omarsar0AI 评分2323 
@WorkBuddy_AI@WorkBuddy_AIAI 评分1010 @WorkBuddy_AI@WorkBuddy_AIAI 评分55 @OpenBMB@OpenBMBAI 评分5353 @OpenBMB@OpenBMB精选AI 评分6868 
推荐理由:面壁智能随 MiniCPM5-2B 一并公开数据、训练配方和 RL 框架,读者可了解小模型训练栈的完整构成。
@OpenBMB@OpenBMBAI 评分4545 

@OpenBMB@OpenBMBAI 评分2424 
@OpenBMB@OpenBMBAI 评分5353 面壁智能开源 2B 参数语言模型 MiniCPM5-2B,在 Artificial Analysis 智能指数上以 23 分位列 4B 以下开源模型第一,Agentic Index 得 20 分。




@kimmonismus@kimmonismusAI 评分3232 
@jxnlco@jxnlcoAI 评分2222 太疯狂了,因为我们已经到了这样一个地步:我看着 vibe coding 做出来的东西,心想“我完全不知道这是怎么做出来的”
@WorkBuddy_AI@WorkBuddy_AIAI 评分2626 
@jxnlco@jxnlcoAI 评分1515 @lifesinger@lifesingerAI 评分1212 现有的各种 AI 应用产品 一定程序上只有两种: 一种是百度,比如豆包 另一种是 Sora2,比如 L... Codex 都不是 AI 应用 而是做 AI 应用的模型工具
@lifesinger@lifesingerAI 评分66 @lifesinger@lifesingerAI 评分1414 @alexandr_wang@alexandr_wangAI 评分1111 @Baidu_Inc@Baidu_IncAI 评分2222 
@kimmonismus@kimmonismusAI 评分4545 
@jxnlco@jxnlcoAI 评分55 @ArtificialAnlys@ArtificialAnlysAI 评分5858 Artificial Analysis 公布了 MiniCPM5-2B 的完整评测结果,并给出其权重下载入口。该模型权重以 Apache 2.0 许可开放。
@ArtificialAnlys@ArtificialAnlysAI 评分2929 
@ArtificialAnlys@ArtificialAnlysAI 评分2525 
@ArtificialAnlys@ArtificialAnlysAI 评分3737 
@ArtificialAnlys@ArtificialAnlysAI 评分3737 
@ArtificialAnlys@ArtificialAnlysAI 评分3737 
@ArtificialAnlys@ArtificialAnlysAI 评分5959 OpenBMB 的 MiniCPM5-2B 在 Artificial Analysis Intelligence Index v4.2 上得分 15,是 4B 以下开源权重模型中的最高分。
