x.com/i/article/205852961337…
X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(535)
@berryxia@berryxiaAI 评分5858 引用Lisan al Gaib (@scaling01)@scaling01@berryxia@berryxiaAI 评分6060 腾讯HY实验室联合四家机构发布Chronicles-OCR基准,用2800张专家标注图像测试AI对3000年中国古文字的识别能力,28个前沿多模态模型几乎全线失败。
引用ModelScope (@ModelScope2022)@ModelScope2022The best VLLM scores only 14% on oracle bone script recognition. Chronicles-OCR, a new ancient Chinese character benchmark from Tencent HY and 4 institutions, just put 28 frontier models to the test across 3,000 years of Chinese writing. 🏺🗂️ modelscope.cn/datasets/Virtu… 2,800 expert-annotated images across 7 scripts: oracle bone, bronze, seal, clerical, regular, running, and cursive. Three findings worth knowing: 🔍 End-to-end detection collapses on ancient scripts. Best H-mean: 16.5. GPT-5 and Gemini 2.5 Pro near 0. 🧠 Reasoning makes it worse. Thinking mode hurts performance across almost all models — chain-of-thought amplifies hallucination when perception fails. 👁️ Models recognize the carrier, not the script. Ancient script classification hits 96.7% by spotting turtle shells and bronze vessels, not the characters themselves.
@kimmonismus@kimmonismusAI 评分99 来源 Phoronix:phoronix.com/review/nvidia-v…
@kimmonismus@kimmonismusAI 评分1717 @kimmonismus@kimmonismusAI 评分6161 Phoronix 公布 NVIDIA 新款 ARM 架构数据中心 CPU Vera 的公开基准测试,作者通读 11 页评测后总结了其性能表现。


@berryxia@berryxiaAI 评分2525 开源地址:github.com/freestylefly/Code…
@berryxia@berryxiaAI 评分5353
引用苍何 (@canghe)@canghex.com/i/article/205957789644…
@berryxia@berryxia精选AI 评分6767
引用OpenRouter (@OpenRouter)@OpenRouterToday we’re announcing our $113M Series B led by @CapitalGVC. Over the last 6 months, weekly volume on OpenRouter grew from 5T to 25T tokens as AI rapidly shifts from experimentation into production. We’re excited for what comes next.
推荐理由:OpenRouter 的 B 轮融资与 token 增长数据,为观察多模型基础设施在生产环境中的需求变化提供参照。
@berryxia@berryxiaAI 评分1111 @berryxia@berryxiaAI 评分4949
引用RyanLee (@RyanLeeMiniMax)@RyanLeeMiniMaxRecently, we took time to consolidate all of the work behind M2 and published it here: our M2 paper on arXiv It’s been just over six months since we first open-sourced M2 on December 23 last year. During that time, a number of our ideas and systems have been broadly adopted by the open-source community — including CISPO, Forge RL System, Self-Evolution. Over the past six months, we’ve felt incredible enthusiasm from the open-source community. Nearly every model release reached the #1 spot on the Hugging Face leaderboard. Now it’s time for a new chapter. We’re getting ready for M3. MSA paper is on the road. arxiv.org/abs/2605.26494
@AYi_AInotes@ayi_ainotesAI 评分3030 我补充一句:最让我震惊的不是 GPT-5.5 有多强,而是一个随便搭的 mini-agent 居然能和官方调了半天的工具打平🤯 兄弟们这说明啥?说明很多公开榜上的成绩,水分可能比我们想象的大的多
@AYi_AInotes@ayi_ainotesAI 评分4141
引用Theo - t3.gg (@theo)@theoThis is the first code bench that actually aligns with how it feels to use these models coding.
@op7418@op7418AI 评分00 
@op7418@op7418AI 评分55 @op7418@op7418AI 评分4949 
引用歸藏(guizang.ai) (@op7418)@op7418藏师傅的小红书图文排版 Skill 预览 完全靠 HTML 和实拍图片,不会被标注 AI AI 会去高质量图片网站帮你寻找对应的主题图片,让你的图文告别只有生硬文字的尴尬情况
@alibaba_cloud@alibaba_cloudAI 评分3434 用 Qwen3.7-Max 驱动 Hermes Agent。去看看 @NousResearch 🚀 (引用推文:Hermes Agent 现已支持 Qwen 3.7 Max)
引用Nous Research (@NousResearch)@NousResearchQwen 3.7 Max is now supported in Hermes Agent
@kimmonismus@kimmonismus精选AI 评分7777 DeepSeek 将 V4-Pro 降价 75% 的举措永久化,小米 MiMo 则把 V2.5 价格最高下调 99%,即日生效。



推荐理由:文章把两家公司的降价归因到注意力架构的改动,读者可据此理解长上下文推理成本为何能结构性下降。
@kimmonismus@kimmonismusAI 评分77 来源 Axios:axios.com/2026/05/26/deepmin…
@kimmonismus@kimmonismusAI 评分1212 Source AGI "Einstein" 测试:teddit.net/r/singularity/com…
@kimmonismus@kimmonismusAI 评分4040 
@vista8@vista8AI 评分1414 让GPT 5.5 Pro调研短剧讨论,写了个短剧剧本生成Skill。 等我测试下效果,再生成几个短片,看看效果。

@vista8@vista8AI 评分2323 已经很少用 Terminal 了,基本都用 Codex App 开发。 连朋友送的 API 都用的少了,不然还要折腾装插件,开启 OpenAI 订阅账号才能有的功能。
@berryxia@berryxiaAI 评分1616 
@kimmonismus@kimmonismusAI 评分2323 主动式 AI 智能体似乎在 ChatGPT 中越来越多了! 我刚在德国查了一下。这里似乎还用不了。 那绝对是重大进步,会非常实用。
引用Max Weinbach (@mweinbach)@mweinbachSo this seems to work and not give me a once an hour hadn’t shipped alert This seems like a big feature in ChatGPT?
@vista8@vista8AI 评分4242 
@vista8@vista8AI 评分3838 @vista8@vista8AI 评分1313 Suno生成了一首很痞的歌曲,很像gala 哈哈哈 Video

@xicilion@xicilionAI 评分11 

@alibaba_cloud@alibaba_cloudAI 评分2222 阿里云被 Omdia 的 Agentic AI Market Radar 评为领导者。Omdia 重点肯定了阿里云在每一层的全栈能力,认可其为首个围绕 Agent 范式来构建整个平台的云服务商。

@berryxia@berryxiaAI 评分4343 引用向阳乔木 (@vista8)@vista8说好不熬夜的,但 AI Coding 太上瘾! 昨晚开发了个 Chrome 新窗口插件,超方便。 1. 番茄钟、音乐播放、Todo、便签、天气、换背景等,独立开发者多件套整合到了一起 😂 2. 支持谷歌搜索,ChatGPT跳转官网带提示词发送。 3. 支持Command + K唤起,快速设置、搜索一切。 已开源,见评论区。 Video
@alibaba_cloud@alibaba_cloudAI 评分3838 1M 上下文。更智能的推理。更多可能性。很高兴看到 Qwen3.7 Max 现已在 Go 中上线,支持 @opencode 🚀
引用OpenCode (@opencode)@opencodeQwen3.7 Max now available in Go - text only - 1M context - smartest model in the Qwen family to date
@vista8@vista8AI 评分2424 乔木Tab Chrome插件开源地址: github.com/joeseesun/qiaomu-… 等后续有空上架。
@vista8@vista8AI 评分3737 
@berryxia@berryxiaAI 评分2929 Typeless感觉每天都在更新啊,有时候中英文识别会瞎识别,目前遇到的就是我说的是seedance,给我直接变成了keling 😁
引用Typeless (@typelessdotcom)@typelessdotcomTypeless 1.5.0 is live for macOS & Windows! ✨ Bringing custom shortcuts to external keyboards. ⌨️ Your favorite setup, working your way. 💙 Go seamless: ⚡️ typeless.com/ #Typeless
@berryxia@berryxiaAI 评分44 
@alibaba_cloud@alibaba_cloudAI 评分3838 
@swyx@swyxAI 评分77 @SenseTime_AI@sensetime_aiAI 评分3636 我们的 SenseSmart Go AI 店员正在重新定义上海的便利店服务。来体验一下吧,祝你有美好的一天!🥰
引用China Xinhua News (@XHNews)@XHNewsMeet Xiaomai - the only staff member running a convenience store in Shanghai. From greeting customers and recommending products to processing payments and handing over purchases, this robot does it all on its own. #Shanghai #Robot #AI #Tech #Future #Innovation #SmartRetail #ChinaTech #Automation #Robotics Video
@vista8@vista8AI 评分4343 这样做完,会生成一个复盘经验文档,非常实用,贴合自己的开发设计审美偏好。
引用向阳乔木 (@vista8)@vista8如何让你的Codex变的越来越聪明,越来越懂你? 上周跟 @HiTw93 直播时,很多人可能没注意他的一段话,他说他的开发Skill waza,每周都能无痛更新。 因为他会让Codex扫描本周对话记录,让AI提炼他的开发经验、审美偏好并写入Skill,从而让它越来越强。 建议人人都试试,做法和提示词见评论第一条。
@vista8@vista8AI 评分5151