关注持久化智能体和记忆。 @coworkerapp 刚刚为你的公司推出了一款出色的持久记忆解决方案。 极其高效、便宜且快速。 https://t.co/cl8yr3QFIr
X:Elvis Saravia
@omarsar0 · X
切换来源
@omarsar0@omarsar0AI 评分2828 @omarsar0@omarsar0AI 评分1515 @omarsar0@omarsar0AI 评分2323 @omarsar0@omarsar0AI 评分5959 @omarsar0@omarsar0AI 评分6363 
@omarsar0@omarsar0精选AI 评分9292 引用@JensenHuang@JensenHuangExciting day for NVIDIA and @huggingface. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. They allow every developer, startup, university, industry and country to build with, customize and benefit from AI. Thank you @ClementDelangue for coming to me. NVIDIA is going to be a great home for Hugging Face, its community and the future of open models. 🤗 https://t.co/q8Om2Xc5ye
推荐理由:转述黄仁勋关于 NVIDIA 将接手 Hugging Face 的表态,并补充对该公司开源模型投入的观察。
@omarsar0@omarsar0AI 评分4444 一篇论文提出将智能体拆分为持久本体与可替换基础设施两半:身份、私有记忆和带版本历史的代码属于智能体本身,推理模型、运行框架、托管服务器和交互入口(chat、API、UI)则可替换。

@omarsar0@omarsar0AI 评分2424 @omarsar0@omarsar0AI 评分2020 只有我这么觉得吗,还是说我们正处在RSI的早期阶段? 过去一个月里,有些事情发生了剧变。 你最近有目睹模型发布的速度吗?
@omarsar0@omarsar0AI 评分1212 @omarsar0@omarsar0AI 评分5454 引用@finkd@finkdMuse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter. This is the biggest jump we've made so far on coding and agentic work. Try it in Muse Code and our API. Next up 🍉 and Muse Spark open weights releases coming soon. https://t.co/XQQEDEJGD7
@omarsar0@omarsar0AI 评分55 @omarsar0@omarsar0AI 评分2626 @omarsar0@omarsar0AI 评分2020 @omarsar0@omarsar0AI 评分2727 非常漂亮的图表,展示了这个模型的效率。天哪! https://t.co/VZ0Equ5R5K
引用@OfficialLoganK@OfficialLoganKGemini 3.8 Flash on DeepSWE 1.1, scores 73.7%! https://t.co/JW7fhVy4He
@omarsar0@omarsar0AI 评分2323 我测试了这个提示词。它确实有效。 它应该有助于减少 Fable 5.1 的"claudese"。 https://t.co/HMPwfvh96q https://t.co/JfJkKf3FpH

@omarsar0@omarsar0精选AI 评分6767 
引用@GoogleDeepMind@GoogleDeepMindTwo new Gemini models are here to help scale your AI agents and secure code: 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. 🔘 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level vulnerability detection and automated patching.
推荐理由:随推附上的对比表把 Gemini 3.8 Flash 的定价与多项基准成绩与 Claude、GPT-5.6 并列,便于横向比较。
@omarsar0@omarsar0AI 评分1818 @omarsar0@omarsar0AI 评分4747 
@omarsar0@omarsar0AI 评分2626 @omarsar0@omarsar0AI 评分55 @omarsar0@omarsar0AI 评分2929 @omarsar0@omarsar0AI 评分5252 一篇论文把每轮对话提取为类型化节点和带属性边、从两跳子图作答,并在固定五个检索根的候选预算下测试图记忆。

@omarsar0@omarsar0AI 评分3939 Meta 提出 AI Research Preference Models,用冻结的预训练 LLM 预测哪个候选方案最值得投入 GPU 时间,解决研究智能体"想法多、算力少"的筛选难题。

@omarsar0@omarsar0AI 评分1515 我只庆幸他没把它们称为“智能体文明”。 想象一下如果他真这么说了,头条会怎么写。 说正经的,这是我们作为一个行业需要弄清楚的重大隐患。https://t.co/LIl0RvNbjt
@omarsar0@omarsar0精选AI 评分6969 
推荐理由:论文报告 openJiuwen 在固定模型策略下靠运行时可适应机制取得基准提升,读者可比对静态与动态 harness 的设计差异。
@omarsar0@omarsar0AI 评分1717 @omarsar0@omarsar0AI 评分1919 @omarsar0@omarsar0AI 评分5555
引用@claudeai@claudeaiWe’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They're the world’s most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3
@omarsar0@omarsar0AI 评分3333
@omarsar0@omarsar0AI 评分1313 定制化前沿智能新时代的早期迹象。当行业其他玩家追上定制模型的潜力与指数级应用时,那将是史诗级的。 拥有你自己的智能栈不只是关于所有权;它关乎你作为一家在智能前沿运营的未来公司如何保持相关性。
@omarsar0@omarsar0AI 评分5353 
@omarsar0@omarsar0AI 评分3131 @omarsar0@omarsar0AI 评分6060 
@omarsar0@omarsar0AI 评分2525 @omarsar0@omarsar0AI 评分88 这也让我想起了 @karpathy 几个月前的这条推文。 不完全是交互式神经视频/模拟,但我觉得大概在 3 到 n 之间(更接近 n)。https://t.co/Yu1mSflOjE

@omarsar0@omarsar0AI 评分4545 
@omarsar0@omarsar0AI 评分2323 @omarsar0@omarsar0AI 评分1414