#编码
#编码
今日 23 条
赵纯想@chunxiangaiAI 评分2727Hugging Face Daily PapersAI 评分3838 AgSpec:突破检索式推测解码在编码智能体流水线中的极限
AgSpec 框架为编码智能体流水线补齐检索式推测解码缺失的语料库与草稿长度策略,从会话、工作区和全局语料检索,并按智能体离线画像设定草稿长度上限、在线自适应调整。在两个仓库级多智能体编码基准上,AgSpec 在多数设置下优于五种检索式草稿器及 EAGLE-3,批大小 1 和 16 时生成吞吐量较自回归解码分别提升最高 4.37 倍和 4.76 倍。
Latent SpaceAI 评分5959 Pi 1.0 与 Pi Durable 发布,AINews 汇总 AI 工程动态
Latent Space 期刊报道 Pi 1.0 与 Pi Durable 同时登上 HN 首页,Pi 1.0 增加 MCP 原生支持、虚拟模型扩展、延迟工具加载和会话中系统消息;Pi Durable 将 Pi 移植到 TypeScript 并外置全部有状态组件,支持检查点崩溃恢复、可插拔存储后端、并行分支对话和热更新工具代码。
Dongxi 东锡 NLP@dongxi_nlpAI 评分3838Hacker News popular via buzzing.cc精选AI 评分7676 DeepSeek Harness 开启全球公开预览并开源
DeepSeek Harness 进入全球公开预览并开源,基于 Cordis 的“一切皆插件”架构,可作为桌面应用运行或从代码启动 Web UI。它支持日常办公、编码、研究、后台任务,可通过“Creator mode”在聊天中创建插件,用 npx @deepseek-ai/dsh web 一条命令启动,源码在 github.com/deepseek-ai/deepseek-harness。
推荐理由:原文给出 DeepSeek Harness 的插件架构、安装方式和适用场景,读者可据此评估是否纳入自己的工作流。
Thariq@trq212AI 评分2929一直在尝试提升我游戏原型里动画的质量,所以让 Claude 教我并帮我找参考。 我让 Claude 做了一个动画编辑器,方便我们迭代跳跃动作,效果让我非常满意。 这是对比视频

Claude Blog精选AI 评分7373 Claude Code 推出 mods:用 TypeScript 函数定制和替换内置功能
Anthropic 为 Claude Code 推出 mods,即小型 TypeScript 函数,可改写提示词、拦截工具调用、审批权限请求或添加新 UI。Mods 随插件分发,可在 CLI 和桌面应用使用,内置的 /diff 功能已改为 mod。
推荐理由:官方公布了 mods 的事件机制、企业管控方式和使用入口,开发者可据此决定是否用 mods 定制自己的 Claude Code 工作流。
Hacker News popular via buzzing.ccAI 评分6767 Earendil 发布 Pi 1.0 稳定版及实验性包 Pi Durable
Earendil 发布 Pi 1.0,一个极简、可扩展的 agent 编码工具,已有每周数十万人使用,MIT 协议开源,提供 curl 或 powershell 安装脚本。
TechCrunch · AIAI 评分6363 Google 发布 Gemini 4 Argon,称其为迄今最强模型
Google(Alphabet)发布新模型 Gemini 4 Argon,主打防御性网络安全,称其可自主发现、验证并修复关键软件漏洞,目前仅通过 Fairwind 安全计划向部分网络安全合作伙伴开放。该模型也用于编码、调试和代码库迁移等日常工程工作,并称在多项基准上显著领先 GPT-6 Astra 与 Anthropic 的 Fable 和 Opus。
ginobefun@hongming731AI 评分2626LangChain 在 Open SWE 智能体 Harness 内构建模型路由器,按任务分类选择最低成本合适模型,将编码线程中位数成本降低 64%,质量变化可忽略不计。
引用ginobefun@hongming731https://x.com/i/article/2105830283580989440
elvis@omarsar0AI 评分5050引用Tapa Ghosh@semiDLExcited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: https://volantissemi.ai/news-insights/our-88m-series-a-demolishing-the-memory-wall-with-photonics-post
Latent SpaceAI 评分5252 Latent Space 访谈 MIT 的 Alex Zhang:RLM、harness 设计与研究品味
Latent Space 播客访谈 MIT 博士生 Alex Zhang,围绕其 Recursive Language Models(RLM)研究展开。
Artificial Analysis@ArtificialAnlys精选AI 评分7070
推荐理由:原文给出三大新模型组合在 Coding Agent Index 的得分与单任务成本对比,读者可以据此在性能和价格之间做选型权衡。
Hacker News popular via buzzing.ccAI 评分6464 生成式 AI 冲击下 Web 开发教育的没落
molily.de 的博客文章汇总了 Baldur Bjarnason、Axel Rauschmayer、Salma Alam-Naylor、Josh W. Comeau、Kyle Cook 和 Rachel Andrew 等多位从业者的陈述,描述生成式 AI 对 Web 开发教育和技术出版的冲击。
OpenAI Developers@OpenAIDevsAI 评分1616听说你们喜欢 @modretro Chromatic。 所以我们往支线任务里加了一台。 还有 48 小时可以加入。 https://side-quests.openai.chatgpt.site/
引用OpenAI Developers@OpenAIDevsReady for a DevDay side quest? 10 builder challenges. 5 Codex Micros up for grabs. Complete any 5 side quests to earn 1 raffle ticket for each. 48 hours. Join from anywhere 🌎
Charlie Holtz@charlieholtzAI 评分2121引用Thomas Paul Mann@thomaspaulmann@charlieholtz Slide to merge!!!
Boris Cherny@bcherny精选AI 评分7070引用ClaudeDevs@ClaudeDevsYou can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
推荐理由:原文介绍了用提示词自定义 Claude Code 行为和界面的新机制,以及通过插件安装和分享的方式。
Artificial Analysis@ArtificialAnlysAI 评分4040Artificial Analysis 的 Coding Agent Index 新增安全拒绝报告,可查看拒绝发生在任务提示词阶段还是智能体已开始工作之后,以及拒绝后切换到了哪个回退模型。

Thariq@trq212AI 评分1616我通常为开发者写作,但下一篇帖子我打算面向那些试图应对 AI 编程智能体变革的公司领导者。 你希望你的领导层能理解智能体的哪些方面?或者如果你正在经营一家公司——你遇到了什么问题?
Thariq@trq212AI 评分6363引用ClaudeDevs@ClaudeDevsYou can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
Claude Code GitHub ReleasesAI 评分5555 Claude Code v2.1.287 发布:新增 Claude Mods 与多项修复
Claude Code 发布 v2.1.287,新增 Claude Mods 插件机制和内置 mod "You should know",Opus 4.7+ 与 Fable 在 Bedrock、Vertex、Foundry 等默认使用 1M 上下文窗口。另修复大量问题,包括 rm 危险命令保护、屏幕阅读器模式和 VSCode 扩展的多项缺陷,并改进 Bash 速度与 MCP 权限提示展示。
ClaudeDevs@ClaudeDevs精选AI 评分6666
推荐理由:官方给出 mods 的可改动范围和安装方式,读者可以据此判断是否用 TypeScript 定制自己的 Claude Code 工作流。
Josh Woodward@joshwoodwardAI 评分4646引用Stitch by Google@stitchbygoogleWe have a Stitch MCP. We have a Stitch SDK. But, we don't have a Stitch CLI. Well, not until now. Introducing the @google/stitch CLI: 🔷 Connect to your local coding agents 🔷 Generate screens and design systems 🔷 Send a local dev server snapshot to Stitch Do it all without leaving the terminal or better yet, ask your favorite harness like @antigravity 😎 Learn more 👇
OpenRouter@OpenRouterAI 评分4545
jason@jxnlcoAI 评分1919
dex@dexhorthyAI 评分3131引用Hari@HarivanshRathion ai psychosis: as i watch engineers fall deeper into the belief that an amalgamation of mathematical probabilities somehow understands their codebase better than they do, i find myself thinking back to a time when software wasn’t built for hypergrowth, but simply to do x without inventing y. ai seems almost fundamentally opposed to this philosophy. ask it to do x and it will eagerly invent y and z before it has even tried to understand x. the danger isn’t that ai writes bad code- it’s that it makes writing unnecessary code 100% free and the human condition is such that some of us will always prefer the fast, steep gains of ai, even when it does a bajillion unrelated things to accomplish something that could have been done without changing anything else. so, somewhat paradoxically, the quality of software may keep declining for as long as ai keeps getting better. the cheaper complexity becomes to create, the less incentive there is to understand or avoid it. try to preserve this craft created by our ancestors write a LOC by hand today
ginobefun@hongming731AI 评分3939引用ginobefun@hongming731https://x.com/i/article/2105450924286353408
jason@jxnlcoAI 评分3636引用/ (Kuramoto) Ko@ko_kuramotoこの時代にゲームボーイの新作を開発しました。 しかも、カートリッジをWi-Fi接続可能な形に魔改造。 ネット経由でゲームボーイ上で生成AIが動きます。 AIと会話して誰が殺人犯か、謎を解け。 #スーパーゲ制デー
jason@jxnlcoAI 评分3939OpenAI 在 Dev Day 上向每位参会者赠送了一台可运行 AI 的 Game Boy,灵感来自 Ko Kuramoto 及其 nanu 团队的硬件改造探索。
引用/ (Kuramoto) Ko@ko_kuramotoということでWi-Fiに接続して、AIをゲームボーイで駆動させるDaydream、再生産を行います! 開発当初チーム内で話していた「グリードアイランドみたいなゲーム作りたいね」という話にちなんで、800台限定。これ以上は再生産なしです。
OpenRouter Announcements精选AI 评分7474 OpenRouter 教程:如何在 CI 中用 LLM eval 门禁拦截 Pull Request
OpenRouter 发布教程,讲解用固定的 eval 集合在 GitHub Actions 中门禁 PR,当 pass rate 低于阈值时阻止合并。
推荐理由:OpenRouter 官方给出可落地的 CI eval 门禁方案,包含阈值测量和 provider 路由等易踩坑细节,方法可直接迁移到自己的仓库。
Apple Machine Learning ResearchAI 评分4343 Apple 提出 RLTL;DR:通过内化自生成反馈实现自我提升
Apple 研究团队提出 RLTL;DR,让策略在每次失败后根据验证器输出自行撰写一条 TL;DR 式反馈,并将上下文中的洞察反向传播以内化“任务→洞察”映射。
Apple Machine Learning ResearchAI 评分5555 Apple 与 EPFL 研究:强模型做自主 ML 工程时复杂 harness 不占优势
Apple 与 EPFL 研究者发表论文,比较开源 SOTA MLE agent harness 与最小 harness 的单会话 coding agent。在相同时间预算和同一前沿 LLM 骨干下,复杂 harness 并无优势,性能主要取决于模型骨干;大规模消融实验显示多层机制在 coding agent 场景中变得冗余。
AK@_akhaliqAI 评分2525LEGO-Anything 面向 3D 场景重建的编码智能体 论文:https://huggingface.co/papers/2609.36380

Latent SpaceAI 评分7474 Latent Space DevDay 访谈:Ari Weinstein 谈 Computer Use 进展,Nikunj Handa 解析 OpenAI 新 API 栈
Latent Space 在 OpenAI DevDay 当天发布访谈,Computer Use 负责人 Ari Weinstein 表示 Computer Use 已与数月前 180 度不同。
fofr@fofrAIAI 评分3333非常激动地分享,Gemini 4 Argon 即将到来。迫不及待想尽快跟大家分享更多内容。
引用Sundar Pichai@sundarpichaiLots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:
ClaudeDevs@ClaudeDevsAI 评分3030
Ars Technica · AIAI 评分6565 Google 发布 Gemini 4 Argon 模型,暂未开放使用
Google 发布 Gemini 4 Argon,称其在编码、知识工作和网络安全方面性能领先,但模型仍处有限测试,普通用户暂无法使用。DeepSWE v1.1 达 77.9%,高于 GPT-6 Astra、Fable 5.1 和 Opus 5.5;API 定价为每百万输入 token $2、输出 $10,输出上限提升至 100 万 token(此前为 64,000)。
Google DeepMind@GoogleDeepMindAI 评分3939推出 Gemini 4 Argon——我们的全新前沿模型。 它专为编码、企业知识工作和网络安全防御等复杂工作流打造——今天起通过我们的 Fairwind Program 向一批受信任的测试者开放。

Google DeepMind精选AI 评分7474 Google DeepMind 发布 Gemini 4 Argon,输出上限扩至 1M tokens
Google DeepMind 发布前沿模型 Gemini 4 Argon,先向 Fairwind Program 的可信网络防御者开放,后续将面向开发者、企业和消费者推出。
推荐理由:官方博客给出定价、输出上限和多个基准成绩,可帮读者评估该模型在编码与安全防御场景的实际定位。
eric zakariasson@ericzakariassonAI 评分4848来看看我们 Grok Bot 市场里的工程类机器人! https://x.ai/bot/marketplace/engineering
引用Grok Bot@botGrok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.