全部AI 动态
全部动态
今日 2 条
Baidu Inc.@Baidu_IncAI 评分2121
inclusionAI Hugging Face modelsAI 评分4747 inclusionAI 发布 Llama-3.2-1B-Instruct-singprobe 流式安全探针
inclusionAI 推出 SingProbe,一个基于 meta-llama/Llama-3.2-1B-Instruct 的流式安全探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,仅增加不到 0.5% 解码开销。
inclusionAI Hugging Face modelsAI 评分4545 inclusionAI 发布 gemma-4-26B-A4B-it-singprobe:基于 Gemma 的流式安全探针
inclusionAI 发布 SingProbe,一个基于 google/gemma-4-26B-A4B-it 的流式安全探针,复用基座模型隐藏状态逐 token 打分,仅增加不到 0.5% 解码开销。
inclusionAI Hugging Face modelsAI 评分5656 inclusionAI 发布 SingProbe 流式安全探测,基于 Qwen3.6-35B-A3B
inclusionAI 在 Hugging Face 发布 SingProbe,一个复用 Qwen3.6-35B-A3B 隐层状态的轻量流式安全探测,逐 token 输出 8 类查询意图、不安全与幻觉风险分数,探测参数仅 4.2M。
inclusionAI Hugging Face modelsAI 评分5151 inclusionAI 发布基于 Qwen3-0.6B 的 SingProbe 流式安全探针
inclusionAI 发布 SingProbe,一个构建在 Qwen/Qwen3-0.6B 之上的流式安全探针,复用基座模型隐藏状态在每个 token 上评估查询意图、回复不安全性和幻觉风险。
inclusionAI Hugging Face modelsAI 评分5656 inclusionAI 发布 SingProbe:基于 Qwen3.5-122B-A10B 的流式安全护栏探针
inclusionAI 在 Hugging Face 发布 SingProbe,一个构建在 Qwen/Qwen3.5-122B-A10B 上的内在流式安全护栏探针,复用基座模型隐藏状态,在每个 token 上输出 8 类查询意图、回复不安全度和幻觉风险分数,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4545 inclusionAI 发布 Qwen3.5-0.8B-singprobe:基于隐藏状态的流式安全探针
inclusionAI 发布 Qwen3.5-0.8B-singprobe,一个基于 Qwen/Qwen3.5-0.8B 隐藏状态的流式安全探针,可在生成过程中逐 token 输出查询意图、回复不安全和幻觉风险评分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4646 inclusionAI 发布 Qwen3.5-397B-A17B-singprobe 流式安全探针
inclusionAI 发布基于 Qwen/Qwen3.5-397B-A17B 的流式安全探针 SingProbe,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4646 inclusionAI 发布 Qwen3.8-27B-singprobe:基于隐藏状态的流式安全探针
inclusionAI 发布 Qwen3.8-27B-singprobe,这是一个构建在 Qwen/Qwen3.8-27B 上的内在流式护栏探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4141 inclusionAI 发布 SingProbe:基于 gemma-4-E4B-it 的流式安全探针
inclusionAI 推出 SingProbe,一个基于 google/gemma-4-E4B-it 的流式安全探针,复用基座模型隐藏状态,在生成每个 token 时对查询意图、回复安全性和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4444 inclusionAI 发布 SingProbe:基于 Qwen3-4B-Instruct-2507 的流式安全探针
inclusionAI 发布 SingProbe,一个构建在 Qwen/Qwen3-4B-Instruct-2507 上的内置流式护栏,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4444 inclusionAI 发布 GLM-5.3-singprobe 流式安全探针
inclusionAI 发布基于 GLM-5.3 的流式安全探针 GLM-5.3-singprobe,复用基座模型隐藏状态逐 token 打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4545 inclusionAI 发布 Qwen3.5-4B-singprobe:基于 Qwen3.5-4B 的流式安全探针
inclusionAI 发布 Qwen3.5-4B-singprobe,这是一个构建在 Qwen/Qwen3.5-4B 上的流式安全探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4242 inclusionAI 发布 Qwen3.5-9B-singprobe:基于隐藏状态的流式安全探针
inclusionAI 发布 Qwen3.5-9B-singprobe,这是一个构建在 Qwen/Qwen3.5-9B 上的内置流式护栏,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4242 inclusionAI 发布 gemma-4-31B-it-singprobe:基于 Gemma 4 的流式安全探针
inclusionAI 发布 SingProbe,一个基于 google/gemma-4-31B-it 的流式护栏探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4242 inclusionAI 发布 gpt-oss-120b-singprobe:基于 gpt-oss-120b 的流式安全探针
inclusionAI 发布 gpt-oss-120b-singprobe,这是构建在 openai/gpt-oss-120b 上的流式护栏探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4646 inclusionAI 发布 Step-3.7-Flash-singprobe:基于 Step-3.7-Flash 的流式安全探针
inclusionAI 发布 SingProbe,一个构建在 stepfun-ai/Step-3.7-Flash 上的内在流式护栏,复用基座模型隐藏状态,在生成每个 token 时对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4444 inclusionAI 发布 Hy3-singprobe:基于腾讯 Hy3 的流式安全探针
inclusionAI 发布 Hy3-singprobe,一个构建在 tencent/Hy3 上的流式安全探针,复用基座模型隐藏状态逐 token 打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4545 inclusionAI 发布 Llama-3.1-8B-Instruct-singprobe 流式安全探针
inclusionAI 发布 Llama-3.1-8B-Instruct-singprobe,一个基于 meta-llama/Llama-3.1-8B-Instruct 的流式安全探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4848 inclusionAI 发布 Qwen3-8B-singprobe:基于 Qwen3-8B 的流式安全探针
inclusionAI 发布 Qwen3-8B-singprobe,这是构建在 Qwen/Qwen3-8B 上的内在流式护栏,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分4545 inclusionAI 发布 gpt-oss-20b-singprobe:基于 gpt-oss-20b 的流式安全探针
inclusionAI 发布 gpt-oss-20b-singprobe,这是构建在 openai/gpt-oss-20b 上的流式护栏探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
inclusionAI Hugging Face modelsAI 评分5050 inclusionAI 发布基于 MiniMax-M2.7 的流式护栏探针 SingProbe
inclusionAI 发布 SingProbe,一个构建在 MiniMaxAI/MiniMax-M2.7 之上的流式护栏探针,在生成过程中复用基座模型隐藏状态,对每个 token 输出查询意图、回复不安全度和幻觉风险评分。
OpenAI NewsAI 评分4343 Perplexity 用 GPT-6 Astra 端到端接管系统
Perplexity 正使用 GPT-6 Astra 撰写沟通内容、修改软件并监控生产系统,且相比早期模型,人工检查频率大幅降低。
OpenRouter Announcements精选AI 评分6262 OpenRouter Ori Eval 教程:用 LLM-as-a-Judge 自动评估 AI Agent 输出
OpenRouter 发布教程,讲解如何用独立的 judge 模型按书面评分标准自动给 Agent 输出打分,弥补确定性测试无法检查最终回答质量的问题。
推荐理由:原文给出可复用的评测方法和 Ori Eval 具体代码,覆盖评分模式选择与 judge 校准的实操细节。
Mustafa Suleyman@mustafasuleymanAI 评分4040这是一个非常直白且符合常识的观点:技术的目的是服务人类,加速人类繁荣。 任何无法实现这一目标的技术都是失败的,应当被拒绝。 我们还没有到那一步。但开始为这种可能性做准备是正确的。
引用Satya Nadella@satyanadellaAny pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive. And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control. So, in this context, we welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like "embedded evaluators" and the broader efforts to develop the mechanisms to make this more than just talk. The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia. This is the approach we are taking: broad access and choice at every layer of the AI stack; enterprise control of learning loops and models; and the “Code of Conduct” that underlies our own first party MAI models that we’ll publish tomorrow for public consultation.
elsewhere articlesAI 评分3838 当具身智能走到十字路口:苏度、蚂蚁灵波、自变量、破壳谈四种一线判断
苏度科技韩铮、蚂蚁灵波沈宇军、自变量王潜、破壳许华哲在 2026 Inclusion 外滩大会圆桌中,围绕具身智能的数据来源、模型路线与落地场景展开了一场未收敛的路线级分歧讨论。对话聚焦 GPT-6 Astra 的能力边界及其对具身行业的冲击,并探讨高成功率与泛化性、客户持续付费等跨越泡沫的指标。嘉宾还就五年后被高估与低估的领域给出各自判断。
Claude Platform release notes精选AI 评分6060 Claude API Messages 支持按需压缩对话,compact-2026-09-04 beta 上线
Claude API 的 Messages API 新增按需压缩对话功能,以 compact-2026-09-04 beta 头开启。发送顶层 compaction 参数后返回带签名的压缩块概括所发消息,后续请求可先发该块替代原消息,支持后台运行、保留最近轮次原文,且在支持保留 thinking 的模型上已保留轮次的 thinking 可保持有效。
推荐理由:原文给出了 compaction 的调用方式和 beta 头参数,开发者可据此评估在长对话场景里如何按需压缩上下文。
LangChain BlogAI 评分4949 医疗与生命科学领域的智能体规模化:来自 Madrigal Pharmaceuticals、Abridge 和 Vizient 的经验
LangChain 博客总结医疗与生命科学行业智能体规模化经验:76% 受访机构将 tracing、评估和支出可见性列为扩大智能体自主权的前提,49% 正在建设公司级智能体平台或“agent factory”,33% 聚焦已有纸质记录的受监管文档与后台流程,26% 在运行或构建面向患者和会员的对话式智能体(含语音),另有 26% 尝试让非工程师在中央护栏内构建智能体。
Sakana AI BlogAI 评分5050 Sakana AI 提出 PC-ALM:无需反向传播训练 1000 层网络
Sakana AI 提出 PC-ALM,一种仅靠局部动力学、无需反向传播即可训练 1000 层神经网络的局部学习替代方案。该方法将预测编码推广为使用增广拉格朗日,引入对偶神经元(拉格朗日乘子)参与层内动力学,使每层成为 PI 反馈控制系统以最小化局部预测误差。PC-ALM 能将信号传播到近乎任意深度,在标准预测编码难以学习的深层窄网络中尤其有效,并可能为神经形态硬件上的深度学习提供参考。
Demis Hassabis@demishassabisAI 评分6262引用Dario Amodei@DarioAmodeiWe Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
Peter McCrory@PeterMcCrory精选AI 评分6969引用Dario Amodei@DarioAmodeiWe Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
推荐理由:Anthropic 首席经济学家推荐 Dario Amodei 新文,提出给第三方评估者永久员工级访问权以核验安全措施,可了解行业自律的具体动作。
Claude Code GitHub ReleasesAI 评分2828 Claude Code v2.1.270 修复只读 git 命令误请求权限问题
Claude Code 发布 v2.1.270,修复了会话运行一段时间后 Bash 中只读 git 命令意外请求权限的问题,该问题为 2.1.269 引入的回归。
Aidan Gomez@aidangomezAI 评分4747引用Sam Altman@samaI agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
LangChain Blog精选AI 评分6666 LangChain 发布付费广告 Agent 并开源,分享构建经验
LangChain 开源其 Paid Media Agent 并复盘构建过程,该 Agent 驻留 Slack,用于管理跨五个付费渠道的广告活动。六个月内付费媒体从 0 做到占营销管道的 20%,CPL 从 6 月到 8 月下降 30%。
推荐理由:原文给出了 Agent 工程的具体设计决策和成本数据,读者可以迁移其上下文分层和权限设计方法到自己的 Agent 项目。
Logan Kilpatrick@OfficialLoganKAI 评分88ByteByteGoAI 评分3232 为什么 Git revert 会产生冲突?
git revert 不重写历史,而是新建一个提交来撤销早先提交的改动,因此当后续提交修改了同一批代码行时就会触发冲突。例如 C2 添加功能、C3 又改了这些行,撤销 C2 便会与 C3 的改动相撞,Git 无法判断哪个版本正确。解决方式是运行 git revert C2,在冲突处暂停后手动修改文件、暂存并继续,最终生成一个干净撤销 C2 且保留 C3 的新提交。
ViggleAI@ViggleAIAI 评分4242你肯定没见过这个: GPT-6 Astra 用于场景与道具建模 + PINOC mcp 用于可动画的高斯泼溅角色
引用PINOC@Viggle_PINOCGPT-6 Astra can now generate animatable Gaussian Splat characters. We connected it to the PINOC MCP and asked for a backrooms-style, Exit 8-ish game. We described the character we wanted and the motions. [MCP link in the comment 👇] Astra generated the character and every motion through PINOC through free preset animations and text to animation, and wrote the loop and the anomaly logic itself, and shipped the whole thing in a few sessions.
Peter McCrory@PeterMcCroryAI 评分4646这是该模型的一个重要局限。我们聚焦于 AI 转型的供给侧(AI 能做什么、扩散多快、劳动者转岗多快)。 价格是灵活的,总需求等于经济体的产出能力。 更多思考见 🧵
引用modest proposal@modestproposal1Anthropic's economic scenario analysis is interesting. But this is not something you can ignore, this is the most important consideration! "the model cannot generate the negative feedback in which disruption depresses demand and amplifies its own labor-market consequences"
Claude Code GitHub ReleasesAI 评分4242 Claude Code v2.1.269 发布:新增插件 eval 套件与输出样式切换
Claude Code v2.1.269 新增 `claude plugin eval`,可对插件运行 eval 套件并输出 JSON 与 HTML 评分报告,同时加入 `/output-style [name]` 用于列出和切换输出样式。
GitHub Blog · AI & ML精选AI 评分6060 GitHub 日本韩国营销负责人分享如何用 GitHub Copilot 把活动运营自动化为代码
GitHub 日本与韩国营销负责人 Tomoko Tanaka 撰文分享如何不写代码,用 GitHub Copilot 把活动运营从策划到会后跟进全流程自动化。
推荐理由:作者以自身营销流程为例,展示如何用 runbook 加 Copilot 在 GitHub 上搭建自动化,方法可迁移到其他重复性工作。