推荐理由:报道给出 DeepSeek 的 API 毛利率与基础设施投入倍数,读者可对照其涨价幅度判断成本优势还能保留多少。
DeepSeek
全部主题DeepSeek(深度求索)的模型发布、开源权重与技术报告——开源大模型价格与性能双卷的风向标。
当前仅显示精选新闻最新精选
第 41–60 条 · 共 106 条@rohanpaul_ai@rohanpaul_ai精选AI 评分6868 
AI前线 · 微信公众号精选AI 评分7777 OpenAI 公布自研推理芯片 Jalapeño 实测结果,多项指标超过英伟达 GB300
OpenAI 公布自研推理芯片 Jalapeño 的实测结果,在 GPT-OSS 120B、DeepSeek R1 670B 和 Kimi K2.5 1T 上,峰值每瓦吞吐量达到对比系统的 1.5 至 1.9 倍,端到端延迟优于对手 1.7 至 3.6 倍。
推荐理由:OpenAI 公布自研推理芯片的实测对标数据,可看到模型公司自研算力对现有 GPU 与软件生态的实际冲击。
@kimmonismus@kimmonismus精选AI 评分7171
引用@kimmonismus@kimmonismusHoly: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency. Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency in OpenAI’s own InferenceX testing! For highly interactive workloads, OpenAI reports 2.1–4.1× higher performance. The chip is rated at 700 watts, but remained at or below 550 watts during the tested workloads. OpenAI plans to begin deploying Jalapeño by the end of 2026. Gen 2 is already deep in development, with Gen 3 taking shape. Probably thats why Tibo said that in 1-2 years 750token/s will be the default
推荐理由:原文给出 GPT-Astra 用 Codex 生成 kernel 的具体加速数据,读者可据此了解 AI 编写底层算子的表现。
@Alibaba_Qwen@Alibaba_Qwen精选AI 评分6767 引用@natolambert@natolambertOver the weekend I had Codex parse 500K arXiv AI/ML papers since ChatGPT to understand which open models are used for research. In 2024, ~30% of papers mentioned an American open model and only 10% a Chinese model. Today, ~40% of papers mention a Chinese (open) LLM, and only 25-30% an American one. Chinese models are the default for research. Chinese mentions are still growing while American open models are stagnating. When looking at this data it's important to remember that papers substantially lag model releases, as research takes a long time. Qwen's steady growth is reflective of this, but so is Llama's lasting power. Some more observations: 1. Qwen has been steadily growing, and today 1/3 of papers which mention any LLM mention qwen. OpenAI's closed models are the highest overall, at ~37%. 2. Llama peaked around April of 2025 at 30% of papers which mention any LLM (including ChatGPT etc). Llama 4 was released at about the same time, and Llama has been declining since. 3. Gemini and Claude are less common than the leading open models, mentioned in 10-15% of papers puts them behind all of Qwen, Llama, and DeepSeek. Open models should be and are the foundations of open research. The % of papers mentioning any LLM have been steadily climbing since 2023. | Year | January | April | July | October | | 2023 | 10.43% | 15.39% | 18.69% | 32.18% | | 2024 | 29.70% | 33.93% | 35.70% | 44.25% | | 2025 | 39.23% | 45.28% | 44.94% | 53.52% | | 2026 | 55.49% | 57.26% | 53.14% | TBD Now over 50% of AI papers, from 10% in 2023. Other notes: - Gemma and Mistral hover around 5-10%. - Our beloved fully-open Olmo models have been ~1% since the first release in Jan. 2024. - DeepSeek has a clear jump after R1 in Jan. 2025 - Data derived from the most popular ML arXiv categories: cs. AI, cs. CL, cs. CV, cs. LG, stat. ML Just like our downloads and derivative model data, this is updated daily on the Interconnects Open Model Dashboard.
推荐理由:引用数据呈现了近三年论文提及开源模型的份额变化,读者可据此观察中美开源模型在研究社区中的位置。
MarkTechPost精选AI 评分7777 FreeToken:单张工作站 GPU 运行 753B GLM-5.2 的边缘 MoE 推理引擎
UC Berkeley 与 UT Austin 研究团队提出边缘原生 MoE 推理引擎 FreeToken,把个人机器当作统一弹性推理平台,可在单张工作站 GPU 上运行 753B 的 GLM-5.2。
推荐理由:文中给出多档消费级与工作站 GPU 的实测吞吐和缓存命中率,可作本地 Agent 推理的选型参考。
@kimmonismus@kimmonismus精选AI 评分7171 

推荐理由:报道给出的资金与团队规模,让读者能看清英伟达自研开放权重模型的投入量级及其双线竞争位置。
@AYi_AInotes@AYi_AInotes精选AI 评分7272 墨问西东创始人池建强体验 DeepSeek Harness 一天后,叫停刚完成两个多月重构的客户端开发,开始论证全面迁移。
推荐理由:借一位创始人停掉两年自研客户端的决定,呈现插件化 Agent 底座对应用开发与迁移选择的影响。
@omarsar0@omarsar0精选AI 评分6767 引用@deepseek_ai@deepseek_aiDeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
推荐理由:官方基准对比显示该实验模型在多模态智能体任务上接近 Opus-4.8,读者可据此了解多模态智能体的当前水平。
@AYi_AInotes@AYi_AInotes精选AI 评分7272 

引用@deepseek_ai@deepseek_aiDeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
推荐理由:逐条列出这个实验性多模态模型的单图 token 上限、分级定价与免费 Files API 额度,可供评估视觉 Agent 的成本与接入方式。
@kimmonismus@kimmonismus精选AI 评分6767
引用@deepseek_ai@deepseek_aiDeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
推荐理由:基准对比显示 Flash 小模型在多模态智能体评测上已接近 Opus-4.8,可供判断轻量模型的能力边界。
@testingcatalog@testingcatalog精选AI 评分7171
引用@deepseek_ai@deepseek_aiMultimodal API support 🔌 🔹 Set model='deepseek-v4-flash-vision-exp' 🔹 Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing 🔹 Supports Chat Completions, Messages & Responses 🔹 Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API. Docs: https://t.co/USZ5gZ3wWB 3/n
推荐理由:材料给出新模型与 Opus 4.8 及自家文本版本的基准对比和计费细节,便于判断实际可用性。
DeepSeek API updates精选AI 评分7070 DeepSeek 发布实验版多模态模型 DeepSeek-V4-Flash-Vision-Exp
DeepSeek 在 API 平台上线实验性多模态视觉理解模型 DeepSeek-V4-Flash-Vision-Exp,通过设置 model='deepseek-v4-flash-vision-exp' 即可调用。
推荐理由:官方公布了多模态 Agent 基准成绩和调用方式,读者可据此判断其视觉 Agent 能力在现有模型中的位置。
IT Home精选AI 评分8181 宇树科技科创板上市首日收涨 460%,开盘一度涨 629%、市值突破 4400 亿
宇树科技 8 月 19 日正式登陆科创板,发行价 150.80 元/股,收盘报 845 元/股,上涨 460.34%,成交额超 230 亿元,市值超 3418 亿元。
推荐理由:宇树科技上市首日收涨 460%,发行市盈率远超行业平均,可对照其 5500 台人形机器人出货量看赛道估值。
硅星人Pro · 微信公众号精选AI 评分7777 DeepSeek V4 Pro 与 V4 Flash 涨价生效,下游订阅与 Agent 平台成本承压
DeepSeek V4 Pro 和 V4 Flash 的新价格于8月17日正式生效,V4 Pro 峰时缓存命中价格从每百万 Token 0.003625 美元提高到 0.044 美元,为原来的 12.14 倍,输出价格从 0.87 美元提高到 3.96 美元;V4 Flash 峰时输出从 0.28 美元提高到 1.32 美元,为原来的 4.71 倍。
推荐理由:文章用峰谷价差和订阅额度变化,量化这次涨价对下游开发者与Agent平台的实际成本冲击。
虎嗅APP · 微信公众号精选AI 评分7777 宇树科技登陆科创板,成 A 股第一家人形机器人上市公司
宇树科技于2026年8月19日登陆科创板,发行4044.64万股,发行价150.80元,募集资金总额60.99亿元,发行市值609.93亿元,成为A股市场第一家人形机器人上市公司。
推荐理由:以上市为节点复盘宇树十年硬件路线,并对比 Figure、Boston Dynamics 等海外公司的不同选择。
机器之心 · 微信公众号精选AI 评分8181 DeepSeek Harness 背后的论文:让 Agent 在运行中改写自己
机器之心解读了 DeepSeek Harness 背后的八十页论文《A Programming Paradigm for Spatiotemporal Composability》。
推荐理由:文章把八十页论文里的效应与余效应机制讲成运行时插件的卸载设计,便于理解自修改智能体如何干净回收组件。
十字路口Crossing · 微信公众号精选AI 评分7878 DeepSeek 开源 Harness DSH:这不是产品,而是一场战略
DeepSeek 于 8 月 13 日发布首个 Harness 版本 DSH,以 MIT 协议在 GitHub 开源,核心口号是一切皆插件。
推荐理由:文章把 DSH 的插件化架构放进 Harness 范式尚未收敛的背景里,解释其开源策略为何是保留选择权而非交付成品。
elsewhere articles精选AI 评分7575 DeepSeek 开源 Harness 框架 DSH:这不是产品,而是一个战略
DeepSeek 于 8 月 13 日发布首个 Harness 版本 DSH,以 MIT 协议开源在 GitHub 上,仓库版本号仍停留在 0.1.0-rc,官方称其为开发者预览版且未来会出现破坏兼容性的变更。
推荐理由:从插件化架构与开源生态两个角度拆解 DSH,帮助读者理解 DeepSeek 把 Harness 探索交给社区的思路。
硅星人Pro · 微信公众号精选AI 评分8080 玩了一夜 DeepSeek Harness,我发现它想用《我的世界》的方式干掉 Claude Code
DeepSeek 在 8 月 13 日晚发布 DeepSeek Harness v0.1 开发者预览版,采用 MIT 协议开源,用 npx @deepseek-ai/dsh web 即可在本地浏览器启动。
推荐理由:文章以一夜实测加插件生态盘点,对比 Claude Code 的封闭外壳路线,呈现 harness 被开源后的竞争变化。
量子位 · 微信公众号精选AI 评分7878 深度体验 DeepSeek Harness:开源编码 Agent 的插件化架构与实测对比
DeepSeek Harness 正式发布并开源,作者内测半个月后把自己的 Vibe Coding 项目从 Codex 迁移到 DSH。DSH 采用一切皆插件的 Cordis 架构,内置 100 多个插件,提供标准、PTC、极简、创造四类 Agent 预设,并有可直接查看原始事件的轨迹回放视图。
推荐理由:作者用半个月内测体验对比 Codex,呈现 DSH 的插件化架构与轨迹回放等设计,便于判断它与现有编码 Agent 的差异。