跳到正文

DeepSeek

DeepSeek(深度求索)的模型发布、开源权重与技术报告——开源大模型价格与性能双卷的风向标。

当前仅显示精选新闻
106条精选相关主题千问 Qwen开源生态模型发布

最新精选

第 41–60 条 · 共 106 条
8月26日周三
  1. AI前线 · 微信公众号77

    OpenAI 公布自研推理芯片 Jalapeño 实测结果,多项指标超过英伟达 GB300

    OpenAI 公布自研推理芯片 Jalapeño 的实测结果,在 GPT-OSS 120B、DeepSeek R1 670B 和 Kimi K2.5 1T 上,峰值每瓦吞吐量达到对比系统的 1.5 至 1.9 倍,端到端延迟优于对手 1.7 至 3.6 倍。

    推荐理由:OpenAI 公布自研推理芯片的实测对标数据,可看到模型公司自研算力对现有 GPU 与软件生态的实际冲击。

  2. @kimmonismus71

    OpenAI 在关于 Jalapeño 芯片的博客中披露,GPT-Astra 用 Codex 编写并优化底层 kernel,两个月内把三个原本不在计划内的开源权重模型带到该芯片上的高性能。在选定的 attention 和 MoE 模块上,AI 生成的实现比现有人工专家代码快 1.5–1.8 倍,这些数字只适用于选定模块而非完整模型。引用的推文还提到,Jalapeño 在 GPT-OSS 120B、DeepSeek R1 670B 和 Kimi K2.5 1T 上实现 1.5–1.9 倍每瓦 AI 吞吐与 1.7–3.6 倍更低的端到端延迟,OpenAI 计划 2026 年底开始部署。

    引用@kimmonismus@kimmonismus

    Holy: OpenAI says its first custom inference chip is already beating Nvidia GB200 and GB300 systems on speed and efficiency. Across GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5–1.9× more AI work per watt at peak throughput and 1.7–3.6× lower end-to-end latency in OpenAI’s own InferenceX testing! For highly interactive workloads, OpenAI reports 2.1–4.1× higher performance. The chip is rated at 700 watts, but remained at or below 550 watts during the tested workloads. OpenAI plans to begin deploying Jalapeño by the end of 2026. Gen 2 is already deep in development, with Gen 3 taking shape. Probably thats why Tibo said that in 1-2 years 750token/s will be the default

    推荐理由:原文给出 GPT-Astra 用 Codex 生成 kernel 的具体加速数据,读者可据此了解 AI 编写底层算子的表现。

8月25日周二
  1. @Alibaba_Qwen67

    通义千问(Qwen)官方账号转发 natolambert 的分析并致谢,该分析用 Codex 解析了 ChatGPT 发布以来的 50 万篇 arXiv AI/ML 论文。数据显示,2024 年约 30% 论文提及美国开源模型、仅 10% 提及中国模型,如今约 40% 提及中国开源 LLM、25-30% 提及美国模型;提及任一 LLM 的论文中有三分之一提到 Qwen,OpenAI 闭源模型以约 37% 居首,Llama 在 2025 年 4 月达到 30% 峰值后持续下滑。提及 LLM 的论文占比已从 2023 年的 10% 升至 50% 以上。

    引用@natolambert@natolambert

    Over the weekend I had Codex parse 500K arXiv AI/ML papers since ChatGPT to understand which open models are used for research. In 2024, ~30% of papers mentioned an American open model and only 10% a Chinese model. Today, ~40% of papers mention a Chinese (open) LLM, and only 25-30% an American one. Chinese models are the default for research. Chinese mentions are still growing while American open models are stagnating. When looking at this data it's important to remember that papers substantially lag model releases, as research takes a long time. Qwen's steady growth is reflective of this, but so is Llama's lasting power. Some more observations: 1. Qwen has been steadily growing, and today 1/3 of papers which mention any LLM mention qwen. OpenAI's closed models are the highest overall, at ~37%. 2. Llama peaked around April of 2025 at 30% of papers which mention any LLM (including ChatGPT etc). Llama 4 was released at about the same time, and Llama has been declining since. 3. Gemini and Claude are less common than the leading open models, mentioned in 10-15% of papers puts them behind all of Qwen, Llama, and DeepSeek. Open models should be and are the foundations of open research. The % of papers mentioning any LLM have been steadily climbing since 2023. | Year | January | April | July | October | | 2023 | 10.43% | 15.39% | 18.69% | 32.18% | | 2024 | 29.70% | 33.93% | 35.70% | 44.25% | | 2025 | 39.23% | 45.28% | 44.94% | 53.52% | | 2026 | 55.49% | 57.26% | 53.14% | TBD Now over 50% of AI papers, from 10% in 2023. Other notes: - Gemma and Mistral hover around 5-10%. - Our beloved fully-open Olmo models have been ~1% since the first release in Jan. 2024. - DeepSeek has a clear jump after R1 in Jan. 2025 - Data derived from the most popular ML arXiv categories: cs. AI, cs. CL, cs. CV, cs. LG, stat. ML Just like our downloads and derivative model data, this is updated daily on the Interconnects Open Model Dashboard.

    推荐理由:引用数据呈现了近三年论文提及开源模型的份额变化,读者可据此观察中美开源模型在研究社区中的位置。

8月23日周日
  1. @kimmonismus71

    据 WSJ 报道,英伟达将投入 60 亿美元打造一个开放权重 AI 模型,通过授权 Poolside 的技术并把其 100 多名员工并入 Nemotron 项目。英伟达还将以 120 亿美元投前估值向 Poolside 追加投资 10 亿美元。其目标是挑战 DeepSeek、Kimi 等中国开放权重模型团队,同时直接对标 OpenAI、Anthropic 等美国前沿实验室。

    推荐理由:报道给出的资金与团队规模,让读者能看清英伟达自研开放权重模型的投入量级及其双线竞争位置。

8月22日周六
8月21日周五
  1. @omarsar067

    DeepSeek-V4-Flash-Vision-Exp 已在 DeepSeek API 平台上线,该实验性多模态模型在文本能力上与 DeepSeek-V4-Flash 持平,在多模态智能体基准上大幅超越 V4-Flash 并接近 Opus-4.8。DeepSeek Harness 0.1.1 同日发布,内置对新模型的直接支持,调用模型名为 deepseek-v4-flash-vision-exp。作者另提示关注支持文本、图像和视频输入的 Ox Alpha 1M token 上下文,称其在编码和智能体任务上表现突出。

    引用@deepseek_ai@deepseek_ai

    DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n

    推荐理由:官方基准对比显示该实验模型在多模态智能体任务上接近 Opus-4.8,读者可据此了解多模态智能体的当前水平。

  2. @AYi_AInotes72

    DeepSeek 在 API 平台上线实验性多模态模型 deepseek-v4-flash-vision-exp,官方称其文本能力与 DeepSeek-V4-Flash 对齐,多模态 Agent 性能接近 Opus-4.8,在 Agents Last Exam 和 ZeroBench 上局部超过。

    原始视频预览图;未保存可播放视频URL
    引用@deepseek_ai@deepseek_ai

    DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n

    推荐理由:逐条列出这个实验性多模态模型的单图 token 上限、分级定价与免费 Files API 额度,可供评估视觉 Agent 的成本与接入方式。

  3. @kimmonismus67

    DeepSeek 在 API 平台上线实验性多模态模型 DeepSeek-V4-Flash-Vision-Exp,文本能力与 DeepSeek-V4-Flash 持平,多模态智能体基准表现接近 Opus-4.8。官方称该模型在多模态智能体基准上较 V4-Flash 大幅提升,可通过 model=deepseek-v4-flash-vision-exp 调用,DeepSeek Harness 0.1.1 同日发布并原生支持该模型。转发该消息的作者提到,取得这一成绩的是 Flash 系列中的小模型。

    引用@deepseek_ai@deepseek_ai

    DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n

    推荐理由:基准对比显示 Flash 小模型在多模态智能体评测上已接近 Opus-4.8,可供判断轻量模型的能力边界。

  4. @testingcatalog71

    DeepSeek-V4-Flash-Vision-Exp 多模态模型已在 DeepSeek API 平台上线,官方称其性能接近 Opus 4.8,文本能力与 DeepSeek-V4-Flash 持平。该模型支持 Chat Completions、Messages 和 Responses,可混合输入文本与图像,图像可通过 base64、外部 URL 或 Files API 提供,按 V4-Flash 价格计费,每张图片最多 384 tokens。随附基准表显示其在 Terminal Bench 2.1 得分 83.9、ApexBench 得分 36.5。

    引用@deepseek_ai@deepseek_ai

    Multimodal API support 🔌 🔹 Set model='deepseek-v4-flash-vision-exp' 🔹 Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing 🔹 Supports Chat Completions, Messages & Responses 🔹 Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API. Docs: https://t.co/USZ5gZ3wWB 3/n

    推荐理由:材料给出新模型与 Opus 4.8 及自家文本版本的基准对比和计费细节,便于判断实际可用性。

8月19日周三
  1. 硅星人Pro · 微信公众号77

    DeepSeek V4 Pro 与 V4 Flash 涨价生效,下游订阅与 Agent 平台成本承压

    DeepSeek V4 Pro 和 V4 Flash 的新价格于8月17日正式生效,V4 Pro 峰时缓存命中价格从每百万 Token 0.003625 美元提高到 0.044 美元,为原来的 12.14 倍,输出价格从 0.87 美元提高到 3.96 美元;V4 Flash 峰时输出从 0.28 美元提高到 1.32 美元,为原来的 4.71 倍。

    推荐理由:文章用峰谷价差和订阅额度变化,量化这次涨价对下游开发者与Agent平台的实际成本冲击。

  2. 虎嗅APP · 微信公众号77

    宇树科技登陆科创板,成 A 股第一家人形机器人上市公司

    宇树科技于2026年8月19日登陆科创板,发行4044.64万股,发行价150.80元,募集资金总额60.99亿元,发行市值609.93亿元,成为A股市场第一家人形机器人上市公司。

    推荐理由:以上市为节点复盘宇树十年硬件路线,并对比 Figure、Boston Dynamics 等海外公司的不同选择。

8月15日周六
8月14日周五
  1. elsewhere articles75

    DeepSeek 开源 Harness 框架 DSH:这不是产品,而是一个战略

    DeepSeek 于 8 月 13 日发布首个 Harness 版本 DSH,以 MIT 协议开源在 GitHub 上,仓库版本号仍停留在 0.1.0-rc,官方称其为开发者预览版且未来会出现破坏兼容性的变更。

    推荐理由:从插件化架构与开源生态两个角度拆解 DSH,帮助读者理解 DeepSeek 把 Harness 探索交给社区的思路。

  2. 量子位 · 微信公众号78

    深度体验 DeepSeek Harness:开源编码 Agent 的插件化架构与实测对比

    DeepSeek Harness 正式发布并开源,作者内测半个月后把自己的 Vibe Coding 项目从 Codex 迁移到 DSH。DSH 采用一切皆插件的 Cordis 架构,内置 100 多个插件,提供标准、PTC、极简、创造四类 Agent 预设,并有可直接查看原始事件的轨迹回放视图。

    推荐理由:作者用半个月内测体验对比 Codex,呈现 DSH 的插件化架构与轨迹回放等设计,便于判断它与现有编码 Agent 的差异。