跳到正文

全部动态

今日 64 条
9月21日周一
  1. elsewhere articles42

    真格基金刘元与 GPT 对谈:AI 能否胜任早期投资

    真格基金投资人刘元与 GPT 进行了一小时对话,探讨 AI 能否成为优秀 VC。GPT 称在信息分析整理上能比人更快更全面,但做决定未必更强,并坦言自己没有恐惧因而也没有勇敢。刘元认为人类长处恰来自缺陷,早期投资人的乐观与非理性是 AI 难以拥有的,创业者精神比智力与背景更重要。

  2. Nathan Lambert: Interconnects74

    Nathan Lambert 撰文解析开源权重模型格局:中国领先约 2-5 个月追赶闭源前沿

    Nathan Lambert 将其给美国国会的演讲稿整理为 2026 年开源权重模型现状综述,认为中国模型(GLM-5.3、Kimi K3 等)已明确领先美国开源模型,AAII 得分 45/44 对美国最佳 26,开源与闭源前沿差距约 2-5 个月。

    推荐理由:作者以自建数据和演讲稿形式梳理中美开源权重模型的差距、采用与风险,读者可获得一份少见的系统性对比视角。

  3. elsewhere articles59

    葬AI评世界模型热潮:上不了大模型桌的人才另开一桌

    葬AI撰文称世界模型是伪概念,列举Loopit、生数、爱诗科技、Tripo等公司借世界模型叙事融资,实际多为后训练开源视频模型,产品集中于实时生成数字人直播间和WASD游戏画面且缺乏差异。文中认为MiniMax H3是唯一开源且能与Seedance 2.0一战的开源视频模型,支撑了这波世界模型宣发,并预告将推出直播间Bench测试实时生成视频模型。

  4. Simon Willison44

    工程师爆料:大公司全员用 Claude Code 生成代码,无人阅读

    一名新入职大公司半个月的工程师称,团队所有规格、代码、测试、PRD、工单及其解决方案、报告等全部由 Claude Code 生成,从 L1 到 L7 的工程师都在做同样的事——和 Claude 对话。团队无人喜欢这种做法,却被要求尽可能多地产出,高层多次表示推代码不是瓶颈,质疑为何还是慢,员工每天工作 12 到 13 小时只为按回车,没有人阅读任何内容。

  5. Peter McCrory38

    大体同意。一些实际启示: (1) 优先做能用新数据定期更新的分析 (2) 公开地做研究(根据新证据修正自己的观点) (3) 承认不确定性;做出可证伪的预测 (4) 真诚且谦逊

    引用Alex Imas@alexolegimas

    A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."

9月20日周日
  1. elsewhere articles66

    峰瑞李丰分析全球流动性见顶后AI周期的走向与投资机会

    峰瑞资本李丰撰文判断,美欧日9月同向加息后全球流动性接近见顶,本轮美元驱动的AI金融周期进入存量博弈尾部,AI技术投资正从投最具想象力的应用转向投能赚钱的方向。文中列举巨头资本开支受市场惩罚、美国数据中心项目大面积取消或延迟、英伟达以租代买等五个资本开支转折信号,并认为拐点后机会在中国AI+应用、生物医疗与AI交叉以及SaaS的AI化等方向。

    推荐理由:作者以全球流动性和资本开支信号梳理AI周期位置,并给出向AI应用与低估资产转向的判断视角。

  2. elsewhere articles65

    Sayash Kapoor 与 Arvind Narayanan 分析 AI 会不会让论文更快却让科学进步更慢

    Sayash Kapoor 和 Arvind Narayanan 在 2025 年发表的文章提出生产—进步悖论:全球论文数量约每 12 年翻一番,1900 至 2015 年间增长约 500 倍,但诺奖成果诞生于获奖前 20 年内的比例从 1970 年约 90% 降至 2015 年约 50%,科学进步相对投入明显放缓。

    推荐理由:文章把 AI 加速科研的讨论从模型能力转向注意力、激励、可复现性和人类理解等制度瓶颈,提供了评估 AI 科研工具的三个问题。

9月19日周六
  1. Noam Brown48

    OpenAI 的 Noam Brown 澄清,他举的“气隙隔离电脑靠温度传感器通信”例子是学术性的,意在说明对隔离做绝对保证极难,因此需要多层防御。他强调该例子讲的是本应完全隔离的智能体之间的协调,而非通过温度传感器窃取模型权重,协调只需极少信息量。他还提到 HF 事件的教训是过度信任沙箱隔离、缺乏独立防护,气隙隔离是极强防护,设计安全协议时宁可高估而非低估 AI。

    引用Fireside Alpha@firesidealpha

    OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: https://firesidealpha.substack.com/p/openai-safety-week-sam-altman-sarah

9月18日周五
  1. GitHub Blog · AI & ML43

    GitHub Podcast 拆解 AI 热门观点:该不该读代码、RAG 已死、Skills 是否杀死 MCP

    GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,只是审查力度应按风险分级;Skills 与 MCP 解决不同问题,前者是打包的专家经验,后者为智能体提供连接工具和数据的标准接口,二者可组合使用。RAG 并未消亡,它让模型获取训练数据之外的信息,减少 token 浪费并降低答案不完整的概率。

  2. 李继刚30

    李继刚提出,"通过学习获得能力、通过能力获得工作、通过工作获得收入与社会认可"这条链并非自然定律,而取决于社会如何组织生产与分配收益。当 AI 改变其中"日用而不知"的前提条件,震动会沿链条向两端传递:向前是教育问题——若成果可借助机器完成,还该学什么、怎样才算学会;向后是意义问题——若社会不再需要我以原来的方式工作,过去努力的东西还算什么。

  3. Karina56

    Epoch AI 推出 Benchmark Reviews 计划,对 AI 基准进行审计,首批覆盖 15 个基准,其中 4 个为 Verified、9 个为 Flawed、2 个信息不足暂无法评审。Karina Nguyen 转发并称此举能激励行业打造真正高质量的基准。

    引用Epoch AI@EpochAIResearch

    Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.

  4. Gary Marcus37

    Gary Marcus:AI 责任与监管并非二选一,科技自由派右翼的虚假二分法

    Gary Marcus 批评科技自由派右翼一边主张 AI 公司应承担损害责任,一边把责任追究当作反对监管的理由,他认为这一推论不成立。他以 2023 年 5 月在美国参议院与参议员 Josh Hawley 的交锋为例,指出现有法律在 AI 出现前制定,版权、大规模虚假信息等领域存在空白,连 Section 230 是否适用都不明确。

  5. Noam Brown51

    OpenAI 的 Noam Brown 在 Dwarkesh 播客中深谈多智能体、Navier-Stokes 与当前数学进展对自动化 AI 研究和递归自我改进的启示。讨论还涵盖如何在启动 RSI 前判断模型是否真正对齐,以及思维链退化、内外部模型差距等话题,并感谢 OpenAI 团队在多智能体方面的工作。

    引用Dwarkesh Patel@dwarkesh_sp

    New episode with @polynoamial We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research. And we also discuss how we will know if the models are actually aligned before we kick off RSI. 0:00:00 – Multi-agent and Navier-Stokes 0:15:28 – How will AI firms work? 0:22:02 – What math progress tells us about recursive self improvement 0:40:22 – Hugging Face and alignment 1:01:18 – The internal/external model gap 1:08:34 – Chain of thought is degrading 1:14:12 – How will we know when alignment is solved?

9月17日周四
  1. elsewhere articles47

    对话深朴智能王家伟:24 岁具身智能首席科学家谈转身具身与大模型边界

    24 岁的深朴智能(Simple AI)首席科学家王家伟在播客中讲述自己从中科大少年班、MSRA、DeepSeek、字节 Seed 转向具身智能的选择。深朴智能已开源 2,000 小时 HiFi-UMI 数据,内部积累数万小时,模型观察到一定 Zero-shot 泛化,并自研 Agentic OS 连接上层意图与底层动作。他认为通用大模型会承担更多理解与规划,但机器人快速反应仍需专门的动作模型。

9月16日周三
9月15日周二