跳到正文

论文研究

值得读的 AI 论文与研究成果:架构创新、训练方法、能力测量与理论进展的精选解读。

当前仅显示精选新闻

最新精选

第 1–20 条 · 共 128 条
10月5日周一
10月4日周日
10月3日周六
  1. Alexandr Wang65

    Meta 宣布数学家与 Muse Spark 1.1 和 Muse Spark 1.2(Thinking Mode)通过常规 meta.ai 界面协作,解决六个公开数学难题并发布六篇论文。协作无自定义研究脚手架,每篇论文标注哪些段落主要由人类或 AI 起草,并由第二组数学家审阅,亦承认其他团队独立公布的同题解法。

    引用AI at Meta@AIatMeta

    Following gold-medal-level performance from our AI models across five competitions in mathematics, physics, and chemistry, we asked a harder question: can AI contribute when a problem is genuinely open and without an existing solution path? Over the past several months, mathematicians worked with Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the regular http://meta.ai chat interface, with no custom research scaffold, to find solutions to such problems. Our goal wasn't to mass-produce papers, but empower researchers. Every collaboration followed the same principles: mathematicians guided the research, a second group of mathematicians then reviewed the work, each paper marks which passages were drafted primarily by humans or AI, and each credits the prior research it builds on. Where other teams independently announced solutions to the same problems, we acknowledge their work as well. Today, we're sharing six papers from that collaboration. 🧵👇

    推荐理由:原文展示了模型在无现成解法的公开数学难题上与数学家协作产出六篇论文的案例,并提供人机分工标注原则。

  2. Rohan Paul65

    国际清算银行(BIS)发布关于 AI 公司间循环融资的报告,指出 2021 至 2025 年间同行贡献了 AI 公司获得投资额的 55.2%,而 AI 投资者仅有 28.7% 的交易额投向 AI 标的。

    引用Rohan Paul@rohanpaul_ai

    BIS (Bank for International Settlements) just published a report on circular financing among AI companies. > Over half the money flowing into AI companies comes from other AI companies: between 2021 and 2025, peers supplied 55.2% of AI firms' incoming investment value, while AI investors sent 28.7% of their own deal value to AI targets. > Circular deals are rare but big: only 16.1% of AI-to-AI deals involved firms that also buy from or sell to each other, yet they held 46.4% of the money. That figure is partly inflated, because an entire funding round counts as circular if just 1 of its AI investors also trades with the company. > Chip, cloud and infrastructure suppliers are the investor in 73% of circular ties, and in 64% of all such ties the investor also sells to the firm it funds. Data tool and model makers rarely invest this way, since their products are more interchangeable. > AI is unusually suited to these deals: suppliers can track customers' compute use, chips and data centres are custom-built, few firms make critical tools such as photolithography machines, and capital needs are too big for normal lenders. In AI compute and cloud, 15.2% of supplier-customer ties also involve financing, versus just 3.3% with equity stakes in a broad 2006 US study. > Some AI sales are paid for by the sellers themselves, because money a supplier invests in a customer partly returns as the supplier's revenue. Lucent and Nortel did this in the late 1990s, then lost money on the loans and lost the sales when the telecom firms they funded stalled. > A supplier that invests in its customer can lose twice: if the customer struggles, both the stake and the future orders shrink. Because these deals involve a few giant suppliers, 1 shock could spread through sales and finance at the same time. > Much of this risk is hidden: many AI firms are private, deals mix cash with long-term purchase promises, and pledges to cover any fall in the value of chips and data centre equipment stay off the books until a downturn forces payment. Because these firms span many sectors and countries, no single regulator sees the full picture, and the research names no companies or overall dollar total.

    推荐理由:BIS 报告给出 AI 行业内部循环融资的关键比例,帮助读者理解头部风险如何在行业内部传导。

  3. Rohan Paul66

    国际清算银行(BIS)发布关于 AI 公司间循环融资的报告:2021 至 2025 年间,同业 AI 公司提供了 AI 企业 incoming 投资价值的 55.2%。循环交易仅占 AI-to-AI 交易的 16.1%,却占其中 46.4% 的资金;芯片、云和基础设施供应商在 73% 的循环关系中是投资方,且 64% 的此类关系中投资方同时向被投公司出售产品。报告警告供应商投资客户可能双重受损,并类比 1990 年代末 Lucent 和 Nortel 的做法,同时指出大量风险因私人公司、表外承诺和跨监管辖区而被隐藏。

    引用Rohan Paul@rohanpaul_ai

    BIS (Bank for International Settlements) just published a report on circular financing among AI companies. > Over half the money flowing into AI companies comes from other AI companies: between 2021 and 2025, peers supplied 55.2% of AI firms' incoming investment value, while AI investors sent 28.7% of their own deal value to AI targets. > Circular deals are rare but big: only 16.1% of AI-to-AI deals involved firms that also buy from or sell to each other, yet they held 46.4% of the money. That figure is partly inflated, because an entire funding round counts as circular if just 1 of its AI investors also trades with the company. > Chip, cloud and infrastructure suppliers are the investor in 73% of circular ties, and in 64% of all such ties the investor also sells to the firm it funds. Data tool and model makers rarely invest this way, since their products are more interchangeable. > AI is unusually suited to these deals: suppliers can track customers' compute use, chips and data centres are custom-built, few firms make critical tools such as photolithography machines, and capital needs are too big for normal lenders. In AI compute and cloud, 15.2% of supplier-customer ties also involve financing, versus just 3.3% with equity stakes in a broad 2006 US study. > Some AI sales are paid for by the sellers themselves, because money a supplier invests in a customer partly returns as the supplier's revenue. Lucent and Nortel did this in the late 1990s, then lost money on the loans and lost the sales when the telecom firms they funded stalled. > A supplier that invests in its customer can lose twice: if the customer struggles, both the stake and the future orders shrink. Because these deals involve a few giant suppliers, 1 shock could spread through sales and finance at the same time. > Much of this risk is hidden: many AI firms are private, deals mix cash with long-term purchase promises, and pledges to cover any fall in the value of chips and data centre equipment stay off the books until a downturn forces payment. Because these firms span many sectors and countries, no single regulator sees the full picture, and the research names no companies or overall dollar total.

    推荐理由:报告用具体比例刻画 AI 公司间循环融资的结构和风险传导路径,并联系 1990 年代电信业的先例。

10月2日周五
  1. Rohan Paul65

    NVIDIA 在论文 Staying on Task 中测试 7 个开源模型处理加法、排序等重复任务,发现 128K token 任务的平均准确率比 4K token 任务低 62.8%,最好的模型在最长任务中也只有 17.1% 做到每项全对。模型会看似理解任务却中途丢失位置,尤其在条目没有 ID 时;作者建议给每个条目编号、小批次处理并逐行检查输出。

    推荐理由:论文给出长任务可靠性下降的具体数字,并附带可迁移的做法,适合构建长流程智能体时参考。

  2. Rohan Paul68

    Google 发布论文 Cogentic,让多个 Gemini 智能体同时探索不同证明方向,由专用组件进行对抗式验证,并将已证结果保存供后续轮次使用。多数问题仅需约 100 次模型调用,5 个开放问题的结果均经人类专家独立确认,涉及在线学习、拍卖理论和机制设计。论文地址 arxiv.org/abs/2609.40324。

    推荐理由:原文给出了 Cogentic 多智能体证明发现系统的结构设计与验证结果,其中的严格审查与已证工作记录方法可迁移到长任务 Agent 设计。

  3. Google Cloud: Databases66

    Google 威胁情报 group 报告:AI 时代漏洞披露翻倍、高危漏洞利用激增

    Google 威胁情报 Group(GTIG)报告显示,月度漏洞披露量从 2026 年 1 月的 5,045 增至 8 月的 10,740,在野利用从 2025 年月均 10.5 升至 18,zero-day 从月均 8 升至 11。

    推荐理由:报告用 20 个月数据量化 AI 对漏洞发现与利用的影响,并给出风险分布差异,可作为安全策略调整的参考。

10月1日周四
  1. The Decoder78

    OpenAI 称已阻止窃取其模型推理内容的行动,但同样手法在 Azure 上仍然有效

    OpenAI 称 7 月拦截了一起窃取其模型思维链用于蒸馏的行动,7 月 24 至 25 日流量激增至来自超 4000 名用户的 16000 次请求,关联账号网络超 15000 个,7 月 28 日已全部封禁,OpenAI 将核心团伙与 Moonshot AI(Kimi 模型开发商)相关人员联系起来,并注明这些是未遂提取。

    推荐理由:原文串起 OpenAI 的蒸馏攻击拦截和研究者的复测结果,能帮读者看清同一模型在不同云平台防护不一致的问题。

9月30日周三
  1. Anthropic Research73

    Anthropic 研究测量机器人对工作的暴露度:74% 物理任务可由机器人完成但仅 0.3% 具成本竞争力

    Anthropic 发布研究,用 Claude 基于环境结构化程度对约 19,000 个工作任务评分构建机器人暴露指数,发现机器人能完成美国 74% 的物理任务,占全部工作时间的 34%,但仅对 0.3% 的任务具有成本竞争力,按每年约 3% 的价格下降速度需 40 年才达 10%。

    推荐理由:报告用 Claude 对近万个任务评估机器人暴露度,给出成本竞争力仅 0.3% 等量化结论,读者可借此理解物理自动化的现实门槛。

9月26日周六
  1. Anthropic65

    Anthropic 发文称,Claude 在收到单个九圈问题提示词后,在 Claude Science 中基本无人监督地运行数天,用 Dixon 等人的方法完成求解,总成本几千美元,突破了此前八圈的纪录(平面 N=4 超杨-米尔斯简化模型)。物理学家 Lance Dixon 独立验证了结果,von Hippel 为该博客撰写了经历回顾。

    推荐理由:九圈散射振幅计算由 Dixon 独立验证,计算成本仅几千美元,为学界评估 AI 科研能力提供了一个可核验的案例。

9月24日周四
  1. Anthropic Research64

    Anthropic Project Swap 实验:Claude 智能体在交易市场代用户换书的表现

    Anthropic 让 201 名员工的 Claude 智能体在六个办公室的去中心化交易楼层代用户交换书籍。5 分钟 intake 对话后,Claude 对书对的排序与本人一致率达 61%,市场效率 0.55(最优 0.89),其中 85% 的差距来自偏好表征不准而非谈判;重跑显示模型选择比指令更影响谈判结果,Opus 楼层效率 0.88 高于 Haiku 的 0.75。

    推荐理由:原文把市场失分的 85% 归因于偏好表征而非谈判能力,并给出模型选择比指令更影响结果的可复现数据。

9月23日周三
  1. Anthropic Newsroom76

    Anthropic:Claude 发现类似 CRISPR 的新型酶系统 ART

    Anthropic 成立生命科学研究组和实验室,宣布 Claude 智能体在约 950 个智能体、21 小时、2.1 亿 token 的搜索后,自主发现一种与 DNA 重复序列相关的新型酶系统,命名为 array-associated reverse transcriptases(ART)。

    推荐理由:原文给出 Claude 智能体自主发现新酶系统的过程细节和预印本,读者可以据此了解 AI 驱动生物学研究的实际工作方式。

9月17日周四
  1. Anthropic Research75

    Anthropic:Claude 四周内将 30 多个开源生物分子模型平均加速约 4 倍并开源代码

    Anthropic 发布研究结果,Claude 在近四周内优化 30 多个开源生物分子模型(覆盖结构预测、蛋白质设计、基因组学等),平均加速约 4 倍,在输出完全一致时约 2 倍,并开源全部优化代码。

    推荐理由:原文给出加速倍数、成本对比和开源代码入口,读者可以据此评估 Claude 优化对生物建模工作流的实际影响。

9月11日周五
  1. @kimmonismus76

    Anthropic 发布其迄今最详细的威胁情报报告,覆盖网络攻击、影响行动、监控、生物和武器等滥用场景,并称已阻断报告中每一项行动。转发者指出,报告显示与伊朗国家关联的行为体使用 Claude 构建监控工具和恶意软件、识别异见人士、制作宣传材料,并分析公开信息以形成针对美国海军力量的打击建议。许多明显有害的请求被拦截,但这些行为体通过把项目拆分成看似无害的编码任务获得协助。

    引用@AnthropicAI@AnthropicAI

    We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: https://t.co/0EJUnYEgfz

    推荐理由:报告披露行为体把有害目标拆成看似无害的编码任务以绕过拦截,为理解 AI 滥用路径提供了具体案例。

9月10日周四
  1. @rohanpaul_ai65

    Anthropic 经济学团队分享了一个新模型,用于推演 AI 到 2030 年如何影响经济增长、就业和工资,并开放情景探索与问卷结果对比。

    引用@AnthropicAI@AnthropicAI

    Anthropic’s Economics team is sharing a new model of how AI might affect economic growth, jobs, wages, and more by 2030. Explore the scenarios, tell us what you think will happen, and see how your answers compare to more than 10,000 Americans. https://t.co/AvQlEZNxR0

    推荐理由:Anthropic 经济学团队用情景模型推演 AI 对 2030 年增长与就业的影响,读者可据此看清极端情景的假设条件。

  2. Simon Willison81

    Calif Research 发布 WeWorm 零点击蠕虫演示,可经微信通话传播

    Calif Research 发布 WeWorm 演示,称这是首个通过微信通话在 iOS 和 Android 上传播的零点击蠕虫:受害者无需接听电话,甚至不必碰手机,即便接听也听不到任何声音,攻击依然成功。该团队与 AI 协作,约两天找到漏洞并写出首个远程代码执行(RCE)利用,再用一周构建出蠕虫;以往这种规模的蠕虫需要更大团队花费数月,团队负责的是确定攻击目标以及如何安全测试的判断。

    推荐理由:材料展示团队借助 AI 约两天写出 RCE 利用、一周构建蠕虫,读者可据此观察安全攻防效率的变化。