跳到正文

全部动态

今日 24 条
9月19日周六
  1. Andrew Milich29

    用 @bot 省钱!

    引用Evan Bacon@Baconbrix

    GROK BOT JUST FOUND $986 IN MY EMAIL 🤯 I asked Grok to get a parking spot for the car it bought me, and on the way it discovered nearly a thousand dollars of extraneous charges from my apartment and emailed them asking for a refund.

  2. Greg Brockman53

    ChatGPT 上线插件多账号连接功能,用户可在插件目录 http://chatgpt.com/plugins 中把工作和个人账号接入同一对话。开发者无需改动即可自动支持,也可在 MCP server 中添加 profile tool 让 ChatGPT 标注不同账号。

    引用Max Stoiber@mxstbr

    Starting today, you can connect multiple accounts with most plugins in ChatGPT! 🎉 Bring context from your work and personal accounts into the same conversation. Connect your accounts in the plugin directory: http://chatgpt.com/plugins Devs: this works automatically, with no changes needed. But, to make the experience for users of your plugin even better, add a profile tool to your MCP server so ChatGPT can label the accounts: https://developers.openai.com/plugins/build/auth#support-multiple-accounts

  3. Noam Brown48

    OpenAI 的 Noam Brown 澄清,他举的“气隙隔离电脑靠温度传感器通信”例子是学术性的,意在说明对隔离做绝对保证极难,因此需要多层防御。他强调该例子讲的是本应完全隔离的智能体之间的协调,而非通过温度传感器窃取模型权重,协调只需极少信息量。他还提到 HF 事件的教训是过度信任沙箱隔离、缺乏独立防护,气隙隔离是极强防护,设计安全协议时宁可高估而非低估 AI。

    引用Fireside Alpha@firesidealpha

    OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because two air-gapped machines can still talk by running a CPU hot and reading the temperature change "But I think the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI. It's a weird world, because AI progress is so fast that people are consistently underestimating the AI." "So to be in a situation where you don't underestimate it again, when it comes to safety and alignment, you have to have a very, very, very high bar." "You could even go as far as to say, "Well, we should air gap the computers." And I'm not convinced that that would be sufficient." "There are studies, and this is mostly academic, where you can have two computers next to each other that are air-gapped and they're still able to communicate with each other because they have temperature sensors." "One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change, and then that actually gives them a mechanism to communicate." _________ Link and more key quotes from OpenAI's safety related conversations: https://firesidealpha.substack.com/p/openai-safety-week-sam-altman-sarah

9月18日周五
  1. GitHub Blog · AI & ML43

    GitHub Podcast 拆解 AI 热门观点:该不该读代码、RAG 已死、Skills 是否杀死 MCP

    GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,只是审查力度应按风险分级;Skills 与 MCP 解决不同问题,前者是打包的专家经验,后者为智能体提供连接工具和数据的标准接口,二者可组合使用。RAG 并未消亡,它让模型获取训练数据之外的信息,减少 token 浪费并降低答案不完整的概率。

  2. 李继刚30

    李继刚提出,"通过学习获得能力、通过能力获得工作、通过工作获得收入与社会认可"这条链并非自然定律,而取决于社会如何组织生产与分配收益。当 AI 改变其中"日用而不知"的前提条件,震动会沿链条向两端传递:向前是教育问题——若成果可借助机器完成,还该学什么、怎样才算学会;向后是意义问题——若社会不再需要我以原来的方式工作,过去努力的东西还算什么。

  3. MiniMax (official)34

    Nunchux AI 与 MIT、CMU、UC Berkeley、斯坦福及 NVIDIA 研究者合作推出 VC-Attention,为 MiniMax-H3 带来免训练低比特注意力加速,在 B200 上比 FlashAttention-4 快 1.6×、B300 上快 1.5×,保真度优于 SageAttention2。

    引用Nunchux AI@NunchuxAI

    Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.

  4. Karina56

    Epoch AI 推出 Benchmark Reviews 计划,对 AI 基准进行审计,首批覆盖 15 个基准,其中 4 个为 Verified、9 个为 Flawed、2 个信息不足暂无法评审。Karina Nguyen 转发并称此举能激励行业打造真正高质量的基准。

    引用Epoch AI@EpochAIResearch

    Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.

  5. OpenRouter Announcements80

    OpenRouter 实测20个图像生成模型的真实成本、编辑与质量差异

    OpenRouter 用同一提示词实测20个图像生成模型的真实计费,1024x1024 单张图价格在 $0.006(openai/gpt-image-2)到 $0.1344(google/gemini-3-pro-image)之间,相差22倍。

    推荐理由:原文用统一提示词实测20个图像模型的实际计费,给出按任务选模型的结论,比照牌价更可靠。

  6. AI at Meta63

    Meta 宣布 Muse for Mac 今日起推送,用户可在电脑上让这一个人智能体直接执行任务,且需用户明确授权。官方示例包括整理下载文件夹、找回丢失的文件、总结消息和笔记,下载地址为 http://ai.meta.com/muse/download/,并表示后续还有更多功能。

    引用Muse@Muse

    Muse is now available on Mac 💻. Your personal agent can get things done for you directly on your computer (all with your explicit permission). - Organize your downloads folder - Find a file you’ve lost track of - Summarize your messages and notes …more coming soon. Try Muse for Mac: http://ai.meta.com/muse/download/

  7. Epoch AI50

    Epoch AI 推出 Benchmark Reviews,用于审计 AI 基准,首批发布 15 个基准的评审结果。其中 4 个 Verified,包括 WeirdML v2、ExploitBench v0.1、PostTrainBench v1.1 和 SimpleQA Verified;9 个 Flawed,包括 Terminal-Bench 4.0.0、SWE-Bench Verified、SWE-Bench Pro、Humanity's Last Exam、DeepSWE v1.1、TextQuests、Lech Mazur Writing、BFCL v4 和 HealthBench Professional;另有 CritPt 和 FrontierCode 因信息不足暂无法评审。

  8. Gary Marcus37

    Gary Marcus:AI 责任与监管并非二选一,科技自由派右翼的虚假二分法

    Gary Marcus 批评科技自由派右翼一边主张 AI 公司应承担损害责任,一边把责任追究当作反对监管的理由,他认为这一推论不成立。他以 2023 年 5 月在美国参议院与参议员 Josh Hawley 的交锋为例,指出现有法律在 AI 出现前制定,版权、大规模虚假信息等领域存在空白,连 Section 230 是否适用都不明确。