

Today we're sharing our work on interaction models. A new class of model trained from scratch to handle real-time interaction natively, instead of gluing it onto a turn-based one. piped.video/A12AVongNN4
关注 AI 研究者、开发者与机构的动态


Today we're sharing our work on interaction models. A new class of model trained from scratch to handle real-time interaction natively, instead of gluing it onto a turn-based one. piped.video/A12AVongNN4
如果这听起来有意思,请与我们联系! openai.com/daybreak/
20 位开发者。8 周中的第 3 周。不再空想,只有交付。 本周三,看谁在追逐自己的第一桶金。 《Race to Revenue》第 3 集 ⠕ 视频
情绪板一直是最棒的部分。现在它只是起点。 上传你的参考图。设定方向。Luma Agents 从那里把情绪板变成成品广告。 把它做成广告 → lumalabs.ai/app 视频
Flow Motion @PixVerse_ Video
Codex 现在可以通过 OpenAI Developers 插件,借助 OpenAI API 帮你更快构建 AI 应用和智能体。 Video
Meet Replit Parallel Agents Build faster by running up to 10 agents in parallel Each agent gets its own copy of your app They work on their own computer Then merge their work agentically Video
Luma Agents 现在可以用 Kling Omni 生成。 更多模型。更广范围。同一工作流。 今天就来试试 → lumalabs.ai/app
Replit 发布 Parallel Agents,可同时并行运行最多 10 个智能体来加快构建速度。每个智能体拥有自己的应用副本,在各自电脑上独立工作,最后以智能体方式合并各自的成果。
下一代模型不只是生成图像——它们将理解世界、运动、交互和动作。我们为此已经构建了一段时间。 视觉智能正在变得实时。@stephenbtl 在 @aiDotEngineer 上谈到了我们的方向: 视频
找到你的下一次逃离。 只需让 Gemini 把你的梦想目的地、喜欢的活动和过去的预订结合起来,为未来的旅行做一次独一无二的头脑风暴。 视频
不用再纠结“到了那儿该玩什么?” Gemini 能根据你过去的聊天记录、偏好以及已连接的应用,为你策划个性化的旅行灵感。 视频
不用再翻收件箱找了。 不用再翻找确认号,Gemini 可以直接从 @Gmail 中找到你的旅行详情,把航班、航站楼和酒店信息汇总到一处。 视频
推荐理由:官方给出跨 Gmail、Photos、搜索与 YouTube 数据生成行程的能力,读者可借此判断个人数据整合的产品边界。
Today we’re launching the OpenAI Deployment Company to help businesses build and deploy AI. It's majority-owned and controlled by OpenAI. It brings together 19 leading investment firms, consultancies, and system integrators to help organizations deploy frontier AI to production for business impact. openai.com/index/openai-laun…
推荐理由:OpenAI 以 $4 billion 和 19 家合作方组建 AI 部署公司,读者可了解前沿模型落地企业的资本与组织方式。
小北部分认同 Karpathy 关于 AI 输出将从 markdown 演进到 HTML、最终走向扩散模型直出交互视频的判断,认为 HTML 在仪表盘、对比和小交互上确实是质变。
This works really well btw, at the end of your query ask your LLM to "structure your response as HTML", then view the generated file in your browser. I've also had some success asking the LLM to present its output as slideshows, etc. More generally, imo audio is the human-preferred input to AIs but vision (images/animations/video) is the preferred output from them. Around a ~third of our brains are a massively parallel processor dedicated to vision, it is the 10-lane superhighway of information into brain. As AI improves, I think we'll see a progression that takes advantage: 1) raw text (hard/effortful to read) 2) markdown (bold, italic, headings, tables, a bit easier on the eyes) <-- current default 3) HTML (still procedural with underlying code, but a lot more flexibility on the graphics, layout, even interactivity) <-- early but forming new good default ...4,5,6,... n) interactive neural videos/simulations Imo the extrapolation (though the technology doesn't exist just yet) ends in some kind of interactive videos generated directly by a diffusion neural net. Many open questions as to how exact/procedural "Software 1.0" artifacts (e.g. interactive simulations) may be woven together with neural artifacts (diffusion grids), but generally something in the direction of the recently viral nitter.net/zan2434/status/2046982… There are also improvements necessary and pending at the input. Audio nor text nor video alone are not enough, e.g. I feel a need to point/gesture to things on the screen, similar to all the things you would do with a person physically next to you and your computer screen. TLDR The input/output mind meld between humans and AIs is ongoing and there is a lot of work to do and significant progress to be made, way before jumping all the way into neuralink-esque BCIs and all that. For what's worth exploring at the current stage, hot tip try ask for HTML.
x.com/i/article/205279610060…
森马将AI用于自身服装业务,带来确收回款几个亿、节省成本几千万。推文借此强调AI要真正用到业务里赚钱,而非自嗨式重复造轮子,并附视频供逐字学习。
忍不住夸冷酸灵的极光感这款牙膏设计。 还是让codex的Chrome自动下单买的,没想到这么好用。 按压挤出,牙膏还可以立起来放洗漱台。 优秀的产品设计,祝大卖。
Anthropic 真的惊为天人 直接把金融服务行业的 AI 工作流模板全开源了 投资银行 / 股票研究 / 私募 / 财富管理 / 基金管理 / KYC 风控 七大业务线的参考 agent / 技能包 / 数据连接器 全部公开 这超出了 demo 的范畴,是把「金融行业 AI 落地」的完整 SOP 摆出来 / 让全行业照抄 · 打开仓库 你会看到这些东西 10 个开箱即用的端到端 agent - Pitch Agent / 自动做 pitch deck(comps + 先例 + LBO → 出品牌排版的 deck) - Meeting Prep Agent / 客户会前自动出 briefing pack Market Researcher / 行业或主题 → 行业概览 + 竞争格局 + peer comps + 标的清单 - Earnings Reviewer / 财报会议 + filings → 模型更新 → 研报草稿 - Model Builder / DCF / LBO / 三表 / comps / 直接在 Excel 里建模 - Valuation Reviewer / 私募估值 + LP 报告 - GL Reconciler / 总账核对 + 找差异 + 路由审批 - Month-End Closer / 月末关账 / accruals / 滚动 / 偏差解读 - Statement Auditor / LP statement 审计 - KYC Screener / KYC 文档解析 + 规则引擎跑 + 标记缺口 每个 agent 都是 self-contained 的,带自己用到的全部 skill,clone 下来直接装 · 7 个垂直行业插件 - financial-analysis(核心):comps / DCF / LBO / 三表 / deck QC / Excel 审计 - investment-banking:CIM / teaser / 流程信 / 买方名单 / 并购模型 / 项目跟踪 - equity-research:财报笔记 / 首次覆盖 / 模型更新 / 投资逻辑跟踪 - private-equity:寻源 / 筛选 / 尽调 / IC memo / 投后监控 - wealth-management:客户复盘 / 财务规划 / 调仓 / 报告 / 税损收割 - fund-admin:GL 核对 / 差异追踪 / NAV 校验 - operations:KYC + 规则引擎 直接用 / 想改也行 / 全是 markdown + json / 没有 build 步骤 · 11 家金融数据商的 MCP 连接器 Daloopa / Morningstar / S&P Global / FactSet / Moody's MT Newswires / Aiera / LSEG / PitchBook / Chronograph / Egnyte 这一行单独值得讲 意味着 Anthropic 已经跟全球最重要的金融数据基础设施都谈完了,你接进 Claude / 这些数据全是开箱即用的,没有这些连接器 / 你自己接每一家的 API / 光人天就要好几个月 LSEG 和 S&P Global 还各自做了 partner-built 的高级插件 一个跑债券相对价值 / 利率曲线 / FX carry / 期权波动率 / 宏观利率监控 一个跑 tear sheet / 财报预览 / 融资 digest · 部署方式两种,一个仓库 方式一:Claude Cowork 插件 装在分析师电脑上 / 个人工作流 方式二:Claude Managed Agents API 跑在公司自己的工作流引擎后面 / 整个公司用 同一份 system prompt / 同一份 skill / 你选在哪儿跑 还附带一个 Microsoft 365 安装工具,让公司 IT admin 把 Claude 部署进 Excel / PowerPoint / Word / Outlook 而且可以走你公司自己的云(Vertex AI / Bedrock / 内部 LLM gateway),不强制走 Anthropic API · 金融是企业 AI 落地最大也最难啃的市场,合规要求高 / 数据敏感 / 流程复杂 Anthropic 这一手 直接把整个行业的 AI 落地 SOP 写明白了 谁照这套搭 / 谁就在 Claude 的轨道上长 谁不照这套搭 / 谁要从零开始踩半年坑 最值得对比的是 OpenAI OpenAI 上周刚上线广告平台 Anthropic 这周直接开源金融行业全套 agent 两家公司的路线分化在这种地方又看得清清楚楚 OpenAI = 大众消费 + 广告 Anthropic = 企业场景 + 开源行业模板 链接 github.com/anthropics/financ…
推荐理由:Anthropic 开源金融行业 agent 模板与数据商 MCP 连接器,可看到企业级落地的整套结构与部署选项。
博主给博客增加了 AI 对话侧边栏,支持随时对话配图、生成标题等操作。例如输入“给第一节配信息图,科普风格”即可生成并自动插入,标题生成后说“选第一个”就能自动替换。功能完善后将同步到开源版本。





We’ve also agreed to acquire Tomoro, which will bring 150 experienced Forward Deployed Engineers and Deployment Specialists to the OpenAI Deployment Company from day one.
推荐理由:OpenAI 用收购直接补入 150 名部署工程师,展示了前沿实验室做企业落地交付的一种组织选择。
The human-perceived RGB is image 1 and the Tesla AI photon count reconstruction is image 2. This is why Tesla FSD can see so well at night or through extreme glare.



Scott Wu is the co-founder of Cognition AI, one of the fastest-growing companies in history. He’s also the greatest competitive programmer the US has ever produced. You may have seen him doing impossible card tricks and mental math. You’ve never seen him asked about weed, Michael Jordan, cancer, and human consciousness over a punnet of strawberries. That is what Colossus editor-in-chief Jeremy Stern did on a recent visit to San Francisco. For those less familiar with @ScottWu46: In 2nd grade, he entered a math competition for 7th graders, lost, and was so furious he still fumes about it 20 years later. The next year he entered the 9th-grade division as a 3rd-grader and got a perfect score. Then he won first place at the US national middle-school math competition and three straight gold medals at the International Olympiad in Informatics, where he became the greatest American gold-medalist and coach in history. Most of the people running the biggest AI companies met as teenagers, competing for their countries on international math and science teams. OpenAI’s Greg Brockman, Anthropic’s Dario Amodei, Meta’s Alexandr Wang, to name just a few. Most agree that the von Neumann among them was Scott Wu. In November 2023, a few weeks after his mother died of lung cancer, on the day Sam Altman was fired from OpenAI, Wu founded his own AI company: Cognition. He was 26 and saw earlier than almost anyone that AI would converge on agents that work in the background, 24/7, like coworkers. He shipped Cognition’s AI software engineer Devin in March 2024. It worked poorly, and he took intense public criticism for it. Now, in its first 18 months of service, Devin has generated $445 million of revenue run rate and usage has doubled every eight weeks. The US Army, Goldman Sachs, and Mercedes-Benz are all customers. Cognition is raising at a valuation around $25 billion. @JeremySternLA sat down with Wu, the emperor of the nerds, to ask the questions we’d all ask one of the smartest people in America—building the most consequential technology of our generation—if we ever got the chance. As well as MJ and weed, they talk about the cluster of competitive math prodigies behind so much of AI, what makes us human when AGI arrives, and why Wu believes he was put on this earth to teach AI how to code. Read the piece below.
Introducing Pareto Code: a new, free, experimental coding router Set `min_coding_score` in your request and route to the cheapest code-capable model that clears your bar, ranked by @ArtificialAnlys. See the Pareto frontier shifting in real time👇
推荐理由:原文给出按编程分数门槛路由到最便宜模型的做法,读者可据此了解低成本编码模型的选型方式。
AI降低内容生产成本 -> 拼选题和审美 -> 拼信任和分发渠道。