跳到正文

#大佬观点

今日 32 条
今天10月2日周五
  1. Rohan Paul35

    黑石集团总裁兼首席运营官乔恩·格雷几个月前也表达了同样的观点。 任何基于规则的业务,如会计、法律、金融,都将被AI彻底颠覆。🎯 例如:摩根大通在股东投票中弃用代理顾问,改用AI替代。

    引用Rohan Paul@rohanpaul_ai

    Rule-based work isn't a career anymore. It's a prompt. In every rule-based profession, humans are now the slow, expensive, error-prone option. Claude Opus 5 got 20 for 20 at 100% and wrapped each task in minutes, while 12 licensed CPAs landed anywhere from 0% to about 90%, with several running out the 3-hour clock.

  2. Rohan Paul42

    AI 可能会自动化初级律师、顾问、银行家和广告从业者目前从事的大量日常工作。 “金字塔底层的人在做的那部分工作——会变成什么样?你唯一能说的是,如果你从未在律所或贝恩、BCG、麦肯锡工作过,你大概不会清楚这到底是怎么回事。” ——班尼迪克·埃文斯(科技分析师、作家、前 a16z 合伙人) ---- 来自 a16z YouTube 频道(链接见评论)

    引用Rohan Paul@rohanpaul_ai

    Blackstone President and COO Jon Gray made this same point few months back. Any rule-based businesses, like accounting, legal, finance, will be completely disrupted by AI. 🎯 e.g. JPMorgan dropped proxy advisors for shareholder votes, replacing them with AI. https://x.com/BloombergTV/status/2016932349737410876/video/1

  3. Thomas Wolf45

    Kevin Buzzard(IMO 满分、数论学家、Lean 形式化数学先驱)写了一篇非常深刻的文章。 如果数学不只是关于“人类理解”,那它又关乎什么? 如果 AI 能力持续指数级增长,而“数学是无限的”,那会发生什么?

    引用Bartosz Naskręcki@nasqret

    I cannot agree more. Kevin Buzzard made so many points I agree with. But the best one is this "I thus believe that in the future we will reach a new “natural boundary” in mathematics, beyond (and perhaps way beyond) where we are now, but where machines are going to get stuck and where it is not viable to expend any more resources to make the next big leap. (...) I believe that the optimal thing to do (...) is to let the machines loose, see what happens, and then begin the journey to where they have stopped." https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-not-to-grieve/

  4. Elon Musk38

    超级智能(前身为 AI)如今在会计测试中表现出色。

    引用Andrew Curran@AndrewCurran_

    'More striking is how fast AI took the lead. Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score. Today, models ace those same tasks.' 'These results are provocative. So much so that we considered not publishing them for fear of misinterpretation. But we think transparency about the findings matters as people and institutions prepare for rapidly advancing AI.'

  5. AYi41

    Meta 首席 AI 官、Scale AI 创始人 Alexandr Wang 首次公开分享创业最绝望的心理死穴,称 YC 创业淘汰比《饥饿游戏》残酷一万倍,90% 的公司不会立刻死掉,而是苟延残喘好几年。他给出两条生存法则:靠第一性原理倒推终局确定性,以及把所有不可控的恐惧置换成高频动作,用带着恐慌的机械执行填满时间。

    引用AYi@AYi_AInotes

    如何从零想出一个估值百亿的创业点子? Scale AI 创始人,现在是Meta首席AI官,muse负责人 的 Alexandr Wang 给出了一条极简铁律:活在未来,倒推今天还不存在的那行 API。 从深夜抢注域名,到肉身坐在客服气泡后死磕每一个访客,这段 4 分钟的复盘,讲透了科技商业里最硬核的起步真相。 很多人可能不知道,现在估值接近 140 亿美元的 AI 数据霸主 Scale AI,在刚起步的前半年,创始人每天也在经历极度严重的精神内耗。 这是 Scale AI 创始人, Alexandr Wang 在 SPC 闭门访谈里,第一次毫无保留地复盘自己在 YC 期间最痛苦的负一阶段。 一句话概括这段分享最值钱的本质: 所有伟大企业的起点,并不是算无遗策的天才顿悟,而是在漫长的游荡期里,靠着第一性原理把脏活做透,硬生生把一个看似不起眼的点子熬成了超级基础设施。 现在几乎所有想做点事、想做个人项目或创业的人,都在经历同一种心理折磨: 打开文档写满了各种点子,却总觉得每一个都不够好; 看着身边的人都在飞速推进,总觉得自己从第一天起就落后了别人半年; 每天在强烈的存在焦虑里打转,不知道自己到底在折腾什么。 Alexandr Wang 当年也是一模一样的处境。 我把他在视频里拆解出的三个底层认知,整理成最干货的复盘讲透👇 ① 选方向的第一性原则:活在未来,倒推缺失的那行 API 当年他在 YC 每天写点子文档,直到读了 Paul Graham 的那篇经典文章:Live in the future, and build what's missing. 他当时看到了一个未来的必然趋势: 未来的人类算力(Human Compute)一定会像计算机算力一样,被极度动态地编排和调用。 但当时整个互联网上,根本没有一个能够像调用服务器一样直接调用人工标注与处理的 API。 于是他花了一整晚买下 ScaleAPI 这个域名,这就是百亿帝国的最初原点。 ② 拆穿创业最大的心理陷阱:起步即落后的虚妄焦虑 在负一阶段,最致命的不是没点子,而是同行压力带来的动作变形。 Alexandr 提到:当你刚萌生一个新点子时,环顾四周,总觉得别人已经跑了很久,自己一开局就落后了。 但事实是,绝大多数人都在各自的迷雾里摸索。 真正的差距从来不是谁先动手两星期,而是谁能在漫长的游荡期里顶住内耗,把方向压力测试到底。 ③ 穿越死亡谷的唯一解法:做无法规模化的笨活与脏活 在 Product Hunt 上线拿到第一波热度后,Scale 经历了整整 4 到 6 个月的空白游荡期。 当时没有爆发式增长,能不能成完全是未知数。 Alexandr 采取的最硬核策略只有一个:当客服。 他在官网挂了 Intercom 聊天气泡,每一个点进网页、发消息咨询的真实访客,背后亲自敲键盘回复的人就是他自己。 正是靠着跟每一个早期客户在泥潭里死磕,直到半年后才终于等来了第一个真正想做大的核心客户。 历史与商业演进的硬核印证: → 硅谷最经典的创业定律: 保罗·格雷厄姆提倡的 Do things that don't scale(做无法规模化的事),在 Scale AI 身上得到了最彻底的验证。世界上最顶级的自动化数据管道,最开始也是创始人靠肉身当客服一点点抠出来的。 → 游荡期是所有顶级公司的必修课: 从 Airbnb 早期靠卖麦片还信用卡债,到 Stripe 创始人亲自跑去客户电脑上敲命令行装插件,没有一家基础设施级巨头能跳过这至少半年的迷茫摸索。 站在另一个更理性的视角来看,这件事给普通人的启发极其锋利: → 不要把摸索期的焦虑误判为失败: 从负一阶段到零的这段时间,内心动荡和怀疑是系统的标配属性,而不是你能力不足的证明。 → 别在战术的勤奋里逃避真正的思考: 想点子不是在文档里盲目堆数量,而是敢于逼问自己:未来五年哪件事一定会发生,而今天还缺了关键工具? → 离真实用户再近一点: 当你不知道下一步该做什么时,去跟每一个点了聊天气泡的真实访客聊半个小时,远比关在屋子里改一百遍商业计划书管用得多。 最后收个尾: 世上从来没有一开局就清晰无比的百亿蓝图。 伟大往往就藏在那份写满废案的文档里,藏在深夜无人问津的客服窗口背后。 熬过负一阶段的迷茫,把未来的缺失变成今天的行动,你才算真正站在了起跑线上。

  6. Dongxi 东锡 NLP67

    Karpathy 发文认为人们将花更多时间理解语言模型的输出,建议让 LLM 用 ASD-STE100 受控语言写作、生成图表、输出 HTML 交互网页,以及用 ElevenLabs 配音生成定制讲解视频。引用者回忆当年求教复杂代码被工程师一句“哦,忘了”回绝,感慨如今 LLMs 能以文字、图表、视频耐心解答问题。

    引用Andrej Karpathy@karpathy

    We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.

    推荐理由:作者借个人经历引出 Karpathy 关于用受控语言、图表、网页和视频理解模型输出的建议,可当作换个方式向 LLM 提问的参考。

  7. Yuchen Jin67

    Yuchen Jin 转引 Andrej Karpathy 关于理解语言模型输出的建议,并表示希望 AI 能直接生成一段 Karpathy 风格的视频,但如今没有 AI 能做到。Karpathy 在引用内容中提出几种输出形式,包括让 LLM 用航空维护文档的受控语言规范 ASD-STE100 解释概念、生成图表和交互式 HTML 网页,以及用 ElevenLabs API key 或本地免费方案生成 3b1b 风格的讲解视频;他认为 LLM 会承担更多工作,人类的工作将上升为监督和理解。Yuchen Jin 还提到 Karpathy 已超过一年没有在 YouTube 上传视频。

    引用Andrej Karpathy@karpathy

    We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.

  8. Suno18

    无论你做什么,只要把你音乐里那种合成的难听声音去掉就行。如果 Suno 保持现在的质量,它走不远……它需要听起来精致、经过母带处理,而不是像在用 Circuit City 倒闭前你能买到的最便宜音箱播放一样。

    引用𝐌𝐫. 𝐖𝐢𝐜𝐤 🇺🇸 🦍@SoonMrWick

    Whatever you do just get rid of that synthetic nasty sound in your music. Suno isnt going anywhere if it remains at the current quality... it needs to sound polished, mastered, not like its playing through the cheapest speakers you could buy at circuit city before they went out of business.

  9. Dongxi 东锡 NLP47

    Tavus 推出 Griffin,号称首个通过视频图灵测试的模型,48% 的实时对话者认为它是真人,此前系统通过率不足 3%,并在 NVIDIA 全双工 AI 视频基准上排名第一。它是首个 Human Interaction Model(HIM)。推文作者借此调侃:用 Agent 干活、Griffin 开会甚至面试,就能同时接成百上千个远程职位。

    引用Tavus@tavus

    Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).

  10. Alexandr Wang25

    muse:"真的感觉像活在未来" ⚡️

    引用Comfortably Smug@ComfortablySmug

    I am absolutely stunned by how good @Muse is, it genuinely feels like living in the future. Pretty clear why Meta stock is up 10% just this week. I think Muse has surprised a lot of people. Hats off @alexandr_wang, your team has shocked the world

  11. Peter Steinberger 🦞23

    我有太多疑问了

    引用yingchao@baggiiiie

    @GergelyOrosz at least we know it thinks it's part of the gpt family and somehow coderabbit once 😆

  12. Elon Musk48

    试试 Grok @Bot!

    引用Beff (e/acc)@beffjezos

    Grok Bots have been life-changing for someone like me with ADHD who has no patience for context switching / navigating slow interfaces to retrieve information We're seeing the beginnings of personal superintelligence that augments each humans to realize their full potential

  13. Chubby♨️27

    今日最佳消息:@GoogleDeepMind 的一位 Senior Staff Research Engineer 称 Bloomberg 的报道是“胡说八道”。 Bloomberg 声称,虽然 Gemini 4 “在广泛用于衡量模型效能的基准测试中表现良好,但当员工真正将其投入工作时,表现就不那么好了。” 所以是的,这让我们更有理由期待一次出色的发布。

    引用Aditya Timmaraju@tadityasrinivas

    This @Bloomberg story is BS. Argon has been my daily driver for a while and it's been a great experience. It's particularly awesome at agentic debugging besides day-to-day coding tasks.

  14. AI as Normal Technology60

    Arvind Narayanan 论 AI 安全运动应选大帐篷还是小帐篷

    Arvind Narayanan 提出 AI 安全存在两种叙事之外的第三种可能:x-risk 警告是真诚但错误的,且对安全政策适得其反。文章以疫情防范不足和网络安全系统性风险为例,论证世界确实对 AI 放大的灾难性和累积性风险投入不足,但 x-risk 框架会加剧党派极化并把资源导向不可行的禁令式政策。作者主张建设包容具体风险防御、社会韧性与透明度、问责政策的大帐篷安全运动。

10月1日周四
  1. Dwarkesh Patel46

    Si Sheppard 谈几百名西班牙士兵如何推翻阿兹特克与印加两大帝国

    军事历史学家 Si Sheppard 在播客中讲述 Cortés 与 Pizarro 如何以数百名征服者推翻阿兹特克和印加帝国:Cortés 在两年半内征服 600 万人口的阿兹特克,Pizarro 十年后征服约 1000 万人口的印加。征服者常在 100 倍甚至 1000 倍兵力劣势下取胜,且多达 99% 的兵力由被征服的原住民组成。

  2. AYi58

    Scale AI 创始人、现任 Meta 首席 AI 官的 Alexandr Wang 在 SPC 闭门访谈中复盘创业负一阶段:他受 Paul Graham 启发预判人类算力会被动态调用,注册 ScaleAPI 域名起步。

    引用AYi@AYi_AInotes

    整个科技圈风向都在卷大企业 AI,
结果 @Meta 的 @AIatMeta AI 掌门人 @alexandr_wang 刚刚发推说他们后台发现,把 Muse 用得最猛的,
竟然是修水管的、开农场的、开杂货店和开小餐馆的老板…… 上线两周,早期用户里竟然有 1/3 直接把它绑定了商业账号。 今天 Meta 顺水推舟,直接上线了 Muse for Small Business,
一口气甩出 22 个官方连接器:
Shopify、Stripe、QuickBooks、Canva、Zoom、Slack……
把小商家的全套数字牛马直接做实了。 我把背后的逻辑和搞钱启示用大白话给大家拆一下👇 ━━ 1|最反直觉的现象:用 AI 最狠的不是大厂,是个体户 大厂做 AI 喜欢去找世界 500 强做大单子,
但真实世界里,最缺时间的是那些“一人身兼八职”的小老板。 官方案例里有个爱荷华州开杂货店的大叔:
每周工作 65 个小时,店里进货是他、管账是他、做促销是他、处理退款还是他。
他缺的根本不是点子,是时间。 以前大家以为 Muse 只是个帮宅男比价订外卖的玩具,
结果这些小老板早就拿它去回客户邮件、查库存、算账了。 2|这 22 个连接器,到底解决了啥? 做过独立电商、自媒体或者小生意的朋友都懂这个痛点:
每天要在 8 个后台之间来回切——
看销量去 Shopify,查退款去 Stripe,做海报去 Canva,记账去 QuickBooks,开会去 Zoom,内部沟通去 Slack。 现在 Muse 把它全收拢到一个对话框里:
▫️查账退款:不用登录 Stripe 后台,直接对它说“把昨天争议的那笔单子退了”,它核对完调接口执行
▫️内容与营销:识别你的产品库存和品牌调性,直接调用 Canva 生成下周促销海报,一键推到 Instagram 商业主页
▫️客户与日程:给它配个专属邮箱,客户发来的询价、改期邮件,它在后台看懂日历直接起草回复 一句话:它不是让你去学一个新软件,而是把你手上所有的 SaaS 变成听话的后台。 3|Meta 这步棋最狠的地方在哪? Facebook 和 Instagram 上有整整 2 亿小微商家。
这是全球最大的个体老板聚集地。 Meta 根本不需要跟微软、谷歌去抢那些复杂的企业级大单,
它只要让这 2 亿小老板在手机上:
“少雇一个客服、少花 2 小时对账、少切 5 个软件”。 而且最聪明的是它的安全机制:
所有涉及扣费、退款、对外公开发布的操作,AI 负责跑腿起草,最后一步依然要人点确认。
这就是我们常说的“人管方向和钱包,AI 管繁琐和流程”。 4|对我们普通人和超级个体意味着什么? 一人公司(One-Person Company)过去最大的瓶颈,不是你不会做核心业务,而是被杂事活活拖死。
你懂做视频,但你不想花时间回商务邮件;
你懂写代码,但你讨厌天天去处理发票和对账。 当 AI Agent 开始长出连接现实商业系统的手(Connectors):
以前需要 3 个人支撑的微型工作室,
以后可能真的 1 个人加一个配置好的 Personal AI 就能跑起来。 5|照例泼盆冷水 ▫️这套连接器目前深度绑定的是海外生态(Shopify、Stripe、QuickBooks 等),国内主流的微店、淘宝、微信支付等生态目前还没打通;
▫️多平台授权意味着你的商业数据、客户往来都在被 AI 扫描,哪些权限给、哪些不给,依然得有边界;
▫️复杂业务逻辑的容错率低,账目核算依然需要人工定期复核,别做甩手掌柜。 ━━ 过去我们讨论 AI,总觉得它是高高在上的算力和算法。 
但今天 Alexandr Wang 这条推说明了一件事:
AI 落地最快、最扎实的地方,
永远是帮街角那家店的老板,把今晚下班的时间提前两个小时。 开源了自定义接入平台(http://muse.ai/platform)。
如果国内也有一个 AI 能接通你所有的工作软件,你最想让它替你干掉哪个日常琐事? https://x.com/alexandr_wang/status/2104925780547399986/video/1

  3. Alexandr Wang30

    强者识强者 @signulll:Muse 的热度是真的。我可以说它是市面上最易用、最聚焦的消费级 AI 产品。极其慷慨的使用限制也让人无法忽视。 它确实能帮普通人应对日常生活中的复杂事务。 如果你回看我四月在 TBPN 的节目,我谈到过 AI 生成的信息流会成为新 AI 体验的基础原语,而 Muse 已经有了一个相当有说服力的版本。你也能开始想象多人层会变得多强大。 我现在的组合是工作用 Claude,生活用 Muse。Facebook 这次执行得很棒。该表扬就表扬。

    引用signüll@signulll

    the muse hype is real. i can def say that it is the easiest to use & most focused consumer ai product on the market. the extremely generous limits also make it impossible to ignore. it actually helps the avg person navigate the day to day complexity of personal life. if you go back to my april tbpn appearance, i talked about how an ai generated feed would become a baseline primitive for new ai experiences & muse has a pretty compelling version of that idea already. & you can start to imagine how powerful the multiplayer layer could become. my stack right now is claude for work & muse for life. great execution from facebook here. credit where credit is due.