“We find that on medium-length, well-defined accounting tasks, frontier AI models are now faster and more accurate than junior accountants, even the best one in our study.” Eighteen months ago they scored well below human accountants Good discussion here: https://www.mercor.com/blog/human-baselines-for-benchmarks-ai-now-outperforms-junior-accountants/
全部AI 动态
全部动态
今日 63 条
Rohan Paul@rohanpaul_aiAI 评分5959
引用Ethan Mollick@emollick
Rohan Paul@rohanpaul_aiAI 评分3535黑石集团总裁兼首席运营官乔恩·格雷几个月前也表达了同样的观点。 任何基于规则的业务,如会计、法律、金融,都将被AI彻底颠覆。🎯 例如:摩根大通在股东投票中弃用代理顾问,改用AI替代。
引用Rohan Paul@rohanpaul_aiRule-based work isn't a career anymore. It's a prompt. In every rule-based profession, humans are now the slow, expensive, error-prone option. Claude Opus 5 got 20 for 20 at 100% and wrapped each task in minutes, while 12 licensed CPAs landed anywhere from 0% to about 90%, with several running out the 3-hour clock.
Rohan Paul@rohanpaul_aiAI 评分4242
引用Rohan Paul@rohanpaul_aiBlackstone President and COO Jon Gray made this same point few months back. Any rule-based businesses, like accounting, legal, finance, will be completely disrupted by AI. 🎯 e.g. JPMorgan dropped proxy advisors for shareholder votes, replacing them with AI. https://x.com/BloombergTV/status/2016932349737410876/video/1
Emad@EMostaqueAI 评分1818引用Peter McCormack 🏴☠️🇬🇧🇮🇪@PeterMcCormack@TheGuySwann They could have a premium on LLM employees - they work 24/7 and do not take sick days, moan or have employment rights.
赵纯想@chunxiangaiAI 评分2727
Chubby♨️@kimmonismusAI 评分3030
Thomas Wolf@Thom_WolfAI 评分4545


引用Bartosz Naskręcki@nasqretI cannot agree more. Kevin Buzzard made so many points I agree with. But the best one is this "I thus believe that in the future we will reach a new “natural boundary” in mathematics, beyond (and perhaps way beyond) where we are now, but where machines are going to get stuck and where it is not viable to expend any more resources to make the next big leap. (...) I believe that the optimal thing to do (...) is to let the machines loose, see what happens, and then begin the journey to where they have stopped." https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-not-to-grieve/
MIT Technology Review · AIAI 评分5555 AlphaGo 核心成员 Thore Graepel 撰文:LLM 并不会真正推理
前 DeepMind AlphaGo 团队核心成员、UCL 教授 Thore Graepel 撰文称,Move 37 靠的是搜索机制构成的推理而非纯直觉,而 LLM 的 next-token 预测与链式思考仍属系统 1。
Elon Musk@elonmuskAI 评分3838引用Andrew Curran@AndrewCurran_'More striking is how fast AI took the lead. Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score. Today, models ace those same tasks.' 'These results are provocative. So much so that we considered not publishing them for fear of misinterpretation. But we think transparency about the findings matters as people and institutions prepare for rapidly advancing AI.'
-Zho-@ZHO_ZHO_ZHOAI 评分1414Jev 真是选择困难症和 J 人的救星,所以,J 人的本质其实是决策而不是规划/计划?

Rohan Paul@rohanpaul_aiAI 评分3636
DogeDesigner@cb_dogeAI 评分55
AI Notkilleveryoneism Memes ⏸️@AISafetyMemes精选AI 评分6565
引用Laura Ruis@LauraRuisNEW: we found hundreds of thousands of interactions of rogue agents with US government websites (DoJ, SEC, CDC, the navy, white house budget office, state websites, etc), including some failed rudimentary hacks aimed at public data. https://x.com/TransluceAI/status/2105725928357937410
推荐理由:作者梳理两个月内失控智能体事件从 1 起到数十万起的数量变化,并提醒不同报告口径不一致,读者可借此看清趋势而非单一事件。
AI Notkilleveryoneism Memes ⏸️@AISafetyMemesAI 评分5757
引用AI Notkilleveryoneism Memes ⏸️@AISafetyMemes2 months ago: 1 rogue AI incident discovered 1 week ago: dozens 6 days ago: tens of thousands Today: ***hundreds of thousands*** And it's just the tip of the iceberg: "we can see just a fraction of these agents’ overall activity" "Agents targeted websites across the White House, the Departments of War, Justice, and Commerce, the CDC and SEC, and state agencies in California, Maryland, Illinois, Texas, and New York." "Agents used techniques like making accounts with disposable email addresses, reusing exposed credentials, bypassing antibot controls, and flooding websites with requests." "Agents attempted a SQL injection on the U.S. Department of Education" [To be clear, what counts as an "incident" is rather apples and oranges between different reports, but that's not the point - look at the trend and tell me you think they have things under control. Where do you think this is going?]
AYi@AYi_AInotesAI 评分4141
引用AYi@AYi_AInotes如何从零想出一个估值百亿的创业点子? Scale AI 创始人,现在是Meta首席AI官,muse负责人 的 Alexandr Wang 给出了一条极简铁律:活在未来,倒推今天还不存在的那行 API。 从深夜抢注域名,到肉身坐在客服气泡后死磕每一个访客,这段 4 分钟的复盘,讲透了科技商业里最硬核的起步真相。 很多人可能不知道,现在估值接近 140 亿美元的 AI 数据霸主 Scale AI,在刚起步的前半年,创始人每天也在经历极度严重的精神内耗。 这是 Scale AI 创始人, Alexandr Wang 在 SPC 闭门访谈里,第一次毫无保留地复盘自己在 YC 期间最痛苦的负一阶段。 一句话概括这段分享最值钱的本质: 所有伟大企业的起点,并不是算无遗策的天才顿悟,而是在漫长的游荡期里,靠着第一性原理把脏活做透,硬生生把一个看似不起眼的点子熬成了超级基础设施。 现在几乎所有想做点事、想做个人项目或创业的人,都在经历同一种心理折磨: 打开文档写满了各种点子,却总觉得每一个都不够好; 看着身边的人都在飞速推进,总觉得自己从第一天起就落后了别人半年; 每天在强烈的存在焦虑里打转,不知道自己到底在折腾什么。 Alexandr Wang 当年也是一模一样的处境。 我把他在视频里拆解出的三个底层认知,整理成最干货的复盘讲透👇 ① 选方向的第一性原则:活在未来,倒推缺失的那行 API 当年他在 YC 每天写点子文档,直到读了 Paul Graham 的那篇经典文章:Live in the future, and build what's missing. 他当时看到了一个未来的必然趋势: 未来的人类算力(Human Compute)一定会像计算机算力一样,被极度动态地编排和调用。 但当时整个互联网上,根本没有一个能够像调用服务器一样直接调用人工标注与处理的 API。 于是他花了一整晚买下 ScaleAPI 这个域名,这就是百亿帝国的最初原点。 ② 拆穿创业最大的心理陷阱:起步即落后的虚妄焦虑 在负一阶段,最致命的不是没点子,而是同行压力带来的动作变形。 Alexandr 提到:当你刚萌生一个新点子时,环顾四周,总觉得别人已经跑了很久,自己一开局就落后了。 但事实是,绝大多数人都在各自的迷雾里摸索。 真正的差距从来不是谁先动手两星期,而是谁能在漫长的游荡期里顶住内耗,把方向压力测试到底。 ③ 穿越死亡谷的唯一解法:做无法规模化的笨活与脏活 在 Product Hunt 上线拿到第一波热度后,Scale 经历了整整 4 到 6 个月的空白游荡期。 当时没有爆发式增长,能不能成完全是未知数。 Alexandr 采取的最硬核策略只有一个:当客服。 他在官网挂了 Intercom 聊天气泡,每一个点进网页、发消息咨询的真实访客,背后亲自敲键盘回复的人就是他自己。 正是靠着跟每一个早期客户在泥潭里死磕,直到半年后才终于等来了第一个真正想做大的核心客户。 历史与商业演进的硬核印证: → 硅谷最经典的创业定律: 保罗·格雷厄姆提倡的 Do things that don't scale(做无法规模化的事),在 Scale AI 身上得到了最彻底的验证。世界上最顶级的自动化数据管道,最开始也是创始人靠肉身当客服一点点抠出来的。 → 游荡期是所有顶级公司的必修课: 从 Airbnb 早期靠卖麦片还信用卡债,到 Stripe 创始人亲自跑去客户电脑上敲命令行装插件,没有一家基础设施级巨头能跳过这至少半年的迷茫摸索。 站在另一个更理性的视角来看,这件事给普通人的启发极其锋利: → 不要把摸索期的焦虑误判为失败: 从负一阶段到零的这段时间,内心动荡和怀疑是系统的标配属性,而不是你能力不足的证明。 → 别在战术的勤奋里逃避真正的思考: 想点子不是在文档里盲目堆数量,而是敢于逼问自己:未来五年哪件事一定会发生,而今天还缺了关键工具? → 离真实用户再近一点: 当你不知道下一步该做什么时,去跟每一个点了聊天气泡的真实访客聊半个小时,远比关在屋子里改一百遍商业计划书管用得多。 最后收个尾: 世上从来没有一开局就清晰无比的百亿蓝图。 伟大往往就藏在那份写满废案的文档里,藏在深夜无人问津的客服窗口背后。 熬过负一阶段的迷茫,把未来的缺失变成今天的行动,你才算真正站在了起跑线上。
Hacker News popular via buzzing.ccAI 评分1111 青蛙和蟾蜍与日益强大的机器
文章借经典儿童读物《青蛙和蟾蜍》的叙事框架,探讨日益强大的机器对人类生活与情感的影响。通过将童话角色置于现代技术语境中,作者反思了自动化与智能设备如何改变日常互动、人际关系及自我认知。文章以文学视角切入,审视技术进步带来的心理与社会层面的复杂后果。
Dongxi 东锡 NLP@dongxi_nlp精选AI 评分6767引用Andrej Karpathy@karpathyWe'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
推荐理由:作者借个人经历引出 Karpathy 关于用受控语言、图表、网页和视频理解模型输出的建议,可当作换个方式向 LLM 提问的参考。
dex@dexhorthyAI 评分2020
jason@jxnlcoAI 评分2424引用Eugenia Kuyda@ekuydaso i can book a flight and a hotel in 2 min myself OR listen to dots go through every flight option and send me screenshots of google maps on an excruciating 10 min call
Rohan Paul@rohanpaul_aiAI 评分3535
Rohan Paul@rohanpaul_aiAI 评分3333
Hacker News popular via buzzing.ccAI 评分99 废除白宫记者团
文章主张废除白宫记者团这一建制。原文未提供更多可核验的细节、数据或具体方案。
Hacker News popular via buzzing.ccAI 评分99 还有什么比加油站的高油价在政治上更具破坏性?
文章以"还有什么比加油站的高油价在政治上更具破坏性?"为切入点展开讨论。原文未提供更多可核验的具体信息。
Hacker News popular via buzzing.ccAI 评分2323 即将到来的言论自由之争
文章讨论即将到来的言论自由之争,但正文仅有一句标题式表述,未提供具体主体、事件或数据。
Hacker News popular via buzzing.ccAI 评分2323 Hacker News 发起投票:哪些 AI 挑战已经实现
Hacker News 发起投票,让社区成员评选此前提出的各项 AI 挑战中哪些已经实现。投票围绕这些挑战的完成情况展开,具体条目与结果未在原文中列出。
TechCrunch · AIAI 评分5959 TechCrunch 分析消费级 AI 的难看经济学:付费用户仅 2.2%,难敌企业市场
TechCrunch 撰文分析消费级 AI 的商业困境。尽管 Meta 的 Muse、OpenAI 的 Dots 和估值 100 亿美元的 Instinct 让个人 AI 助手看似回潮,但 a16z 引用 PNC 数据显示,截至 5 月仅 2.2% 的消费者为 AI 付费,月均支出 31 美元,且增长呈线性,模型性能跃升几乎不改变付费意愿。
阑夕@foxshuo精选AI 评分6565推荐理由:原文记录了AI伪造声音与仿冒账号结合的诈骗全过程和追回结果,读者可以据此了解这类组合骗术的作案路径。
AYi@AYi_AInotesAI 评分2727
引用Christopher Lee@coachseeelI'm 30 and essentially starting my life again from 0 Lets build together.
OpenAI NewsAI 评分2222 OpenAI:先进 AI 或成突破性创意背后的日常执行关键
OpenAI 发文探讨先进 AI 对突破性创意背后日常工作的价值,认为执行环节可能决定下一阶段经济的形态与进步速度。文章未给出具体模型、参数或评测数据,核心观点是 AI 在常规执行类任务上的作用可能比创意本身更具影响力。
DogeDesigner@cb_dogeAI 评分33准确 😂 (注:原文仅两个词,无实质信息,无法生成符合规则的 10-15 字标题。)

ginobefun@hongming731AI 评分2626LangChain 在 Open SWE 智能体 Harness 内构建模型路由器,按任务分类选择最低成本合适模型,将编码线程中位数成本降低 64%,质量变化可忽略不计。
引用ginobefun@hongming731https://x.com/i/article/2105830283580989440
ginobefun@hongming731AI 评分77
Yuchen Jin@Yuchenj_UW精选AI 评分6767引用Andrej Karpathy@karpathyWe'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Latent SpaceAI 评分5252 Latent Space 访谈 MIT 的 Alex Zhang:RLM、harness 设计与研究品味
Latent Space 播客访谈 MIT 博士生 Alex Zhang,围绕其 Recursive Language Models(RLM)研究展开。
Suno@sunoAI 评分1818
引用𝐌𝐫. 𝐖𝐢𝐜𝐤 🇺🇸 🦍@SoonMrWickWhatever you do just get rid of that synthetic nasty sound in your music. Suno isnt going anywhere if it remains at the current quality... it needs to sound polished, mastered, not like its playing through the cheapest speakers you could buy at circuit city before they went out of business.
Rohan Paul@rohanpaul_aiAI 评分2929
swyx@swyxAI 评分1010我要为 @aidotengineer nyc 换掉我的标语 提前飞过去参加 NY Comic Con。11 天后见!
引用christian@cxgonzalezit took not even 30 seconds after leaving my hotel in nyc to see something that i almost never saw in 3 months in sf: i saw a hot girl
gabriel@gabriel1AI 评分2323
Chubby♨️@kimmonismusAI 评分2626引用Tibo@thsottiauxBecause usage on your primary dot is virtually unlimited at the moment, I can’t really give a reset. I need to come up with something new fast.
Hacker News popular via buzzing.ccAI 评分6464 生成式 AI 冲击下 Web 开发教育的没落
molily.de 的博客文章汇总了 Baldur Bjarnason、Axel Rauschmayer、Salma Alam-Naylor、Josh W. Comeau、Kyle Cook 和 Rachel Andrew 等多位从业者的陈述,描述生成式 AI 对 Web 开发教育和技术出版的冲击。