'More striking is how fast AI took the lead. Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score. Today, models ace those same tasks.' 'These results are provocative. So much so that we considered not publishing them for fear of misinterpretation. But we think transparency about the findings matters as people and institutions prepare for rapidly advancing AI.'
#xAI
#xAI
今日 3 条
Elon Musk@elonmuskAI 评分3838引用Andrew Curran@AndrewCurran_
DogeDesigner@cb_dogeAI 评分55
Elon Musk@elonmuskAI 评分4848引用Beff (e/acc)@beffjezosGrok Bots have been life-changing for someone like me with ADHD who has no patience for context switching / navigating slow interfaces to retrieve information We're seeing the beginnings of personal superintelligence that augments each humans to realize their full potential
Gary MarcusAI 评分5151 Gary Marcus 批评白宫《超级智能协议》是弱约束的自查承诺
Gary Marcus 评论白宫《超级智能协议》,认为其实质是签署企业承诺不受监管、不让公众发声、只靠自律的弱约束文本。他质疑协议中独立外部审计的独立性可能受大公司选择和业务关系影响,并指出两周前业内谈论的 AI 发展限速已不见踪影,称 Dario、Sam 和 Elon 都退缩了。
Yuchen Jin@Yuchenj_UWAI 评分2525我 3 周前试了 Grok Bot。 2 周前装了 Instint。 上周装了 Muse。 现在显然我还得试试 Dots。 个人 AI 助手之战开始了。

a16z NewsAI 评分6464 AI 代客购物时代,电商平台利润池归属谁
a16z 分析 AI 购物助手对电商利润池的冲击:Amazon 封禁 Muse,而 Instacart 与 Shopify 选择接入。文章指出 2025 年 Amazon 广告收入达 690 亿美元,超过除 AWS 外的 340 亿美元经营利润,助手若接管购买决策将动摇广告与佣金模式,关键在于平台能带来多少新增需求、以及是否只截流本会发生的订单。
karminski-牙医@karminski3AI 评分3535
a16z NewsAI 评分4646 a16z 图表周报:AI 代码生成催生大量应用,但下载量停滞、爆款占比骤降
AI 代码生成工具推动应用数量激增,iOS、Android 和 Chrome 每月新增应用数量翻倍甚至翻四倍,但下载量(iOS 为评分)基本停滞,达到 10+ 评分、100+ 下载等增长门槛的应用占比大幅下滑。
Lee Robinson@leerobAI 评分4242引用Peng Zheng@pengzheng_wrote down some of the design thinking behind Grok Bot. persistent roles, clear state, scoped context, coordinated teams — an interface designed to move you from operating AI to delegating work. https://x.ai/news/designing-grok-bot
Lee Robinson@leerobAI 评分5454引用Lee Robinson@leerobGrok @Bot has made a few simple yet powerful technical decisions that I believe make it easy and enjoyable to use. 1. The best UI is none at all. The product interface is dramatically simpler than alternatives without sacrificing functionality. How is this possible? It's one of the first products designed for current frontier model capabilities and has a UI restrained enough to remain easy to use as models improve exponentially. Everyone knows how to text. 2. A thin harness for the client, a thick harness for the server. You might have noticed the app feels very fluid to use, even for a beta product. This is primarily because of everything we didn't have to build. The app harness is essentially a single tool to send messages between the client and server. The complexity moves to the server, where you can still use the coding agent harness with specialized tools as needed. This helps make the UI fast and responsive on desktop and mobile. 3. An always-on computer. Most coding agents and assistants today start fresh with every question you ask. Some of these sessions are on your local machine and others happen in the cloud. We believe strongly that cloud is the future, which is why it's the only option. Further, rather than spinning up virtual machines for every conversation, your bots connect to their own computer. This means you can still run agents on the bot's persistent filesystem. It's closer to what programmers have been doing by using Tailscale from their phones to connect to a remote computer and run an agent TUI. You get those capabilities without the hassle. 4. Your bots can use the browser. Coding agents have shown that most work on a computer can be expressed and run as code. You can ask for a task in natural language and the agent will decide to write a script to complete it. This is amazing, but there's still many tasks which can't be completed without logging into a website and clicking around the browser. Models and harnesses are now good enough to reliably handle this. The combination of writing code and using browsers means you can automate almost any task on a computer. Further, you can ask Grok Bot to record you doing the task, and then turn it into something repeatable.
@elonmusk@elonmuskAI 评分55
Karina@karinanguyenAI 评分3737Grok 4.6 在 DiligenceBench 金融测试中以约 52–53% 位列第二,与 Claude Opus 5 基本持平,Sonnet 5 以 46.2% 落后。

@frxiaobei@frxiaobeiAI 评分2222 不到一年时间,大模型御三家从 ChatGPT、Claude、Gemini 变成了 ChatGPT、Claude、Grok。 不过还好还是 CCG。