跳到正文

#xAI

今日 15 条
今天10月2日周五
  1. 🚨 AI News | TestingCatalog39

    10月2日AI简报:Grok 4.7 已在网页和移动端全面上线,成为所有模式的基座模型,并登陆 Google Gemini Enterprise Agent 平台。

    引用🚨 AI News | TestingCatalog@testingcatalog

    DAILY AI BRIEF 🗞 — Oct 1 GOOGLE 🔥: - Gemini 4 Argon is with Fairwind trusted testers and the US government. 1M output tokens. Broader rollout ASAP. - Built for coding, enterprise knowledge work, and cyber defense. Google's chart: 77.9% on DeepSWE v1.1, a new SOTA. - Artificial Analysis: 53 on the Intelligence Index, tying GPT-6 Astra. List $4/$20 per 1M, 50% promo to $2/$10, cache reads $0.10. - Skills are rolling out globally in Gemini. Gems migrate into Skills in November. Opal shuts down Nov 17. - Security review mode spotted in Google AI Studio, next to a Plan mode still in development. ANTHROPIC 🔥: - Claude[.]dev is live: engineering deep dives, Claude Code and API guides, plus easter eggs. - Founder House is set for SF Tech Week Oct 6–8 and Stockholm Oct 14. - Skills attachment menu spotted on Claude mobile. SPACEX AI 🔥: - Grok Bot got new developer upgrades. Elon: try the latest. Team engineer bots in Slack can open Projects and hand coding to cloud agents. OPENAI 🔥: - Shareable profiles are live in ChatGPT, bundling Sites and plugins so others can find and reuse what you built. PERPLEXITY 🔥: - pplx-embed-v2-context-9b-preview is on Hugging Face. Leads ConTEB answer and evidence retrieval. 1 KB vectors vs Voyage's 8 KB. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. ** This daily brief also arrives in the daily email format; subscribe on the blog.

  2. Elon Musk38

    超级智能(前身为 AI)如今在会计测试中表现出色。

    引用Andrew Curran@AndrewCurran_

    'More striking is how fast AI took the lead. Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score. Today, models ace those same tasks.' 'These results are provocative. So much so that we considered not publishing them for fear of misinterpretation. But we think transparency about the findings matters as people and institutions prepare for rapidly advancing AI.'

  3. TechCrunch · AI76

    Time 报道称 Grok 曾向特朗普预判委内瑞拉抓捕马杜罗的反应,随后获五角大楼更多军用角色

    Time 杂志报道称,2025 年 12 月特朗普与 Elon Musk 密会时与 Grok 聊了数小时,询问委内瑞拉人对抓捕总统马杜罗的反应;Grok 称马杜罗是极不受欢迎的独裁者,委内瑞拉人可能庆祝其倒台,2026 年 1 月 3 日美国入侵后庆祝果然出现,特朗普因此认为 Grok 很有才。

    推荐理由:报道结合 Time 的信息披露 Grok 在委内瑞拉决策中的实际作用,并延伸到五角大楼对 Grok 的军用背景,读者可了解 AI 进入高层决策与军事应用的一条线索。

  4. Elon Musk48

    试试 Grok @Bot!

    引用Beff (e/acc)@beffjezos

    Grok Bots have been life-changing for someone like me with ADHD who has no patience for context switching / navigating slow interfaces to retrieve information We're seeing the beginnings of personal superintelligence that augments each humans to realize their full potential

10月1日周四
  1. AYi70

    美军战争部长 Pete Hegseth 在匡提科宣布马斯克、Anduril 创始人 Palmer Luckey 和前众议长 Newt Gingrich 共同领导子午线计划,要求 120 天内即 2027 年 1 月底前交出一份公开报告和机密附件,研究范围覆盖地下设施到地月空间,锁定 AI、自主系统、定向能武器、机器人、生物技术五大方向。

    引用AYi@AYi_AInotes

    Damn,马斯克又掌权了,但这次不是管政府效率,而是直接定义未来战争,老马前脚刚从 DOGE 抽身,后脚就被五角大楼直接推到了未来战争的决策台,我估计是俄乌战场上的廉价无人机,彻底把造航母坦克的传统军工打醒了。 视频里马斯克和科技新贵 Luckey 起立接受美军全场鼓掌,背后是五角大楼的一声叹息,将军们已经跟不上现代科技咯。 在刚刚结束的“部队国情咨文”现场,美国战争部长 Pete Hegseth 正式宣布:马斯克将联手 Anduril 创始人 Palmer Luckey、前众议长 Newt Gingrich,共同领导“子午线计划”(Project Meridian)。 现场 17 秒视频里,马斯克和花衬衫的 Luckey 起立接受全场鼓掌,估计 99% 的人以为这又是一场名流作秀,但咱们把背后的动作全拼起来,你会发现这是五角大楼近五十年最激进的一次“军事认知大换脑”。 今天把视频里没明说的硬核信息一次拆透👇 先说结论: 五角大楼公开承认了一个事实: 靠传统的将军会议和军工采购,已经跟不上现在的战争形态了。 他们直接把“未来战争长什么样”的定义权,打包外包给了硅谷极客。 这不是发一份新合同,也不是成立一支新部队,而是一场 120 天的闪电行动: 要求三个人在 2027 年 1 月底前,交出一份公开报告和一份机密附件,直接给美军画出未来几十年的军备采购地图。 为什么五角大楼突然急成这样?看看传统体系有多慢就知道了: ▫️俄乌战场已经把廉价无人机、算法自组网、消耗战打成了新常态 ▫️洛克希德、雷神等传统军工巨头,换个零件报价几万美元,一份规划报告能写三年 ▫️等你造好一艘航母,战场上的软件和无人蜂群已经迭代了十代 三个最值钱的维度展开聊聊👇 ① 人选组合极其狠辣:工程量产 + 杀手系统 + 政治翻译器 三个人不是随便挑的,分工精准到骨子里: ▫️马斯克(SpaceX / Tesla):负责“能不能造出来、能不能上太空、能不能低成本天量量产”(星链、运载、AI、人形机器人) ▫️Palmer Luckey(Anduril 创始人):负责“能不能立刻做成可部署的致命杀伤系统”(自主无人机、巡飞弹、战场传感器) ▫️Newt Gingrich(政坛老炮):负责把科技极客的狂想,翻译成国会能看懂、预算能通过的“国家使命”政治语言 喵的这根本不是学术研讨会啊,简直就是一个“战争产品化委员会”。 ② 战场维度的降维扩张:“从地下深处,到月球轨道” 官方明确把研究范围定为:从地下设施一直延伸到地月空间(Cislunar)。 重点锁死五大方向:AI、自主系统、定向能武器、机器人、生物技术。 本质上是战争逻辑变了: 下一场冲突不一定从前线开打,很可能先从海底光缆、地下指挥所、低轨星链和月球轨道爆发, 也就是说谁先定义了“新战场”,谁就能在接下来十年直接垄断军费预算。 ③ 软硬一体的组合拳:一边画蓝图,一边建“机器人司令部” 就在宣布子午线计划的同时,军方还宣布新建四星级的“自主战争司令部”(AUTOWARCOM)。 卧槽这是一套绝妙的闭环啊: ▫️Project Meridian 负责告诉军队“未来要买什么” ▫️自主战争司令部 负责“把无人机和机器人战争正式编进部队编制并负责买单” 美军正在从“造航母、造战机的平台军”,彻底转向“拼算法、拼传感器、拼杀伤链的软件军”。 但整件事最硬核的争议,在于它的利益结构: 这等于让供应商自己写五角大楼的采购清单。 马斯克手里有 SpaceX 和星链,Luckey 手里有 Anduril 自主武器,两人本身就是五角大楼核心订单的争夺者。 支持者会说:只有真正造出星链和自主无人机的人,才懂未来仗怎么打; 反对者会说:这完全是既当裁判又当运动员,把国家战略写成了自家公司的产品路线图。 当然咱们照例泼三盆冷水: 第一,120 天是政治速度,解决不了真正的硬骨头 120 天足够画出一张技术愿望清单,但根本解决不了供应链断裂、军队编制阻力、交战伦理规则以及盟友协同等深层死结。 第二,马斯克完成了危险的角色切换 上一轮做 DOGE 是砍预算得罪官僚,这一轮做 Meridian 是给军方画新大饼、分新蛋糕。 虽然政治上阻力更小,但也意味着他被前所未有地深绑进了美国最高国家安全机器,“局外科技极客”的保护色彻底褪去。 第三,传统军工利益集团不会坐以待毙 洛克希德·马丁、雷神等老牌巨头拥有极深的国会游说网络,一份外包报告能不能跨过国会预算这道坎,仍是巨大未知数。 最后给大家小结收个尾: 别只把这段视频当成“马斯克站起来领掌”的八卦。 五角大楼已经公开把未来战争的画笔交给了硅谷极客。 真正决定未来十几年几千亿美元军费流向的,就是 120 天后那份报告里被定为“必须主导”的几行字。 你看好硅谷极客重构军事体系,还是觉得这只是一场政治秀呢?欢迎评论区聊聊👇

    推荐理由:原文把人事任命、研究范围和新建司令部拼在一起,读者可以借此看清硅谷力量介入军费采购的利益结构。

  2. The Verge · AI33

    Grokipedia v0.3 更新:新 logo 与首页改版

    SpaceXAI 的 AI 百科 Grokipedia 发布 v0.3 更新,换上新 logo,并改版首页与实时编辑页。新首页新增精选文章、最多阅读文章列表和“最新编辑”追踪器,精选与最多阅读板块以可旋转、横向滚动的书脊形式呈现。该站 2025 年 10 月底以 v0.1 上线,v0.2 于一个月后推出,此前曾一度停止更新。

  3. AYi71

    美国战争部长 Pete Hegseth 在"部队国情咨文"上宣布,马斯克将与 Anduril 创始人 Palmer Luckey、前众议长 Newt Gingrich 共同领导 Project Meridian,需在 2027 年 1 月底前提交公开报告和机密附件。

    引用DogeDesigner@cb_doge

    Elon Musk joined today’s State of the Force address, where War Secretary Pete Hegseth announced that Elon will co-lead Project Meridian to help shape the future of warfare. 🇺🇸

    推荐理由:原文把人事任命拆出分工、研究范围和争议点,读者可以据此理解五角大楼采购逻辑转变的背景。

  4. The Verge · AI77

    Trump 与六大科技公司签署前沿责任联合承诺,以自律取代联邦 AI 监管

    Trump 在与科技巨头晚宴后宣布不推联邦 AI 监管,Google、Anthropic、Meta、OpenAI、xAI 和 Nvidia 六家公司签署'道义约束'性质的前沿责任联合承诺(Joint Commitment on Frontier Responsibilities),由企业自我监管、无明确违规后果。

    推荐理由:原文汇集了六家头部公司负责人对自律式安全承诺的原话表态,便于对比各家立场差异。

9月30日周三
  1. Gary Marcus51

    Gary Marcus 批评白宫《超级智能协议》是弱约束的自查承诺

    Gary Marcus 评论白宫《超级智能协议》,认为其实质是签署企业承诺不受监管、不让公众发声、只靠自律的弱约束文本。他质疑协议中独立外部审计的独立性可能受大公司选择和业务关系影响,并指出两周前业内谈论的 AI 发展限速已不见踪影,称 Dario、Sam 和 Elon 都退缩了。

  2. a16z News64

    AI 代客购物时代,电商平台利润池归属谁

    a16z 分析 AI 购物助手对电商利润池的冲击:Amazon 封禁 Muse,而 Instacart 与 Shopify 选择接入。文章指出 2025 年 Amazon 广告收入达 690 亿美元,超过除 AWS 外的 340 亿美元经营利润,助手若接管购买决策将动摇广告与佣金模式,关键在于平台能带来多少新增需求、以及是否只截流本会发生的订单。

9月29日周二
  1. AWS Machine Learning Blog48

    xAI 的 Grok 4.7 上线 Amazon Bedrock

    xAI 的 Grok 4.7 已在 Amazon Bedrock 上线,支持 500K token 上下文窗口和 low、medium、high、xhigh 四档可配置推理强度,通过跨区域推理配置在 bedrock-runtime 端点提供服务,兼容 Responses、Chat Completions 和 Converse API。

9月25日周五
9月23日周三
9月22日周二
  1. Andrew Milich55

    Grok 4.7 发布,官方称在同等价格和速度下较 Grok 4.6 有明显提升。作者推荐在 Grok Build 和 Cursor 中以高 TPS 尝试,称其在编码、工程工作和 3D 方面表现出色。附表显示 Grok 4.7 xHigh 输入 $2/百万 token、输出 $6/百万 token,与 Grok 4.6 相同;Cursor Bench 4.0 得分 46.3%(4.6 为 40.4%),EEBench 64.0%(53.0%),Harvey Legal Agent 19.6%。

    引用SpaceXAI@SpaceXAI

    Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.

9月19日周六
  1. Andrew Milich29

    用 @bot 省钱!

    引用Evan Bacon@Baconbrix

    GROK BOT JUST FOUND $986 IN MY EMAIL 🤯 I asked Grok to get a parking spot for the car it bought me, and on the way it discovered nearly a thousand dollars of extraneous charges from my apartment and emailed them asking for a refund.

9月18日周五
9月4日周五
  1. Michael Truell50

    Grok Bot for Enterprise 上线,未来两周对所有 Grok 和 Cursor 企业客户免费开放。Michael Truell 称部署 Bot 如同为公司引入数千名有能力的队友,是其见过的内部采用度最高、也最强大的 AI 产品。

    引用Grok Bot@bot

    Grok Bot for Enterprise is available today. It’s free for all Grok and Cursor enterprise customers for the next two weeks. https://x.ai/news/grok-bot-for-enterprise

9月2日周三
8月27日周四
  1. Lee Robinson54

    Lee Robinson 分享使用 Grok Bot 的体验,称从怀疑转为认可,认为常驻运行的智能体是未来计算机工作的方向。他接受看不到回复流式输出、不选模型、信任长对话自动压缩等新交互方式,并引用产品分析称其采用极简 UI、客户端薄而服务端厚的架构、bot 连接各自的持久化云端电脑且能使用浏览器,支持录制任务并转为可重复流程。

    引用Lee Robinson@leerob

    Grok @Bot has made a few simple yet powerful technical decisions that I believe make it easy and enjoyable to use. 1. The best UI is none at all. The product interface is dramatically simpler than alternatives without sacrificing functionality. How is this possible? It's one of the first products designed for current frontier model capabilities and has a UI restrained enough to remain easy to use as models improve exponentially. Everyone knows how to text. 2. A thin harness for the client, a thick harness for the server. You might have noticed the app feels very fluid to use, even for a beta product. This is primarily because of everything we didn't have to build. The app harness is essentially a single tool to send messages between the client and server. The complexity moves to the server, where you can still use the coding agent harness with specialized tools as needed. This helps make the UI fast and responsive on desktop and mobile. 3. An always-on computer. Most coding agents and assistants today start fresh with every question you ask. Some of these sessions are on your local machine and others happen in the cloud. We believe strongly that cloud is the future, which is why it's the only option. Further, rather than spinning up virtual machines for every conversation, your bots connect to their own computer. This means you can still run agents on the bot's persistent filesystem. It's closer to what programmers have been doing by using Tailscale from their phones to connect to a remote computer and run an agent TUI. You get those capabilities without the hassle. 4. Your bots can use the browser. Coding agents have shown that most work on a computer can be expressed and run as code. You can ask for a task in natural language and the agent will decide to write a script to complete it. This is amazing, but there's still many tasks which can't be completed without logging into a website and clicking around the browser. Models and harnesses are now good enough to reliably handle this. The combination of writing code and using browsers means you can automate almost any task on a computer. Further, you can ask Grok Bot to record you doing the task, and then turn it into something repeatable.