跳到正文

全部动态

今日 64 条
8月30日周日
8月27日周四
  1. Lee Robinson54

    Lee Robinson 分享使用 Grok Bot 的体验,称从怀疑转为认可,认为常驻运行的智能体是未来计算机工作的方向。他接受看不到回复流式输出、不选模型、信任长对话自动压缩等新交互方式,并引用产品分析称其采用极简 UI、客户端薄而服务端厚的架构、bot 连接各自的持久化云端电脑且能使用浏览器,支持录制任务并转为可重复流程。

    引用Lee Robinson@leerob

    Grok @Bot has made a few simple yet powerful technical decisions that I believe make it easy and enjoyable to use. 1. The best UI is none at all. The product interface is dramatically simpler than alternatives without sacrificing functionality. How is this possible? It's one of the first products designed for current frontier model capabilities and has a UI restrained enough to remain easy to use as models improve exponentially. Everyone knows how to text. 2. A thin harness for the client, a thick harness for the server. You might have noticed the app feels very fluid to use, even for a beta product. This is primarily because of everything we didn't have to build. The app harness is essentially a single tool to send messages between the client and server. The complexity moves to the server, where you can still use the coding agent harness with specialized tools as needed. This helps make the UI fast and responsive on desktop and mobile. 3. An always-on computer. Most coding agents and assistants today start fresh with every question you ask. Some of these sessions are on your local machine and others happen in the cloud. We believe strongly that cloud is the future, which is why it's the only option. Further, rather than spinning up virtual machines for every conversation, your bots connect to their own computer. This means you can still run agents on the bot's persistent filesystem. It's closer to what programmers have been doing by using Tailscale from their phones to connect to a remote computer and run an agent TUI. You get those capabilities without the hassle. 4. Your bots can use the browser. Coding agents have shown that most work on a computer can be expressed and run as code. You can ask for a task in natural language and the agent will decide to write a script to complete it. This is amazing, but there's still many tasks which can't be completed without logging into a website and clicking around the browser. Models and harnesses are now good enough to reliably handle this. The combination of writing code and using browsers means you can automate almost any task on a computer. Further, you can ask Grok Bot to record you doing the task, and then turn it into something repeatable.

8月25日周二
8月24日周一
  1. Andrew Ng50

    在捍卫 AI 开放性的斗争中,Marin 项目是模型训练开放性的一份珍贵示范——开放代码、数据、配方,甚至实验结果。公开分享 AI 研究曾是常态;我很感激 @percyliang 的开放实验室做法。

    引用Percy Liang@percyliang

    🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.

  2. elsewhere articles49

    22 岁 RoboParty 创始人黄一:一年 5 轮融资超 1 亿美元,谈具身智能创业

    RoboParty 萝博派对创始人兼 CEO 黄一在一年内完成 5 轮融资、累计超 1 亿美元,股东包括知名 VC 及小米、宁德等产业方。他将具身智能比作 42 公里马拉松:机器人本体已跑完 1/4,小脑约跑了一半,大脑可能才跑了一两公里。公司从第一代原型机 RPO 走向 RP1,近 150 人团队采用“蜂巢结构”管理,RP1 先服务科研教育场景。

8月23日周日
8月21日周五
  1. jietang44

    精彩评论:FLOPs 是智能;参数是知识!

    引用Liam Fedus@LiamFedus

    An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!

8月20日周四
8月19日周三
8月18日周二
8月17日周一
8月15日周六
  1. Nathan Lambert: Interconnects71

    Nathan Lambert 解析 GLM-5.3 为何能紧跟前沿

    Z.ai 发布 GLM-5.3,目前仅在 coding plan 提供,两周内将开放权重到 Hugging Face。作者认为其与 GLM-5.2 同底座、靠大幅扩展后训练提升成绩,在部分智能体编码基准上超越 Kimi K3 甚至个别超越 Claude Fable 5 或 GPT-5.6-Sol,参数约 750B。

    推荐理由:作者给出了对 GLM-5.3 成绩来源的解释框架,包括发布节奏、后训练策略和 RL 数据产业等背景,可用于理解中美前沿模型竞争的成因。

8月14日周五
8月13日周四
  1. Jensen Huang39

    强大的 A100 集群从 2020 年到 2029 年都具备任务能力。NVIDIA 计算不只是芯片。CUDA 为开发者和 NVIDIA 工程师提供了共同平台,让 Ampere、Hopper 和 Blackwell 在整个使用寿命期内持续升级。 CUDA 让 NVIDIA 计算具备通用性。通用性让它可互换。可互换性驱动利用率并延长耐用性,使 NVIDIA 算力成为一项生产性资产:可租用、耐用且可融资。

    引用Business Insider@BusinessInsider

    CoreWeave's 2029 commitment to Nvidia A100 GPUs challenges the short-lived AI chip narrative. https://bit.ly/4wkKn8t

8月12日周三
8月11日周二
  1. Google Developers Blog44

    为什么 Go 是 AI 辅助软件工程的理想语言

    Google 发文论证 Go 是 AI 辅助软件工程的理想语言:当 AI 智能体可秒级生成数百行代码,开发者重心从编写转向审查与维护,语言的可读性和工具链一致性变得更重要。Go 自带格式化、测试框架、依赖管理与安全工具,能让 AI 更快、更便宜、更可靠地处理代码,并减少上下文窗口污染与 token 成本。

8月10日周一
  1. Every latest articles58

    一次 vibe coding 如何把安全漏洞写进线上应用

    作者讲述自己 vibe coding 应用 Tastemaker 并用 Claude 添加 MCP 连接器后,被 GPT-5.6 Sol 审查发现线上连接器存在不应公开的注册路由,随后下线功能并撤销会话,未发现用户数据被访问。文章将其归因于'解释深度错觉'和只测试预期路径的倾向,并总结出学习领域基本原理、请人类专家把关、不把 AI 自我保证当唯一证据等规则,文末给出需要付费订阅解锁的五步提示词。