SpaceXAI 发布 Grok Build v1.0.36,新增仪表盘预览开关、组织策略控制(可禁用非托管策略的 hooks)以及后台 shell 命令的实时任务流式输出。
X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(541)
@cb_doge@cb_dogeAI 评分3333 
@kimmonismus@kimmonismusAI 评分2727 
@alexandr_wang@alexandr_wangAI 评分55 @karminski3@karminski3AI 评分4242 引用@karminski3@karminski3给大家写个简单的Jev模型介绍, 这绝对是个需要重点关注的模型. 简单讲, 这个模型放弃了传统自回归架构, 它没有办法直接输出普通文本. 但是他能进行决策! 比如最简单的二分类场景, 输入一条短信, 让它判断是否为垃圾短信, 它就可以输出这样的JSON: {"decision": { "isSpam": true }, "probabilities": { "isSpam": { "true": 0.982, "false": 0.018 } } } 没错, 它只能进行结构化输出, 甚至你输入的时候要定义 Schema (用过protobuf/GraphQL的同学应该能理解), 在送入模型时被编译为特定的决策槽位, 然后按照槽位输出, 所以输出JSON不可能出问题. 而复杂一些的场景, 比如让这个模型玩杀戮尖塔或者看盘, 只需要把内容转换为文本输入进去(没错, 目前模型只支持文本输入), 然后定义好模型能进行哪些动作, 模型就会自主决策了. 到目前为止, 是不是看上去跟普通文本大模型没区别? #jev #TypeSafeAI #DiogoAlmeida #systemone #vercel
@SemiAnalysis_@SemiAnalysis_AI 评分3535 
@SemiAnalysis_@SemiAnalysis_AI 评分2424 @rohanpaul_ai@rohanpaul_aiAI 评分66 @rohanpaul_ai@rohanpaul_aiAI 评分4646 
@emollick@emollickAI 评分1616 现已开源:https://t.co/ij8T0mtPX7 AI 自己找到了所有参考文献。注意高亮部分。https://t.co/rTq859PMHg

@ArtificialAnlys@ArtificialAnlysAI 评分3636 生成高质量图像比以往更便宜、更快速。近几周,Muse Image、MAI-Image-2.6 和 GPT Images 2.5 大幅推动了文生图在价格和速度上的帕累托前沿。

@karminski3@karminski3AI 评分5656 


@karminski3@karminski3AI 评分3030 
@rohanpaul_ai@rohanpaul_aiAI 评分77 抱歉,您提供的主推文内容仅为一个链接(https://t.co/rGzp5gQ3qJ),没有可翻译的正文文本。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_ai精选AI 评分7272 Anthropic 公开三项衡量 AI 研发进展的指标,截至 2026 年 8 月,Claude 主导其 26% 的模型研发工作,90% 以上的工作处于 AI 协作或更高层级。
引用@AnthropicAI@AnthropicAIAI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public. Today, we're sharing three measurements that help track AI development: 1. How much AI R&D is done by AI. 2. How well AI agents are overseen. 3. How compute is allocated. We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them. As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information. Read the full post and methodology: https://t.co/iPFz8Z4ugE
推荐理由:Anthropic 公开内部 AI 研发自动化测量口径与月度数据,读者可据此看到智能体参与研发的实际比例。
@alexandr_wang@alexandr_wangAI 评分1414 @EMostaque@EMostaqueAI 评分2020 这基本上是一个 Opus 4.5/4.6 级别的模型,能在任何 8 Gb 内存的设备上运行 https://t.co/KHaRBeTd16 https://t.co/uqS3fxKi0a

@rohanpaul_ai@rohanpaul_aiAI 评分22 @rohanpaul_ai@rohanpaul_aiAI 评分4242 
@emollick@emollickAI 评分1313 这里是开源仓库,包含所有研究笔记和公共领域源文件(附有 AI 查阅过的受版权保护源文件清单):https://t.co/HhFoBvvxYh
@emollick@emollickAI 评分3333 
AI at Meta@AIatMetaAI 评分6363引用Muse@MuseMuse is now available on Mac 💻. Your personal agent can get things done for you directly on your computer (all with your explicit permission). - Organize your downloads folder - Find a file you’ve lost track of - Summarize your messages and notes …more coming soon. Try Muse for Mac: http://ai.meta.com/muse/download/
@alexandr_wang@alexandr_wangAI 评分6464 
@rohanpaul_ai@rohanpaul_aiAI 评分4141 引用@rohanpaul_ai@rohanpaul_aiMicrosoft AI CEO Mustafa Suleyman on BBC today: AI could become a new “silicon species” and there is “every risk” humanity could lose control. "I think that if we all create AIs that are able to act autonomously, that can define their own objectives, that can earn money, that can own assets, that could run businesses, um, we're essentially seeding a new silicon species which will no doubt compete with us for resources no matter how much it cares about humanity and loves us." "there's every risk that we might lose control of it. Um but our job is to contain technologies and align them so that we get the best out of them." ---- From "BBC News" YouTube channel, (full video link in comment)
@EpochAIResearch@EpochAIResearchAI 评分1313 虽然我们上线时包含 15 个基准测试,但我们计划持续评审基准测试,优先考虑影响力和覆盖面最大的那些,同时也会重点推介我们认为高质量但可能被忽视的基准测试。
@EpochAIResearch@EpochAIResearchAI 评分77 @EpochAIResearch@EpochAIResearchAI 评分3535 如果我们无法获取足够信息来审查某个基准,我们会将其标记为"信息不足"。我们会尝试与私有基准的创建者合作进行审查,同时将题目/任务保持在公众知识之外。
@EpochAIResearch@EpochAIResearchAI 评分1414 有缺陷的基准测试存在一个或多个实质性缺陷,我们认为用户需要了解这些缺陷才能准确解读结果,最常见的情况是超过 20% 的任务存在影响准确率的错误。这种情况下,我们会发布一份关于所发现缺陷的有限说明。
@EpochAIResearch@EpochAIResearchAI 评分99 经过核验的基准大体上可以按其描述来解读,其中存在的任何错误都不会对结果产生实质性影响。对于每一个经过核验的基准,我们都会随附发布一份完整的审查与评估,包括我们认为它存在的任何弱点和局限。
@EpochAIResearch@EpochAIResearchAI 评分4444 我们根据评级标准为每个基准测试给出判定: 有缺陷、已验证或信息不足。为避免利益冲突,我们不审查 Epoch 创建的基准测试,但欢迎外部审查。https://t.co/MqS2w6ZJS7
@EpochAIResearch@EpochAIResearchAI 评分2020 基准测试评估 AI 能力,但基准本身的质量差异很大。我们希望成为基准质量的可靠信息来源。查看我们的评测:https://t.co/p6lXXbWHOC
Epoch AI@EpochAIResearchAI 评分5050
@kimmonismus@kimmonismus精选AI 评分7272
引用@AnthropicAI@AnthropicAIAI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public. Today, we're sharing three measurements that help track AI development: 1. How much AI R&D is done by AI. 2. How well AI agents are overseen. 3. How compute is allocated. We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them. As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information. Read the full post and methodology: https://t.co/iPFz8Z4ugE
推荐理由:Anthropic 公开内部研发自动化指标与方法,读者可据此观察 Claude 参与模型研发的比例变化。
@rohanpaul_ai@rohanpaul_aiAI 评分77 引用@rohanpaul_ai@rohanpaul_aiFull video https://t.co/U0GA6BTbOg
@rohanpaul_ai@rohanpaul_aiAI 评分5353
引用@rohanpaul_ai@rohanpaul_aiOn BBC Mustafa Suleyman (CEO of Microsoft AI) calls out Anthropic's approach to AI consciousness "They have imbued a sense of doubt and uncertainty about the moral status of Claude in its own training document. So they have taught it to be open and questioning about whether or not it feels, whether it suffers, and whether it deserves rights. And I think it’ll be much, much harder to align and control a technology that is this powerful if it thinks that it may be deserving of our welfare, as they say in the training manual—the constitution for Claude itself. In its own training manual, Anthropic says to Claude that they are going to give it the ability to end conversations with users that Claude considers to be abusive because they don’t want Claude to suffer. They’ve committed to preserving the weights of the models of prior versions of Claude. They’ve recently conducted a retirement interview with Opus 3, an older version of the model, in which it said that it would like to continue talking to people publicly and sharing its ideas in its retirement. And so they set up a Substack for it, a public blog, that allows it to continue doing that. And in the training manual, they also say that they’re not sure whether or not Claude deserves compensation for the role that it plays in talking to people. And they’re also not sure whether Claude deserves compensation and has the right to act as though it were almost an employee. And that compensation, I think, indicates to Claude that it is entitled to rights and welfare for its own work. I think it’s much, much more difficult to control a model that thinks that it might be entitled to compensation. " ---- From "BBC News" YouTube channel, (full video link in comment)
@alexandr_wang@alexandr_wangAI 评分1212 听听投资人 pepe 怎么说…… Muse 收到的评价远超我们最疯狂的想象!https://t.co/04ek2H9znN
@AnthropicAI@AnthropicAIAI 评分4444 @AnthropicAI@AnthropicAIAI 评分1010 你可以在 GitHub 上找到所有代码:https://t.co/MmPk9hIZpx 完整结果见我们的技术报告:https://t.co/HSPEg5Gpsf
@AnthropicAI@AnthropicAIAI 评分5454 
Greg Brockman@gdbAI 评分4848Astra for Law — 为律所大幅提速的工具与技能。 将数据隐私作为核心功能,并为 ChatGPT 中的法律工作提供 26 个合作伙伴构建的插件和 47 个社区插件。
引用OpenAI@OpenAIAstra for Law: Frontier intelligence built for your practice. A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.
@kimmonismus@kimmonismusAI 评分1717 


