ZHO 分享《牛来》的观影体验,称正如 6 个月前所说,AI 用得越久,越喜欢人类笨拙的真诚、物体无所谓的存在与自然肆意的变幻。多视角的存在让大脑不再被单一形式束缚,回到作为一个人的纯粹体验。
X
关注 AI 研究者、开发者与机构的动态
按账号或来源筛选(536)
@ZHO_ZHO_ZHO@ZHO_ZHO_ZHOAI 评分1313 @rohanpaul_ai@rohanpaul_aiAI 评分2525 
@rohanpaul_ai@rohanpaul_aiAI 评分3838 
@rohanpaul_ai@rohanpaul_aiAI 评分2424 
@rohanpaul_ai@rohanpaul_aiAI 评分3131 
@rohanpaul_ai@rohanpaul_aiAI 评分1414 
@rohanpaul_ai@rohanpaul_aiAI 评分3030 
@rohanpaul_ai@rohanpaul_aiAI 评分5858 
@rohanpaul_ai@rohanpaul_aiAI 评分22 @rohanpaul_ai@rohanpaul_aiAI 评分6161 据彭博报道,Anthropic 在计划公开上市前寻求超过 100 亿美元的循环信贷额度,目前银行承诺额已高于约 100 亿美元的目标,但 Anthropic 也可能将额度控制在目标水平或以下。


@jxnlco@jxnlcoAI 评分2626 @testingcatalog@testingcatalog精选AI 评分6666 
引用@OpenAI@OpenAIAs models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. https://t.co/ecbMMmVoox
推荐理由:OpenAI 说明为安全与对齐暂停两周前沿 RL 训练,读者可了解头部实验室推进部署前如何调整训练节奏。
@MiniMax_AI@MiniMax_AIAI 评分2424 @rohanpaul_ai@rohanpaul_aiAI 评分77 抱歉,您提供的主推文内容仅包含一个链接(https://t.co/MQeYBfXr8r),没有可翻译的正文文字。请提供推文的实际文字内容,我将为您翻译。
@rohanpaul_ai@rohanpaul_ai精选AI 评分7878
引用@OpenAI@OpenAIAs models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. https://t.co/ecbMMmVoox
推荐理由:OpenAI 因 Astra 评测可能触及自主零日攻击的 Critical 阈值而暂停前沿 RL 训练,读者可看到其研究环境安全门槛的变化。
@kimmonismus@kimmonismusAI 评分3838 
@kimmonismus@kimmonismusAI 评分1313 @kimmonismus@kimmonismus精选AI 评分7373
引用@OpenAI@OpenAIAs models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. https://t.co/ecbMMmVoox
推荐理由:OpenAI 暂停最新部署模型的 RL 训练并披露 Astra 或触及 Critical 网络安全阈值,可据此了解前沿模型发布前的安全审查流程。
@rohanpaul_ai@rohanpaul_aiAI 评分4242 
@rohanpaul_ai@rohanpaul_aiAI 评分5050 Rohan Paul 在 X 上分享了 Nori V1 的详情链接,称其为面向表格的开源权重基础模型,代码与权重采用 Apache 2.0 许可,可免费用于商业用途。
@suno@sunoAI 评分6060 
@OpenAI@OpenAI精选AI 评分6565 推荐理由:OpenAI 披露暂停 RL 训练两周并扩大监控,读者可借此了解前沿模型内部开发阶段的安全管控安排。
@OpenAI@OpenAIAI 评分4949 @kimmonismus@kimmonismusAI 评分3737 
@kimmonismus@kimmonismusAI 评分2727 
@kimmonismus@kimmonismusAI 评分2828 
@kimmonismus@kimmonismusAI 评分2222 
@elonmusk@elonmuskAI 评分66 @mustafasuleyman@mustafasuleymanAI 评分2525 上周我们在 Text-to-Image 基准上测试了该模型,排名第 2 本周我们在图像编辑排行榜上运行了同一模型,得分排名第 3 MAI-Image-2.6 现已成为任何图像生成任务的全面高性能选手。
@jxnlco@jxnlcoAI 评分3232 期待见到大家!你们在哪?https://t.co/JOugVwCNHm
引用@OpenAIDevs@OpenAIDevsOpenAI DevDay is going global. Starting this October, DevDay Exchange is bringing builders together to swap build notes, share real projects, and meet the teams building OpenAI tools in: Bengaluru Tokyo Seoul Berlin Paris London São Paulo Mexico City We're ready to see what developers around the 🌎 are pushing to prod.
@krea_ai@krea_aiAI 评分1111 谁想要预览? 首场展示将于 8 月 27 日在 NYC 举行。 下方 RSVP 👇 https://t.co/MZIZdnzWA2

@krea_ai@krea_aiAI 评分2424 我们正在筹办一场私密活动,揭晓即将登陆 Krea 的新产品和模型。 通过以下链接申请:https://t.co/tlXfT9gWg2
@mustafasuleyman@mustafasuleymanAI 评分2424 我们在图像编辑基准测试中排名第三!……击败了 Google 的 Nano Banana 和 Meta 的 Muse Image 模型。https://t.co/Mm5ImJKvho
@OpenAIDevs@OpenAIDevsAI 评分44 @OpenAIDevs@OpenAIDevsAI 评分3434 
@omarsar0@omarsar0AI 评分5656 引用@ArtificialAnlys@ArtificialAnlysAnnouncing the Artificial Analysis Search Index, benchmarking how search API providers perform on quality, cost, and speed when used by an agent. We are initiating coverage with Parallel, Exa, Firecrawl, You (dot) com, Tavily, Keenable, and Brave Search is one of the most important tools for agents. Search providers make different choices about how they search, rank, and package results, and those choices change what the model reads and how it acts. We are expanding our benchmarking coverage to search APIs, so developers can pick a search provider on measured quality, cost, and speed. Each provider result pairs a search API provider with the same model, GPT-5.6 Luna (medium). The model runs inside Stirrup, our open-source agent harness, with tools for searching and fetching pages from the web - only the search provider behind the search tool changes. At launch, the leaderboard covers 11 results across 7 search providers, and we’ll keep expanding coverage as we look to provide the most accurate and comprehensive benchmarking of search providers for AI agent usage. Key elements of the Artificial Analysis Search Index: ➤ Three equally weighted benchmarks: the Search Index is the average of DeepSearchQA (900 broad research questions that need many searches, graded with an F1 score over answer items), BrowseComp (a 200-sample hard subset of facts that need multi-hop browsing), and AA-Omniscience (a 600 question private subset, balanced across 6 domains) ➤ Same agent, different search provider: the agent has 25 turns available to complete each task. Its web search tool returns the search provider's native response payload (with content modes standardized to snippets), with a maximum of 10 results and contamination sources filtered out ➤ Model-only baseline: we compare search agent results to the same model answering single-shot without tools, showing how much each provider lifts the model above its internal knowledge ➤ Cost and Time per Task: we aggregate the time and cost spent on both model inference and search. This is key - search APIs have different cost and latency structures, but these can be offset where they help an agent use fewer turns and save on costly language model inference Key results: ➤ Parallel, Exa, and Firecrawl have the strongest overall performance, with Artificial Analysis Search Index scores of 75, 74, and 73 respectively at launch ➤ All search providers tested substantially improve knowledge-based benchmark performance: the model only baseline scores 33 on the Search Index, while search-included provider results score between 65 and 75 ➤ Focused search results reduce spend on model inference: Parallel Search (advanced) search costs more per task than Parallel Search (basic) but less per task in total ($0.084 vs $0.11). Higher quality results cut the model's token use by over 40% in this case, more than offsetting increased search costs while reaching higher benchmark scores ➤ Fast search calls do not guarantee fast tasks: Parallel Search (turbo) has the fastest average search calls among Parallel's tiers (0.51s per query vs 1.03s for Parallel Search (basic)) but the basic tier scores higher on quality (73 vs 67) and the two land close on total time per task
@omarsar0@omarsar0AI 评分2323 @dexhorthy@dexhorthyAI 评分1414 Syncs and A/B Testing 200 Agents: 🦄 AI That Works #70 https://t.co/iRclUZOH9C
@kimmonismus@kimmonismusAI 评分4444 
@gabriel1@gabriel1AI 评分55