X:Elvis Saravia
@omarsar0 · X
切换来源
@omarsar0@omarsar0AI 评分1212 @omarsar0@omarsar0AI 评分6464 引用@opencode@opencodeOx Alpha (stealth model) is free for the next week - 1M Context - Multi-modal - Zero Data Retention Generous rate limits, near unlimited usage We have capacity for 100T tokens per day, lets see what you can do
@omarsar0@omarsar0精选AI 评分6767 引用@deepseek_ai@deepseek_aiDeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
推荐理由:官方基准对比显示该实验模型在多模态智能体任务上接近 Opus-4.8,读者可据此了解多模态智能体的当前水平。
@omarsar0@omarsar0AI 评分4444 
@omarsar0@omarsar0AI 评分4848 现代AI智能体将经验积累在harness(提示词、记忆、工具、技能、路由规则)中而非模型权重里,更新任一harness组件都可能导致原有可靠行为失效,论文将此定义为harness级遗忘并给出度量方法。

@omarsar0@omarsar0AI 评分2727 @omarsar0@omarsar0AI 评分2121 @omarsar0@omarsar0AI 评分3232
@omarsar0@omarsar0AI 评分5050 
@omarsar0@omarsar0AI 评分3737 @omarsar0@omarsar0AI 评分2323 现实生活比封闭基准测试更难。 用户意图逐步显现。任务分阶段展开。外部条件不断变化。 在模拟的咖啡电商场景中,智能体跟踪库存、采购、履约和客户需求——并随着条件变化修订其计划。

@omarsar0@omarsar0AI 评分3535 
@omarsar0@omarsar0AI 评分2929 
@omarsar0@omarsar0精选AI 评分7171 
推荐理由:推文展示自评审在训练循环之外同样奏效,并给出 IMO 2026 官方认证的 42/42 成绩作为依据。
@omarsar0@omarsar0AI 评分2222 
@omarsar0@omarsar0AI 评分3535 
@omarsar0@omarsar0AI 评分2424 TEMPO 就是被提出的答案。 它把一条长轨迹划分为多个宏观步骤。在每一步,同一个模型从演员切换为评论家。 评论家对当前状态进行推理,调用工具,并在整个任务完成之前估算预期的剩余回报。
@omarsar0@omarsar0AI 评分2020 核心问题是信用分配。 一次长时程 rollout 可能耗时数十小时,而单一的终端奖励必须归因到数千次交互上。 这正是 GRPO 这类无价值(value-free)RL 方法开始吃力的地方。
@omarsar0@omarsar0AI 评分1818 我很享受深入研究这个循环是如何构建的。如果你正把智能体从原型推向生产,这个仓库值得你花时间。 给这个仓库点个 Star。https://t.co/Vy7Wuu1Boz
@omarsar0@omarsar0精选AI 评分6666 
推荐理由:作者给出同一基准下的成本对比和多模型路由配置,便于判断自托管 agent harness 的实际取舍。
@omarsar0@omarsar0AI 评分2727 @omarsar0@omarsar0AI 评分6060 引用@Replit@ReplitReplit Free Mode, powered by @OpenAI GPT-5.6 Luna. Let’s make intelligence accessible to everyone. https://t.co/UDcrYl5HZL
@omarsar0@omarsar0AI 评分2222
@omarsar0@omarsar0AI 评分3535 
@omarsar0@omarsar0AI 评分2727 @omarsar0@omarsar0AI 评分5151 微软提出 Agent Lightning v1.0,用约 3,500 行代码通过 LLM endpoint proxy 把任意 agent harness 接入强化学习。

@omarsar0@omarsar0AI 评分55 @omarsar0@omarsar0AI 评分2929 @omarsar0@omarsar0AI 评分2727 @omarsar0@omarsar0AI 评分3636 我看好简单的智能体框架。很喜欢 Vercel 的这个想法。 fx 是一个新的编码智能体框架,用于模型基准测试、沙箱、评估和 gym。https://t.co/7YiOPGzyow
@omarsar0@omarsar0AI 评分4444 
@omarsar0@omarsar0AI 评分1414 @omarsar0@omarsar0AI 评分1111 太对了!借助 AI 智能体做研究的方式太多了。有太多研究问题值得探索。有太多实验要跑。有太多要学。有太多知识要构建。 研究是我 tokenmaxxing 最多的地方。
@omarsar0@omarsar0AI 评分5656 引用@ArtificialAnlys@ArtificialAnlysAnnouncing the Artificial Analysis Search Index, benchmarking how search API providers perform on quality, cost, and speed when used by an agent. We are initiating coverage with Parallel, Exa, Firecrawl, You (dot) com, Tavily, Keenable, and Brave Search is one of the most important tools for agents. Search providers make different choices about how they search, rank, and package results, and those choices change what the model reads and how it acts. We are expanding our benchmarking coverage to search APIs, so developers can pick a search provider on measured quality, cost, and speed. Each provider result pairs a search API provider with the same model, GPT-5.6 Luna (medium). The model runs inside Stirrup, our open-source agent harness, with tools for searching and fetching pages from the web - only the search provider behind the search tool changes. At launch, the leaderboard covers 11 results across 7 search providers, and we’ll keep expanding coverage as we look to provide the most accurate and comprehensive benchmarking of search providers for AI agent usage. Key elements of the Artificial Analysis Search Index: ➤ Three equally weighted benchmarks: the Search Index is the average of DeepSearchQA (900 broad research questions that need many searches, graded with an F1 score over answer items), BrowseComp (a 200-sample hard subset of facts that need multi-hop browsing), and AA-Omniscience (a 600 question private subset, balanced across 6 domains) ➤ Same agent, different search provider: the agent has 25 turns available to complete each task. Its web search tool returns the search provider's native response payload (with content modes standardized to snippets), with a maximum of 10 results and contamination sources filtered out ➤ Model-only baseline: we compare search agent results to the same model answering single-shot without tools, showing how much each provider lifts the model above its internal knowledge ➤ Cost and Time per Task: we aggregate the time and cost spent on both model inference and search. This is key - search APIs have different cost and latency structures, but these can be offset where they help an agent use fewer turns and save on costly language model inference Key results: ➤ Parallel, Exa, and Firecrawl have the strongest overall performance, with Artificial Analysis Search Index scores of 75, 74, and 73 respectively at launch ➤ All search providers tested substantially improve knowledge-based benchmark performance: the model only baseline scores 33 on the Search Index, while search-included provider results score between 65 and 75 ➤ Focused search results reduce spend on model inference: Parallel Search (advanced) search costs more per task than Parallel Search (basic) but less per task in total ($0.084 vs $0.11). Higher quality results cut the model's token use by over 40% in this case, more than offsetting increased search costs while reaching higher benchmark scores ➤ Fast search calls do not guarantee fast tasks: Parallel Search (turbo) has the fastest average search calls among Parallel's tiers (0.51s per query vs 1.03s for Parallel Search (basic)) but the basic tier scores higher on quality (73 vs 67) and the two land close on total time per task
@omarsar0@omarsar0AI 评分2323 @omarsar0@omarsar0AI 评分3535
@omarsar0@omarsar0AI 评分44 在我们的学院追踪更多热门 AI 论文:https://t.co/qF2b2uvKf1 论文:https://t.co/RsfC1mckzH
@omarsar0@omarsar0精选AI 评分6868 
推荐理由:用 1902 次运行把智能体协作画成时序网络,读者能看到协调开销如何随团队规模与任务形态变化。
@omarsar0@omarsar0AI 评分2828 推荐资源。这是我见过的最大的智能体技能数据集。非常适合为你的智能体挖掘酷炫的想法和模式。https://t.co/z00rGoYPaT
@omarsar0@omarsar0AI 评分2121 