Introducing flow-1, our new model trained with RL to find errors in agent traces. It matches GPT-6-sol in trace intelligence while being 23x cheaper. It also costs 25% less to run than GPT-6-luna. flow-1 finally makes it possible to monitor and understand every agent run, without sampling. 1/6
X:Elvis Saravia
@omarsar0 · X
切换来源
elvis@omarsar0AI 评分4242引用Robert@skull8888888888
elvis@omarsar0AI 评分5151作者用 @pidotdev 的 Pi Durable 快速搭起一个个人智能体,作为主动式个人智能体的实验场,之后计划开源。

elvis@omarsar0AI 评分5656引用Ofir Ehrlich@OfirEhrlichYou can’t safely test an enterprise agent on a real company’s data. So we built a company for it to work in. Era is live today, and it’s free. It generates a complete simulated enterprise that behaves like a real one across Salesforce, Slack, Jira, Zendesk, Gong, Deel and more, along with cloud databases and storage. Agents interact with it through live MCP and API interfaces. People leave. Deals change. Records get duplicated. Permissions differ across systems. And because Era generated the company, it knows the exact ground truth. Test, benchmark and improve agents against realistic enterprise workloads, use the failures for targeted post-training, then rerun the same environment to measure the impact. Huge thanks to our research partners @NVIDIA, @Decart, @Composio, @openlayerco, @Deel, @Eragon and @Plurai, with more coming soon.
elvis@omarsar0AI 评分4545推荐。LLM 智能体喜欢结构,所以语料库能改善智能体搜索并不意外。
引用DAIR.AI@dair_aiBanger paper from Microsoft and colleagues. If you run agents that search a large document collection, this one is worth your time. (bookmark it) They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it. The original documents stay in place. The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again. Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%. It also beats an LLM Wiki layer and three other navigation layers. The map can be built without LLM calls and updated as new documents arrive. Paper: https://arxiv.org/abs/2609.37226 Chat with Paper: https://academy.dair.ai/papers/follow-the-entities-a-corpus-map-for-agentic-search-2609.37226
elvis@omarsar0AI 评分6464引用Luis Wenus@luiswenusToday, Nolla Health became the first organization in the U.S. (and possibly the world) to receive regulatory approval for an AI to issue initial prescriptions. This makes Nolla the first ever actual end-to-end AI doctor.
elvis@omarsar0AI 评分5050
elvis@omarsar0AI 评分4545引用Kun Chen@kunchenguidthere’s something quite awkward about all the personal agents in the current hype cycle muse, grok bot, instinct, dots, and whatever google, anthropic will come up with none of them is “mine” i’d be trusting a vendor for some of the most sensitive data about me and bet on them handling it with extreme responsibility. what if they have a data leakage incident? what if they hired a wrong employee? what if they simply break bad? i’d also be betting on their model. right now i’ve set up a lot of my stuff in grok bot, but what if their model falls behind? what if another model becomes 10x better on pure intelligence? then i’ll either miss out on a better assistant, or have to bite the bullet and do a full migration i’d also have to bet on their product. they may not build the features i want. they may not allow me to customize enough. and one day my assistant may show me ads that feels like way too much betting than what i’d be comfortable with, if i truly want literally everything in my life to go through and be taken care of by an assistant with all that considered, i’m more bullish on open source personal agents that run on people’s own computers. but i think these open source projects have to break out of the assumption that their users are developers and they can just throw a repo at them and ask them to launch a terminal a well polished, fool proof, community maintained, vendor agnostic, free personal agent that everyone can run by themselves without locking into a SaaS - that’s what i think many people will need
elvis@omarsar0AI 评分4545引用Deep Barot@deepcabinwalaEveryone got a coding agent. Nobody got a QA engineer. Until today. Meet Ship, your Autonomous Quality Engineer. It tests deployed PRs, reproduces bugs from Slack and Linear, and hands Claude or Codex the context to fix them. Try it free: https://ship.contextqa.com
elvis@omarsar0AI 评分6464UT Austin 在 SWE-bench Verified 和 Terminal-Bench 0.0 上运行近 35000 次编码 agent 实验,分别变化压缩方式、触发时机和删除量三个决策。

elvis@omarsar0AI 评分5959Elvis Saravia 在引用 Karpathy 关于用 ASD-STE100 写作、图表、网页和讲解视频等方式理解 LLM 输出的帖子时,分享了自己数月来实验的通用人机协作界面。
引用Andrej Karpathy@karpathyWe'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
elvis@omarsar0AI 评分3232引用Yuchen Jin@Yuchenj_UWI was a terminal person for 15+ years. I loved Vim, knew all the shortcuts, and a black terminal made me feel like I was a cool hacker in The Matrix. But terminals assume humans operate computers directly via files, commands, processes. Agents changed all that. You just say your intent. The agent operates the machine, using the same programming languages and Linux commands we spent years learning. I’m still glad I learned computer systems before the AI era though. Understanding what sits beneath the abstraction makes you a better systems thinker. That’s a moat in the AI era.
elvis@omarsar0AI 评分5353Google 发布 VeriHarness 论文,把同一基座模型变成智能体验证器,用于长程任务结果校验。

elvis@omarsar0AI 评分4949
elvis@omarsar0AI 评分3232引用Yuchen Jin@Yuchenj_UWI haven’t touched Claude Code or Codex CLI in a while. The terminal era is over imo. It's the wrong interface for coding agents. Tabs are ephemeral, but context is persistent, and managing 30 tabs is pure cognitive overhead. I don’t really need an IDE like Cursor either. I rarely navigate the whole codebase anymore. The new primitive is the agent, not the file. (Codex desktop app is the best agentic UI for now. But we’re still early.)
elvis@omarsar0AI 评分5959
elvis@omarsar0AI 评分3535
elvis@omarsar0AI 评分5959引用DAIR.AI@dair_aiBanger paper from Meta Superintelligence Labs. They find something super interesting and unexpected. (bookmark it) Base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples. Post-trained models win on pass@1. At large K, base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop. This is because post-training pushes each task toward always solved or never solved. Consistency goes up, and coverage goes down. The authors call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it shows up in most settings, grows with model size, and can be estimated from a few rollouts. Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass@1. Paper: https://academy.dair.ai/papers/sharpening-tax-in-post-training-2610.01509
elvis@omarsar0AI 评分4343
elvis@omarsar0AI 评分4040
elvis@omarsar0AI 评分4949
elvis@omarsar0AI 评分3838引用Kled AI@useKledThis is Kled V3. We've solved data collection for artificial general intelligence. Any consumer dataset that can exist, can now be collected from physical reality in under 72 hours. All powered by the largest and most comprehensive data application layer on the planet. (Thread)
elvis@omarsar0AI 评分6161Google Research 发布 Cogentic,一个基于 Gemini 的多智能体自动证明发现框架,面向无专家提示的理论计算机科学开放问题。

elvis@omarsar0AI 评分4444引用Pi@pidotdevPeople of Pi: We've shipped Pi 1.0 with Pi Durable. Go make them yours. https://earendil.com/posts/pi-1-0/
elvis@omarsar0AI 评分4545
elvis@omarsar0AI 评分4848
elvis@omarsar0AI 评分5050引用Tapa Ghosh@semiDLExcited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: https://volantissemi.ai/news-insights/our-88m-series-a-demolishing-the-memory-wall-with-photonics-post
elvis@omarsar0AI 评分4848引用Soohyun Bae@RealSoohyunBaeYour AI voice sounds human. So why can't it say your product's name? A great AI voice reads "Porsche Taycan" as TAY-can. Porsche says TIE-kahn. It guessed from the spelling, and nobody caught it, because nobody listens to line 1,200. Today we're launching Onepin: the production step after text-to-speech. It checks every line of voiceover before it ships using the voices you already work with. Onepin can: ➤ Check people's and product names against a 4-million-word pronunciation dictionary ➤ Spell out prices and dates before the voice speaks ➤ Score every line of audio for naturalness, clarity and word accuracy ➤ Fix the one wrong word in the same voice, without re-rendering the take Works with your voice subscription on @ElevenLabs, @OpenAI, @Google and 30+ more. No phonetic spellings to type. No re-rolls. No switching providers. Free to start, no credit card required. Hear the before and after in the thread ⬇️
elvis@omarsar0AI 评分4848
引用David Stout@DavidstoutHalf a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸
elvis@omarsar0AI 评分5858引用Tavus@tavusIntroducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human. Previous systems have had a pass rate <3%. It is #1 on NVIDIA's benchmark for full-duplex AI video. It’s the first Human Interaction Model (HIM).
elvis@omarsar0AI 评分5757引用Mac Liu@themacliuI’m excited to announce that @arceuslegal is launching with $17M in funding, led by @greycroftvc, with participation from @craft_ventures, @spc, and others. As a founder, I always hated how helpless I felt working with law firms. I went through four or five different firms and somehow the experience was always the same. I’d be waiting on something important to our business with no idea when I’d hear back. I’d have to re-explain our business over and over again. And I dreaded jumping on calls because I knew every minute was costing me money. We started Arceus because we believe every business deserves a better law firm. One that moves faster, costs less, and puts the client first. And we’re just getting started. ↓
elvis@omarsar0AI 评分5858Meta 及合作机构发布 Context Language Models(CLMs)论文,把上下文作为文件让模型用 Bash 自由读写编辑,自行决定保留、重写或删除内容。
