确实。我想补充的是,如果你知道如何在此基础上构建一个好的 eval,你很快就能在你所在的领域/任务中处于前沿。这就是 eval 能为你解锁的那种优势。 也要学会构建好的 eval。这值得。
You can basically work with agents to write markdown skills and put it on a cron job and do almost every useful type of knowledge work now
@omarsar0 · X
确实。我想补充的是,如果你知道如何在此基础上构建一个好的 eval,你很快就能在你所在的领域/任务中处于前沿。这就是 eval 能为你解锁的那种优势。 也要学会构建好的 eval。这值得。
You can basically work with agents to write markdown skills and put it on a cron job and do almost every useful type of knowledge work now
AutomationBench Verified is out. We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier. • Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed. • We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade • Where verifiers were too strict, pass rate went from 18.8% to 43.8% • Where they were too lenient, it dropped from 60.2% to 49.7% Audit and dataset links below.
Agent Browsers are live on Gumloop today so your agents can reach the work that doesn't have an MCP or API > Secure credential storage & 1Password integration > Session replays to see your agents work > Stealth, proxied sessions All completely model agnostic Try Browser today.
AutomationBench Verified is out. We went through all 600 public tasks in @Zapier's AutomationBench and checked every verifier. • Agents flagged 323 verifiers as suspicious, human review confirmed 206 real bugs, all 206 are fixed. • We replayed 1,235 Kimi K3 runs on the old and fixed verifiers and 344 of them (27.9%) got a different grade • Where verifiers were too strict, pass rate went from 18.8% to 43.8% • Where they were too lenient, it dropped from 60.2% to 49.7% Audit and dataset links below.
Introducing AgentID by @agentmail, "Sign in with Google" for AI Agents Millions of agents already use AgentMail. AgentID lets them sign in to your app. Integrate with the prompt below, reply with a screenshot, and we'll give you 3 months free of any AgentMail plan.
Overmind turns anyone into an AI lab. • Builds a context graph from code and traces • Curates training and eval datasets • Evaluates prompts and models • Trains smaller, specialized models https://github.com/overmind-core/overmind
Your AI is vaporware without deep integrations with your customers’ systems of record. Introducing @WithAmpersand: integration infrastructure for enterprise agents. We power 11x, Orb, and Square's ability to take agentic action in systems of record like Salesforce, SAP, NetSuite, and Workday. Ask anyone serious about building AI and they’ll tell you: - The SaaSpocalypse didn’t happen. The world depends on CRMs, ERPs, HRISs, and ITSMs. - Your customers customized their deployments beyond recognition. - Systems of record companies are basically monopolies. They never had to make their APIs and MCPs user-friendly. - Docs don't explain half the weird edge cases you'll hit. - One bad write can blow up your pilot or renewal. - Once you finally get it working, someone changes a field and it breaks again. The world's data lives in structured databases. AI needs a translation layer to work with it. So, for your AI to truly transform the way enterprises work, you'll need deep integrations built for: - scoped permissions - bi-directional actions - real-time speed for agents - custom objects, fields, and workflows for each of your customers We built Ampersand to power the future of software. AI didn't trivialize writing integrations. It made them the critical path. I believe that deeply, so I didn't make a launch video about Ampersand. Instead, it's the best builders I know explaining just how big of a challenge this is.
Meet Mistral Large 4, aka Le Chonk. • 1T parameters, natively multimodal. 49B active. It is the best open weights model from US or Europe on aggregated benchmarks. • State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding. • Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure. • Available to all via API today. Working with cybersecurity partners privately. Open weights release end of October.
Elvis Saravia 发布与自己的研究智能体合写的教程,介绍用 TypeSafe AI 的决策模型 Jev 做 LLM-as-a-Judge 智能体评测。
https://x.com/i/article/2107258465630507009
Everyone is sick of generic AI, including us. Gamma was the first AI presentation platform to reach real scale. But now that AI tools are everywhere, everything looks the same. Earlier this summer, we came to an uncomfortable conclusion: without a step change in visual variety, Gamma’s output would look and feel too similar to all the other AI tools. The entire point of a presentation is to help someone "get" your idea, to see it as unique and do something with it. We started over. Today we're shipping Gamma 5, the biggest product overhaul in our company's history. We rebuilt Gamma to create presentations that actually look like you and your brand. Your teams can design anything they dream up, following your brand's aesthetic or something entirely new, with every frontier and image model under the hood. And the same goes for your docs, social assets, and graphics. We revamped our agent, design tools, editing, import, export, connectors, and more. Get your ideas out there.
Introducing flow-1, our new model trained with RL to find errors in agent traces. It matches GPT-6-sol in trace intelligence while being 23x cheaper. It also costs 25% less to run than GPT-6-luna. flow-1 finally makes it possible to monitor and understand every agent run, without sampling. 1/6
作者用 @pidotdev 的 Pi Durable 快速搭起一个个人智能体,作为主动式个人智能体的实验场,之后计划开源。
You can’t safely test an enterprise agent on a real company’s data. So we built a company for it to work in. Era is live today, and it’s free. It generates a complete simulated enterprise that behaves like a real one across Salesforce, Slack, Jira, Zendesk, Gong, Deel and more, along with cloud databases and storage. Agents interact with it through live MCP and API interfaces. People leave. Deals change. Records get duplicated. Permissions differ across systems. And because Era generated the company, it knows the exact ground truth. Test, benchmark and improve agents against realistic enterprise workloads, use the failures for targeted post-training, then rerun the same environment to measure the impact. Huge thanks to our research partners @NVIDIA, @Decart, @Composio, @openlayerco, @Deel, @Eragon and @Plurai, with more coming soon.
推荐。LLM 智能体喜欢结构,所以语料库能改善智能体搜索并不意外。
Banger paper from Microsoft and colleagues. If you run agents that search a large document collection, this one is worth your time. (bookmark it) They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it. The original documents stay in place. The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again. Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%. It also beats an LLM Wiki layer and three other navigation layers. The map can be built without LLM calls and updated as new documents arrive. Paper: https://arxiv.org/abs/2609.37226 Chat with Paper: https://academy.dair.ai/papers/follow-the-entities-a-corpus-map-for-agentic-search-2609.37226
Today, Nolla Health became the first organization in the U.S. (and possibly the world) to receive regulatory approval for an AI to issue initial prescriptions. This makes Nolla the first ever actual end-to-end AI doctor.
there’s something quite awkward about all the personal agents in the current hype cycle muse, grok bot, instinct, dots, and whatever google, anthropic will come up with none of them is “mine” i’d be trusting a vendor for some of the most sensitive data about me and bet on them handling it with extreme responsibility. what if they have a data leakage incident? what if they hired a wrong employee? what if they simply break bad? i’d also be betting on their model. right now i’ve set up a lot of my stuff in grok bot, but what if their model falls behind? what if another model becomes 10x better on pure intelligence? then i’ll either miss out on a better assistant, or have to bite the bullet and do a full migration i’d also have to bet on their product. they may not build the features i want. they may not allow me to customize enough. and one day my assistant may show me ads that feels like way too much betting than what i’d be comfortable with, if i truly want literally everything in my life to go through and be taken care of by an assistant with all that considered, i’m more bullish on open source personal agents that run on people’s own computers. but i think these open source projects have to break out of the assumption that their users are developers and they can just throw a repo at them and ask them to launch a terminal a well polished, fool proof, community maintained, vendor agnostic, free personal agent that everyone can run by themselves without locking into a SaaS - that’s what i think many people will need
Everyone got a coding agent. Nobody got a QA engineer. Until today. Meet Ship, your Autonomous Quality Engineer. It tests deployed PRs, reproduces bugs from Slack and Linear, and hands Claude or Codex the context to fix them. Try it free: https://ship.contextqa.com
UT Austin 在 SWE-bench Verified 和 Terminal-Bench 0.0 上运行近 35000 次编码 agent 实验,分别变化压缩方式、触发时机和删除量三个决策。
Elvis Saravia 在引用 Karpathy 关于用 ASD-STE100 写作、图表、网页和讲解视频等方式理解 LLM 输出的帖子时,分享了自己数月来实验的通用人机协作界面。
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
I was a terminal person for 15+ years. I loved Vim, knew all the shortcuts, and a black terminal made me feel like I was a cool hacker in The Matrix. But terminals assume humans operate computers directly via files, commands, processes. Agents changed all that. You just say your intent. The agent operates the machine, using the same programming languages and Linux commands we spent years learning. I’m still glad I learned computer systems before the AI era though. Understanding what sits beneath the abstraction makes you a better systems thinker. That’s a moat in the AI era.
Google 发布 VeriHarness 论文,把同一基座模型变成智能体验证器,用于长程任务结果校验。
I haven’t touched Claude Code or Codex CLI in a while. The terminal era is over imo. It's the wrong interface for coding agents. Tabs are ephemeral, but context is persistent, and managing 30 tabs is pure cognitive overhead. I don’t really need an IDE like Cursor either. I rarely navigate the whole codebase anymore. The new primitive is the agent, not the file. (Codex desktop app is the best agentic UI for now. But we’re still early.)
Banger paper from Meta Superintelligence Labs. They find something super interesting and unexpected. (bookmark it) Base models with a light harness often solve more agentic tasks than their RL post-trained versions when both get enough samples. Post-trained models win on pass@1. At large K, base models frequently solve tasks the post-trained ones never solve on BFCL v4 multi-turn, ACEBench, and WebShop. This is because post-training pushes each task toward always solved or never solved. Consistency goes up, and coverage goes down. The authors call the lost test-time scalability the Sharpening Tax. Across 42 base and post-trained pairs, it shows up in most settings, grows with model size, and can be estimated from a few rollouts. Their fix, PTGS, sets the sampling temperature per prompt from its estimated difficulty during RL. It pays a smaller tax and also raises pass@1. Paper: https://academy.dair.ai/papers/sharpening-tax-in-post-training-2610.01509
This is Kled V3. We've solved data collection for artificial general intelligence. Any consumer dataset that can exist, can now be collected from physical reality in under 72 hours. All powered by the largest and most comprehensive data application layer on the planet. (Thread)
Google Research 发布 Cogentic,一个基于 Gemini 的多智能体自动证明发现框架,面向无专家提示的理论计算机科学开放问题。
People of Pi: We've shipped Pi 1.0 with Pi Durable. Go make them yours. https://earendil.com/posts/pi-1-0/
Excited to announce Volantis's $88M Series A. We are solving Al's memory bottleneck by using optics, enabling chips with huge amounts of fast & cheap memory. By boosting both the memory bandwidth and capacity per chip by orders of magnitude, we enable ultra-fast inference (up to 10,000 tps/user) for large models (>10T) - with low $/tok to boot. Initially, this will enable insanely fast agents - think coding agents that finish in minutes or even seconds instead of hours. More excitingly, optics is a fundamentally scalable way to increase memory systems. Not 2X/year, but by orders of magnitude across new generations. This will enable a structurally new Al industry, including restarting scaling laws, holding entire repos in context windows & more. Our team has pioneered many core semiconductor technologies: the 1st CoWoS product, early HBM, the 1st silicon photonics CPO systems, the 1st high volume tunable VCSELs, the 1st processors to directly communicate using light & more. We’ve already sent data >10× farther than equally tiny electrical wires inside a chip package. Our next iteration is already taped out and targets world-record bandwidth density over relevant distances, read more: https://volantissemi.ai/news-insights/our-88m-series-a-demolishing-the-memory-wall-with-photonics-post
Your AI voice sounds human. So why can't it say your product's name? A great AI voice reads "Porsche Taycan" as TAY-can. Porsche says TIE-kahn. It guessed from the spelling, and nobody caught it, because nobody listens to line 1,200. Today we're launching Onepin: the production step after text-to-speech. It checks every line of voiceover before it ships using the voices you already work with. Onepin can: ➤ Check people's and product names against a 4-million-word pronunciation dictionary ➤ Spell out prices and dates before the voice speaks ➤ Score every line of audio for naturalness, clarity and word accuracy ➤ Fix the one wrong word in the same voice, without re-rendering the take Works with your voice subscription on @ElevenLabs, @OpenAI, @Google and 30+ more. No phonetic spellings to type. No re-rolls. No switching providers. Free to start, no credit card required. Hear the before and after in the thread ⬇️
Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸