

@swyx · X
顺便别忘了"还有一件事……" 他们做 dots 已经有一阵子了
Introducing dots, powered by GPT-6 Astra. Remarkably capable, always-on agents built to handle everything.
我要为 @aidotengineer nyc 换掉我的标语 提前飞过去参加 NY Comic Con。11 天后见!
it took not even 30 seconds after leaving my hotel in nyc to see something that i almost never saw in 3 months in sf: i saw a hot girl
把别人的 logo 裁掉、也不注明出处,就做一个简单的 YouTube 视频,这正常吗?或者说在专业科技媒体里这是怎么运作的?


TypeSafe AI is talking with investors about raising $1 billion or more as enthusiasm builds around its Jev model and potential valuations above $10 billion. Read more: https://thein.fo/4hcSXC4
Flow has raised a $50M Series B at a $750M valuation, co-led by Antonio Gracias (Valor) and Gavin Baker (Atreides) with Sequoia Capital, Roelof Botha (SpaceX, Block), and more. When I became a mechanical engineer, I wanted to invent. In reality, 90% of the day to day execution work (CAD, excel, sim) was digital manual labour. That’s about to flip. The engineer’s role will change more profoundly than at any point since the invention of CAD, with humans focusing on invention, architecture, tradeoffs and judgement calls, and agents taking the grunt work. Anduril, Joby, Stoke Space, Rivian & many more - Flow is now the default platform for requirements and verification for frontier hardware teams. But it’s not just next-gen anymore. Industry giants including General Motors PPU and Volkswagen & Rivian’s joint venture (RV Tech) are reshaping their core development practices with Flow’s AI. The revolution is just starting. Models are getting better quickly at CAD, analysis, simulation, and tool use. In the next year, AI will move from integrating the work to doing the work. One day, this shift won't just enable us to iterate faster, it will enable us to design a class of products that we could not have dreamt of before. We’re hiring.
[AINews] The Future of Latent Space https://www.latent.space/p/ainews-the-future-of-latent-space - Plans for AINews v3 - Plans for a new home! - We are open for business - and @supabase are our first sponsors!
更多会议应该采用这种做法。太多高价值的时间被毫无思考地浪费掉了。
512GB RAM Studios. Apple was good to us. 🦞 https://t.co/NyvtNH6lRa
哈哈哈哈哈 非技术人士简直被烤得外焦里嫩 你能想象在零上下文、零推理、零内部世界模型的情况下报道 AI 吗。那一定快乐得不得了,一切都太棒了,表面价值就够了, 活在这个时代真好
With Codex, @asana finished a frontend test migration from Enzyme to React Testing Library in two calendar weeks—a project expected to take five more years. https://t.co/WcTLX1JZ6v
如果你从这个角度来看 OpenRouter,这是一个正面的 UBB 解读 https://t.co/rOtLC5UkC5
祝贺我们的朋友 @ona_hq 加入 @openai! 在这里看他们的分享,了解 Codex 下一步的 alpha 内幕 👀
Three layers you need to run agent swarms at scale: - Runtime: solved. - Orchestration and triggers: solved. - Coordination (how agents pick up tasks from each other, verify they have cleared a stage, and proceed): not solved. Stripe called theirs Minions. RAMP called theirs Inspect. Both are internal infrastructure for running fleets of background agents. Both built from scratch. @loujaybee says this shouldn't keep happening. GitHub is a poor coordination layer for agents. Noisy, designed for humans, not built for hundreds of parallel pull requests. Lou's fix: a CLI gateway that any local coding agent can invoke to check whether it has cleared its current SDLC stage and can proceed. piped.video/watch?v=5Sui_OnS…
You have Claude Fable for only a few days. Here's how to make the most of it. Introducing /improve: use your most capable model to audit your codebase and write plans for cheaper models to execute later. Studies your code, figures out bugs, perf, tech debt, missing tests, what to build and writes plans any agent can run.
just finished rerunning FC Diamond on my historical charts. none of the official tables/charts are capturing the degree of takeoff. nitter.net/karpathy/status/206440… its this same chart all the way down difficulty classes (below) breaks every curve fit because Fable is a diffferent CLASS of model, with beeeeeg model smell.
Anthropic is partnering with @SpaceX to run models on their compute, exciting!



It's finally out!!! @METR_Evals found that more than half of SWEBench results is unmergeable slop. FrontierCode represents over 1000+ hours of maintainer validated software engineering work most frontier models cannot yet solve, much less solve with high quality. Cog had IOI Gold medalists and top code maintainers Look At The Data — FrontierCode includes 3000+ rubrics covering code quality and anticheat reward hacking plaguing other benchmarks. FC Diamond is so hard that Opus 4.8 scores 13.8%. Three eras of AI coding : Three eras of benchmarks 2021 • Autocomplete : HumanEval 2023 • Passing Tests: SWEBench, TerminalBench 2026 • Maintainable Code: FrontierCode to me the most beautiful chart when I requested a special historical run into all extant old models, the data was finding that the easiest third of FC tasks (in FC Extended) were rapidlly and suddenly solved over late 2025 - Opus almost doubled from a 41% pass rate to 74% in 4 months. This describes the "WTF happened in Dec 2025" vibe shift that a lot of folks from @dhh to @karpathy have called out: it is the difference between getting 95% success in 2 rerolls vs 6, making it finally feasible to go up the next layer of abstraction in agentic coding, eg @GeoffreyHuntley's ralph loops or @bcherny's /goals or @steipete's "loops that prompt your agents" without fearing too much that things go off the rails. My guess: as AI accelerates from here, each FrontierCode tier will saturate in sequence, hopefully ~annually. I've already asked the team to prepare FrontierCode 2027.... The old mountains will be destroyed. Their rubble becomes regolith. And from that regolith, the next model forest grows. Circle of life.
This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *qualitatively* also, this is a major-version-bump-deserving step change forward (imo of the same order as Claude 4.5 was in November), peaking especially for long problem-solving sessions on very difficult problems. You can give it a lot more ambitious tasks than what you're used to, the model "gets it" and it will just go, and it's never felt this tempting to stop looking at the code at all (but don't do this in prod!). The model still has quirks that people will run into and the safeguards are configured to be a little too trigger happy for launch, which can hopefully be tuned over time. I feel a lot of things changing as working software increasingly comes out on a tap. The Jevon's paradox kicks in and I feel my own demand for software growing substantially. You can ask for anything - explainers, visualizers, dashboards, bespoke single-use apps (e.g. a full wandb that is hyper-specific just for your project), you can 10X your test suite, auto-optimize code, run giant research projects with custom HTML for the results, anything! "Free your mind" (Matrix ref). Really looking forward to all the things people build!
推荐理由:转发的评测者认为官方榜单未体现 Fable 5 的进步幅度,可作为判断这次模型跃迁的定性参考。
Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. Its capabilities exceed those of any model we’ve ever made generally available. Video