Prediction: universities will get priced out of AI evals work We just spent $50k yesterday benchmarking 40 long horizon coding tasks As we design harder tasks for agents, trajectories are getting longer and need more tokens. That means substantially more inference and compute. How is this economically feasible for academics working off university grants?
X:Karina Nguyen
@karinanguyen · X
切换来源
Karina@karinanguyenAI 评分4343引用Anand Kannappan@anandnk24
Karina@karinanguyenAI 评分3434
引用Thoughtful@thoughtfullabPostTrainBench v1.2 is out! A few updates: 1. Cloud GPU support. You can now run the benchmark with identical settings through Harbor + Modal using our new Harbor adapter. 2. New leaderboard leaders. Fable 5.1 takes #1 at 44.6%, followed by Opus 5.5 at 43.8% and GPT-6 (Astra) at 41.9%. 3. Evaluation fixes. Removed BFCL, fixed HumanEval and remote-code scoring, added averaging across multiple seeds, and switched contamination checks to majority vote.
Karina@karinanguyenAI 评分5252引用Google DeepMind@GoogleDeepMindIntroducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
Karina@karinanguyenAI 评分3131
引用Rehan Sheikh@rehan_sheiinteractive media will be huge in the next year or two! this is incredible
Karina@karinanguyenAI 评分3131推出 ACTx486,一种全新交互式媒介的研究演示。 如果你能和任何视频对话,想问什么就问什么,会怎样? 我们的系统将一档现有播客变成了能倾听、能回应、能适应的东西。 研究来自 @jakubzeg:

Karina@karinanguyenAI 评分5656引用Epoch AI@EpochAIResearchIntroducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.

