跳到正文
@omarsar0· @omarsar0 · X·· 26 天前AI 评分31
AI 导读

一项长时程基准显示,Astra 开局领先并保持优势长达 19 小时,但 Fable 5.1 在最后几小时追了上来,原因尚不明确。Qwen3.8 Max、Gemini 3.8 Flash 和 Grok 4.6 因更具成本效益而进入成本-性能帕累托前沿。

正文

This is a really neat benchmark. It's super interesting that Astra starts strong and holds the lead for up to 19 hours, but Fable 5.1 catches up in the final hours.

Not entirely clear why.

Different models solve long-horizon problems with different strategies. It's no surprise that we see different trends.

And then there is also the question of cost-performance efficiency. Qwen3.8 Max, Gemini 3.8 Flash, and Grok 4.6 are more economical for research, earning a place on the cost–performance Pareto frontier.

For AI research, things get complicated not because of the duration of the task but because agents (even with the best models) tend to stagnate due to low-quality exploration. Might be due to OOD. Or simply that agents are just simply not "creative" enough (well, at least not at human levels), which the top researchers are exceptional at, from years of experience building intuition and deep expertise.

The other question is: where exactly are research agents spending their compute? Are they using it efficiently? Earlier work from @intology found that coding agents mostly spend compute on hyperparameter tuning, rarely attempting algorithmic research. https://t.co/BmjrGu9mYB Not sure where the research stands on that today, but all of these are interesting research questions.

来源:@omarsar0 · x.com