AI 导读
对 AI 系统进行基准测试是一个持续的过程,与模型共同演进。新的基准测试用新兴问题挑战 AI 能力,从而塑造研究过程的方向和反馈信号。随后它们随着模型的进步而调整,瞄准 AI 与人类智能之间的残余差距。 我们仍在开发 ARC-AGI-4,这是我们在今年早些时候发布 ARC-AGI-3 之后开始开发的。它将于 2027 年 Q1 推出。我们认为它会非常特别。
正文
Benchmarking AI systems is a continual process that co-evolves with the models. New benchmarks challenge AI capabilities with emerging questions to shape the directions and feedback signal of the research process. Then they adapt as models progress, targeting the residual between AI and human intelligence.
We are still working on ARC-AGI-4, which we started developing after releasing ARC-AGI-3 earlier this year. It is coming Q1 2027. We think it's going to be really special.
来源:@fchollet · x.com