AI 导读
GPT-6 Astra 现已上线 SimpleBench。 Astra Pro 得分 86.5%,几乎与 Claude Fable 5.1 的 86.6% 持平。两者均超过该基准的人类基线 83.7%。 SimpleBench 测试日常推理,涵盖空间逻辑、社交情境和陷阱题。这正是其意义所在:AI 在那些看似简单、却屡屡难倒即使高能力模型的问题上正取得进展。 所以是的,看来我们得到了完好无损的通用(!)人工智能
正文
GPT-6 Astra is now on SimpleBench.
Astra Pro scores 86.5%, almost tied with Claude Fable 5.1 at 86.6%. Both beat the benchmark’s human baseline of 83.7%.
SimpleBench tests everyday reasoning, from spatial logic to social situations and trick questions. That’s what makes this meaningful: AI is making progress on the seemingly simple questions that have repeatedly tripped up even highly capable models.
So yeah, looks like we got intact artificial general (!) intelligence
来源:@kimmonismus · x.com