跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-18AI 评分31
AI 导读

AI 编程智能体可以在一个任务上花费数小时,却没有对时间流逝的校准感知。 因此,长时程评估可能需要直接测量持续时间遵循能力,而不是将持续的任务表现视为智能体知道何时停止的证据。

正文

AI coding agents can spend hours on a task without a calibrated sense of time passing.

Long-horizon evaluations may therefore need to measure duration-following directly instead of treating sustained task performance as evidence that an agent knows when to stop. https://t.co/tH9yBn0Nfi https://t.co/i092jyyyPv

来源:@rohanpaul_ai · x.com