跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-23AI 评分40
AI 导读

斯坦福与CMU提出Task Model Induction,将真实工作屏幕录制拆解为独立任务并重建为保留循环结构的目标树,恢复74.9%的实际发生内容,超过现有最佳摘要器两倍以上。基于该方法构建的技能在未见任务上表现提升30%。该方法在错误修复和死胡同等混乱会话中表现最弱,建议挖掘演示数据时输入目标结构而非原始记录。

正文

New Stanford + Carnegie paper.

Screen recordings of real work look like free training data for agents. They aren't, because nobody works in one clean sequence of steps.

People switch between unrelated tasks and redo the same fix until it passes. Most tools flatten that into one list, so the agent learns a workflow nobody performed.

Task Model Induction, from Stanford and CMU, first pulls a recording apart into the separate tasks that were running, then rebuilds each as a goal tree with its loops intact.

It works out the goals and the running order in separate passes and merges them, because asking one model for both at once flattens the detail.

That recovers 74.9% of what actually happened, more than double the best current summarizer, and skills built from it did 30% better on unseen tasks.

It's weakest where sessions get messy, on error fixing and dead ends.

So if you're mining demonstrations for agent skills, feed the goal structure, not the transcript.

– arxiv. org/abs/2608.20319

Title: "Inducing Task Models from Computer-Use Traces"

来源:@rohanpaul_ai · x.com