跳到正文
@rohanpaul_ai· @rohanpaul_ai · X·· 2026-08-24AI 评分52
AI 导读

DataSpace 是一个面向异构工作区的数据智能体基准,包含 410 个跨语言任务和 7439 条文件,要求智能体整合数据库、CSV/JSON、长文档和视频后返回指定表格。

正文

A strong LLM does not automatically make a reliable data agent: DataSpace shows that the harness, cross-source joins, and final-table handling all materially affect whether the job actually gets done.

It tests 410 tasks where agents must combine databases, CSV/JSON files, long documents, and video, then return the exact requested table.

The best model gets only 66.34% right.

And with the model fixed, simply changing the agent harness moves accuracy from 30.98% to 46.34%.

Across every tested model, joins and mixing multiple data types are consistent weak spots.

But many failures happen even later: the agent has already found or computed the right information, then misunderstands the requested result or submits the wrong columns.

– arxiv. org/abs/2608.03451

Title: "DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces"

来源:@rohanpaul_ai · x.com