跳到正文
@omarsar0· @omarsar0 · X·· 2026-09-02AI 评分52
AI 导读

一篇论文把每轮对话提取为类型化节点和带属性边、从两跳子图作答,并在固定五个检索根的候选预算下测试图记忆。

正文

Finally, a good paper testing if graph memory actually beats flat retrieval for long-term agents.

(bookmark this one)

Researchers extract each conversational turn into typed nodes and attributed edges, answer from a two-hop subgraph, and hold the candidate-generation budget fixed at five retrieval roots.

On LongMemEval the graph gets token F1 0.42 against 0.47 for a flat vector baseline, and a paired bootstrap over 500 questions puts the gap at -0.050 (95% CI -0.085 to -0.016).

The damage concentrates on questions that require recalling a specific prior assistant turn, where judged correctness falls from 0.911 to 0.607. Splitting a turn into entities discards the surface form those questions depend on.

The forgetting module fares much better. One pruning pass over a persistent 27,021-node graph, scored on recency, access frequency, degree centrality and age, removes 9.8% of nodes and 9.5% of stored bytes with token F1 unchanged.

Paper: https://t.co/KDUecWNGTH

Chat with Paper: https://t.co/b661ajV2ri

来源:@omarsar0 · x.com