IBM 提出生成式检索方法 STAIR,用文档自带的目录作为寻址方案,解决传统检索按长度切块丢弃文档层级结构的问题。在 SearchTome 上,STAIR 的 Recall@1 达 82.6%,高于微调后的 Differentiable Search Index 的 76.9%,BM25 和 DPR 分别为 59.5% 和 68.7%。
Great RAG paper from IBM.
There are some really good ideas on how to solve common RAG issues.
It's well known that retrievers chunk long documents by length, which discards the hierarchy the document already has.
So they propose using a table of contents.
A table of contents helps to encodes exactly the global structure that chunking throws away.
STAIR uses that table of contents as the addressing scheme for a generative retriever, so the model stores and retrieves information from its own parameters against a structure the corpus supplies.
On SearchTome, it reaches Recall@1 of 82.6 percent against 76.9 percent for a fine-tuned Differentiable Search Index, a statistically significant gap, with BM25 at 59.5 percent and DPR at 68.7 percent.
Hallucination stays below 0.05 percent, which is the standing objection to generative retrieval and the reason grounding the address space in a real hierarchy is worth the extra structure. The ablations also show it generalizes where very few training samples exist.
Paper: https://t.co/zQbpGz3Ojd
来源:@omarsar0 · x.com