A PDF page contains a table where the column header supplies the unit and the footnote changes the meaning of a value. A text extractor returns the numbers in the wrong order. The language model may produce a fluent answer from a representation that already lost the evidence.
Document RAG needs to decide what information the retrieval system can see. OCR text, layout-aware extraction, page images, and mixed representations preserve different parts of the document.
Retrieval and answering are separate evaluations
ColPali, published at ICLR 2025, studies retrieval over visually rich document pages and introduces the ViDoRe benchmark. It uses multi-vector page-image representations with late interaction rather than relying only on extracted text.
The paper's retrieval results do not by themselves establish accurate end-to-end answers from every page. Finding the correct page is one stage; reading the right cell, preserving its unit, and applying a qualification are additional tasks.
Choose the page or region as an evidence unit
For a synthetic procurement question, the answer might need a product row, a capacity column, and a footnote excluding one operating temperature. Mark those elements in the evaluation reference. A page-level hit can be relevant while still requiring precise region interpretation.
Preserve document ID, page number, version, extraction method, and any region coordinates. A citation to a whole 200-page PDF is much less useful than a source view that identifies the table and condition.
Compare representations on the actual documents
Evaluate text extraction plus lexical retrieval, text embeddings, visual retrieval, and a mixed path on the same task set. Include scanned pages, dense tables, charts, multi-column layouts, low-resolution images, and languages the product supports.
Keep the indexing and query cost visible. Multi-vector representations can change storage and retrieval work. A system that improves table retrieval but increases latency beyond the user's decision window has a trade-off to report.
Check the answer against the structure
A generative model may confuse a row label, repeat a number from a nearby column, or overlook a footnote even after the correct page is retrieved. Use field-level checks where the task can be expressed structurally, and source-aware review where interpretation is required.
Test whether the answer preserves units and time periods. A price per case is not a price per item. A chart axis can be logarithmic. A footnote may limit the result to a selected sample. These are semantic conditions, not decorative page details.
Keep sensitive content within permission boundaries
A page image may contain more information than the requested field. Sending it to a model can expose neighboring personal or confidential material. Apply source access rules, use region-level processing when appropriate, and document the provider and retention path.
Do not assume that replacing OCR with images removes prompt-injection risk. A page can display instructions directed at the agent. Treat the content as evidence and retain tool authority outside the document.
Show a useful recovery
If extraction fails, expose that state rather than replacing unreadable content with a confident guess. Ask for a better source, show the page for manual review, or route the task to a validated alternative. Record which representation failed and whether the fallback completed the task.
Multimodal retrieval is promising when visual structure carries necessary evidence. Its value is demonstrated by improved task completion with traceable sources, not merely by a higher probability of finding a related page.