Home / Blog / RAG pipeline
Applied AI and Search Systems

RAG Without Mystery: From Documents to Reliable Retrieval

Retrieval-augmented generation is not “put PDFs in a vector database.” It is a data pipeline that turns changing source material into traceable evidence, finds the right pieces for a question and gives a model enough context to answer with grounded citations. When answers are wrong, the failure may be in parsing, access filtering, chunk boundaries, retrieval ranking or generation. Treat each stage as a measurable component instead of reaching first for a longer prompt.

Follow the evidence through the system

Every answer should be traceable from source version to retrieved passage
1 / SourceOwner, version, permissions and freshness.
2 / ParseExtract text, tables and layout faithfully.
3 / StructureSplit by meaning; keep section metadata.
4 / IndexEmbed chunks and retain provenance.
5 / RetrieveSearch, filter, rank and select evidence.
6 / AnswerGenerate with citations and abstention.

Start with document identity and access

Give each source a stable document ID, version, content hash, owner, effective date, ingestion time and access policy. A re-upload should update or supersede the right version instead of silently creating duplicate evidence. Deletions and permission changes must propagate to derived chunks and indexes; hiding a source in the file store while leaving its chunks searchable is a data exposure.

Make ingestion repeatable. Record parser version, extraction warnings and chunking configuration for each run. Keep the original file or a controlled reference when retention policy allows, so a bad answer can be traced to the exact source and transformation. Treat untrusted document text as data, not instructions: retrieved content must not override system policy or authorize tool actions.

Parse for meaning, not just text volume

PDF extraction can scramble columns, omit headers or detach a table from its units. HTML may contain navigation and repeated boilerplate. OCR introduces uncertainty. Preserve headings, page numbers, table labels and source links as metadata; flag low-quality extraction for review. For tables, store row/column context or render a concise textual representation that does not lose the relationship between a value and its heading.

Choose chunk boundaries for the corpus

Content shapeUseful starting boundaryFailure to watch
Policies and manualsHeading plus coherent section, with parent contextExceptions separated from the rule they modify
Code and API docsSymbol, endpoint or complete exampleSignature detached from parameters or version
Tables and catalogsRows with repeated column labels and unitsNumbers retrieved without their category
Support conversationsIssue, resolution and product/version metadataOne reply stripped of the question it answers

Fixed token windows are reproducible baselines, not universal truth. Overlap can protect context at boundaries but increases index size and duplicate results. Structure-aware chunks with parent/child retrieval can return a focused passage while preserving the larger section for generation. Compare strategies on your documents and questions rather than selecting a size by folklore.

Embeddings are one retrieval signal

An embedding maps text to a vector so semantically related passages can be found even when wording differs. It does not guarantee relevance, factual correctness or permission. Store embedding model/version, dimensions and chunk identity; changing models usually requires a planned re-index. For product codes, names, dates and exact error strings, lexical search can outperform semantic similarity. Hybrid retrieval can combine both, followed by a reranker when the latency and cost are justified.

Apply authorization filters before content reaches the model. Use tenant, audience, document status and effective-date metadata as filters, and test cross-tenant queries explicitly. Security is not fixed by asking the model to ignore unauthorized passages after retrieval.

Evaluate retrieval separately from the answer

Build a small, representative set of real questions with expected supporting documents or passages. Measure whether the right evidence appears in the top results, whether irrelevant passages crowd it out and whether access filters hold. Then evaluate the generated answer for groundedness, relevance, completeness and citation correctness. Track retrieval and generation separately so a fluent answer does not conceal a broken search stage.

Include adversarial and ordinary cases: no answer in the corpus, conflicting versions, stale policy, ambiguous terms, tables, multilingual questions and prompt injection inside documents. Define when the system should abstain, ask a clarifying question or route to a human. Review failures by stage and rerun the same evaluation set after changes.

Make citations useful to the reader

Return source title, version or date, section/page and a stable link alongside the answer. A citation should point to the passage actually retrieved, not merely a document with a similar title. Preserve mappings from each chunk to its source offsets so the interface can show context and the user can verify the claim. If the retrieval is weak, do not manufacture a confident answer from general model knowledge.

In summary

Reliable RAG is governed document ingestion plus retrieval you can inspect and generation you can evaluate. Preserve identity, permissions and provenance; parse structure; test chunking and hybrid search; measure retrieval separately from answer quality; and make abstention a valid outcome. The vector database is one component in that system, not the system itself.

References