Khachatur Pepanyan ← All notes

Notes · RAG · open source · 5 min read

Refuse before you hallucinate

DocuQuery is a small open-source RAG service built around one rule: answer only from the retrieved documents, cite every claim, and when the documents do not contain the answer — refuse, without ever calling the LLM.

0LLM calls on a refusal
6document formats
160tests in CI

Three ways naive RAG fails

DocuQuery addresses each one by design rather than by prompt wording alone.

Refusal is a code path, not a prompt

Retrieved chunks are filtered by a similarity score threshold. If nothing passes, the service returns a fixed refusal message immediately: no model call, zero tokens, zero chance of inventing an answer. The system prompt still forbids using knowledge outside the context, but the cheapest hallucination is the call you never make.

A hard token budget

A token-budget manager packs chunks greedily by descending relevance score until it reaches a configured ceiling, counted with tiktoken’s cl100k_base encoding. No query ever sends unbounded context, so cost per query is predictable — and the most relevant evidence always makes it in first.

Citations you can follow

Every factual statement carries a [Source: file, Section: heading] citation, and every citation maps back to an exact chunk with its source, section and chunk id. Ingestion keeps provenance native to each format: page numbers for PDF, headings for DOCX and Markdown, sheet names for XLSX, plus CSV and cleaned-up HTML.

POST /api/v1/query
  1. embed the question
  2. retrieve top-k chunks (vector store, cosine)
  3. drop chunks below the score threshold
  4. nothing left?  → return the refusal, no LLM call
  5. pack chunks into the token budget (best first)
  6. build the grounded prompt, temperature 0
  7. generate (or stream over SSE)
  8. attach one citation per chunk used
  9. log the query to the audit trail

The boring engineering that makes it production-shaped

Takeaways

  1. Make “I don’t know” a first-class, deterministic outcome.
  2. Budget tokens explicitly; relevance decides what gets in.
  3. If an answer cannot be traced to a chunk, it should not ship.