Notes · RAG · open source · 5 min read
Refuse before you hallucinate
DocuQuery is a small open-source RAG service built around one rule: answer only from the retrieved documents, cite every claim, and when the documents do not contain the answer — refuse, without ever calling the LLM.
Three ways naive RAG fails
- Hallucinated answers — the model fills gaps with facts that are not in the retrieved context.
- Unbounded token costs — whole documents get stuffed into every prompt.
- Answers nobody can audit — no way to tell which source a sentence came from.
DocuQuery addresses each one by design rather than by prompt wording alone.
Refusal is a code path, not a prompt
Retrieved chunks are filtered by a similarity score threshold. If nothing passes, the service returns a fixed refusal message immediately: no model call, zero tokens, zero chance of inventing an answer. The system prompt still forbids using knowledge outside the context, but the cheapest hallucination is the call you never make.
A hard token budget
A token-budget manager packs chunks greedily by descending relevance score until it reaches a configured ceiling, counted with
tiktoken’s cl100k_base encoding. No query ever sends unbounded context, so cost per query is predictable — and the most relevant
evidence always makes it in first.
Citations you can follow
Every factual statement carries a [Source: file, Section: heading] citation, and every citation maps back to an exact
chunk with its source, section and chunk id. Ingestion keeps provenance native to each format: page numbers for PDF, headings for
DOCX and Markdown, sheet names for XLSX, plus CSV and cleaned-up HTML.
POST /api/v1/query 1. embed the question 2. retrieve top-k chunks (vector store, cosine) 3. drop chunks below the score threshold 4. nothing left? → return the refusal, no LLM call 5. pack chunks into the token budget (best first) 6. build the grounded prompt, temperature 0 7. generate (or stream over SSE) 8. attach one citation per chunk used 9. log the query to the audit trail
The boring engineering that makes it production-shaped
- Clean architecture: routes depend on interfaces, so swapping the vector store touches one adapter.
- Fully async I/O; the one synchronous driver is offloaded with
asyncio.to_thread. - Server-Sent Events streaming with proxy-buffering disabled, ending in an explicit
[DONE]. - Every query lands in an audit trail; Docker Compose for one-command runs; 160 tests in CI.
Takeaways
- Make “I don’t know” a first-class, deterministic outcome.
- Budget tokens explicitly; relevance decides what gets in.
- If an answer cannot be traced to a chunk, it should not ship.