End-to-end reference for the 5-stage RAG retrieval pipeline that turns a user query into LLM-ready context.
For ingestion (how documents enter Qdrant), see ingestion.md. For operational commands, see pipeline-commands.md. For environment variable reference, see configuration.md.
ChatController.stream()
-> ChatService.buildStructuredPromptWithContextOutcome()
-> RetrievalService.retrieve()
|-- 1. QueryVersionExtractor -- version detection + query boosting
|-- 2. HybridSearchService -- parallel dense+sparse search with RRF fusion
|-- 3. dedupeByHashThenUrl() -- UUID -> SHA-256 hash -> URL dedup
+-- 4. RerankerService -- LLM reranking with Caffeine caching
-> SearchQualityLevel -- quality assessment for LLM calibration
-> 5. PromptTruncator -- priority-based truncation
-> OpenAIStreamingService.streamResponse()
| Stage | File | Entry point |
|---|---|---|
| Version extraction | util/QueryVersionExtractor.java |
boostQueryWithVersionContext() |
| Hybrid search | service/HybridSearchService.java |
search() |
| Deduplication | service/RetrievalService.java |
dedupeByHashThenUrl() |
| LLM reranking | service/RerankerService.java |
rerank() |
| Prompt assembly | application/prompt/PromptTruncator.java |
truncate() |
All paths are relative to src/main/java/com/williamcallahan/javachat/.
No LLM calls -- all enhancement is deterministic regex and string manipulation.
QueryVersionExtractor (util/QueryVersionExtractor.java) detects Java version mentions in queries using the regex pattern \b(?:java\s*se|javase|java|jdk)[\s-]*(\d{1,3})\b (case-insensitive).
Matches: Java 25, JDK 24, java25, jdk-25, Java SE 24, JavaSE 25.
boostQueryWithVersionContext() prepends synonym expansions to the raw query, one per requested release:
JDK 25 Java SE 25 Java 25 release documentation: <original query>
For multi-release comparison queries such as Java 21 vs 24, each requested release contributes its own synonym prefix separated by ; . This improves dense-embedding recall by injecting version-related terms the embedding model can anchor on.
extractVersionNumbers() returns every requested Java release in encounter order, including comparison shorthand such as Java 21/22 or JDK 21 vs 17. Extraction is independent of the indexed release inventory so a corpus gap cannot erase the learner's requested version before adjacent-evidence resolution.
For an indexed requested release, RetrievalService builds a RetrievalConstraint whose docVersions is an any-of docVersion keyword filter pushed to Qdrant server-side via QdrantRetrievalConstraintBuilder (matchKeywords). A Java release missing from the corpus resolves to the nearest indexed release below and above it; a request outside the indexed range uses the nearest available side. For example, Java 22 searches Java 21 and Java 25 documentation. Versioned dependency queries use the same-family identities from DocsSourceRegistry to select exact or adjacent indexed versions and label the requested/evidence relationship in each source record. Repository policy for version gaps is owned solely by AGENTS.md [VG1]. There is no client-side URL/title post-filter fallback.
HybridSearchService (service/HybridSearchService.java) performs parallel hybrid search across 4 Qdrant collections using the direct gRPC client (io.qdrant:client), not Spring AI VectorStore abstractions.
Each query is encoded into two vectors:
| Vector | Named vector | Model | Dimensions | Source |
|---|---|---|---|---|
| Dense | dense |
Qwen3-Embedding-4B | 2560 | EmbeddingClient.embed() |
| Sparse | bm25 |
Murmur3 feature-hashed TF | variable | LexicalSparseVectorEncoder.encode() |
LexicalSparseVectorEncoder (service/LexicalSparseVectorEncoder.java) builds BM25-style sparse vectors:
- Normalize text via
AsciiTextNormalizer.toLowerAscii()+toLowerCase(Locale.ROOT) - Tokenize with Lucene
StandardAnalyzer; discard tokens shorter than 2 characters - Hash each token with Murmur3-32 (seed=0), sign-extend to unsigned long
- Count term frequencies in
Map<Long, Integer> - Cap at 256 unique tokens (sorted by count descending, tie-break by index)
- Return
SparseVector(indices, values)sorted by index ascending
IDF weighting is applied server-side by Qdrant's modifier=idf on the sparse vector config (set during collection creation by QdrantIndexInitializer). The encoder only sends raw term counts.
For each of the 4 collections, a QueryPoints request is built with two prefetch stages fused by RRF:
QueryPoints {
prefetch: [
{ nearest: denseVector, using: "dense", limit: prefetchLimit, filter: ... },
{ nearest: sparseVector, using: "bm25", limit: prefetchLimit, filter: ... }
],
query: rrf(k=rrfK),
limit: topK
}
All 4 requests fire asynchronously via qdrantClient.queryAsync() (returns ListenableFuture, converted to CompletableFuture). Each future is awaited with the configured timeout (default 10s).
The sparse prefetch is skipped if the sparse vector has no indices (empty query after tokenization).
RRF formula (server-side Qdrant): score = Sum(1 / (k + rank_i)) where k defaults to 60.
As collection results merge, mergePoints() deduplicates by Qdrant point UUID in a LinkedHashMap<String, ScoredResult>. If the same UUID appears in multiple collections, the higher score wins.
Any collection query failure throws HybridSearchPartialFailureException; retrieval never returns silently degraded context.
| Collection | Default name | Property |
|---|---|---|
| Books | java-chat-qwen3-embedding-4b-2560-books |
app.qdrant.collections.books |
| Docs | java-chat-qwen3-embedding-4b-2560-docs |
app.qdrant.collections.docs |
| Articles | java-chat-qwen3-embedding-4b-2560-articles |
app.qdrant.collections.articles |
| PDFs | java-chat-qwen3-embedding-4b-2560-pdfs |
app.qdrant.collections.pdfs |
Each Qdrant point is converted to a Spring AI Document with typed metadata fields: url, title, package, hash, docSet, docPath, sourceName, sourceKind, docVersion, docType, chunkIndex, pageStart, pageEnd, score, collection.
Three-layer dedup, applied in sequence:
| Layer | Location | Key | Winner |
|---|---|---|---|
| UUID | HybridSearchService.mergePoints() |
Qdrant point UUID | Highest score |
| Content hash | RetrievalService.dedupeByHashThenUrl() |
docVersion + hash (SHA-256) |
First seen |
| URL | RetrievalService.dedupeByHashThenUrl() |
url metadata |
First seen |
Both hash and URL dedup use LinkedHashMap.putIfAbsent to preserve reranker ordering. Content-hash dedup is keyed by docVersion + hash so the same chunk can appear once per requested Java release in a multi-release comparison. Documents with neither hash nor URL are kept unconditionally (with a warning log).
File: RetrievalService.java lines 182-213.
RerankerService (service/RerankerService.java) selects and orders search results by relevance using an LLM call; documents the LLM judges unrelated are dropped.
The reranking prompt instructs the LLM to consider:
- Java-specific context and domain relevance
- Version relevance to the query
- Source authority -- official docs preferred over blogs/third-party
- Stable vs. early-access -- stable release docs preferred over preview content
- Learning value for the user
Each document is presented as [index] title | url followed by the first 500 characters of content.
- Model: the configured chat provider and chat model
- Temperature:
0.0(deterministic) - Timeout: configurable via
app.rag.reranker-timeout(default 8s) - Response format:
{"order": [0, 3, ...]}(0-based indices, most relevant first; subset or empty when not all documents are relevant)
The response must be exactly one JSON object with exactly one order field. The parser does not
unwrap Markdown fences, extract an embedded object from prose, or accept trailing JSON or other text.
It rejects duplicate keys, unknown fields, floating-point or string index coercions, and malformed JSON.
The order array may be a strict subset of the source indices: only documents the LLM judges
relevant appear, most relevant first, and the array may be empty when no document is relevant
(off-topic queries yield no context and no citations). Duplicate, null, negative, or
out-of-range indices, or an array larger than the source set, fail reranking rather than being
skipped. A valid selection is then limited to searchReturnK (default 6).
Results are cached in a Caffeine cache (reranker-cache) keyed by:
query + ":" + docsHash + ":" + returnK
where docsHash is Integer.toHexString() of concatenated document URLs (or text hashCodes).
On any failure (timeout, parse error, LLM unavailable), RerankerService throws RerankingFailureException. There is no fallback to original ordering.
Chat and embeddings share OPENAI_BASE_URL and OPENAI_API_KEY, but embedding model selection is independent
of OPENAI_MODEL: retrieval always embeds with qwen/qwen3-embedding-4b. User-facing retrieval sends
X-Tier: production-z; ingestion, probes, and warmups send X-Tier: batch.
After retrieval completes, the prompt is assembled and truncated to fit the model's token budget.
StructuredPrompt (domain/prompt/StructuredPrompt.java) renders segments in this order:
- System prompt (with appended SEARCH CONTEXT quality note)
- Context documents, each prefixed with
[CTX N] <normalized URL> - Conversation history turns (role-prefixed)
- Current user query
ChatService.buildStructuredPromptWithContextOutcome() (line 186-195) appends a quality note to the system prompt based on SearchQualityLevel.describeQuality():
| Level | Trigger | Message injected into system prompt |
|---|---|---|
NONE |
No documents retrieved | "No relevant documents found. Using general knowledge only." |
KEYWORD_SEARCH |
Any document URL contains local-search or keyword |
"Found N documents via keyword search (embedding service unavailable). Results may be less semantically relevant." |
HIGH_QUALITY |
All documents have text longer than 100 characters | "Found N high-quality relevant documents via semantic search." |
MIXED_QUALITY |
Some documents below 100 character threshold | "Found N documents (M high-quality) via search. Some results may be less relevant." |
When the quality message contains "less relevant" or "keyword search", an additional low-quality search prompt is appended from SystemPromptConfig.getLowQualitySearchPrompt(). This calibrates LLM confidence -- the model is told explicitly when its context may be unreliable.
PromptTruncator (application/prompt/PromptTruncator.java) fits the assembled prompt within a model-specific token budget:
| Priority | Segment type | Truncation behavior |
|---|---|---|
| CRITICAL | System prompt | Never truncated |
| HIGH | Current user query | Never truncated |
| HIGH | Authoritative context documents (e.g. curated lessons) | Retained before conversation history |
| MEDIUM | Conversation history | Oldest turns removed first |
| LOW | Ordinary retrieved context documents | Least relevant removed first |
Algorithm (lines 49-101):
- Reserve tokens for system prompt + current query (non-negotiable)
- If those alone exceed the budget, return minimal prompt (system + query only)
- If authoritative (HIGH-priority) context is present, throw
AuthoritativeContextDoesNotFitExceptionwhen no segment fits the remaining budget - Fit HIGH priority context documents (e.g. curated lesson context) into remaining budget
- Fit conversation history newest-first into remaining budget
- Fit LOW priority context documents in reranker order (most relevant first) into remaining budget
- Re-index surviving documents with sequential
[CTX N]markers
The GPT-5.4 prompt budget is application-owned:
| Provider/model | Token budget | Constant |
|---|---|---|
| Shared gateway GPT-5.4 | 100,000 | GPT54_INPUT_TOKEN_BUDGET |
Token estimation uses a conservative (text.length() / 4) + 1 approximation (~4 characters per token for English text).
GPT-5.4 RAG retrieval is reduced upstream to max 3 documents (RAG_LIMIT_CONSTRAINED)
with max 600 tokens each (RAG_TOKEN_LIMIT_CONSTRAINED), defined in ModelConfiguration.java.
All properties are bound via AppProperties (@ConfigurationProperties(prefix = "app")). See configuration.md for environment variable names.
| Property | Default | Description |
|---|---|---|
app.qdrant.dense-vector-name |
dense |
Named vector key for dense embeddings |
app.qdrant.sparse-vector-name |
bm25 |
Named vector key for sparse BM25 tokens |
app.qdrant.prefetch-limit |
14 |
Per-stage candidate count for each dense/sparse prefetch before RRF fusion |
app.qdrant.rrf-k |
60 |
RRF k parameter: score = Sum(1 / (k + rank)) |
app.qdrant.query-timeout |
10s |
Timeout for hybrid search fan-out across all collections |
app.qdrant.ensure-payload-indexes |
true |
Create payload indexes on startup for metadata filtering |
app.qdrant.ensure-collections |
false |
Validate existing collections without creating missing collections on normal startup |
| Property | Default |
|---|---|
app.qdrant.collections.books |
java-chat-qwen3-embedding-4b-2560-books |
app.qdrant.collections.docs |
java-chat-qwen3-embedding-4b-2560-docs |
app.qdrant.collections.articles |
java-chat-qwen3-embedding-4b-2560-articles |
app.qdrant.collections.pdfs |
java-chat-qwen3-embedding-4b-2560-pdfs |
All four must be non-blank and distinct (validated on startup).
| Property | Default | Validation | Description |
|---|---|---|---|
app.rag.search-top-k |
8 |
Must be > 0 | Candidates fetched from hybrid search before reranking |
app.rag.search-return-k |
6 |
Must be > 0, must be <= search-top-k |
Results returned to the LLM after reranking |
app.rag.reranker-timeout |
8s |
Must be positive and less than 20s |
Timeout for LLM reranking call |
app.rag.search-citations |
3 |
Must be >= 0 | Citation references included in the response |
app.rag.search-mmr-lambda |
0.5 |
Must be in [0.0, 1.0] | MMR lambda (higher = relevance, lower = diversity) |
app.rag.chunk-max-tokens |
900 |
Must be > 0 | Max tokens per ingested chunk (see ingestion.md) |
app.rag.chunk-overlap-tokens |
150 |
Must be >= 0, must be < chunk-max-tokens |
Overlap between consecutive chunks |
| Property | Default | Description |
|---|---|---|
app.embeddings.model |
qwen/qwen3-embedding-4b |
Embedding-owned gateway model; independent from OPENAI_MODEL |
app.embeddings.dimensions |
2560 (application and code defaults) |
Must match the active embedding model output size |
For the shared gateway and explicit local development mode, see configuration.md#embeddings.