Production-Grade RAG with Spring AI 2.0 in Java: Chunking, Reranking and Hallucination Guards
A Spring AI 2.0 RAG pipeline on Spring Boot 4.1, taken apart one stage at a time on a real PostgreSQL with pgvector: chunking, once-only PDF ingestion, the similarity threshold that is off by default, injection-proof tenant filters, reranking and answer checks. Every number comes from a committed transcript, and the page says plainly what was scripted.
Ask a language model what your company’s annual leave policy is and it will answer, fluently, without ever having seen your handbook. Retrieval-augmented generation, RAG, is the fix that most teams reach for: before the question goes to the model, find the few paragraphs of your documents that are most likely to contain the answer, paste them into the prompt, and tell the model to answer from those and nothing else. The idea takes ten lines to demonstrate. What takes the rest of the page is the places it quietly goes wrong once real PDFs, real users and more than one customer arrive.
This page builds one pipeline on Spring Boot 4.1 and Spring AI 2.0 and takes it apart a stage at a time: cutting documents into chunks, getting a PDF in exactly once, deciding what counts as a match, keeping one tenant’s documents away from another’s, re-rating candidates, and checking the answer after it is written. Every number and every block of output below came out of a test in asmhatre/spring-ai, rag module, running against a real PostgreSQL with pgvector and real PDF files, and every code block links to the file it was taken from. The complete code tour is the companion post, Spring AI RAG in Java: Complete Runnable Code. This page replaces an earlier version whose code did not build; the section “What the earlier version got wrong” says exactly how.
Versions. Spring Boot 4.1.1, Spring AI 2.0.1 (Maven Central dates 2.0.0 at 12 June 2026 and 2.0.1 at 20 August 2026), JDK 25 (Temurin 25.0.4.1), PostgreSQL 16.13 with pgvector 0.6.0. API names were read from the jars with javap (transcript 12). Retrieved 21 September 2026.
What is real and what is scripted. The vector store, the PDFs, the metadata filters, the HTTP layer and every Spring AI class run for real. The chat model and the embedding model are scripted stand-ins, so that anyone can run the tests with no API key. That means nothing here shows how a real model answers, what similarity scores a real embedding model gives, whether reranking improves answers, or how well a real judge catches an invented claim. Where a number depends on that, the text says so. The application was never run against the OpenAI API while writing this.
A model only knows what is in its prompt
A language model has two sources of knowledge: what it absorbed during training, and what you put in the prompt. Your documents are not in the first, so RAG works on the second. The pipeline has two lanes. In the ingestion lane, run once per file, you read the document, cut it into pieces called chunks, turn each chunk into a list of numbers called an embedding (texts about similar things get similar numbers), and store the chunk with its embedding in a vector store. In the query lane, run on every question, you embed the question the same way, ask the store for the nearest chunks, and put them in the prompt with the question.
The picture is the map for the page. Every arrow is a hand-off between two pieces of code, and every hand-off has a way to fail without raising an error: a chunk cut in the wrong place, a duplicate stored, a threshold that accepts everything, a filter that lets another customer’s text through, a rating that silently becomes zero. The vector store here is PostgreSQL with the pgvector extension, which Spring AI drives through PgVectorStore. If embeddings are new to you, Vector Embeddings and Semantic Search in Java explains what they are and how nearest-neighbour search works.
The stages and which Spring AI class each one stands on: chapter 1.
Spring AI packages the query lane as an advisor: a piece of code that wraps every call a ChatClient makes, and can change the prompt on the way out. QuestionAnswerAdvisor does the whole job: it searches the store with the user’s question and appends the matching chunks to the prompt.
That is the whole pipeline, from SimpleAdvisorTest.java. The search asks for the two nearest chunks and ignores anything scoring below 0.3 (a number that only means something for the embedding model used here; the retrieval section returns to it). Here is the prompt the model receives for “How many days of annual leave do employees get?”, captured from that test (transcript 11):
How many days of annual leave do employees get?
Context information is below, surrounded by ---------------------
---------------------
4.1 Annual Leave Entitlement. Full-time employees are entitled to 20 working days
of annual leave per calendar year. Part-time employees receive leave pro rata.
4.2 Leave Carryover. Unused annual leave may be carried over for a maximum of 5 days
into the next calendar year and must be used by 31 March.
4.3 Requesting Leave. All leave requests must be submitted through the HR portal
at least two weeks in advance for absences longer than three days.
4.7 Sick Leave. Sick leave is separate from annual leave and is not deducted from it.
A medical certificate is required after three consecutive days.
---------------------
Given the context and provided history information and not prior knowledge,
reply to the user comment. If the answer is not in the context, inform
the user that you can't answer the question.
The retrieved chunks sit between two lines of dashes, and the template ends with the only instruction that stops the model from making something up: if the answer is not in the context, say so. Whether the model obeys is up to the model. Now the same advisor with a question the handbook cannot answer:
What is the capital of Mongolia?
Context information is below, surrounded by ---------------------
---------------------
---------------------
Given the context and provided history information and not prior knowledge,
reply to the user comment. If the answer is not in the context, inform
the user that you can't answer the question.
The empty block is the fingerprint of a missing safety net. Nothing found, and the prompt is the same template with nothing between the dashes. No code decided that there was no answer; the only thing between the user and an invented reply is the model reading the last sentence of the template. Everything from here on is about not depending on that.
What the two prompts show and what the advisor does not do: chapter 1.
Advisors in general, and the tools QuestionAnswerAdvisor shares with RetrievalAugmentationAdvisor: Spring AI RAG reference.
QuestionAnswerAdvisor ships in spring-ai-vector-store-advisor and the larger RetrievalAugmentationAdvisor in spring-ai-rag; the dependency tree is in transcript 13.
Chunking decides what the model can be shown
A chunk is the unit that gets embedded, retrieved and pasted into the prompt. Cut too large and the prompt fills with text that has nothing to do with the question. Cut too small and an answer is split across two chunks, each of which looks half relevant. Spring AI’s TokenTextSplitter is the standard tool, and it is worth knowing what it does before you tune it. To find out, the repository cuts one fixed document three ways: forty numbered sentences, 4,268 characters, 880 tokens.
The top bar is the default splitter: it treats 800 tokens as a chunk, so the document becomes two chunks and a short handbook page becomes exactly one. The middle bar sets withChunkSize(100). The splitter did not cut at 100 tokens; it cut back to the last full stop, which gave ten chunks of 88 tokens. The bottom bar is a recursive chunker written for this repository (RecursiveChunker.java), which repeats about 80 characters at the start of each chunk so a sentence that straddles a boundary is whole in one of them. The measurements (transcript 02, transcript 03):
chunks: 2, sizes in tokens: [792, 88]
chunks: 10
boundaries where the first 30 characters of a chunk already appear in the chunk before it: 0 of 9
minChunkSizeChars=350: 10 chunks, 9 of the first 9 end on a full stop
minChunkSizeChars= 50: 10 chunks, 9 of the first 9 end on a full stop
chunks: 14, longest: 395 characters
boundaries where the next chunk opens with words the previous one ended with: 13 of 13
Three facts in there are worth carrying away. TokenTextSplitter has no overlap: not one of nine boundaries repeated any text. The setting minChunkSizeChars changed nothing between 350 and 50, and reading the splitter’s bytecode shows why: it cuts an over-long chunk back to its last sentence-ending mark only if that mark is more than minChunkSizeChars characters into the chunk, and in text with a full stop every hundred characters that is always true. And the overlap has a price: the recursive chunker’s chunks often begin mid-sentence, because the overlap starts at a word boundary.
Here is the bean that chooses the chunker. Everything downstream takes a DocumentTransformer, so swapping strategy is a one-bean change:
(From RagConfig.java.) A third strategy, a semantic chunker that embeds every sentence and starts a new chunk where the topic changes, is in SemanticChunker.java. On twelve sentences about three topics it produced three chunks, and it sent the twelve sentences to the embedding model in one batch (transcript 03). Treat that result with suspicion: the scripted embedding model is exactly the kind that separates topics by shared words.
No chunking strategy here has been benchmarked. The earlier version of this article carried a table of retrieval-precision percentages for each strategy. Nothing in this repository measures retrieval quality, so the table is gone. What you can rely on is what the transcripts show: what each splitter does to a known document. Which one retrieves best on your documents is an experiment you have to run with your own questions.
Every chunker measured, with the caveats: chapter 2.
The recursive chunker’s separators and the hard-cut fallback: RecursiveChunker.java.
The constructor new TokenTextSplitter(512, 128, 5, 10_000, true) from the earlier version compiles against Spring AI 1.1.0 and not against 2.0.1; the builder is the way now (transcript 14).
Token counts in the transcript come from jtokkit’s cl100k_base encoding, the library TokenTextSplitter itself uses; the test that produced them is ChunkingTest (ChunkingTest.java).
Getting a PDF in, once
Ingestion has three jobs that tutorials skip: keep the page number (a citation needs it), clean what the PDF reader returns, and make sure uploading the same file twice does not store it twice. PagePdfDocumentReader with withPagesPerDocument(1) returns one document per page, and the only metadata it adds is page_number. Every chunk cut from that page inherits it.
The reader also pads text. Page 2 of the sample handbook came back as 862 characters, with a run of 133 spaces at the end of a line. Padding costs tokens and changes what gets embedded, so the service collapses it (IngestionService.java):
as read : 862 characters, longest run of spaces 133
tidied : 304 characters, longest run of spaces 1
The padding belongs to these PDFs. The sample PDFs are generated with PDFBox inside the repository, so the exact amount of padding is an artefact of that generator and the reader. Real PDFs pad differently, or not at all. Print what your own files produce before assuming the same.
Now the part that matters most. Upload a file, upload it again, edit one page and upload that, and count the rows in PostgreSQL. The service remembers, for each file name, the SHA-256 of the bytes and the ids of the chunks it wrote, and it decides what to do with each upload:
That flow is the ingest method (IngestionService.java). The heart of it, with the delete-by-id that stops the old version lingering:
String hash = sha256(pdf);
var known = tracker.find(filename);
if (known.isPresent() && known.get().hash().equals(hash)) {
metrics.counter("rag.ingestion.skipped").increment();
return IngestionResult.skipped(filename);
}
List<Document> pages = new PagePdfDocumentReader(pdf, PdfDocumentReaderConfig.builder()
.withPagesPerDocument(1)
.build()).get();
List<Document> tagged = pages.stream().map(page -> {
Map<String, Object> metadata = new HashMap<>(page.getMetadata());
metadata.put(SOURCE_FILE, filename);
metadata.put("source_hash", hash);
metadata.putAll(callerMetadata);
return new Document(tidy(page.getText()), metadata);
}).toList();
List<Document> chunks = chunker.apply(tagged);
int replaced = 0;
if (known.isPresent()) {
// The file changed: remove what the old version wrote, by id, before adding the new chunks.
replaced = known.get().chunkIds().size();
vectorStore.delete(known.get().chunkIds());
}
vectorStore.add(chunks);
tracker.record(filename, hash, chunks.stream().map(Document::getId).toList());
metrics.counter("rag.chunks.ingested").increment(chunks.size());
return known.isPresent()
? IngestionResult.updated(filename, chunks.size(), replaced)
: IngestionResult.ingested(filename, chunks.size());
Here is what happened on real PostgreSQL (transcript 05), followed by the naive alternative that most tutorials show, calling vectorStore.add() on every upload with nothing remembered:
1. first upload -> status=ingested chunksWritten=4 chunksReplaced=0 | rows in table=4, texts embedded so far=4
2. same bytes again -> status=skipped chunksWritten=0 chunksReplaced=0 | rows in table=4, texts embedded so far=4
3. page 2 edited (20 -> 22 days) -> status=updated chunksWritten=4 chunksReplaced=4 | rows in table=4, texts embedded so far=8
rows still saying "20 working days": 0, rows saying "22 working days": 1
4. same text, exported again -> status=updated chunksWritten=4 chunksReplaced=4 | rows in table=4, texts embedded so far=12
the two PDFs have identical text and different bytes: true
the file hash is a hash of bytes, so a re-export counts as a change and is re-embedded
--- the naive version: vectorStore.add() on every upload, nothing remembered ---
the same 4 chunks added on two more uploads: rows in table 4 -> 12
rows saying "22 working days" now: 3
The naive version tripled the rows, and the same page was stored three times, so retrieval returned three copies of it and pushed other passages out of the top results. The tracked version skipped identical bytes, spent no embedding calls doing so, and after an edit deleted the four old chunks before adding four new ones, leaving zero rows that still said “20 working days”.
The hash is of bytes, not of text. The fourth line above is a limit, not a feature: two PDFs with identical text but different bytes count as a change and are re-embedded. The PDFBox generator in the tests produces different bytes on every run for the same text. Hashing the extracted text instead would avoid that, at the cost of reading the file first. And the tracker is in memory: after a restart it has forgotten every file, the next upload of each one is re-embedded, and the old chunks are not deleted, because their ids are gone. Persist the tracker, or delete by a source_file filter (the chunks carry one; no test here exercises that path).
Page metadata, the tidy function, tenant tags and the tracker’s limits: chapter 3.
What one stored row looks like, read back from PostgreSQL: transcript 10.
The schema the application expects and never creates itself: init.sql.
Retrieval: the threshold that is not there by default
Retrieval asks the store for the topK chunks nearest to the question. Nearest is not the same as relevant: a store with four chunks will happily return all four for any question at all, ranked by how little they resemble it. What separates “the answer is here” from “this is the closest thing I have” is a similarity threshold, a minimum score below which a chunk is not a candidate. Spring AI’s retriever has one, and by default it accepts everything.
The bars are the four chunks of one handbook scored against a leave question. With the threshold at 0.3 two chunks survive; with the default of 0.0 all four do. The second half of the picture is the one that bites: for a question about Mongolia, every chunk scored exactly 0.0000, and the default threshold still accepted all four. So the safety net that Spring AI’s larger advisor offers for “nothing found” never fires while the threshold accepts everything. The numbers (transcript 06):
--- question: "How many days of annual leave do employees get?" ---
similarityThreshold 0.0 (the default): 4 chunk(s)
score 0.5891 page 2 "4.1 Annual Leave Entitlement. Full-time empl..."
score 0.4077 page 3 "4.3 Requesting Leave. All leave requests mus..."
score 0.2535 page 1 "Acme Employee Handbook 2026 3.5 Probationar..."
score 0.2023 page 4 "6.1 Expenses. Receipts are required for ever..."
similarityThreshold 0.3: 2 chunk(s)
--- question: "What is the capital of Mongolia?" ---
similarityThreshold 0.0 (the default): 4 chunk(s)
score 0.0000 page 1 "Acme Employee Handbook 2026 3.5 Probationar..."
score 0.0000 page 2 "4.1 Annual Leave Entitlement. Full-time empl..."
score 0.0000 page 3 "4.3 Requesting Leave. All leave requests mus..."
score 0.0000 page 4 "6.1 Expenses. Receipts are required for ever..."
similarityThreshold 0.3: 0 chunk(s)
Do not copy the 0.3. These scores come from the scripted hashing embedding model, which scores by shared words. A real embedding model produces a different range, and the threshold that separates “answerable” from “not” has to be found for your model and your corpus, for example by asking a set of questions you know the documents cannot answer and looking at the highest score each one gets.
The larger advisor, RetrievalAugmentationAdvisor, is wired in RagConfig.java. The part to read is the queryAugmenter line: allowEmptyContext(false).
/** Stages 2 to 4: retrieve, rerank, and put the survivors into the prompt. */
@Bean
RetrievalAugmentationAdvisor retrievalAdvisor(VectorStore vectorStore, LlmReranker reranker,
RagProperties properties) {
return RetrievalAugmentationAdvisor.builder()
.documentRetriever(VectorStoreDocumentRetriever.builder()
.vectorStore(vectorStore)
.topK(properties.retrieval().topK())
.similarityThreshold(properties.retrieval().similarityThreshold())
.build())
.documentPostProcessors(reranker)
.queryAugmenter(ContextualQueryAugmenter.builder().allowEmptyContext(false).build())
.build();
}
What that flag does when nothing is retrieved, captured from the same test:
--- nothing retrieved, allowEmptyContext(false): the prompt the model receives ---
The user query is outside your knowledge base.
Politely inform the user that you can't answer it.
--- nothing retrieved, allowEmptyContext(true): the prompt the model receives ---
What is the capital of Mongolia?
With false the model is told the question is outside its knowledge base and to say so politely. With true the model receives the bare question, which for a document service means it answers from whatever it was trained on. That is the failure RAG exists to prevent, and it is what you get if you set the flag the other way to “be more helpful”. QuestionAnswerAdvisor has no such switch, as the earlier section showed.
Threshold, prompt shape and the empty-context path in full: chapter 4.
As soon as two customers share a store, a question from one of them can retrieve the other’s text. The fix is metadata: tag every chunk with the tenant when it is ingested (the controller does, and the ingestion service passes the tag through to every chunk), and pass a filter with every search. Two fictional companies, Acme and Globex, each uploaded a handbook with a section 4.1 on annual leave: Acme says 20 working days and Globex says 25.
How you build the filter matters more than that it exists. The obvious way is a string. The safe way is the builder (TenantFilterTest.java):
// The mistake: build the filter as text from a value the caller controls.
String glued = "tenant_id == '" + HOSTILE_SINGLE + "'";
t.line("filter string built by concatenation: %s", glued);
List<Document> injected = store.similaritySearch(
SearchRequest.builder().query(QUESTION).topK(5).filterExpression(glued).build());
var acmeOnly = new FilterExpressionBuilder().eq("tenant_id", "acme").build();
List<Document> scoped = store.similaritySearch(
SearchRequest.builder().query(QUESTION).topK(3).filterExpression(acmeOnly).build());
The string version puts whatever the caller sent into the filter’s syntax. A caller who belongs to Globex and sends the tenant id globex' || tenant_id == 'acme turns the filter into “tenant is globex or tenant is acme”. Here is what that did on PostgreSQL, alongside the two correct calls (the in-memory store gave the same results, in the same transcript, transcript 07):
no filter, top 3: 3 chunk(s)
tenant=acme page=2 "4.1 Annual Leave Entitlement. Full-time ..."
tenant=acme page=3 "4.3 Requesting Leave. All leave requests..."
tenant=globex page=1 "Globex Staff Manual 2026 4.1 Annual Lea..."
eq("tenant_id", "acme") built with FilterExpressionBuilder, top 3: 3 chunk(s)
tenant=acme page=2 "4.1 Annual Leave Entitlement. Full-time ..."
tenant=acme page=3 "4.3 Requesting Leave. All leave requests..."
tenant=acme page=1 "Acme Employee Handbook 2026 3.5 Probati..."
filter string built by concatenation: tenant_id == 'globex' || tenant_id == 'acme'
result, top 5: 5 chunk(s)
tenant=acme page=2 "4.1 Annual Leave Entitlement. Full-time ..."
tenant=acme page=3 "4.3 Requesting Leave. All leave requests..."
tenant=globex page=1 "Globex Staff Manual 2026 4.1 Annual Lea..."
tenant=acme page=1 "Acme Employee Handbook 2026 3.5 Probati..."
tenant=acme page=4 "6.1 Expenses. Receipts are required for ..."
same text passed to FilterExpressionBuilder.eq(), top 5: 0 chunk(s)
double-quote variant passed to FilterExpressionBuilder.eq(), top 5: 0 chunk(s)
Unfiltered, the top three included Globex’s chunk for an Acme question. With eq("tenant_id", "acme") only Acme’s came back. With the concatenated string the hostile input returned both companies’ chunks. And the same hostile text handed to the builder as a value returned nothing, with a single-quote or a double-quote version, on both stores. The controller does it that way:
// The tenant goes in as a value of an expression, never as text glued into a filter string.
var filter = request.tenantId() == null ? null : new FilterExpressionBuilder()
.eq("tenant_id", request.tenantId()).build();
return queries.ask(request.question(), filter);
(From RagController.java.) The two hostile strings tried are the ones in the test; this is not a security audit of the filter parser.
Not tested: what the filter costs. pgvector applies a WHERE clause after scanning an approximate (HNSW) index, so a selective filter can return fewer than topK rows; pgvector 0.8.0 added iterative index scans to deal with it (see the pgvector README, “Filtering”). This repository ran pgvector 0.6.0, which predates them, and never ran a filtered query on enough rows to see the effect. If one tenant is a small slice of your table, test that before you ship.
The tenant id in the request body is trusted in this demo, which is a demonstration and not a design; where a real one comes from is the subject of the security posts on this blog, starting with Spring Security 7.1 JWT Authentication.
Vector search is fast and approximate: it finds chunks that resemble the question. A common second step is to retrieve generously (say twenty chunks) and then have a model re-rate each candidate for how well it answers the question, keeping only the best few. Spring AI ships the hook for this, DocumentPostProcessor, and no reranker: spring-ai-rag contains no class with “rerank” in its name (transcript 12). This one is written for the repository:
(From LlmReranker.java.) Each candidate gets its own rating call, so twenty candidates cost twenty model calls before the answer is even requested. Those calls go out on virtual threads, one per candidate, so they wait together instead of one after another:
With 200 milliseconds of scripted latency per call, twenty calls one after another take 4,000 milliseconds or more, and through the reranker they finished in under 1,000 (the transcript records both as true or false, not as raw timings, because timings change every run). That shows the structure is right. It does not model a real API, which adds network variance and rate limits (Virtual Threads vs Platform Threads covers why virtual threads suit blocking waits like this). The transcript (transcript 08) shows the order before and after, the timing checks, and what happens when a reply is not a bare integer:
candidates in: 4, model calls made: 4, chunks out: 2
order from the vector search, best first:
similarity 0.5891 page 2
similarity 0.4077 page 3
similarity 0.2535 page 1
similarity 0.2023 page 4
order after reranking, best first:
rerank_score 8 page 2
rerank_score 6 page 3
--- latency: 20 candidates, each rating call takes 200 ms (a Thread.sleep in the fake model) ---
the 20 calls one after another take 4000 ms or more: true
LlmReranker, one virtual thread per candidate, takes under 1000 ms: true
a real API adds its own rate limits, which this test cannot show
reply "Score: 8" for every candidate -> failures counted: 4 of 4
scores assigned: [0, 0]
pages kept, in order: [2, 3] (the vector-search order, because every score is 0)
reply " 9\n" (padded) -> score 9
The failure is silent. The reranker expects the model to reply with a bare integer. A reply of Score: 8 failed to parse for every candidate, every score became 0, and the four survivors were simply the vector-search order. Nothing threw; the only sign is the rag.rerank.failures counter. A padded " 9\n" parsed fine. Watch that counter, or ask the model for structured output instead of parsing text.
One more limit, stated plainly: the scripted rater is word overlap and agrees with the scripted embeddings, so this repository cannot show that reranking improves which chunks reach the prompt. It shows the mechanics and the cost.
Even with a strict prompt, a model sometimes says something the chunks do not. The last stage asks a second model call, a judge, whether the answer is supported by the chunks that were in the prompt. Spring AI has one, FactCheckingEvaluator, and the query service wraps it and reports one of three statuses:
The picture shows why no_context is its own status: when nothing was retrieved there is nothing to check against, so the judge is not called at all. When chunks were retrieved and the judge says “yes”, the status is answered. Anything else is ungrounded, including a judge call that fails, because an unchecked answer is not a supported one. The service:
private boolean isGrounded(String question, String answer, List<Document> context) {
try {
return factChecker.evaluate(new EvaluationRequest(question, context, answer)).isPass();
} catch (RuntimeException e) {
// If the judge itself fails, the answer is unchecked, which is not the same as supported.
metrics.counter("rag.faithfulness.judge_failures").increment();
return false;
}
}
(The isGrounded method from RagQueryService.java; the rest of ask() is in the code tour.) The five cases the test drives, with the scripted model deciding both the answer and the verdict (transcript 09):
--- 1. the answer is in the chunks ---
status=answered grounded=true sources=4
answer: Full-time employees are entitled to 20 working days
of annual leave per calendar year.
--- 2. the model answers with something the chunks do not say ---
status=ungrounded grounded=false sources=4
answer: Employees get 30 days of annual leave.
the check the judge model was given:
Evaluate whether or not the following claim is supported by the provided document.
Respond with "yes" if the claim is supported, or "no" if it is not.
--- 3. the judge says "Yes." instead of "yes" ---
status=ungrounded grounded=false sources=4
answer: Full-time employees are entitled to 20 working days
of annual leave per calendar year.
with the reply "YES": grounded=true
--- 4. the judge call itself fails ---
status=ungrounded grounded=false sources=4
answer: Full-time employees are entitled to 20 working days
of annual leave per calendar year.
rag.faithfulness.judge_failures = 1
--- 5. nothing is retrieved (threshold 0.3, off-topic question) ---
status=no_context grounded=false sources=0
answer: I don't have enough information in the provided documents.
fact-check calls made: 0
“Yes.” is not “yes”. The evaluator compares the judge’s reply to the word yes ignoring case but not punctuation: YES counted as grounded and Yes. did not. A real judge model that likes to end its answer with a full stop would fail every answer, and your dashboard would show a wall of ungrounded. Tell the judge to reply with one word, and watch the rate.
What this stage does not tell you: how good a real judge is at catching an invented claim. The tests script the verdict, so they prove the wiring and the four outcomes, not the judgement. It also adds one more model call to every question that finds any context. Measure the judge on questions where you know the answer is wrong before you let it decide anything.
Every stage above fails without an exception, so the only way to see it is to count. The service records a handful of Micrometer meters and exposes them through Actuator’s Prometheus endpoint (Spring Boot Actuator in Production covers exposing and securing endpoints). This is the part of the query method that records the outcome of every question, whatever happened:
(From RagQueryService.java.) The end-to-end test uploaded two files, uploaded one of them again, asked two questions over real HTTP, and read the endpoint back (transcript 10):
--- GET /actuator/prometheus (only the rag_ series; the timer's sum and max are left out because they change every run) ---
rag_chunks_ingested_total 5.0
rag_context_chunks_count 2
rag_context_chunks_sum 5.0
rag_context_chunks_max 4.0
rag_ingestion_skipped_total 1.0
rag_queries_total{status="answered"} 2.0
rag_query_duration_seconds_count{status="answered"} 2
Two files, five chunks written, one re-upload skipped, two questions answered, and the context-size summary says the retriever handed over four chunks for one question and one for the other. Read against the earlier sections, those meters are the alarms:
Meter
What a change means
Section
rag_queries_total{status=...}
A rising share of ungrounded or no_context: the judge is disagreeing with the model more often, or the threshold is rejecting more questions.
Checking the answer; Retrieval
rag_context_chunks
Staying at its maximum on every question means the threshold accepts everything.
Retrieval
rag_ingestion_skipped_total against rag_chunks_ingested_total
Chunks written on every restart means the tracker is not remembering anything.
Getting a PDF in
rag_rerank_failures_total
Any increase: ratings that did not parse are counting as 0.
Reranking
rag_faithfulness_judge_failures_total
Any increase: answers are being marked ungrounded because the judge call failed, not because they were wrong.
Checking the answer
Two of those never appeared in the transcript. Micrometer registers a counter when it is first incremented, so rag_rerank_failures_total and rag_faithfulness_judge_failures_total are absent from the endpoint until the first failure. A dashboard that plots a missing series as “no data” and an alert that fires on “increase > 0” behave differently on it; test yours.
The largest gap between this repository and a production system is not a meter. It has no evaluation set: a list of real questions with known correct answers that you run after every change to the chunker, the threshold or the model, and score. Every quality claim above is about mechanics for that reason, and nothing here can tell you whether your answers are getting better.
The meters, the alert suggestions and a ten-line production checklist: chapter 6.
This page replaces an article that was published as “Production-Grade RAG with Spring AI 1.0” and said Spring AI 1.1.0 in its body. The reason for the rewrite was that its code no longer worked on Spring AI 2.0. Only part of that turned out to be true. The repository keeps the earlier code as LegacyIngestion.java and runs it against both versions (transcript 14), and it reads its configuration keys against the jars’ metadata (transcript 14, transcript 15). The findings, worst first:
spring-ai-openai-spring-boot-starter latest: 1.0.0-M6
spring-ai-pgvector-store-spring-boot-starter latest: 1.0.0-M6
the ids that replaced them:
spring-ai-starter-model-openai latest: 2.0.1
spring-ai-starter-vector-store-pgvector latest: 2.0.1
## mvn validate on the article's dependency block with spring-ai-bom 2.0.1
'dependencies.dependency.version' for org.springframework.ai:spring-ai-openai-spring-boot-starter:jar is missing. @ line 31, column 17
'dependencies.dependency.version' for org.springframework.ai:spring-ai-pgvector-store-spring-boot-starter:jar is missing. @ line 35, column 17
First, the two starters in the earlier project file, spring-ai-openai-spring-boot-starter and spring-ai-pgvector-store-spring-boot-starter, were last published as 1.0.0-M6, a milestone. Neither Spring AI BOM manages them, so the earlier project could not be resolved with 1.1.0 or with 2.0.1. The names that replaced them are spring-ai-starter-model-openai and spring-ai-starter-vector-store-pgvector.
legacy-1x/src/LegacyIngestion.java:17: error: incompatible types: ExtractedTextFormatter is not a functional interface
.withPageExtractedTextFormatter(text -> text.replaceAll("s{3,}", " "))
^
legacy-1x/src/LegacyIngestion.java:22: error: no suitable constructor found for TokenTextSplitter(int,int,int,int,boolean)
TokenTextSplitter splitter = new TokenTextSplitter(512, 128, 5, 10_000, true);
^
Second, two lines of the earlier ingestion code do not compile (the compiler output above is from transcript 14, run against 2.0.1). The PDF page-cleanup lambda fails on both versions, because the method takes a final class, ExtractedTextFormatter, not a function. (Look closely at the quoted line: the regular expression reads s{3,}, where \s{3,} was clearly meant. The backslashes had been lost from the published article, which is why the repository copy has that too.) The other error is the one that is a Spring AI 2.0 change: the five-argument TokenTextSplitter constructor is now six arguments, and the builder is the way to configure it.
## the 1.x configuration keys in the metadata of spring-ai-autoconfigure-model-openai
1.1.0 spring.ai.openai.chat.options.model current
1.1.0 spring.ai.openai.chat.options.temperature current
1.1.0 spring.ai.openai.embedding.options.model current
2.0.1 spring.ai.openai.chat.options.model deprecated, use spring.ai.openai.chat.model
2.0.1 spring.ai.openai.chat.options.temperature deprecated, use spring.ai.openai.chat.temperature
2.0.1 spring.ai.openai.embedding.options.model deprecated, use spring.ai.openai.embedding.model
--- keys written the 1.x way ---
spring.ai.openai.chat.options.model DEPRECATED, use spring.ai.openai.chat.model
spring.ai.openai.chat.options.temperature DEPRECATED, use spring.ai.openai.chat.temperature
spring.ai.openai.embedding.options.model DEPRECATED, use spring.ai.openai.embedding.model
spring.ai.openai.chat.optoins.model UNKNOWN
Third, the earlier application.yml (the two blocks above are from transcript 14 and transcript 15) used chat.options.model, chat.options.temperature and embedding.options.model. In the Spring AI 1.1.0 configuration metadata those keys are current; in 2.0.1 they are deprecated, with flat replacements (chat.model, chat.temperature, embedding.model). That is the second 2.0 change in this list. This repository read the metadata and did not start an application with the old keys, so “deprecated” is all it can say. A misspelt key, the last line above, is reported as unknown; Spring Boot binds by name, and a key that nothing binds to is normally ignored without an error (also not tested here). The migration guide (Spring AI 1.x to 2.0 Migration Guide) says the .options prefix is gone for some model types and does not mention chat; on this evidence, for chat it is deprecated and not removed.
Fourth, what was not code but was wrong or unsupported, and is gone from this page: a table of retrieval-precision percentages for three chunking strategies that nothing measured; the claim that Spring AI supports a semantic response cache through SemanticSearchCacheAdvisor (a scan of every jar on the project’s classpath finds no class with that name; the count is the line beginning “classes in any jar on this project’s classpath” in transcript 12); sample outputs that no run had produced; a tenant filter built by string concatenation, the injectable version shown in the tenant section; a reranker that used parallelStream() for blocking model calls; and specific scale and latency figures for vector stores that were quoted without a source.
The premise was half right. The request that started this rewrite was “the code has been broken since 2.0 GA”. Two of the four problems are 2.0 changes: the splitter constructor and the deprecated keys. The other two, the starter names and the lambda, were broken before 2.0 shipped, and the article’s own text disagreed with its title about which 1.x it meant. The lesson worth keeping is the method: every Java block on this page is a slice of a file that the repository’s build compiles, so a snippet that goes stale stops compiling instead of quietly staying on the page.
Every stage above costs something, in code, in latency or in model calls. This table is the decision in one place. It is not a ranking of quality, because nothing here measured quality.
Your situation
What to use
What it costs
A prototype, a handful of documents, every question is answerable from them
QuestionAnswerAdvisor with a similarity threshold
One embedding call per question
Users will ask things the documents cannot answer
RetrievalAugmentationAdvisor with allowEmptyContext(false) and a threshold you measured
The same, plus finding the threshold for your embedding model
More than one customer, team or access level shares the store
A tenant tag at ingestion and a FilterExpressionBuilder filter on every search
Test filtered queries on a table of production size
Files are uploaded again, or edited, or the service restarts
Track a hash and the chunk ids per file; persist the tracker
A table to keep, or a delete by source_file filter
The right chunk is retrieved but ranks below the wrong ones
Retrieve more, re-rate, keep the best few
One model call per candidate, and a failure counter
A wrong answer is expensive
A judge after the answer, shown to the user as such
About one more model call per question, and an evaluation set to trust it
The production checklist that goes with the table: chapter 6.
This page ran PostgreSQL with pgvector, because it needs no new infrastructure if you already run PostgreSQL, and the pgvector starter is one dependency (transcript 13). Spring AI has starters for many other stores and the interface (VectorStore) is the same, so the code in the ingestion and query sections would not change. Nothing here compares stores, and the scale and latency figures the earlier version quoted for pgvector, Qdrant and Weaviate had no source, so they are gone. Test with your own volume.
How do I keep one tenant’s documents away from another’s?
Tag every chunk with the tenant when it is ingested, and pass a filter built with FilterExpressionBuilder on every search, never a string built from the request. The hostile-input results are in transcript 07, and the untested cost of filtering an HNSW index is in the callout in that section.
How do I handle documents that change?
Remember, per file, a hash of its bytes and the ids of the chunks you wrote; skip identical bytes; on a change delete the old ids before adding the new chunks. The rows and embedding counts for each case are in transcript 05. Persist what you remember, or delete by a source_file filter, so that a restart does not orphan the old chunks.
What chunk size should I use?
The earlier version answered this for code documentation with a specific range and a rule about splitting at functions. Nothing in this repository tested either, so that answer is gone. What is measured is how the splitters behave on a known document (transcript 02, transcript 03); the right size for your documents is what a set of your own questions retrieves best.
Is Spring AI ready for production?
Maven Central dates Spring AI 2.0.0 at 12 June 2026 and 2.0.1 at 20 August 2026, so 2.0 is a released major version. It also changed enough that code written for 1.x needs edits (Spring AI 1.x to 2.0 Migration Guide). What this repository shows is that the pieces can be wired together and observed; it never called a real model, so it is not evidence about how the whole behaves under production load or with a real model’s answers.
Conclusion
A RAG service works when nothing between the document and the answer is trusted to behave well by default. The defaults here lean the other way: the splitter cuts without overlap, the retriever accepts every chunk, the simple advisor sends an empty context, the filter takes whatever string it is given, the judge is picky about a full stop, and the reranker written here turns a reply it cannot parse into a zero. Each of them was found by running the code and printing what it did, which is also how the earlier article’s broken dependencies and lambda turned up. Set the threshold, refuse to answer on an empty context, build filters as values, remember what you have ingested, count the failures, and build an evaluation set, which is the one thing this repository does not have.
Should you build this yourself? If you have a few dozen documents and questions they can all answer, the single advisor in the second section is enough and the rest is machinery you do not need yet. If several customers share the data, or wrong answers cost money, the tenant filter, the empty-context rule and the ingestion tracker are not optional, and the reranker and judge are worth trying once you have an evaluation set to see whether they help. If you do not want to run any of this, a managed retrieval service does the same job with a different set of trade-offs; nothing on this page compares them. The complete code is walked through in Spring AI RAG in Java: Complete Runnable Code.
No Comments yet!