Skip to main content

Testing LLM Apps in Java: Spring AI Evaluators and LLM-as-Judge in JUnit 6

How to test an LLM application in Java with Spring AI’s RelevancyEvaluator and FactCheckingEvaluator: what they send to the judge, why “Yes.” fails, golden datasets, pass-rate gates and a free deterministic judge for CI.

You change a prompt, or the model version, or the retrieval settings, and the application still compiles, still starts, and still returns a fluent answer. Whether the answers got worse is the one thing the compiler cannot tell you. For ordinary code you would write a test. For an LLM application the output is different every run and “correct” is a judgement about meaning, so a plain assertEquals does not work. Spring AI ships a small answer to this: evaluators, which ask a second model to judge the first one’s output. This article takes them apart. It shows exactly what the two built-in evaluators send to the judge, the one-word trap in how they read its reply, how to turn them into a regression suite with a golden dataset, how to stop a noisy judge from making your build flaky, and how to run the whole thing in continuous integration without spending money. Depth is in expandable sections, so you can read straight through or open only what you need.
Versions, and an honest limit. Spring Boot 4.1.1, Spring AI 2.0.1, JUnit 6.0.3 (the version Boot manages; 6.1.3 is the newest on Maven Central) and Java 25. All the code is in the evaluation module of asmhatre/spring-ai, and every console block below is quoted from a file under its output/ directory, written by a test that asserts the same lines. No live model was used. The application’s answers are scripted, and the judge is a stub, a deterministic word-overlap judge I wrote, or a seeded noise simulation, each labelled where it appears. So this article is evidence about how the evaluators and your test suite behave. It is not evidence about how well any real model judges; the one test that would call one (LiveJudgeTest) is skipped here for lack of an API key, and I have not run it.

An evaluator is a test where a second model does the assertion

Start with the problem in your own words. You have an assistant that answers support questions from a few retrieved documents. A customer asks “how long do I have to request a refund on an annual plan?” and the documents say 30 days. Three different things can go wrong: the assistant can answer a different question (relevance), say something the documents do not support (a hallucination, which the field calls a groundedness or fact-checking failure), or leave out the one number that mattered. The wording can be perfect in all three cases. Spring AI’s answer is the Evaluator interface. It has one method: you give it an EvaluationRequest (the question, the documents the assistant was shown, and the assistant’s answer) and you get an EvaluationResponse (a pass flag, a score and some feedback). The two built-in implementations, RelevancyEvaluator and FactCheckingEvaluator, do their work by sending a prompt to a judge: another ChatModel, often the same one you are testing, asked to answer yes or no. That is what “LLM-as-judge” means. Everything in this article is a consequence of that sentence.
Golden casequestion + docsYour appanswers from docsEvaluatorbuilds the promptJudge model“yes” or “no”Verdictpass, scoreGatepass rate ≥ threshold?Two models are involved, and either can be wrong. The suite you are about to build is mostly about not being fooled by the second one.
Read the diagram left to right. The golden case and your application are the familiar half: input in, answer out. The evaluator and judge are the new half, and the gate on the right is what makes it a test rather than a report. A single verdict is rarely what you act on; the fraction of cases that pass is, and a later section explains why.
Going deeper: the three types you will touch
Evaluator is a one-method interface in spring-ai-commons (package org.springframework.ai.evaluation) with a default doGetSupportingData that maps the request’s documents to their text, applies a filter I did not decode, and joins what is left with System.lineSeparator() (so the prompt’s line breaks are platform-dependent). EvaluationRequest has three constructors: question plus answer, documents plus answer, or all three. EvaluationResponse holds pass, a float score, a feedback string and a metadata map. I read those signatures with javap against the 2.0.1 jars rather than from the docs. The two built-in evaluators live in spring-ai-client-chat (org.springframework.ai.chat.evaluation). Both take a ChatClient.Builder, not a ChatModel, so you can give the judge its own system prompt, advisors and options.

What the two built-in evaluators actually send

The application under test is deliberately small, so the evaluators are the interesting part. SupportAssistant.java takes a question and the retrieved documents, and asks the model to answer using only that context:
    static final String SYSTEM = "Answer using only the context. If the context does not contain the answer, say you do not know.";

    private final ChatClient client;

    public SupportAssistant(ChatClient.Builder builder) {
        this.client = builder.defaultSystem(SYSTEM).build();
    }

    public String answer(String question, List<Document> context) {
        String joined = context.stream().map(Document::getText).collect(Collectors.joining("\n"));
        return client.prompt()
                .user(u -> u.text("Context:\n{context}\n\nQuestion: {question}")
                        .param("context", joined)
                        .param("question", question))
                .call()
                .content();
    }
In a real application a QuestionAnswerAdvisor would fetch those documents from a vector store. Passing them in as an argument means a test controls precisely what the model was shown, which is what lets an evaluator check the answer against it. Now the part the documentation describes only loosely. I pointed both evaluators at a recording stub that replies “yes” and printed what it was sent (01-evaluator-prompts.txt, from EvaluatorAnatomyTest.java):
--- RelevancyEvaluator: prompt sent to the judge (1 message(s)) ---
	Your task is to evaluate if the response for the query
	is in line with the context information provided.

	You have two options to answer. Either YES or NO.

	Answer YES, if the response for the query
	is in line with context information otherwise NO.

	Query:
	How long do I have to request a refund on an annual plan?

	Response:
	You can request a refund on an annual plan within 30 days of purchase.

	Context:
	Annual plans can be refunded within 30 days of purchase.
	Monthly plans are not refundable.

	Answer:

--- RelevancyEvaluator: judge said "yes" ---
pass=true score=1.0 feedback='' metadata={}
Two things stand out. First, RelevancyEvaluator does not check whether the answer is on topic; its prompt asks whether the response is “in line with the context information provided”. That is a groundedness check, so the two built-in evaluators overlap far more than their names suggest. Second, the documents are joined with a line break and nothing else: no numbering, no source names, so a judge cannot tell you which document disagreed. The fact-checking evaluator sends a shorter prompt built from a document and a claim, and returns a verdict the same way (same 01-evaluator-prompts.txt):
--- FactCheckingEvaluator: prompt sent to the judge ---
	Evaluate whether or not the following claim is supported by the provided document.
	Respond with "yes" if the claim is supported, or "no" if it is not.

	Document:
	Annual plans can be refunded within 30 days of purchase.
	Monthly plans are not refundable.

	Claim:
	You can request a refund on an annual plan within 30 days of purchase.

--- FactCheckingEvaluator: judge said "yes" ---
pass=true score=0.0 feedback='' metadata={}

--- RelevancyEvaluator: judge said "no" ---
pass=false score=0.0 feedback=''
Two quiet surprises in the response. FactCheckingEvaluator returns pass=true score=0.0 for a pass: it builds its response without a score, so the field defaults to zero. If you chart or average getScore() across fact-checking results you will average zeros. RelevancyEvaluator does set a score, but only ever 1.0 or 0.0. And on a fail, feedback is an empty string from both: the judge is asked for one word, so there is no explanation to return. Use isPass(), and build your own message if you need a reason.
Going deeper: reading the evaluators’ bytecode
I read the evaluator classes with javap -c -constants on spring-ai-client-chat-2.0.1.jar rather than from the documentation. RelevancyEvaluator.evaluate renders its prompt template with the parameters query, response and context, makes one call, strips the reply and compares it with "yes" using equalsIgnoreCase; a match gives (true, 1.0), anything else gives (false, 0.0). FactCheckingEvaluator does the same with document and claim and the three-argument EvaluationResponse constructor, which has no score. Its static forBespokeMinicheck factory swaps in a shorter prompt (just the document and the claim) intended for models fine-tuned for fact-checking; I confirmed the prompt text in the bytecode but did not run that variant.

The judge has to say exactly “yes”, or your test fails

The comparison in the last section is the bug waiting for you. The evaluators strip whitespace from the judge’s reply and require it to equal yes, ignoring case. A model that is helpful and polite says “Yes.” or “Yes, the response is in line with the context.” Those are failures, with no error and no warning, just a false negative that you will read as the application regressing. I fed the evaluators a range of replies (02-verdict-parsing.txt):
judge reply                                          relevancy fact-chk  wrapped  
'yes'                                                true      true      true     
'YES'                                                true      true      true     
'  yes\n'                                            true      true      true     
'Yes.'                                               false     false     true     
'yes!'                                               false     false     true     
'Yes, the response is in line with the context.'     false     false     true     
'**Yes**'                                            false     false     true     
'no'                                                 false     false     false    
'No.'                                                false     false     false    
''                                                   false     false     false    
The first two columns are the two built-in evaluators; they agree on every row. Only the bare word, in any case and with any surrounding whitespace, passes. The third column is a thirty-line fix, VerdictNormalizingModel.java, which wraps the judge ChatModel and reduces the reply to its first word with the punctuation removed:
    @Override
    public ChatResponse call(Prompt prompt) {
        String reply = delegate.call(prompt).getResult().getOutput().getText();
        String first = reply == null ? "" : reply.strip().split("\\s+", 2)[0].replaceAll("[^A-Za-z]", "").toLowerCase(Locale.ROOT);
        String verdict = first.equals("yes") || first.equals("no") ? first : reply;
        return new ChatResponse(List.of(new Generation(new AssistantMessage(verdict))));
    }
Wrap the judge before you hand it to ChatClient.builder(...) and “Yes.”, “yes!”, “**Yes**” and the sentence form all pass, while “No.” and the empty reply still fail. It deliberately leaves any reply that does not begin with yes or no untouched, so a confused judge still produces a failure rather than a guess.
Prefer fixing the prompt and normalising. A stricter instruction (“answer with exactly one word”) reduces chatty replies, and the built-in prompts already ask for yes or no. But models, versions and temperature all change how closely they obey, and the failure mode is silent. The wrapper costs nothing and makes the dependence on exact wording disappear.
Going deeper: why the normaliser is a ChatModel and not a string helper
The evaluators call ChatClient.prompt().user(...).call().content() internally, and there is no hook for post-processing the string. Wrapping the model keeps the built-in evaluators untouched, so you keep their prompts and their upgrades. An advisor on the judge’s ChatClient.Builder would also work and is the more idiomatic Spring AI route; I used a model wrapper because it makes the test of the wrapper a three-line unit test with no client setup.

A golden dataset turns one-off checks into a regression suite

One evaluated answer is a demo. A golden dataset is what makes it a test: a fixed list of questions, each with the documents retrieval should supply and the facts a good answer must contain. You run the whole list on every change and compare. The dataset here has twelve rows about a fictional product, kept as JSON next to the tests:
  {"id": "refund-annual", "question": "How long do I have to request a refund on an annual plan?",
   "context": ["Annual plans can be refunded within 30 days of purchase. Monthly plans are not refundable."], "expectedFacts": ["30 days"]},
That is the first row of support-golden.json. EvalRunner.java walks the rows: it asks the assistant, then applies three checks to each answer. The first two are the built-in evaluators with a judge model. The third, ContainsFactsEvaluator.java, is plain Java that asks whether the answer contains each expected fact as text. It makes no model call and cannot be fooled by a fluent wrong number, which will matter in a moment.
    public EvalReport run(List<GoldenCase> cases) {
        return new EvalReport(cases.stream().map(this::runOne).toList());
    }

    private CaseResult runOne(GoldenCase c) {
        String answer = assistant.answer(c.question(), c.documents());
        EvaluationRequest request = new EvaluationRequest(c.question(), c.documents(), answer);
        boolean relevant = relevancy.evaluate(request).isPass();
        boolean grounded = factChecking.evaluate(request).isPass();
        boolean facts = new ContainsFactsEvaluator(c.expectedFacts()).evaluate(request).isPass();
        return new CaseResult(c.id(), answer, relevant, grounded, facts);
    }
To see the suite do its job I ran it against two versions of the application. The healthy build gives sensible answers. The regressed build simulates a bad prompt change with three wrong answers: a hallucinated 60-day refund window, an off-topic reply, and a plausible but wrong price (12 USD instead of 8). The answers are scripted (Answers.java), so this proves the suite catches those three kinds of failure, not that any real model makes them. The result is 03-golden-run.txt:
pass rate: 1.00   failing: []

== regressed build ==
case               relevant  grounded  has facts passed 
refund-annual      false     false     false     false  
refund-monthly     true      true      true      true   
storage-pro        true      true      true      true   
support-hours      true      true      true      true   
sso-plan           false     false     false     false  
data-region        true      true      true      true   
api-rate-limit     true      true      true      true   
backup-retention   true      true      true      true   
macos-agent        true      true      true      true   
reset-link         true      true      true      true   
extra-seat         true      true      false     false  
trial-card         true      true      true      true   
pass rate: 0.75   failing: [refund-annual, sso-plan, extra-seat]

== the gate ==
healthy build   : requirePassRate(0.90) returned normally
regressed build : java.lang.AssertionError: pass rate 0.75 is below the 0.90 threshold; failing cases: [refund-annual, sso-plan, extra-seat]
The first line of the block is the healthy build’s result: all twelve cases pass. The table below it is the regressed build, which fails three. Look at the three check columns for the failing rows. The hallucinated answer fails all three checks. The off-topic reply fails all three. But the wrong price, 12 USD per month, passes both judge checks and is caught only by the plain-code fact check, because five of its six content words appear in the context. A judge that reads for gist will let a changed number through; an exact check will not. That is the reason the suite uses both kinds. The last two lines of the file show the gate: EvalReport.requirePassRate(0.90) returns quietly for the healthy build and throws an AssertionError naming the three failing cases for the regressed one.
    /** Fails the build (with an {@link AssertionError} naming the failing cases) when the pass rate is below the threshold. */
    public void requirePassRate(double threshold) {
        if (passRate() < threshold) {
            throw new AssertionError("pass rate %.2f is below the %.2f threshold; failing cases: %s".formatted(passRate(), threshold, failedIds()));
        }
    }
Going deeper: one JUnit invocation per row, and when not to
GoldenDatasetTest.java also contains a @ParameterizedTest with a @MethodSource over the case ids, so each golden row appears as its own named invocation in any test report: [1] refund-annual, [2] refund-monthly, and so on. That is the nicest reporting you can get, and it is correct only when the judge is deterministic. With a real model, one flipped verdict on one row fails that invocation and the build, which is exactly the flakiness the next section is about. For a real judge, run the loop in one test method and assert on the aggregate.

Gate on a pass rate, because a model judge is noisy

A deterministic judge gives the same verdict every run. A model does not, even at temperature zero across provider updates, and a judge that is right 97 times in 100 will still turn a twelve-case suite red surprisingly often if you demand that every case passes. The arithmetic is unforgiving: with two judge checks per case and 3 percent of verdicts wrong, the chance that a perfectly healthy build gets through all twelve cases untouched is only about half. I measured that effect, with a clearly marked simulation. FlakyJudge.java wraps a correct judge and flips its verdict with a fixed probability from a seeded random generator, so every run is repeatable. For each noise level I ran the suite 200 times against the healthy build and 200 times against the regressed one, and counted how often each gate said “fail” (04-noisy-judge.txt):
noise-free pass rate: healthy build 1.00, regressed build 0.75

== 3% of judge verdicts flipped (mean pass rate: healthy 0.947, regressed 0.706) ==
gate                   healthy build fails it       regressed build fails it
every case must pass   96 of 200 runs               200 of 200 runs
pass rate >= 0.90      29 of 200 runs               200 of 200 runs
pass rate >= 0.85      29 of 200 runs               200 of 200 runs
pass rate >= 0.80      2 of 200 runs                200 of 200 runs

== 8% of judge verdicts flipped (mean pass rate: healthy 0.859, regressed 0.640) ==
gate                   healthy build fails it       regressed build fails it
every case must pass   166 of 200 runs              200 of 200 runs
pass rate >= 0.90      102 of 200 runs              200 of 200 runs
pass rate >= 0.85      102 of 200 runs              200 of 200 runs
pass rate >= 0.80      47 of 200 runs               200 of 200 runs
every case96rate ≥ 0.9029rate ≥ 0.802healthy builds wrongly failed, out of 200 runs, at 3% simulated judge noise (fewer is better)
The chart is the 3 percent block. Requiring every case to pass failed the healthy build in 96 of 200 runs; a coin flip. A 0.80 gate failed it twice, and still caught the regressed build all 200 times, because the regressed build’s true pass rate of 0.75 sits below the line even after noise. Two details are easy to miss. The 0.90 and 0.85 rows are identical because with twelve cases the only possible pass rates are multiples of 1/12 (0.917, then 0.833), so any threshold between those values is the same gate. And at 8 percent noise even the 0.80 gate raised 47 false alarms, which is the honest answer to “what threshold should I use”: the one you set from the noise you measured. Run your suite repeatedly against a known-good build, look at the spread, and put the line below it and above the regressed build.
This is a simulation, not a measurement. The flip rates are numbers I chose; I do not know how often any real judge model errs on your questions. What the experiment shows is how a gate rule behaves given a noise level, and that a strict all-pass rule is hostile to any noise at all. The way to learn your own noise level is to run the same golden set against the same build many times with your real judge, which is exactly what the nightly layer in the next section is for.
Going deeper: ways to quiet a real judge
Set the judge’s temperature to zero, and give it a model that is not the one being tested (a model grading its own output tends to be lenient, a widely reported effect that I did not test here). Use a larger golden set so one flipped verdict moves the pass rate less, run each case several times and take the majority, and keep a plain-code check beside every judge check so a flip in one cannot hide a real regression in the other. The table above also argues for a mild gate over a strict one: you are monitoring a rate, not proving a property.

Keep pull requests free: a deterministic judge, and what it gets wrong

Every judge call is a model call, and a model call is money and a network dependency. A suite that needs an API key to run is a suite that does not run on a contributor’s laptop, or on a fork’s pull request, or when the provider has an outage. The way out is to treat the judge like any other dependency you replace in tests. Because the built-in evaluators take a ChatClient.Builder, and a builder can wrap any ChatModel, you can hand them a fake judge and still run the real evaluators, the real prompts and the real verdict parsing. RuleBasedJudge.java is such a fake. It reads the claim and the evidence out of the prompt and answers “yes” when at least 60 percent of the claim’s content words and numbers also appear in the evidence:
    public ChatResponse call(Prompt prompt) {
        String text = prompt.getInstructions().getLast().getText();
        String claim;
        String evidence;
        Matcher r = RELEVANCY.matcher(text);
        Matcher f = FACT.matcher(text);
        if (r.find()) {
            claim = r.group(1);
            evidence = r.group(2);
        }
        else if (f.find()) {
            evidence = f.group(1);
            claim = f.group(2);
        }
        else {
            throw new IllegalArgumentException("RuleBasedJudge does not recognise this prompt: " + text);
        }
        return new ChatResponse(List.of(new Generation(new AssistantMessage(overlap(claim, evidence) >= minimumOverlap ? "yes" : "no"))));
    }

    /** Share of the claim's content words that also appear in the evidence. */
    public static double overlap(String claim, String evidence) {
        Set<String> have = words(evidence);
        Set<String> need = words(claim);
        if (need.isEmpty()) {
            return 0;
        }
        long hit = need.stream().filter(have::contains).count();
        return hit / (double) need.size();
    }
It understands nothing, so its limits matter more than its strengths. I labelled five answers to one question by hand and asked it to judge them (06-judge-limits.txt):
answer type                correct?  judge says     verdict
correct, same words        true      pass           right
WRONG: negation dropped    false     pass           WRONG
WRONG: roles swapped       false     pass           WRONG
correct, paraphrased       true      fail           WRONG
correct, but a refusal     true      fail           WRONG
A word-overlap judge accepts a dropped negation (“Monthly plans are refundable” shares every content word with the truth), accepts swapped roles, and rejects a correct paraphrase and an honest refusal. Those are the exact failures a real model judge is supposed to catch, which is the argument for having two layers rather than replacing one with the other.
1. Unit: prompts and wiring — every buildstub model, no evaluator, milliseconds2. Golden run with RuleBasedJudge — every pull requestreal evaluators and prompts, fake judge: 12 app calls + 24 judge calls, 0 network3. Golden run with a real judge — nightly, key requiredsame dataset and evaluators; measures the noise; may cost money
The picture is the pyramid I would use. Layer two is the new idea: the golden dataset and the evaluators run on every pull request, at no cost, and catch the regressions a cheap judge can catch (a swapped number, an off-topic reply). Layer three is the same loop with a real judge, which LiveJudgeTest.java implements. I counted the calls in 07-ci-layers.txt:
layer                                        app calls    judge calls  network
1 unit: prompt and wiring, no evaluator      1            0            none
2 golden run, RuleBasedJudge (pull request)  12           24           none
3 golden run, real judge (nightly)           n/a          n/a          needs OPENAI_API_KEY; skipped here
pass rate at layer 2: 1.00
The live test is annotated so that it can never run by accident. It is skipped unless OPENAI_API_KEY is set, and the skip is recorded by surefire (08-live-judge-skipped.txt):
testcase: "realJudgeAgreesWithTheKnownGoodAndKnownBadBuilds"
skipped : "Environment variable [OPENAI_API_KEY] does not exist"
Be exact about what this proves. The skipped test compiled, and that is all I can claim for it. No result from a real judge model appears anywhere in this article. When you run it with a key, the first thing to read is not the pass rate but the spread across repeated runs.
Going deeper: choosing a threshold for the overlap rule
The 0.6 in new RuleBasedJudge(0.6) is a tuned constant, and tuning it is a small instance of the whole problem. Too low and an unrelated answer sharing a few words passes; too high and a terse correct answer fails. I picked 0.6 because every good answer in the golden set scores above it and each regression scores below it, which is a fact about this dataset. When your dataset changes, rerun it. The point of the layer is to be cheap and repeatable, not right.

Grade from one to five, and combine checks to get a reason

Yes or no throws away information, and an empty feedback field throws away the reason. Two small custom evaluators fix both. GradedEvaluator.java asks the judge for a grade from 1 to 5 and passes at or above a minimum. Because models rarely return a bare digit, it pulls the first standalone 1–5 out of whatever comes back:
    public EvaluationResponse evaluate(EvaluationRequest request) {
        String reply = builder.build().prompt()
                .user(PROMPT.formatted(request.getUserText(), request.getResponseContent(), doGetSupportingData(request)))
                .call().content();
        Matcher m = GRADE.matcher(reply == null ? "" : reply);
        if (!m.find()) {
            return new EvaluationResponse(false, 0f, "judge did not return a grade: " + reply, Map.of());
        }
        int grade = Integer.parseInt(m.group(1));
        return new EvaluationResponse(grade >= minimumGrade, grade / 5f, "grade " + grade + " of 5", Map.of("grade", grade));
    }
CompositeEvaluator.java runs several evaluators on one request. It passes only if all of them do, averages the scores and lists which ones failed in the feedback. Here are both in 05-graded-and-composite.txt:
judge reply                                  pass   score  feedback
'5'                                          true   1.0    grade 5 of 5
'4/5'                                        true   0.8    grade 4 of 5
'Score: 4 - mostly right, misses the date'   true   0.8    grade 4 of 5
'3'                                          false  0.6    grade 3 of 5
'I would say excellent'                      false  0.0    judge did not return a grade: I would say excellent
'10'                                         false  0.0    judge did not return a grade: 10

== CompositeEvaluator on the regressed 'extra-seat' answer ("12 USD" instead of "8 USD") ==
answer   : Extra seats cost 12 USD per month.
pass     : false
score    : 0.5
feedback : contains-facts failed (missing: [8 USD]);
verdicts : {relevancy=true, contains-facts=false}
The grading rows show the tolerant parse: “4/5” and “Score: 4 – mostly right, misses the date” both read as 4, while “10” and a reply with no number are failures that carry the judge’s raw text as feedback, so you can see what went wrong. The composite row is the regressed price answer from earlier: it fails with contains-facts failed (missing: [8 USD]) and a metadata map showing the relevancy judge said yes. That one line is what you want in a failing test report.
Going deeper: using the metadata map
EvaluationResponse carries a Map<String, Object> that the built-in evaluators leave empty. The composite uses it for per-evaluator verdicts; you could add the judge’s raw reply, the prompt, token usage or latency. Be careful with prompts in test reports if your golden set contains customer text.

Should you even do this?

Yes, but start smaller than you think. A golden dataset of fifteen rows you have read, a plain-code check for every fact that matters, and a pass-rate gate with a mild threshold will catch most of the regressions that matter. Add a model judge for the things code cannot check (tone, completeness, paraphrase) and treat its verdicts as a noisy signal to watch, not a proof. Do not let a model judge be the only thing standing between a prompt change and production, and do not make every pull request wait for one. What this article did not test: any real judge model, agreement between a model judge and human labels, bias from a model grading itself, cost per evaluation run, or evaluation of tool-calling and multi-turn conversations. Those are the questions to answer with your own data before you trust the gate.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.