Files

evaluation

Companion code for Testing LLM Apps in Java: Spring AI Evaluators and LLM-as-Judge in JUnit 6, part of the Spring AI series on ankurm.com.

A small retrieval-grounded support assistant, a 12-row golden dataset, and the tests that show what Spring AI's RelevancyEvaluator and FactCheckingEvaluator really send and accept, how to gate a build on a pass rate, and how to keep a model-graded test suite free and repeatable in CI.

No live model is used anywhere. The "application model" is a scripted stub and the "judge" is either a recording stub or RuleBasedJudge, a deterministic word-overlap judge. Judge noise in NoisyJudgeTest is a seeded simulation. The one test that talks to a real model, LiveJudgeTest, is skipped unless OPENAI_API_KEY is set and was never run for the article. So nothing here says how well any real model judges.

Versions

Component Version
Spring Boot 4.1.1
Spring AI 2.0.1 (spring-ai-client-chat)
JUnit 6.0.3 (managed by Boot 4.1.1; 6.1.3 is the latest on Maven Central)
Java 25 (LTS)

Quickstart

scripts/run-all.sh     # runs 21 tests (1 skipped) and regenerates output/01 .. 08

Two consecutive runs produce byte-identical files.

What's here

File What it shows
SupportAssistant.java The application under test
EvalRunner.java, EvalReport.java Run the golden set through three evaluators; requirePassRate is the build gate
RuleBasedJudge.java A deterministic ChatModel that the real evaluators run against in CI
VerdictNormalizingModel.java Turns "Yes." into "yes", which is all the built-in evaluators accept
GradedEvaluator.java A 1-5 judge, because the built-ins only return 0 or 1
ContainsFactsEvaluator.java, CompositeEvaluator.java A no-model check, and a combiner that restores the feedback the built-ins leave empty
src/test/resources/golden/support-golden.json The golden dataset

Output files

File Written by
01-evaluator-prompts.txt EvaluatorAnatomyTest: the exact prompts, and the responses
02-verdict-parsing.txt VerdictParsingTest: which judge replies pass
03-golden-run.txt GoldenDatasetTest: a healthy and a regressed build, and the gate
04-noisy-judge.txt NoisyJudgeTest: simulated noise against five gates
05-graded-and-composite.txt GradedCompositeTest
06-judge-limits.txt JudgeLimitsTest: where the deterministic judge is wrong
07-ci-layers.txt CiLayersTest: model calls per testing layer
08-live-judge-skipped.txt cut from the surefire report by run-all.sh