Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
3.5 KiB
evaluation
Companion code for Testing LLM Apps in Java: Spring AI Evaluators and LLM-as-Judge in JUnit 6, part of the Spring AI series on ankurm.com.
A small retrieval-grounded support assistant, a 12-row golden dataset, and the tests that show what Spring AI's RelevancyEvaluator and FactCheckingEvaluator really send and accept, how to gate a build on a pass rate, and how to keep a model-graded test suite free and repeatable in CI.
No live model is used anywhere. The "application model" is a scripted stub and the "judge" is either a recording stub or RuleBasedJudge, a deterministic word-overlap judge. Judge noise in NoisyJudgeTest is a seeded simulation. The one test that talks to a real model, LiveJudgeTest, is skipped unless OPENAI_API_KEY is set and was never run for the article. So nothing here says how well any real model judges.
Versions
| Component | Version |
|---|---|
| Spring Boot | 4.1.1 |
| Spring AI | 2.0.1 (spring-ai-client-chat) |
| JUnit | 6.0.3 (managed by Boot 4.1.1; 6.1.3 is the latest on Maven Central) |
| Java | 25 (LTS) |
Quickstart
scripts/run-all.sh # runs 21 tests (1 skipped) and regenerates output/01 .. 08
Two consecutive runs produce byte-identical files.
What's here
| File | What it shows |
|---|---|
SupportAssistant.java |
The application under test |
EvalRunner.java, EvalReport.java |
Run the golden set through three evaluators; requirePassRate is the build gate |
RuleBasedJudge.java |
A deterministic ChatModel that the real evaluators run against in CI |
VerdictNormalizingModel.java |
Turns "Yes." into "yes", which is all the built-in evaluators accept |
GradedEvaluator.java |
A 1-5 judge, because the built-ins only return 0 or 1 |
ContainsFactsEvaluator.java, CompositeEvaluator.java |
A no-model check, and a combiner that restores the feedback the built-ins leave empty |
src/test/resources/golden/support-golden.json |
The golden dataset |
Output files
| File | Written by |
|---|---|
01-evaluator-prompts.txt |
EvaluatorAnatomyTest: the exact prompts, and the responses |
02-verdict-parsing.txt |
VerdictParsingTest: which judge replies pass |
03-golden-run.txt |
GoldenDatasetTest: a healthy and a regressed build, and the gate |
04-noisy-judge.txt |
NoisyJudgeTest: simulated noise against five gates |
05-graded-and-composite.txt |
GradedCompositeTest |
06-judge-limits.txt |
JudgeLimitsTest: where the deterministic judge is wrong |
07-ci-layers.txt |
CiLayersTest: model calls per testing layer |
08-live-judge-skipped.txt |
cut from the surefire report by run-all.sh |