# evaluation Companion code for [Testing LLM Apps in Java: Spring AI Evaluators and LLM-as-Judge in JUnit 6](https://ankurm.com/spring-ai-2-0-testing-llm-apps-evaluators-llm-as-judge-junit-6/), part of the [Spring AI series](../README.md) on ankurm.com. A small retrieval-grounded support assistant, a 12-row golden dataset, and the tests that show what Spring AI's `RelevancyEvaluator` and `FactCheckingEvaluator` really send and accept, how to gate a build on a pass rate, and how to keep a model-graded test suite free and repeatable in CI. **No live model is used anywhere.** The "application model" is a scripted stub and the "judge" is either a recording stub or `RuleBasedJudge`, a deterministic word-overlap judge. Judge noise in `NoisyJudgeTest` is a seeded simulation. The one test that talks to a real model, `LiveJudgeTest`, is skipped unless `OPENAI_API_KEY` is set and was never run for the article. So nothing here says how well any real model judges. ## Versions | Component | Version | |---|---| | Spring Boot | 4.1.1 | | Spring AI | 2.0.1 (`spring-ai-client-chat`) | | JUnit | 6.0.3 (managed by Boot 4.1.1; 6.1.3 is the latest on Maven Central) | | Java | 25 (LTS) | ## Quickstart ```bash scripts/run-all.sh # runs 21 tests (1 skipped) and regenerates output/01 .. 08 ``` Two consecutive runs produce byte-identical files. ## What's here | File | What it shows | |---|---| | [`SupportAssistant.java`](src/main/java/com/ankurm/evaluation/SupportAssistant.java) | The application under test | | [`EvalRunner.java`](src/main/java/com/ankurm/evaluation/EvalRunner.java), [`EvalReport.java`](src/main/java/com/ankurm/evaluation/EvalReport.java) | Run the golden set through three evaluators; `requirePassRate` is the build gate | | [`RuleBasedJudge.java`](src/main/java/com/ankurm/evaluation/RuleBasedJudge.java) | A deterministic `ChatModel` that the real evaluators run against in CI | | [`VerdictNormalizingModel.java`](src/main/java/com/ankurm/evaluation/VerdictNormalizingModel.java) | Turns "Yes." into "yes", which is all the built-in evaluators accept | | [`GradedEvaluator.java`](src/main/java/com/ankurm/evaluation/GradedEvaluator.java) | A 1-5 judge, because the built-ins only return 0 or 1 | | [`ContainsFactsEvaluator.java`](src/main/java/com/ankurm/evaluation/ContainsFactsEvaluator.java), [`CompositeEvaluator.java`](src/main/java/com/ankurm/evaluation/CompositeEvaluator.java) | A no-model check, and a combiner that restores the feedback the built-ins leave empty | | [`src/test/resources/golden/support-golden.json`](src/test/resources/golden/support-golden.json) | The golden dataset | ## Output files | File | Written by | |---|---| | [`01-evaluator-prompts.txt`](output/01-evaluator-prompts.txt) | `EvaluatorAnatomyTest`: the exact prompts, and the responses | | [`02-verdict-parsing.txt`](output/02-verdict-parsing.txt) | `VerdictParsingTest`: which judge replies pass | | [`03-golden-run.txt`](output/03-golden-run.txt) | `GoldenDatasetTest`: a healthy and a regressed build, and the gate | | [`04-noisy-judge.txt`](output/04-noisy-judge.txt) | `NoisyJudgeTest`: simulated noise against five gates | | [`05-graded-and-composite.txt`](output/05-graded-and-composite.txt) | `GradedCompositeTest` | | [`06-judge-limits.txt`](output/06-judge-limits.txt) | `JudgeLimitsTest`: where the deterministic judge is wrong | | [`07-ci-layers.txt`](output/07-ci-layers.txt) | `CiLayersTest`: model calls per testing layer | | [`08-live-judge-skipped.txt`](output/08-live-judge-skipped.txt) | cut from the surefire report by `run-all.sh` |