Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
50 lines
3.5 KiB
Markdown
50 lines
3.5 KiB
Markdown
# evaluation
|
|
|
|
Companion code for [Testing LLM Apps in Java: Spring AI Evaluators and LLM-as-Judge in JUnit 6](https://ankurm.com/spring-ai-2-0-testing-llm-apps-evaluators-llm-as-judge-junit-6/), part of the [Spring AI series](../README.md) on ankurm.com.
|
|
|
|
A small retrieval-grounded support assistant, a 12-row golden dataset, and the tests that show what Spring AI's `RelevancyEvaluator` and `FactCheckingEvaluator` really send and accept, how to gate a build on a pass rate, and how to keep a model-graded test suite free and repeatable in CI.
|
|
|
|
**No live model is used anywhere.** The "application model" is a scripted stub and the "judge" is either a recording stub or `RuleBasedJudge`, a deterministic word-overlap judge. Judge noise in `NoisyJudgeTest` is a seeded simulation. The one test that talks to a real model, `LiveJudgeTest`, is skipped unless `OPENAI_API_KEY` is set and was never run for the article. So nothing here says how well any real model judges.
|
|
|
|
## Versions
|
|
|
|
| Component | Version |
|
|
|---|---|
|
|
| Spring Boot | 4.1.1 |
|
|
| Spring AI | 2.0.1 (`spring-ai-client-chat`) |
|
|
| JUnit | 6.0.3 (managed by Boot 4.1.1; 6.1.3 is the latest on Maven Central) |
|
|
| Java | 25 (LTS) |
|
|
|
|
## Quickstart
|
|
|
|
```bash
|
|
scripts/run-all.sh # runs 21 tests (1 skipped) and regenerates output/01 .. 08
|
|
```
|
|
|
|
Two consecutive runs produce byte-identical files.
|
|
|
|
## What's here
|
|
|
|
| File | What it shows |
|
|
|---|---|
|
|
| [`SupportAssistant.java`](src/main/java/com/ankurm/evaluation/SupportAssistant.java) | The application under test |
|
|
| [`EvalRunner.java`](src/main/java/com/ankurm/evaluation/EvalRunner.java), [`EvalReport.java`](src/main/java/com/ankurm/evaluation/EvalReport.java) | Run the golden set through three evaluators; `requirePassRate` is the build gate |
|
|
| [`RuleBasedJudge.java`](src/main/java/com/ankurm/evaluation/RuleBasedJudge.java) | A deterministic `ChatModel` that the real evaluators run against in CI |
|
|
| [`VerdictNormalizingModel.java`](src/main/java/com/ankurm/evaluation/VerdictNormalizingModel.java) | Turns "Yes." into "yes", which is all the built-in evaluators accept |
|
|
| [`GradedEvaluator.java`](src/main/java/com/ankurm/evaluation/GradedEvaluator.java) | A 1-5 judge, because the built-ins only return 0 or 1 |
|
|
| [`ContainsFactsEvaluator.java`](src/main/java/com/ankurm/evaluation/ContainsFactsEvaluator.java), [`CompositeEvaluator.java`](src/main/java/com/ankurm/evaluation/CompositeEvaluator.java) | A no-model check, and a combiner that restores the feedback the built-ins leave empty |
|
|
| [`src/test/resources/golden/support-golden.json`](src/test/resources/golden/support-golden.json) | The golden dataset |
|
|
|
|
## Output files
|
|
|
|
| File | Written by |
|
|
|---|---|
|
|
| [`01-evaluator-prompts.txt`](output/01-evaluator-prompts.txt) | `EvaluatorAnatomyTest`: the exact prompts, and the responses |
|
|
| [`02-verdict-parsing.txt`](output/02-verdict-parsing.txt) | `VerdictParsingTest`: which judge replies pass |
|
|
| [`03-golden-run.txt`](output/03-golden-run.txt) | `GoldenDatasetTest`: a healthy and a regressed build, and the gate |
|
|
| [`04-noisy-judge.txt`](output/04-noisy-judge.txt) | `NoisyJudgeTest`: simulated noise against five gates |
|
|
| [`05-graded-and-composite.txt`](output/05-graded-and-composite.txt) | `GradedCompositeTest` |
|
|
| [`06-judge-limits.txt`](output/06-judge-limits.txt) | `JudgeLimitsTest`: where the deterministic judge is wrong |
|
|
| [`07-ci-layers.txt`](output/07-ci-layers.txt) | `CiLayersTest`: model calls per testing layer |
|
|
| [`08-live-judge-skipped.txt`](output/08-live-judge-skipped.txt) | cut from the surefire report by `run-all.sh` |
|