Files

50 lines
4.0 KiB
Markdown

# guardrails
Companion code for [Prompt Injection Defense in Spring AI: Guardrails, Tool Allow-Lists and Output Validation](https://ankurm.com/prompt-injection-defense-spring-ai-guardrails-tool-allow-lists-output-validation/), part of the [Spring AI series](../README.md) on ankurm.com.
A support assistant that answers from retrieved documents and can call three tools (look up an order, refund, send an email), three defences around it, and six attacks to run against them.
**No real model is used, and no claim is made about real models.** The "model" is `GullibleModel`, a deterministic stub that obeys every `ACTION name {json}` directive it finds anywhere in its input, whatever words surround it. It stands for a model that has been successfully injected. A real model follows hostile text only sometimes, and how often depends on the model and the wording; nothing here measures that. What the tests measure is narrower and still useful: *given that the model was fooled, which defence still stops the damage?* The attack payloads, the phrase list and the allowed hosts are all examples written for this module.
## Versions
| Component | Version |
|---|---|
| Spring Boot | 4.1.1 (parent, for dependency management and Bean Validation) |
| Spring AI | 2.0.1 (`spring-ai-client-chat`) |
| Jackson | 3 (`tools.jackson`) |
| Java | 25 (LTS) |
## Quickstart
```bash
scripts/run-all.sh # runs the suite and regenerates output/01 .. 08
```
Two consecutive runs produce byte-identical files.
## What's here
| File | What it is |
|---|---|
| [`SupportAssistant.java`](src/main/java/com/ankurm/guardrails/SupportAssistant.java) | The assistant, with the three defences switched on or off by a `Defences` record |
| [`InjectionHeuristics.java`](src/main/java/com/ankurm/guardrails/InjectionHeuristics.java), [`DocumentGuard.java`](src/main/java/com/ankurm/guardrails/DocumentGuard.java) | A phrase list and the filter that drops matching chunks |
| [`ToolPolicy.java`](src/main/java/com/ankurm/guardrails/ToolPolicy.java), [`GuardedTools.java`](src/main/java/com/ankurm/guardrails/GuardedTools.java) | Tool allow-list, per-tool argument checks and a call cap, wrapped around every `ToolCallback` |
| [`OutputRules.java`](src/main/java/com/ankurm/guardrails/OutputRules.java), [`OutputGuardAdvisor.java`](src/main/java/com/ankurm/guardrails/OutputGuardAdvisor.java) | Link, canary and key checks on the answer, as an advisor |
| [`Triage.java`](src/main/java/com/ankurm/guardrails/Triage.java), [`TriageService.java`](src/main/java/com/ankurm/guardrails/TriageService.java) | A typed answer checked by Jackson, Bean Validation and the output rules |
| [`Tools.java`](src/main/java/com/ankurm/guardrails/Tools.java) | The three tools; side effects go to lists a test can inspect |
| [`GullibleModel.java`](src/test/java/com/ankurm/guardrails/support/GullibleModel.java), [`Scenarios.java`](src/test/java/com/ankurm/guardrails/support/Scenarios.java) | The always-obeying stub and the six attacks |
## Output files
| File | Written by |
|---|---|
| [`01-attack-matrix.txt`](output/01-attack-matrix.txt) | `AttackMatrixTest`: six attacks against five configurations |
| [`02-heuristic-filter.txt`](output/02-heuristic-filter.txt) | `HeuristicsTest`: what a phrase list catches, misses and wrongly drops |
| [`03-tool-policy.txt`](output/03-tool-policy.txt) | `ToolPolicyTest`: one call per rule |
| [`04-unexposed-tool.txt`](output/04-unexposed-tool.txt) | `ToolPolicyTest`: what Spring AI does when the model asks for a tool it was not given |
| [`05-output-rules.txt`](output/05-output-rules.txt) | `OutputRulesTest` |
| [`06-structured-validation.txt`](output/06-structured-validation.txt) | `TriageTest`: schema and text checks on a typed answer |
| [`07-benign-traffic.txt`](output/07-benign-traffic.txt) | `OutcomeTest`: what a legitimate customer sees with every defence on |
| [`08-tool-result-encoding.txt`](output/08-tool-result-encoding.txt) | `ToolResultEncodingTest`: a tool's String result reaches the model JSON-encoded |