spring-ai

Runnable companion code for the Spring AI articles on ankurm.com. One directory per module; each module is one commit and carries its own README, tests and captured output.

Module What it is Article
getting-started/ One ChatClient bean, three endpoints (plain call, templated system prompt, streaming), and a test proving spring.ai.model.chat switches providers with zero code change. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Spring AI 2.0 in 10 Minutes: ChatClient on Spring Boot 4.1
rag/ Ingest PDFs, chunk, retrieve from pgvector, rerank, answer, check the answer. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Production-grade RAG with Spring AI and the complete example
mcp-server/ An order-lookup service exposed as MCP tools, a resource, and a prompt with @McpTool/@McpResource/@McpPrompt, served over Streamable HTTP (Spring AI 2.0's default MCP server transport). Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Build an MCP Server with Spring AI 2.0
mcp-client/ ChatClient calling tools from two real external MCP servers (filesystem, git) over stdio via defaultToolCallbacks(ToolCallbackProvider...), contrasted with a local @Tool method, with every call logged through one Micrometer ObservationHandler. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Spring AI MCP Client: Calling External MCP Servers from ChatClient
mcp-secure/ The mcp-server article's order-lookup tools behind a real OAuth2 resource server: JWT validation, one scope per tool via @PreAuthorize, unauthenticated tool discovery rejected outright, and every call audit-logged through MDC -- denials included. Spring Boot 4.1.1, Spring AI 2.0.1, Spring Security 7.1.1, Java 25. Securing an MCP Server with Spring Security 7
tool-calling/ @Tool methods, ToolCallingAdvisor (the advisor-layer replacement for Spring AI 1.x's per-model tool loop), returnDirect, ToolContext, and ToolSearchToolCallingAdvisor for progressive disclosure across a 230-tool synthetic library -- every test driven by a hand-written ScriptedChatModel, no live model anywhere. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Tool Calling in Spring AI 2.0
structured-output/ ChatClient.entity() mapping LLM responses to Java records, lists and maps; StructuredOutputValidationAdvisor retrying non-conforming JSON with a real enum-constrained schema, including a captured run that exhausts every retry without throwing. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Structured Output in Spring AI 2.0
ollama-local/ Chat and embeddings against a real local qwen2.5:0.5b/all-minilm, no API key, driven by a Testcontainers-managed Ollama container started from a baked image; a confirmed model unload via keep_alive: 0 and /api/ps, not a scripted model anywhere. Spring Boot 4.1.1, Spring AI 2.0.1, Testcontainers 2.0.5, Java 25. Run LLMs Locally with Spring AI and Ollama
chat-memory/ MessageChatMemoryAdvisor, MessageWindowChatMemory, the JDBC and Redis ChatMemoryRepository, per-user conversation IDs and a token-budget memory of our own, with the traps reproduced against a real PostgreSQL 16 and Redis Stack: a 36-character conversation_id, tool messages dropped on save, concurrent writers, a 1.x table under the 2.0 repository, and a Redis repository that silently steps aside for a custom ChatMemory. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Chat Memory in Spring AI 2.0: JDBC, Redis and Windowed Conversations
advisors/ Three custom advisors -- a logger, a PII redactor (with a stream-safe restore) and a per-request / per-user token budget -- and tests for how the chain is ordered, what BaseAdvisor does on a stream, where an advisor sits relative to memory and the tool loop, and what a refusal looks like on a call, a stream and over HTTP (429). A recording stub model, no live model. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Writing Custom Advisors in Spring AI 2.0: Logging, PII Redaction and Token Budgets
vector-stores/ The same 30,000-document dataset behind VectorStore on pgvector, Redis, Qdrant and Elasticsearch: ingest time, recall@10, latency, metadata filtering and running cost, with the defaults that cost recall reproduced (Elasticsearch's quantised mapping, Redis EF_RUNTIME, pgvector post-filtering, Qdrant payload indexes). Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Choosing a Vector Store for Spring AI
evaluation/ Testing an LLM app: RelevancyEvaluator and FactCheckingEvaluator (exactly what they send and which judge replies they accept), a 12-case golden dataset with a pass-rate gate, a deterministic judge for CI, simulated judge noise, a 1-5 graded evaluator and a composite. Stub models only; the one live-judge test is skipped without a key. Spring Boot 4.1.1, Spring AI 2.0.1, JUnit 6, Java 25. Testing LLM Apps in Java: Spring AI Evaluators and LLM-as-Judge in JUnit 6
observability/ What Spring AI 2.0.1 records on its own (model, chat client, advisor and tool meters, spans for a tool call), token usage turned into cost per endpoint, a Grafana dashboard checked against live Prometheus and Grafana, and the traps: histograms are opt-in, a response with no usage looks like a free call, prompt text is logged only if switched on. The real OpenAiChatModel against a local fake server, so counts are approximate and prices illustrative. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Observability for Spring AI: Tokens, Latency and Cost with Micrometer and OpenTelemetry
guardrails/ Prompt injection against a Spring AI assistant with tools: a poisoned document, a poisoned tool result, a markdown-image leak and a system prompt leak, run against a document filter, a tool allow-list with argument policies, and output validation, alone and together (6 of 6 attacks succeed with no defence, 0 of 6 with all three). A deliberately gullible stub model, so it measures what each defence stops when the model is fooled, not how often a real model is. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Prompt Injection Defense in Spring AI
multimodal/ A receipt image through Media and ChatClient.entity(...) into a Java record, on the real OpenAiChatModel, AnthropicChatModel and OllamaChatModel against a local server that OCRs the image it receives (so accuracy figures describe OCR, not any vision model). The same image on three wire formats, arithmetic validation and a repair retry, accuracy under tilt, shrinking and noise, and an image-token estimate from a documented formula. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Multimodal Spring AI: Extract Structured Data from Images
text-to-sql/ A question to SQL to rows, safely, on a real PostgreSQL 16: a schema prompt built through the restricted role, a JSqlParser guard (one statement, SELECT only, listed tables and functions), a read-only role with column grants and a statement timeout, a row cap, and evaluation by comparing results. Sixteen queries against four setups (13 harmful: 13 succeed with neither protection, 0 with both). The model is a script, not a language model. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Text-to-SQL with Spring AI, Done Safely
providers/ One ticket-summarising app on the real OpenAiChatModel, AnthropicChatModel and GoogleGenAiChatModel against a local server in all three wire formats: where each puts the system prompt and what it sends by default, portable versus provider-specific options (and the per-call ChatOptions that crashes two providers and silently resets the third), prompt caching, a cost table from the vendors' price sheets, and failover with the retry multiplication measured. No vendor API called; token counts and cache hits are simulated from documented rules; latency not measured. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. Anthropic Claude vs OpenAI vs Gemini in Spring AI 2.0
llm-gateway/ A gateway service on Spring AI 2.0: model-hint routing, failover under the tool-calling advisor (a provider failing mid tool loop does not re-run the tool), one circuit breaker per provider, dollar caps that reserve before the call, a token-per-hour limiter, and a tenant-keyed semantic cache. Real OpenAI and Anthropic models against a local fake; no vendor API called. Spring Boot 4.1.1, Spring AI 2.0.1, Java 25. The LLM Gateway Pattern for Java Microservices

Upgrading from Spring AI 1.x: migration guide.

S
Description
Runnable companion code for the Spring AI articles on ankurm.com
Readme
734 KiB
Languages
Java 95.8%
Shell 3.6%
Python 0.6%