Semantic Caching for LLM Calls in Spring Boot with Redis Vector Search
Build a semantic cache for LLM calls in Spring Boot with Spring AI and Redis vector search, then measure threshold, hit rate, wrong answers and savings on 600 replayed requests.
Build a semantic cache for LLM calls in Spring Boot with Spring AI and Redis vector search, then measure threshold, hit rate, wrong answers and savings on 600 replayed requests.
Run a DistilBERT sentiment model from Java with ONNX Runtime and DJL: tokenizing, identical answers by two routes, batching, a Spring Boot endpoint, and measured latency.
Jlama runs a 4-bit TinyLlama inside the JVM on the JDK 25 Vector API. Measured tokens per second, streaming, the token-limit surprise, temperature, and a same-machine comparison with Ollama.
Two Java agents talk over the A2A protocol: the agent card, a raw SendMessage call, blocking versus streaming, input-required follow-ups and an agent calling another agent, with every message shown and a table comparing A2A with MCP.
Embabel plans the order of an agent's steps from the types they need and produce. A small research agent shows the plan, a skipped step, a failing condition and a cost choice, against plain Spring AI.
Build a travel planner from small LangChain4j agents: sequence, parallel, loop, supervisor and error recovery, with the exact transcripts each step produced and the traps that cost time.
Write a plain Java interface and LangChain4j implements it. Adds prompt templates, per-user memory, tools, retrieval and @AiService on Spring Boot 4.1.1, with the exact messages sent at each step and two traps that cost an afternoon.
You built a feature on one AI provider. Then someone asks the questions every team asks sooner or later: could we try a different one, what would it cost, and what happens when the one we use has a bad afternoon? In Spring AI the answer to the first question is “mostly yes, it is a configuration change”. The word mostly is where the surprises live. This article takes one small application and runs it on OpenAI, Anthropic Claude and Google Gemini through the three real Spring AI model classes. It looks at what each one sends over the wire, which options carry over and which silently do not, how prompt caching differs, what a batch of requests costs on each price sheet, and how to fail over from one provider to the next without multiplying your retries by accident. The depth is in expandable sections, so you can read straight through or open only what you need. Versions, and an honest limit. Spring Boot 4.1.1, Spring AI 2.0.1 and Java 25. The code is the providers module of asmhatre/spring-ai, and every console block below is quoted from a file under its output/ directory. I did not call any vendor’s API; there are no keys in my build environment. The real OpenAiChatModel, AnthropicChatModel and GoogleGenAiChatModel talk to a local server that answers in the three wire formats, so what each client sends and how it reacts to a reply or an error is real. Three things are simulated, and the article says so wherever they matter: token counts are characters divided by four, the answer is a fixed string, and cache hits follow the rules the vendors document (thresholds, prefix matching), applied to the request that actually arrived. So the cost table is a worked example of the published price sheets, not a bill, and I did not measure latency or answer quality at all. Prices and cache rules are as I read them on 9 October 2026; they change often.
“Which customers haven't ordered this year?” is a question a manager can ask in plain words and a database can answer in one query, if somebody writes the query. A language model can write it. The catch is that you are now running SQL written by something that sometimes gets it wrong, can be talked into things, and has no idea which table holds the secrets. This article builds a text-to-SQL service in Spring AI and spends most of its length on the part around the model: what it is shown, what is checked before the SQL runs, what the database itself refuses, how many rows come back, and how you tell whether the answers are right. Everything runs against a real PostgreSQL, so every refusal quoted is PostgreSQL's own. The depth is in expandable sections, so you can read straight through or open only what you need. Versions, and an honest limit. Spring Boot 4.1.1, Spring AI 2.0.1, JSqlParser 5.4, PostgreSQL 16 and Java 25. The code is the text-to-sql module of asmhatre/spring-ai, and every console block below is quoted from a file under its output/ directory. I did not use a language model, and nothing here measures how well one writes SQL. The “model” is a script that returns a prepared SQL string for each question, some correct, some wrong and some hostile on purpose, so the code around it can be tested. The database, though, is real: results, error messages, timeouts and permission failures are PostgreSQL 16's. The schema, the data, the attack queries and the eight evaluation questions are examples I wrote.
You have a shoebox of receipts, or a phone full of photos of them, and what you want is rows: merchant, date, line items, total. Typing them in is the job nobody wants. A model that can look at an image can do it, and Spring AI lets you hand it the image almost as easily as you hand it a string. The easy part is sending the picture. The part that decides whether you can trust the result is everything after: getting a typed Java record back instead of a paragraph, noticing when the numbers do not add up, and knowing what happens when the photo is tilted, small or grainy. This article builds that, step by step, and measures it. The depth is in expandable sections, so you can read straight through or open only what you need. Versions, and an honest limit. Spring Boot 4.1.1, Spring AI 2.0.1, Jackson 3 and Java 25. The code is the multimodal module of asmhatre/spring-ai, and every console block below is quoted from a file under its output/ directory. I did not run a vision model, and nothing here measures one. There are no API keys in my build environment, so the real OpenAiChatModel, AnthropicChatModel and OllamaChatModel classes talk to a local server that stands in for all three services. That server is not a language model: it decodes the image it receives, runs the tesseract OCR program on the pixels and parses the text with a few regular expressions. So the requests are the real ones and the answers depend on the image, but every accuracy figure below describes OCR plus a parser, and says nothing about GPT, Claude or Gemini. Token figures are arithmetic from a published formula, not a bill. The receipts are ten I wrote by hand.