Versions, and an honest limit. LangChain4j 1.22.0 (the core library is generally available), the Spring Boot starter and embeddings module at 1.22.0-beta32 (these are beta only; no stable release exists), Spring Boot 4.1.1, JUnit 6.1.3 and Java 25. All the code is in theai-servicesandspring-boot-ai-servicemodules of asmhatre/langchain4j-demo, and every console block below is quoted from a file under a module’soutput/directory, written by a test that asserts the same lines. No live model was used. The model is a small scripted stand-in, so every run is identical and the tests can check exactly what LangChain4j sent. That tells you what the library does. It tells you nothing about how any real model answers.
An AI Service is an interface you write and LangChain4j implements
A language model is, from your program’s point of view, a function. You give it a list of messages, and it gives you one message back. The first message in the list can be a system message (instructions about how to behave), followed by user messages (what the person said) and AI messages (what the model said earlier). An AI Service hides that function behind an ordinary Java method. You declare the method you wish you had, such asString answer(String product, String question), and LangChain4j generates a class that implements it. Calling the method builds the messages, calls the model, and returns the text.
ChatModel box. Everything to the left of it is your own code and stays the same, which is why the examples here run against a scripted model and would run unchanged against a real one.
Here is the whole interface. It is quoted from SupportAgent.java, and there is no class that implements it. {{product}} and {{question}} are placeholders that LangChain4j fills from the method arguments.
@SystemMessage("You are a concise support agent for {{product}}. Answer in one sentence.")
@UserMessage("Customer asks: {{question}}")
String answer(@V("product") String product, @V("question") String question);
And this is how you get an object that implements it, from AiServicesTest.java. AiServices.create takes the interface and a model and returns the implementation.
ScriptedChatModel model = new ScriptedChatModel(r -> AiMessage.from("Reset it from Settings > Security."));
SupportAgent agent = AiServices.create(SupportAgent.class, model);
String reply = agent.answer("Acme Vault", "How do I reset my password?");
Running it and printing what the model received gives the transcript below, quoted from 01-prompt-template.txt. The annotations became two messages, with the placeholders replaced.
Messages LangChain4j sent to the model:
SYSTEM: You are a concise support agent for Acme Vault. Answer in one sentence.
USER: Customer asks: How do I reset my password?
Reply: Reset it from Settings > Security.
Implementation is a java.lang.reflect.Proxy: true
Going deeper: what AiServices.create actually returns
AiServices.create(Class, ChatModel) is a shortcut for the builder. The full form, AiServices.builder(SupportAgent.class)...build(), is where every other capability in this article is attached: chatMemory, chatMemoryProvider, tools, contentRetriever and retrievalAugmentor are all methods on that builder (read from the 1.22.0 jar with javap). The returned object is a java.lang.reflect.Proxy, so it is not an instance of any class you wrote. The last line of the transcript above shows Proxy.isProxyClass returning true, and the test asserts it.
If the interface is badly configured, the failure happens at build() time, not on the first call. The @MemoryId trap later in this article is an example.
- LangChain4j reference: AI Services
- ScriptedChatModel — the scripted stand-in used throughout, a single small class.
Prompts are annotations on the method, not strings in your code
Because the prompt lives next to the method that uses it, you can read the contract of a model call the way you read any other method: name, parameters, return type, and the instructions.@SystemMessage sets the standing instructions. @UserMessage is a template for what the user says. @V("name") binds a method parameter to a {{name}} placeholder.
Every method in this article returns String, which gives you the model’s text. I did not test other return types.
Going deeper: when you do not need @V
With a single parameter and no
@V, the template refers to it as {{it}}. The Spring Boot example later in this article uses exactly that: @UserMessage("Greet {{it}}") on a method taking one String, and the transcript in 01-boot-context.txt shows the result, USER: Greet Ankur.
The demo’s parent POM compiles with -parameters (<parameters>true</parameters> in the compiler plugin configuration). Every example here uses @V or {{it}}, so nothing in this article depends on it, and I did not test what happens without it.
Memory means the history is sent again on every call
A model has no memory of its own. If you ask “What is my name?” as a separate call, it cannot know. The only way a conversation works is that you send the whole history again each time. AChatMemory is the object that remembers that history for you and adds it to every request.
You usually want a separate history for each user, so LangChain4j lets you mark a parameter with @MemoryId and give the builder a chatMemoryProvider, which creates a fresh memory the first time it sees a new id.
The interface, quoted from ChatAgent.java:
String chat(@MemoryId String customerId, @UserMessage String message);
And the builder, from AiServicesTest.java. MessageWindowChatMemory.withMaxMessages(4) keeps only the newest four messages.
ChatAgent agent = AiServices.builder(ChatAgent.class)
.chatModel(model)
.chatMemoryProvider(id -> MessageWindowChatMemory.withMaxMessages(4))
.build();
The test makes five calls: three from “alice”, one from “bob” in between, and then a fourth from alice. The transcript, quoted from 02-memory.txt, shows how many messages the model received each time.
Messages the model saw on each call (memory window = 4 stored messages):
call 1: 1 messages
call 2: 3 messages
call 3: 1 messages
call 4: 5 messages
call 5: 5 messages
Second call, in full:
USER: My name is Alice
AI: noted (1 messages seen)
USER: What is my name?
Read it from the top. Call 2 sees three messages, because alice’s first question and the first reply were replayed before her new question. Call 3 is bob, and he sees one message, because his history is separate from alice’s. That is what @MemoryId buys you.
To see it plainly, a second test records, at the moment of each call, how many messages the model received and how many the memory held. It is quoted from 06-window-probe.txt.
Memory window = 4. For each call: messages the model received, messages stored at that moment, messages stored afterwards.
call 1: model received 1, stored 1, stored afterwards 2
call 2: model received 3, stored 3, stored afterwards 4
call 3: model received 5, stored 4, stored afterwards 4
call 4: model received 5, stored 4, stored afterwards 4
call 5: model received 5, stored 4, stored afterwards 4
Largest number of messages the memory object held after any add, used on its own: 4
A window of 4 shows the model 5. Calls 4 and 5 are alice’s, and the model received five messages even though the window is four. The window limits what is stored. The request carries the stored messages plus the new user message that has just arrived. Used on its own, the memory object never holds more than four after an add; inside an AI Service call the model received five while the memory held four. The transcript is 06-window-probe.txt. So size the window with one message to spare when you are budgeting tokens. This is what 1.22.0 does, and it is the kind of detail that could change, which is why the tests assert it.
Going deeper: windows, tokens and where history can live
A message window counts messages, not words. One long reply and one short one use the same slot, so a window of ten messages can still be a large request. LangChain4j also ships
TokenWindowChatMemory next to it, for when tokens are what you are budgeting; I confirmed the class exists in the 1.22.0 jar but did not run it.
Memory in these examples lives in the JVM, so it disappears on restart and is not shared between application instances. For anything real you want the history in a store, and that is a separate decision from the AI Service code.
- AiServicesTest — the memory test asserts the 3, 1 and 5 message counts above, and the window probe asserts all five rows of the second table.
- Chat memory in Spring AI 2.0 — the same problem, solved with a database or Redis behind it.
- LangChain4j reference: Chat Memory
One @MemoryId anywhere breaks the whole interface
Here is a trap that costs people an afternoon. Suppose you put two methods on one interface: a plain one with no memory, and one with@MemoryId. You would expect only the second to need a memory provider. That is not what happens. The test below builds exactly that interface with no provider, and the whole build fails. The interface is MemoryIdMixed in AiServicesTest.java, and the failure is quoted from 05-memoryid-trap.txt.
interface MemoryIdMixed {
String plain(String message);
String remembered(@dev.langchain4j.service.MemoryId String id, @dev.langchain4j.service.UserMessage String message);
}
Building an interface where ONE method has @MemoryId, with no ChatMemoryProvider:
IllegalConfigurationException: In order to use @MemoryId, please configure the ChatMemoryProvider on the 'com.ankurm.lc4j.aiservices.AiServicesTest$MemoryIdMixed'.
The error message mentions @MemoryId and a missing ChatMemoryProvider, and it names the whole interface, not the method. The fix is to keep memory-aware methods on their own interface. That is why this article has SupportAgent (templates, no memory) and ChatAgent (memory) as two separate interfaces rather than one.
Tools: the model asks, and your Java method runs
A model cannot look up an order, and it cannot call your database. What it can do is reply with a structured request that says, in effect, “please runorderStatus with this argument and tell me what came back.” Your program runs the method and sends the result back, and the model then writes its final answer using that result.
The important point is that the model never executes anything. It only asks. LangChain4j receives the request, finds the matching Java method, runs it, and makes a second call to the model.
@Tool carries the description the model reads to decide when to call it, and @P describes the parameter.
@Tool("Looks up the shipping status of an order by its id")
public String orderStatus(@P("the order id, for example A-1001") String orderId) {
calls.add("orderStatus(" + orderId + ")");
return "Order " + orderId + " shipped on 2026-10-08, arriving 2026-10-13";
}
Registering it is one builder call. The memory provider line is included because the example uses the memory-aware interface; the two handler lines are explained in the callout below. From AiServicesTest.java:
.chatMemoryProvider(id -> MessageWindowChatMemory.withMaxMessages(10))
.tools(tools)
// 1.22.0 logs a notice that the defaults will change; set both explicitly
.toolArgumentsErrorHandler(dev.langchain4j.service.tool.ToolArgumentsErrorHandler.sendExceptionMessageToLlm())
.toolExecutionErrorHandler(dev.langchain4j.service.tool.ToolExecutionErrorHandler.failInvocationUnlessVisibleToLlm())
.build();
The transcript, from 03-tools.txt, shows the tool description that was sent with the first request, the second request after the tool ran, the Java methods that actually executed, and the number of model round trips.
Tool specification sent with the first request: [orderStatus - Looks up the shipping status of an order by its id]
Second request (after the tool ran):
USER: Where is order A-1001?
AI: tool calls [orderStatus{"orderId":"A-1001"}]
TOOL_EXECUTION_RESULT: Order A-1001 shipped on 2026-10-08, arriving 2026-10-13
Java methods actually executed: [orderStatus(A-1001)]
Model round trips: 2
Final reply: Your order is on its way and arrives on 2026-10-13.
Set the tool error handlers yourself. On 1.22.0, using tools without configuring error handlers makes the library log a long notice once per JVM. It says that the defaults are planned to change in an upcoming release, and that today an exception thrown by a tool is sent to the model as the exception’s message. Exception messages are written for developers, so they can leak file paths, SQL or other internals to the model provider and into the chat memory. The notice recommends two lines, which the example above uses:toolArgumentsErrorHandler(ToolArgumentsErrorHandler.sendExceptionMessageToLlm())andtoolExecutionErrorHandler(ToolExecutionErrorHandler.failInvocationUnlessVisibleToLlm()). When I re-ran the tests with both set, the notice did not appear. I set them but did not exercise the failure path, so what this article claims is that observation and what the notice itself says.
Going deeper: describing tools well, and what the model sees
The model never sees your Java. It sees a specification: a name, a description and a parameter schema. The first line of the transcript shows what was sent for
orderStatus. The quality of the @Tool and @P descriptions is therefore the quality of your interface to the model, and a vague description leads to a tool being called at the wrong time or not at all. That is a property of real models; the scripted model here always does what it is told, so this article cannot show it.
A tool can signal a failure meant for the model by throwing an exception that implements ToolErrorVisibleToLlm; the library’s own notice shows ToolErrorVisibleToLlm.from("There is no order with this ID.") as the example. Everything else fails the invocation under the recommended setting.
- Tool calling in Spring AI 2.0 — the same idea in the other big Java framework.
- LangChain4j reference: Tools and error handling
Retrieval rewrites the user message before it is sent
A model only knows what it was trained on plus what you put in the request. If your answer lives in your own documents, such as a refund policy, you have to put the relevant part into the request yourself. Retrieval-augmented generation, usually shortened to RAG, automates that: search your documents for the passages closest in meaning to the question, and add them to the message.String ask(String question). Retrieval is attached on the builder, from AiServicesTest.java:
KnowledgeAgent agent = AiServices.builder(KnowledgeAgent.class)
.chatModel(model)
.contentRetriever(EmbeddingStoreContentRetriever.builder()
.embeddingStore(store)
.embeddingModel(embeddingModel)
.maxResults(1)
.build())
.build();
The store holds three short facts, embedded in-process with the all-MiniLM-L6-v2 model, so the demo needs no external service. The transcript, from 04-rag.txt, shows what the model received when the question was about refunds.
What the model received after retrieval (maxResults=1):
USER: How long do refunds take?
Answer using the following information:
Refunds are issued to the original payment method within 5 business days.
Only the refund passage was added, because maxResults(1) asked for one result, and the test asserts that the Windows and macOS fact did not appear. LangChain4j appended the retrieved text after the question under a fixed lead-in sentence, which is what 1.22.0 does by default.
Going deeper: what this tiny example does and does not prove
Three sentences are enough to show the mechanism but nothing about retrieval quality. Real documents need to be split into passages, and how you split them, how many results you ask for and whether you filter by score all change the answers. None of that was tested here.
Pulling in the embedding model also pulls in ONNX Runtime 1.20.0 and the Deep Java Library (DJL) 0.36.0 API and tokenizer as transitive dependencies, as the Maven dependency tree in 07-embedding-dependencies.txt shows. That is worth knowing before you ship it, because it adds native libraries to your application.
- Choosing a vector store for Spring AI — where the passages would live in a real system.
- LangChain4j reference: RAG
With Spring Boot, @AiService replaces the builder code
Everything so far usedAiServices.builder(...), which you would call from a configuration class and expose as a bean. The Spring Boot starter lets you skip that. You put @AiService on the interface, and the starter finds it, builds the implementation and registers it as a bean. It also looks in the application context for a ChatModel (and a ChatMemory, if you want memory) and wires them in.
This is the whole setup, from BootApp.java. The two @Bean methods stand in for the model and memory you would normally get from a provider starter.
@Bean
ScriptedChatModel chatModel() {
return new ScriptedChatModel(r -> AiMessage.from("Reply built from " + r.messages().size() + " message(s)."));
}
@Bean
dev.langchain4j.memory.ChatMemory chatMemory() {
return MessageWindowChatMemory.withMaxMessages(10);
}
/** The whole declaration. No @Bean method, no builder: the scanner creates and registers it. */
@AiService
public interface Greeter {
@SystemMessage("You greet people briefly.")
@UserMessage("Greet {{it}}")
String greet(String name);
}
A real @SpringBootTest starts the application and calls the interface. Its output is quoted from 01-boot-context.txt.
Spring Boot version: 4.1.1
Bean of the interface type exists: [bootApp.Greeter]
Messages sent to the model:
SYSTEM: You greet people briefly.
USER: Greet Ankur
Reply: Reply built from 2 message(s).
A starter built for Spring Boot 3.5 ran on Spring Boot 4.1.1. The 1.22.0-beta32 starter declares Spring Boot 3.5.13 in its own POM, not Boot 4. I did not want to assume that mattered either way, so the test above runs it on Boot 4.1.1 and it works for this scenario: the context starts,@AiServiceis scanned, and the interface becomes a bean. That is the only thing tested. I did not test the provider starters (the ones that turnlangchain4j.open-ai...properties into a model bean), streaming, or the explicit wiring mode on Boot 4, so check those yourself before relying on them. The starter is also still a beta release, which is worth knowing before you build on it.
Going deeper: the attributes of @AiService
Reading the annotation from the 1.22.0-beta32 jar with
javap, it has these attributes: wiringMode (either AUTOMATIC or EXPLICIT), and string attributes naming the beans to use: chatModel, streamingChatModel, chatMemory, chatMemoryProvider, contentRetriever, retrievalAugmentor, moderationModel, toolProvider, toolExecutionErrorHandler and toolArgumentsErrorHandler, plus a tools array.
My reading of this, which I did not run, is that AUTOMATIC picks up beans by type and EXPLICIT makes you name each one, which is how you would use two different models in one application. I only exercised the default, so treat that as the documented intent rather than something this article verified.
- LangChain4j reference: Spring Boot integration
- Spring AI 2.0 ChatClient on Spring Boot 4.1 — the equivalent starting point in Spring AI.
Should you even do this?
A fair answer. AI Services are a good fit when your calls follow a shape you can name: “answer a support question”, “summarise this ticket”, “look up an order”. The interface documents the contract, and swapping providers does not touch it. They are a worse fit when the prompt has to be assembled dynamically from many pieces, because annotations are static and you end up fighting them. Weigh the maturity as well: the core library is stable at 1.22.0, but the Spring starter and the embeddings module are beta releases, and the starter was built against an older Spring Boot than the one used here. What this article did not test: any real model, how tool descriptions affect real tool choices, the tool error path, retrieval quality on real documents, streaming, token usage and cost, structured output beyond String, and the provider starters on Boot 4. The scripted model proves what the library sends, not what any model will do with it.
Further reading
- The code for this article: asmhatre/langchain4j-demo, with the transcripts under
ai-services/output/andspring-boot-ai-service/output/, and the README. - Spring AI 2.0 ChatClient on Spring Boot 4.1 and tool calling in Spring AI 2.0 — the other major Java framework for the same job.
- Chat memory in Spring AI 2.0 and choosing a vector store.
- Testing LLM apps in Java — once the calls work, how to check the answers stay good.
- LangChain4j documentation and the AI Services tutorial
- Spring Boot reference documentation
No Comments yet!