Versions, and an honest limit. LangChain4jlangchain4j-agentic1.22.0-beta32. It has no stable release yet, so every API in this article may change, and the stable core library is at 1.22.0. Java 25 and JUnit 6.1.3. All the code is in theagenticmodule of asmhatre/langchain4j-demo, and every console block below is quoted from a file underagentic/output/, written by a test that asserts the same facts. No live model was used. The model is a small scripted stand-in, so every run is identical and the tests can check exactly what LangChain4j did. That tells you what the library does with the replies it gets. It tells you nothing about how well any real model plans, scores or chooses.
An agent is an AI Service with a name and an output
If you have read the article on AI Services, an agent will look familiar: a Java interface with one method and a prompt, which LangChain4j implements. Two things are new. The method carries@Agent, which gives it a name and a description, and an outputKey, which is the name under which its result is stored.
That stored result is the important part. All the agents in a workflow share one scratchpad, called the AgenticScope: a map from names to values. An agent reads its inputs from the scratchpad and writes its output back to it, so one agent’s result becomes another’s input without you passing anything by hand.
Here is the first agent, quoted from FlightAgent.java. It takes a destination and a traveler, and stores its answer under the key flights.
@UserMessage("List one flight option to {{destination}} for {{traveler}}.")
@Agent(value = "Finds a flight", outputKey = "flights")
String findFlight(@V("destination") String destination, @V("traveler") String traveler);
The hotel and activity agents (HotelAgent, ActivityAgent) have the same shape and write hotels and activities. The fourth agent reads those three keys and writes the plan:
@UserMessage("Write a one-line itinerary for {{destination}}. Flight: {{flights}}. Hotel: {{hotels}}. Activity: {{activities}}.")
@Agent(value = "Writes the itinerary", outputKey = "plan")
String write(@V("destination") String destination,
@V("flights") String flights,
@V("hotels") String hotels,
@V("activities") String activities);
A sequence runs agents in order and passes state by name
A sequence is the simplest workflow: run these agents, one after another. You describe the entry point as another interface, here TripPlannerWithScope, whose method returns the answer together with the scope so that the test can look inside. Then you ask the library to build it, from AgenticTest.java: TripPlannerWithScope trip = AgenticServices.sequenceBuilder(TripPlannerWithScope.class)
.subAgents(flights(model), hotels(model), activities(model), planner(model))
.outputKey("plan")
.build();
The transcript, from 01-sequential.txt, lists the agents in the order they ran, what each read, what each wrote, and the final contents of the scope.
Order the agents ran in:
findFlight read [traveler, destination] and wrote AI-101 Mumbai to Lisbon
findHotel read [traveler, destination] and wrote Hotel Alfama
findActivity read [traveler, destination] and wrote Tram 28
write read [hotels, activities, destination, flights] and wrote Fly AI-101, stay at Hotel Alfama, ride Tram 28.
Shared state at the end:
traveler = Ankur
hotels = Hotel Alfama
activities = Tram 28
destination = Lisbon
flights = AI-101 Mumbai to Lisbon
plan = Fly AI-101, stay at Hotel Alfama, ride Tram 28.
The prompt the planner was sent (last model call):
Write a one-line itinerary for Lisbon. Flight: AI-101 Mumbai to Lisbon. Hotel: Hotel Alfama. Activity: Tram 28.
Model calls: 4
Three things are worth reading off this. The order is the order you passed to subAgents. The first three agents read only traveler and destination, which the caller supplied, while the planner read the three keys the others had written. And the last model call, the planner’s prompt, contains all three results, even though you never wrote code that moved them.
Going deeper: how an agent finds its inputs
The library matches by parameter name, not by what the template mentions. The test in AgenticTest.java builds the same sequence twice with a planner whose template mentions
{{flights}}, {{hotels}} and {{activities}}. In version A the planner declares only @V("destination"). In version B it declares all four parameters. In each version the test leaves one finder out. The transcript is 02-state-by-parameter.txt.
A) Planner declares only @V destination, template mentions {{flights}} {{hotels}} {{activities}}.
omit nothing: Value for the variable 'hotels' is missing
omit flights: Value for the variable 'hotels' is missing
omit hotels: Value for the variable 'hotels' is missing
omit activities: Value for the variable 'hotels' is missing
B) Planner declares @V destination, @V flights, @V hotels, @V activities.
omit nothing: OK: Fly AI-101, stay at Hotel Alfama, ride Tram 28.
omit flights: Missing argument: flights
omit hotels: Missing argument: hotels
omit activities: Missing argument: activities
Version A fails in every case, including the one where nothing was left out, because the template needs values that no parameter supplies. Version B works when everything is present and, when a finder is missing, fails with Missing argument: <key>, which names the exact key. So declare every value an agent needs as a @V parameter, and make the name equal the producer’s outputKey. Note that the template-only error always named hotels, whichever finder was left out, so that message is not a reliable pointer to the cause.
Parallel agents run at the same time
The flight, hotel and activity searches do not depend on each other. Running them one after another wastes time while each waits for a model. A parallel workflow starts several agents together and continues when all have finished. Here the three finders are grouped into one parallel agent, which then acts as the first step of an ordinary sequence. UntypedAgent gather = AgenticServices.parallelBuilder()
.subAgents(agent(FlightAgent.class, slow), agent(HotelAgent.class, slow), agent(ActivityAgent.class, slow))
.build();
ScriptedChatModel fast = travelModel(0);
TripPlanner trip = AgenticServices.sequenceBuilder(TripPlanner.class)
.subAgents(gather, agent(PlannerAgent.class, fast))
.outputKey("plan").build();
To make the timing visible, the scripted model sleeps 300 ms inside every call. The transcript, from 03-parallel.txt, records when each call started and ended, and on which thread.
Three finders, each model call sleeps 300 ms.
#27 start 0 ms end 301 ms
#31 start 1 ms end 302 ms
#29 start 0 ms end 302 ms
Distinct threads: 3
Wall clock for the whole sequence: 320 ms (sequential would be at least 900 ms)
Plan: Fly AI-101, stay at Hotel Alfama, ride Tram 28.
The three calls started within about two milliseconds of each other, ran on three different threads, and ended together, so the whole sequence took roughly the time of one call instead of three. The test asserts three distinct threads and a wall-clock time under 800 ms. The thread names were blank, so the transcript shows only their numeric ids.
Going deeper: when parallel is the wrong choice
Parallel agents must not depend on each other’s output: anything one of them reads must already be in the scope before the group starts. They also share the same scope, so two parallel agents must write different keys. The builder has an
executor method for supplying your own thread pool; I read that from the jar but did not run it.
The speed-up here is what you get from a model that waits 300 ms. A real provider adds rate limits and cost, and calling three models at once triples the burst against those limits. That is a property of real providers, which this article cannot show.
A loop repeats until a condition holds
Sometimes one pass is not good enough. A common pattern is to draft, score, improve and score again, until the score is high enough or you run out of patience. Two small agents do this: ScorerAgent readsplan and writes an integer score, and ImproverAgent reads plan and score and writes a new plan. Because the improver writes the same key it reads, each pass replaces the plan with a better one.
Refiner loop = AgenticServices.loopBuilder(Refiner.class)
.subAgents(agent(ScorerAgent.class, m), agent(ImproverAgent.class, m))
.maxIterations(5)
.outputKey("plan")
.testExitAtLoopEnd(testExitAtLoopEnd)
.exitCondition(scope -> scope.readState("score", 0) >= 8)
.build();
Two settings matter. maxIterations(5) is the safety limit, so that a model that never reaches the target cannot loop for ever. The exit condition looks at the scope and says when to stop. The scripted scorer answers 5, 7, 9 and 10 on successive calls, and the test runs the loop twice, once with testExitAtLoopEnd off and once on. The transcript is 04-loop.txt.
Scorer replies 5, 7, 9, 10. Exit condition: score >= 8. maxIterations 5.
testExitAtLoopEnd = false (default): final plan 'draft 3', scorer ran 3 times, improver ran 2 times, final score 9
testExitAtLoopEnd = true: final plan 'draft 4', scorer ran 3 times, improver ran 3 times, final score 9
With the default, the condition is checked after each agent, so once the scorer returned 9 the loop stopped before the improver ran a third time, and the plan is “draft 3”. With testExitAtLoopEnd(true) the condition is only checked at the end of a full pass, so the improver ran a third time and the plan is “draft 4”, one improvement beyond what was needed.
Set an output key on the loop, or the result is null. My first version of this test returnednullas the final plan, because the loop had nooutputKeysaying which value in the scope was its answer. Adding.outputKey("plan")fixed it. The scope still held the right plan the whole time; only the method’s return value was missing. If an agentic method returns null, check this first.
Going deeper: other workflow shapes
The same builder class has a
conditionalBuilder (run an agent only when a condition on the scope holds), a parallelMapperBuilder (run one agent once per item of a list), a plannerBuilder for your own planning strategy, and humanInTheLoopBuilder. I listed these from the jar with javap and did not run them, so this article makes no claim about how they behave.
A supervisor lets the model choose the order
In a sequence you decide the order. In a supervisor workflow, a model decides. You hand it a set of agents, and for each step it is asked which agent to run next and with what arguments, until it says it is done. That is useful when the right steps depend on the request: “find me a flight” should not book a hotel. The supervisor is itself a model call, so it has a prompt of its own: a list of the agents with their descriptions and arguments, an instruction to answer in JSON, and a reserved name,done, for finishing. The full prompt is in the last details section of this part.
SupervisorAgent s = AgenticServices.supervisorBuilder().chatModel(sup)
.subAgents(agent(FlightAgent.class, workers), agent(HotelAgent.class, workers), agent(ActivityAgent.class, workers))
.maxAgentsInvocations(5).build();
Here the supervisor model is scripted to answer with three JSON decisions: run findFlight, run findHotel, then done. The decisions have to carry every argument the chosen agent declares, which for these agents includes traveler. The transcript is 05-supervisor.txt.
Supervisor model was scripted to reply with three JSON decisions.
Answer: Hotel Alfama
Worker calls: 2
List one flight option to Lisbon for Ankur.
List one hotel option in Lisbon for Ankur.
Supervisor model calls: 3
First prompt the supervisor saw:
The user request is: 'Find me a flight and a hotel for Lisbon'.
The last received response is: ''.
You must answer strictly in the following JSON format: {
"agentName": (type: string),
"arguments": (type: java.util.Map<java.lang.String, java.lang.Object>)
}
What each responseStrategy returned for the same three decisions:
(default) Hotel Alfama
SCORED Flight AI-101 and Hotel Alfama booked options found. (supervisor calls: 4)
The extra SCORED call asked the supervisor model:
You are a response evaluator that is provided with two responses to a user request.
Your role is to score the two responses based on their relevance for the user request.
For each of the two responses, response1 and response2, you will return a score, respectively score 1 and score 2,
between 0.0 and 1.0, where 0.0 means the response is completely irrelevant to the user request,
and 1.0 means the response is perfectly relevant to the user request.
Return only the score and nothing else, without any additional text or explanation.
The user request is: 'Find me a flight and a hotel for Lisbon'.
The first response is: 'Hotel Alfama'.
The second response is: 'Flight AI-101 and Hotel Alfama booked options found.'.
You must answer strictly in the following JSON format: {
"score1": (type: double),
"score2": (type: double)
}
SUMMARY Flight AI-101 and Hotel Alfama booked options found. (supervisor calls: 3)
LAST Hotel Alfama (supervisor calls: 3)
The supervisor chose two of the three agents, so only two worker calls were made, and the activity agent never ran. The first prompt shows the exact JSON shape the supervisor model must reply in. The last lines show that the same three scripted decisions produce different answers depending on the response strategy. The default returned Hotel Alfama, which is the output of the last worker, not the supervisor’s “done” summary. SUMMARY returned that summary, LAST returned the last worker’s output, and SCORED asked the supervisor model to rate the two candidates, which cost a fourth supervisor call. My scripted scores favoured the summary, so that is what it returned. The prompt it used is in the transcript.
Which answer you get is a setting, not a given. In this test the default behaved likeLAST. If you wanted the model’s own recap of what was done, you had to ask forSUMMARY, and if you wanted the better of the two chosen by the model,SCOREDcosts an extra call. I checked this on 1.22.0-beta32 with a scripted model; with a real model the scores would differ.
What can go wrong
The supervisor depends on a model returning valid JSON and eventually sayingdone. Two tests probe what happens when it does not. The transcript is 06-supervisor-limits.txt.
A) Supervisor model never says done, maxAgentsInvocations = 2.
returned: AI-101 Mumbai to Lisbon
worker calls: 2, supervisor calls: 2
B) Supervisor model replies in prose instead of JSON.
threw: JsonParseException: Unrecognized token 'Sure': was expecting (JSON String, Number, Array, Object or token 'null', 'true' or 'false')
at [Source: REDACTED (`StreamReadFeature.INCLUDE_SOURCE_IN_LOCATION` disabled); line: 1, column: 1]
In case A the supervisor model asked for the flight agent every time and never said done, with maxAgentsInvocations(2). The library stopped after two worker calls and two supervisor calls and returned the last worker output, without an error. That is a safe ending, but silent: nothing tells your code that the supervisor ran out of budget instead of finishing. In case B the model replied in prose, and the library failed with a JSON parse error naming the unexpected word. Real models are usually better behaved than this, but the guard rails you rely on are the invocation limit and your own handling of that exception.
Going deeper: two agents with the same method name
Agent names come from the method name by default. Two agents that both have a method called
find would be indistinguishable to the supervisor, so the library renames them. The test builds two such agents (Finders) and records what the supervisor model was shown, in 07-name-collision.txt. The line that matters is:
The comma separated list of available agents is: '{'find$0', 'Finds a train', [destination: String]}, {'find$1', 'Finds a bus', [destination: String]}'.
The suffixes $0 and $1 are meaningless to a model. Give your agents distinct method names, or set the name in @Agent, so that the supervisor sees something it can reason about. That transcript also contains the full supervisor system prompt, if you want to see how the instructions are phrased.
When an agent fails, you decide what happens next
Model calls fail: a provider is down, a request times out. By default the exception just propagates and the whole workflow ends. An error handler on the workflow builder gets the failure and returns one of three recoveries: retry the agent, substitute a result, or throw. The test uses FlakyAgent, whose model throws on its first N calls, and runs four setups. The transcript is 08-error-handling.txt.FlakyAgent's model throws on its first N calls.
no handler, N=1: Attempt[outcome=threw IllegalStateException: model unavailable (attempt 1), modelCalls=1]
handler retry(), N=2: Attempt[outcome=returned 'AI-101 Mumbai to Lisbon', modelCalls=3] (handler invoked 2 times)
handler result(".."), N=5: Attempt[outcome=returned 'no flight found, ask the traveler', modelCalls=1]
handler throwException(), N=5: Attempt[outcome=threw IllegalStateException: model unavailable (attempt 1), modelCalls=1]
Without a handler, one failed call ended the workflow with the original exception. With ErrorRecoveryResult.retry() the handler ran twice, the model was called three times, and the workflow completed with the real answer. With result("...") the agent’s output was replaced by the string you gave, after a single model call, so the rest of the workflow carried on with a fallback value. With throwException() the original exception came out, as if there were no handler.
The retry case is one line plus a handler, from FlowsTest.java:
Attempt retry = flaky(2, ctx -> {
seenByHandler.incrementAndGet();
return ErrorRecoveryResult.retry();
});
An unconditional retry is a loop you did not size. The retry example works because the failure was scripted to stop after two attempts. The handler as written would retry for ever against a provider that stays down, and the test did not exercise that. If you retry, count attempts yourself in the handler and fall back toresult(...)orthrowException()after a limit.
Going deeper: what the handler receives
The handler gets an
ErrorContext, a record with three parts: the failing agent’s name, the AgenticScope as it stood, and the AgentInvocationException that was thrown. Having the scope means a fallback can depend on what was already gathered. I confirmed this shape from the jar; the example here ignores the context and only decides what to return.
Should you even do this?
A fair answer. A fixed sequence of model calls is just a pipeline, and you may not need a framework for that: three method calls and a few strings work. The module earns its place when you want the parts to be independently testable, to run independent steps in parallel, to loop on a score, or to hand the ordering to a model. Be careful with the supervisor: it hands control flow to a model, so it is the least predictable shape and the hardest to test, and the guard rails shown above (invocation limit, JSON handling, response strategy) are yours to set. Weigh the maturity as well: the agentic module is a beta release with no stable version, so the API in this article can change. What this article did not test: any real model, whether a real model produces the JSON the supervisor needs, the conditional, mapper and planner workflows, human-in-the-loop, agents over A2A, persistence of the scope, and behaviour under real provider failures. The scripted model proves what the library does with the replies it is given, not what any model will answer.
Further reading
- The code for this article: asmhatre/langchain4j-demo, with the transcripts under
agentic/output/and the README. - LangChain4j AI Services with Spring Boot 4 — the single-agent foundation this article builds on.
- Testing LLM apps in Java — once the workflow runs, how to check the answers stay good.
- LangChain4j documentation: Agents and agentic AI
No Comments yet!