Versions, and an honest limit. Embabelembabel-agent-api1.5.3, which brings Spring AI 2.0.1 and Spring Boot 4.1.1 with it (see 07-dependencies.txt). Java 25 and JUnit 6.1.3. All the code is in theembabelmodule of asmhatre/java-ai-agents, and every console block below is quoted from a file underembabel/output/, written by a test that asserts the same facts. No live model was used, and the agents were not started as a Spring Boot application. The tests deploy the agent on a test platform that ships insideembabel-agent-apiand replace the LLM withScriptedLlmOperations, a scripted stand-in from the same jar. That tells you what the planner does. It tells you nothing about how well a real model summarises, and nothing about the Spring Boot starter, the shell, or the model provider integrations, none of which I ran.
The vocabulary is your own domain types
The planner reasons about types. So the first thing to write is the set of objects that flow through the job. For a research task there are four: a topic to research, the sources found for it, a summary of those sources, and a finished brief. Plain Java records are enough. From Domain.java: public record Topic(String text) {
}
public record Sources(List<String> items) {
}
public record Summary(String text) {
}
public record Brief(String text) {
}
Actions say what they need, and one of them names the goal
An Embabel agent is a class annotated@Agent. Each @Action method declares its inputs as parameters and its output as the return type. Exactly one marks the end with @AchievesGoal. Here is the research agent from ResearchAgent.java, with the search corpus left out. findSources does not use a model at all, because it is an ordinary lookup, and the other two ask the model through OperationContext:
@Action(description = "Look the topic up in the local corpus")
public Sources findSources(Topic topic) {
return new Sources(CORPUS.getOrDefault(topic.text(), List.of()));
}
@Action(description = "Summarise the sources with an LLM")
public Summary summarise(Sources sources, OperationContext context) {
return context.ai().withDefaultLlm()
.createObject("Summarise these notes in one sentence: " + String.join(" ", sources.items()), Summary.class);
}
@AchievesGoal(description = "A short brief about the topic")
@Action(description = "Write the brief from the summary")
public Brief writeBrief(Topic topic, Summary summary, OperationContext context) {
return context.ai().withDefaultLlm()
.createObject("Write a two-line brief titled '" + topic.text() + "' from: " + summary.text(), Brief.class);
}
No method calls another. summarise needs a Sources, so something must produce one first, and only findSources does. writeBrief needs a Topic and a Summary, and it is the goal. That is all the planner is given.
Going deeper: what a planner does here, and which one
The default planner for
@Agent is GOAP, goal-oriented action planning, an idea from game AI. The agent’s current knowledge is a set of facts (what objects exist, which conditions hold). Each action has preconditions (the types and conditions it needs) and effects (what it adds). The planner searches for the cheapest chain of actions from the current facts to a state where the goal is achieved. Embabel’s enum PlannerType also lists UTILITY, HYBRID and SUPERVISOR (read from the jar with javap); I only ran the default. The transcript in the next section prints the planner class the process used, AStarGoapPlanner.
- Embabel on GitHub
- ResearchAgent.java — the full agent, including the small in-memory corpus used instead of a search engine.
Running it: the planner chose the order
To run an agent you read its annotations into an agent definition, deploy it on a platform, and start a process with the starting objects. The test uses a helper for that, from EmbabelTest.java. The platform is the test platform from the Embabel jar, and the model is a scripted one that hands back prepared objects in order. private static Run run(Object agentInstance, Map<String, Object> inputs, Object... llmReplies) {
ScriptedLlmOperations llm = new ScriptedLlmOperations();
for (Object o : llmReplies) llm.returnObject(o);
AgentPlatform platform = IntegrationTestUtils.dummyAgentPlatform(llm);
Agent agent = (Agent) new AgentMetadataReader().createAgentMetadata(agentInstance);
platform.deploy(agent);
AgentProcess p = platform.runAgentFrom(agent, ProcessOptions.DEFAULT, inputs);
return new Run(p, llm);
}
The caller supplies a Topic under the key it. Everything else is up to the planner. The transcript, from 01-plan.txt, lists the actions in the order they ran, the objects left on the blackboard (the process’s shared memory), and the prompts the model received.
ResearchAgent, input Topic("virtual threads"), planner GOAP.
Planner: AStarGoapPlanner
Status: COMPLETED
Actions in the order they ran:
com.ankurm.agents.embabel.ResearchAgent.findSources
com.ankurm.agents.embabel.ResearchAgent.summarise
com.ankurm.agents.embabel.ResearchAgent.writeBrief
Objects on the blackboard at the end:
Topic
Sources[items=[JEP 444 makes virtual threads final in Java 21., Virtual threads are cheap to block, so thread-per-request scales again.]]
Summary
Brief
Prompts the LLM received (2):
Summarise these notes in one sentence: JEP 444 makes virtual threads final in Java 21. Virtual threads are cheap to block, so thread-per-request scales again.
Write a two-line brief titled 'virtual threads' from: Virtual threads make blocking cheap.
The planner ran findSources, then summarise, then writeBrief, which is the order of the arrows in the diagram. The model was called twice, once for each action that asked it, and never for findSources. Each object that an action returned was added to the blackboard, which is how the next action found its input.
When the data is already known, the planner skips the step
Now give the process aSources object at the start as well as the topic. Nothing in the agent changes. The transcript is 02-skips-known.txt.
Same agent, but the caller also hands over a Sources object.
Actions in the order they ran:
com.ankurm.agents.embabel.ResearchAgent.summarise
com.ankurm.agents.embabel.ResearchAgent.writeBrief
Prompts the LLM received:
Summarise these notes in one sentence: A note we already had.
Write a two-line brief titled 'virtual threads' from: Virtual threads make blocking cheap.
findSources did not run, and the first prompt used the supplied note. In hand-written orchestration this is the “if we already have it” branch you would have to remember to write. Here it falls out of the planner finding that the facts it needs already exist.
Conditions can make the goal unreachable, and that is a feature
Summarising nothing is a waste of a model call. A condition is a named yes-or-no check that an action can require withpre. GuardedResearchAgent adds one, hasSources, and requires it before summarise. Note the first line, which is easy to leave out and which the callout below explains:
@Action(post = "hasSources")
public Sources findSources(Topic topic) {
return delegate.findSources(topic);
}
@Condition(name = "hasSources")
public boolean hasSources(Sources sources) {
return !sources.items().isEmpty();
}
@Action(pre = "hasSources")
public Summary summarise(Sources sources, OperationContext context) {
return delegate.summarise(sources, context);
}
The transcript runs it with a topic that has sources and one that does not. It is 03-condition.txt.
GuardedResearchAgent: summarise has pre = "hasSources".
Topic "virtual threads": status COMPLETED, actions:
com.ankurm.agents.embabel.GuardedResearchAgent.findSources
com.ankurm.agents.embabel.GuardedResearchAgent.summarise
com.ankurm.agents.embabel.GuardedResearchAgent.writeBrief
LLM prompts: 2
Topic "quantum gravity": status STUCK, actions:
com.ankurm.agents.embabel.GuardedResearchAgent.findSources
LLM prompts: 0
With a known topic the plan is the same three steps. With an unknown topic findSources ran, found nothing, and the process stopped with status STUCK before any model call: zero prompts were sent. That is the agent declining to continue rather than summarising an empty list.
A condition nobody establishes makes the goal unreachable from the start. My first version of this agent had the condition and thepre, but notpost = "hasSources"onfindSources. The planner works from the declared effects of actions, and no action claimed to makehasSourcestrue, so it could not build any plan. Even for a topic that has sources the process wasSTUCKand no action ran at all, not evenfindSources. The transcript, 06-condition-trap.txt, comes from ForgetfulGuardedAgent, which is the same apart from its name and that one missing attribute.
ForgetfulGuardedAgent: identical to GuardedResearchAgent except findSources has no post = "hasSources".
Topic "virtual threads" (a topic that has sources): status STUCK, actions run: 0
LLM prompts: 0
Going deeper: what post means, and what it cannot do
post = "hasSources" says that after findSources runs, the condition is expected to hold. The planner uses that to build a plan. The condition method itself is then evaluated for real after the action has run, which is why the unknown topic stops after findSources. I observed this behaviour; I did not read the planner’s source to confirm that is the mechanism, so treat that explanation as my reading of the results.
The condition method takes a Sources parameter. I only used a condition with a domain-object parameter; I did not test conditions with other signatures.
Cost: when two actions can do the same job
Sometimes there are two ways to produce the same type. CostAwareResearchAgent has two methods that both return aSummary: one asks the model and is marked expensive, and one just takes the first source and is marked cheap.
@Action(cost = 0.9, description = "Summarise with an LLM")
public Summary summariseWithLlm(Sources sources, OperationContext context) {
return delegate.summarise(sources, context);
}
@Action(cost = 0.1, description = "Take the first source verbatim")
public Summary summariseLocally(Sources sources) {
return new Summary(sources.items().getFirst());
}
The transcript is 04-cost.txt. The scripted model was given a single object to hand back, the final brief, because the planner was expected to avoid the summarising call.
CostAwareResearchAgent: summariseWithLlm cost 0.9, summariseLocally cost 0.1, both return Summary.
Status: COMPLETED
Actions in the order they ran:
com.ankurm.agents.embabel.CostAwareResearchAgent.findSources
com.ankurm.agents.embabel.CostAwareResearchAgent.summariseLocally
com.ankurm.agents.embabel.CostAwareResearchAgent.writeBrief
Prompts the LLM received (1):
Write a two-line brief titled 'virtual threads' from: JEP 444 makes virtual threads final in Java 21.
The planner picked summariseLocally, and the model was called once instead of twice. The cost numbers are yours to set, and they only mean something relative to each other: here they say “a model call is worth avoiding”. They do not measure real money or latency. Whether the cheaper action is good enough is a question about your task, and the planner cannot judge it.
The same job in plain Spring AI
Here is the equivalent written the ordinary way, with Spring AI’sChatClient and the order written by hand. It is PlainResearch.java and it reuses the same lookup and the same prompts.
public Brief research(Topic topic) {
Sources sources = search.findSources(topic);
Summary summary = new Summary(chat.prompt()
.user("Summarise these notes in one sentence: " + String.join(" ", sources.items()))
.call().content());
return new Brief(chat.prompt()
.user("Write a two-line brief titled '" + topic.text() + "' from: " + summary.text())
.call().content());
}
A test runs both versions against scripted models and compares what each sent. The transcript is 05-plain-spring-ai.txt.
Plain Spring AI (ChatClient, order written by hand):
result: Brief
prompts:
Summarise these notes in one sentence: JEP 444 makes virtual threads final in Java 21. Virtual threads are cheap to block, so thread-per-request scales again.
Write a two-line brief titled 'virtual threads' from: Virtual threads make blocking cheap.
Embabel:
result: Brief
prompts:
Summarise these notes in one sentence: JEP 444 makes virtual threads final in Java 21. Virtual threads are cheap to block, so thread-per-request scales again.
Write a two-line brief titled 'virtual threads' from: Virtual threads make blocking cheap.
Same prompts: true
The two send identical prompts and return the same brief. So for this fixed three-step job the two approaches produce the same behaviour, and the plain version is easier to read straight down. The differences show up when the situation changes. These are the rows demonstrated above, not a general verdict:
| Situation | Plain Spring AI | Embabel |
|---|---|---|
| Order of steps | You write it | The planner derives it from types |
| Data already available | You add an if | The step is skipped (transcript 02) |
| Nothing to summarise | You add a guard | Declared as a condition; the process stops (transcript 03) |
| Two ways to get one result | You choose in code | The planner takes the cheaper (transcript 04) |
| When it goes wrong | A stack trace at the line that failed | A process status, such as STUCK, that you have to interpret (transcript 06) |
Going deeper: what the plain version does not show
The plain version above has no guard for an empty
Sources. I left it out on purpose to keep the two prompt-for-prompt comparable, and a real version would need it, which is the third row of the table. I did not write the plain equivalents of the skip, guard and cost cases, so the “you add an if” cells describe what you would have to do, not code I ran.
- Spring AI 2.0 ChatClient on Spring Boot 4.1 — the plain version’s building block.
- Tool calling in Spring AI 2.0
Should you even do this?
A fair answer. A fixed pipeline of a few model calls does not need a planner, and the plain version above is shorter and easier to debug. Embabel earns its place when the set of steps is large or changes, when steps are optional, when several actions can produce the same type, or when you want guards and costs declared in one place instead of scattered through control flow. The price is a new vocabulary (actions, goals, conditions, blackboard), plans that are harder to follow than a method body, and failures that appear as a process status instead of a stack trace. What this article did not test: any real model; running the agent inside a Spring Boot application with the starter; the other planners; the shell; chaining several agents; and how the planner scales with many actions. The scripted model proves what the planner does with the objects it is given, not what any model will write.
Further reading
- The code for this article: asmhatre/java-ai-agents, with the transcripts under
embabel/output/and the README. - Embabel agent framework on GitHub
- Spring AI 2.0 ChatClient on Spring Boot 4.1 and tool calling in Spring AI 2.0 — the plain approach this article compares against.
- Agentic workflows with LangChain4j — the same problem solved by composing workflows instead of planning.
- Testing LLM apps in Java
No Comments yet!