Skip to main content

Embabel: Goal-Oriented AI Agents on the JVM

Embabel plans the order of an agent’s steps from the types they need and produce. A small research agent shows the plan, a skipped step, a failing condition and a cost choice, against plain Spring AI.

Most programs that use a language model end up with an orchestrator: a method that calls the model, takes the answer, decides what to call next, and so on. You write the order by hand. That is fine for three steps. It gets harder when steps are optional, when some can be skipped because you already have the data, and when two different steps could produce the same result and one is cheaper. Embabel is a JVM framework that takes a different approach. You do not write the order. You describe what each step needs and what it produces, name the goal, and a planner works out the sequence. This article builds a small research agent, shows the plan the planner chose, then changes the situation (data already known, a precondition that fails, two ways to do one step) and shows the plan change. It ends with the same job written in plain Spring AI, so you can judge whether the planner is worth having. Depth is in expandable sections, so you can read straight through or open only what you need.
Versions, and an honest limit. Embabel embabel-agent-api 1.5.3, which brings Spring AI 2.0.1 and Spring Boot 4.1.1 with it (see 07-dependencies.txt). Java 25 and JUnit 6.1.3. All the code is in the embabel module of asmhatre/java-ai-agents, and every console block below is quoted from a file under embabel/output/, written by a test that asserts the same facts. No live model was used, and the agents were not started as a Spring Boot application. The tests deploy the agent on a test platform that ships inside embabel-agent-api and replace the LLM with ScriptedLlmOperations, a scripted stand-in from the same jar. That tells you what the planner does. It tells you nothing about how well a real model summarises, and nothing about the Spring Boot starter, the shell, or the model provider integrations, none of which I ran.

The vocabulary is your own domain types

The planner reasons about types. So the first thing to write is the set of objects that flow through the job. For a research task there are four: a topic to research, the sources found for it, a summary of those sources, and a finished brief. Plain Java records are enough. From Domain.java:
    public record Topic(String text) {
    }

    public record Sources(List<String> items) {
    }

    public record Summary(String text) {
    }

    public record Brief(String text) {
    }
Topicwhat the caller givesSourcesfound for the topicSummaryof the sourcesBriefthe goal
The diagram is the whole design. Each arrow will become one method. Nothing in the code will say “do the arrows in this order”; the order is implied by which type each method needs and which type it returns.

Actions say what they need, and one of them names the goal

An Embabel agent is a class annotated @Agent. Each @Action method declares its inputs as parameters and its output as the return type. Exactly one marks the end with @AchievesGoal. Here is the research agent from ResearchAgent.java, with the search corpus left out. findSources does not use a model at all, because it is an ordinary lookup, and the other two ask the model through OperationContext:
    @Action(description = "Look the topic up in the local corpus")
    public Sources findSources(Topic topic) {
        return new Sources(CORPUS.getOrDefault(topic.text(), List.of()));
    }

    @Action(description = "Summarise the sources with an LLM")
    public Summary summarise(Sources sources, OperationContext context) {
        return context.ai().withDefaultLlm()
                .createObject("Summarise these notes in one sentence: " + String.join(" ", sources.items()), Summary.class);
    }

    @AchievesGoal(description = "A short brief about the topic")
    @Action(description = "Write the brief from the summary")
    public Brief writeBrief(Topic topic, Summary summary, OperationContext context) {
        return context.ai().withDefaultLlm()
                .createObject("Write a two-line brief titled '" + topic.text() + "' from: " + summary.text(), Brief.class);
    }
No method calls another. summarise needs a Sources, so something must produce one first, and only findSources does. writeBrief needs a Topic and a Summary, and it is the goal. That is all the planner is given.
Going deeper: what a planner does here, and which one
The default planner for @Agent is GOAP, goal-oriented action planning, an idea from game AI. The agent’s current knowledge is a set of facts (what objects exist, which conditions hold). Each action has preconditions (the types and conditions it needs) and effects (what it adds). The planner searches for the cheapest chain of actions from the current facts to a state where the goal is achieved. Embabel’s enum PlannerType also lists UTILITY, HYBRID and SUPERVISOR (read from the jar with javap); I only ran the default. The transcript in the next section prints the planner class the process used, AStarGoapPlanner.

Running it: the planner chose the order

To run an agent you read its annotations into an agent definition, deploy it on a platform, and start a process with the starting objects. The test uses a helper for that, from EmbabelTest.java. The platform is the test platform from the Embabel jar, and the model is a scripted one that hands back prepared objects in order.
    private static Run run(Object agentInstance, Map<String, Object> inputs, Object... llmReplies) {
        ScriptedLlmOperations llm = new ScriptedLlmOperations();
        for (Object o : llmReplies) llm.returnObject(o);
        AgentPlatform platform = IntegrationTestUtils.dummyAgentPlatform(llm);
        Agent agent = (Agent) new AgentMetadataReader().createAgentMetadata(agentInstance);
        platform.deploy(agent);
        AgentProcess p = platform.runAgentFrom(agent, ProcessOptions.DEFAULT, inputs);
        return new Run(p, llm);
    }
The caller supplies a Topic under the key it. Everything else is up to the planner. The transcript, from 01-plan.txt, lists the actions in the order they ran, the objects left on the blackboard (the process’s shared memory), and the prompts the model received.
ResearchAgent, input Topic("virtual threads"), planner GOAP.

Planner: AStarGoapPlanner
Status: COMPLETED

Actions in the order they ran:
  com.ankurm.agents.embabel.ResearchAgent.findSources
  com.ankurm.agents.embabel.ResearchAgent.summarise
  com.ankurm.agents.embabel.ResearchAgent.writeBrief

Objects on the blackboard at the end:
  Topic
  Sources[items=[JEP 444 makes virtual threads final in Java 21., Virtual threads are cheap to block, so thread-per-request scales again.]]
  Summary
  Brief

Prompts the LLM received (2):
  Summarise these notes in one sentence: JEP 444 makes virtual threads final in Java 21. Virtual threads are cheap to block, so thread-per-request scales again.
  Write a two-line brief titled 'virtual threads' from: Virtual threads make blocking cheap.
The planner ran findSources, then summarise, then writeBrief, which is the order of the arrows in the diagram. The model was called twice, once for each action that asked it, and never for findSources. Each object that an action returned was added to the blackboard, which is how the next action found its input.

When the data is already known, the planner skips the step

Now give the process a Sources object at the start as well as the topic. Nothing in the agent changes. The transcript is 02-skips-known.txt.
Same agent, but the caller also hands over a Sources object.

Actions in the order they ran:
  com.ankurm.agents.embabel.ResearchAgent.summarise
  com.ankurm.agents.embabel.ResearchAgent.writeBrief

Prompts the LLM received:
  Summarise these notes in one sentence: A note we already had.
  Write a two-line brief titled 'virtual threads' from: Virtual threads make blocking cheap.
findSources did not run, and the first prompt used the supplied note. In hand-written orchestration this is the “if we already have it” branch you would have to remember to write. Here it falls out of the planner finding that the facts it needs already exist.

Conditions can make the goal unreachable, and that is a feature

Summarising nothing is a waste of a model call. A condition is a named yes-or-no check that an action can require with pre. GuardedResearchAgent adds one, hasSources, and requires it before summarise. Note the first line, which is easy to leave out and which the callout below explains:
    @Action(post = "hasSources")
    public Sources findSources(Topic topic) {
        return delegate.findSources(topic);
    }

    @Condition(name = "hasSources")
    public boolean hasSources(Sources sources) {
        return !sources.items().isEmpty();
    }

    @Action(pre = "hasSources")
    public Summary summarise(Sources sources, OperationContext context) {
        return delegate.summarise(sources, context);
    }
The transcript runs it with a topic that has sources and one that does not. It is 03-condition.txt.
GuardedResearchAgent: summarise has pre = "hasSources".

Topic "virtual threads": status COMPLETED, actions:
  com.ankurm.agents.embabel.GuardedResearchAgent.findSources
  com.ankurm.agents.embabel.GuardedResearchAgent.summarise
  com.ankurm.agents.embabel.GuardedResearchAgent.writeBrief
  LLM prompts: 2

Topic "quantum gravity": status STUCK, actions:
  com.ankurm.agents.embabel.GuardedResearchAgent.findSources
  LLM prompts: 0
With a known topic the plan is the same three steps. With an unknown topic findSources ran, found nothing, and the process stopped with status STUCK before any model call: zero prompts were sent. That is the agent declining to continue rather than summarising an empty list.
A condition nobody establishes makes the goal unreachable from the start. My first version of this agent had the condition and the pre, but not post = "hasSources" on findSources. The planner works from the declared effects of actions, and no action claimed to make hasSources true, so it could not build any plan. Even for a topic that has sources the process was STUCK and no action ran at all, not even findSources. The transcript, 06-condition-trap.txt, comes from ForgetfulGuardedAgent, which is the same apart from its name and that one missing attribute.
ForgetfulGuardedAgent: identical to GuardedResearchAgent except findSources has no post = "hasSources".

Topic "virtual threads" (a topic that has sources): status STUCK, actions run: 0
  LLM prompts: 0
Going deeper: what post means, and what it cannot do
post = "hasSources" says that after findSources runs, the condition is expected to hold. The planner uses that to build a plan. The condition method itself is then evaluated for real after the action has run, which is why the unknown topic stops after findSources. I observed this behaviour; I did not read the planner’s source to confirm that is the mechanism, so treat that explanation as my reading of the results. The condition method takes a Sources parameter. I only used a condition with a domain-object parameter; I did not test conditions with other signatures.

Cost: when two actions can do the same job

Sometimes there are two ways to produce the same type. CostAwareResearchAgent has two methods that both return a Summary: one asks the model and is marked expensive, and one just takes the first source and is marked cheap.
    @Action(cost = 0.9, description = "Summarise with an LLM")
    public Summary summariseWithLlm(Sources sources, OperationContext context) {
        return delegate.summarise(sources, context);
    }

    @Action(cost = 0.1, description = "Take the first source verbatim")
    public Summary summariseLocally(Sources sources) {
        return new Summary(sources.items().getFirst());
    }
The transcript is 04-cost.txt. The scripted model was given a single object to hand back, the final brief, because the planner was expected to avoid the summarising call.
CostAwareResearchAgent: summariseWithLlm cost 0.9, summariseLocally cost 0.1, both return Summary.

Status: COMPLETED
Actions in the order they ran:
  com.ankurm.agents.embabel.CostAwareResearchAgent.findSources
  com.ankurm.agents.embabel.CostAwareResearchAgent.summariseLocally
  com.ankurm.agents.embabel.CostAwareResearchAgent.writeBrief

Prompts the LLM received (1):
  Write a two-line brief titled 'virtual threads' from: JEP 444 makes virtual threads final in Java 21.
The planner picked summariseLocally, and the model was called once instead of twice. The cost numbers are yours to set, and they only mean something relative to each other: here they say “a model call is worth avoiding”. They do not measure real money or latency. Whether the cheaper action is good enough is a question about your task, and the planner cannot judge it.

The same job in plain Spring AI

Here is the equivalent written the ordinary way, with Spring AI’s ChatClient and the order written by hand. It is PlainResearch.java and it reuses the same lookup and the same prompts.
    public Brief research(Topic topic) {
        Sources sources = search.findSources(topic);
        Summary summary = new Summary(chat.prompt()
                .user("Summarise these notes in one sentence: " + String.join(" ", sources.items()))
                .call().content());
        return new Brief(chat.prompt()
                .user("Write a two-line brief titled '" + topic.text() + "' from: " + summary.text())
                .call().content());
    }
A test runs both versions against scripted models and compares what each sent. The transcript is 05-plain-spring-ai.txt.
Plain Spring AI (ChatClient, order written by hand):
  result: Brief
  prompts:
    Summarise these notes in one sentence: JEP 444 makes virtual threads final in Java 21. Virtual threads are cheap to block, so thread-per-request scales again.
    Write a two-line brief titled 'virtual threads' from: Virtual threads make blocking cheap.

Embabel:
  result: Brief
  prompts:
    Summarise these notes in one sentence: JEP 444 makes virtual threads final in Java 21. Virtual threads are cheap to block, so thread-per-request scales again.
    Write a two-line brief titled 'virtual threads' from: Virtual threads make blocking cheap.

Same prompts: true
The two send identical prompts and return the same brief. So for this fixed three-step job the two approaches produce the same behaviour, and the plain version is easier to read straight down. The differences show up when the situation changes. These are the rows demonstrated above, not a general verdict:
SituationPlain Spring AIEmbabel
Order of stepsYou write itThe planner derives it from types
Data already availableYou add an ifThe step is skipped (transcript 02)
Nothing to summariseYou add a guardDeclared as a condition; the process stops (transcript 03)
Two ways to get one resultYou choose in codeThe planner takes the cheaper (transcript 04)
When it goes wrongA stack trace at the line that failedA process status, such as STUCK, that you have to interpret (transcript 06)
Going deeper: what the plain version does not show
The plain version above has no guard for an empty Sources. I left it out on purpose to keep the two prompt-for-prompt comparable, and a real version would need it, which is the third row of the table. I did not write the plain equivalents of the skip, guard and cost cases, so the “you add an if” cells describe what you would have to do, not code I ran.

Should you even do this?

A fair answer. A fixed pipeline of a few model calls does not need a planner, and the plain version above is shorter and easier to debug. Embabel earns its place when the set of steps is large or changes, when steps are optional, when several actions can produce the same type, or when you want guards and costs declared in one place instead of scattered through control flow. The price is a new vocabulary (actions, goals, conditions, blackboard), plans that are harder to follow than a method body, and failures that appear as a process status instead of a stack trace. What this article did not test: any real model; running the agent inside a Spring Boot application with the starter; the other planners; the shell; chaining several agents; and how the planner scales with many actions. The scripted model proves what the planner does with the objects it is given, not what any model will write.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.