Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01G8ikz8xdWuTP5yun8DZ1hk
java-ai-agents
Companion code for the ankurm.com posts on agents on the JVM. Every number and transcript quoted in a post
comes from a file in a module's output/ folder, and the tests that write those files also assert the same
facts, so the build fails when a claim stops being true.
| Module | Post | What it shows |
|---|---|---|
embabel |
Embabel: Goal-Oriented AI Agents on the JVM | actions, goals, conditions and cost-based planning with GOAP, compared with plain Spring AI |
a2a |
A2A Protocol in Java: Agents That Talk to Each Other | agent card, tasks, streaming, input-required and agent-to-agent calls with the A2A Java SDK |
jlama |
Local LLM Inference in Pure Java with Jlama | load and run a 4-bit TinyLlama inside the JVM on the Vector API, tokens per second, streaming, and the same measurement against Ollama |
Versions (verified on Maven Central, 2026-10-11)
| Component | Version | Note |
|---|---|---|
Embabel embabel-agent-api |
1.5.3 | brings Spring AI 2.0.1 and Spring Boot 4.1.1 transitively |
A2A Java SDK io.github.a2asdk |
1.0.0.Alpha3 | alpha; targets A2A protocol 1.0. The newest stable release, 0.3.3.Final, targets the 0.3 protocol and its API differs |
Jlama jlama-core (com.github.tjake) |
0.8.4 | needs --add-modules jdk.incubator.vector at compile and run time |
| Model | tjake/TinyLlama-1.1B-Chat-v1.0-Jlama-Q4 |
1.1 GB, downloaded by Jlama into jlama/models/ on first run (git-ignored) |
| JDK | 25 LTS | |
| JUnit | 6.1.3 |
Run
export JAVA_HOME=/path/to/jdk-25
mvn test # everything
mvn test -pl a2a -am # one module
No API key is needed. The A2A agents contain no model call, because the protocol is the subject. The Embabel tests use ScriptedLlmOperations, the scripted stand-in shipped inside
embabel-agent-api, and the Spring AI comparison uses a small scripted ChatModel.
Each module writes its transcripts to <module>/output/ when you run the tests.
The jlama module needs a model and takes minutes
mvn test -pl jlama downloads the 1.1 GB model on the first run and then runs a few minutes of inference on the CPU.
The Ollama comparison is skipped unless you pass the address of a running Ollama that has a model named tl:
# Modelfile: FROM ./tinyllama-1.1b-chat-v1.0.Q4_0.gguf (TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF)
# plus the TinyLlama chat TEMPLATE and PARAMETER stop "</s>"
ollama create tl -f Modelfile
mvn test -pl jlama -Dollama.url=http://127.0.0.1:11434
The committed numbers in jlama/output/ come from one 2-vCPU cloud machine, not a laptop. Run it on yours and expect different figures.