Files
java-ai-agents/README.md
T

3.7 KiB

java-ai-agents

Companion code for the ankurm.com posts on agents on the JVM. Every number and transcript quoted in a post comes from a file in a module's output/ folder, and the tests that write those files also assert the same facts, so the build fails when a claim stops being true.

Module Post What it shows
embabel Embabel: Goal-Oriented AI Agents on the JVM actions, goals, conditions and cost-based planning with GOAP, compared with plain Spring AI
a2a A2A Protocol in Java: Agents That Talk to Each Other agent card, tasks, streaming, input-required and agent-to-agent calls with the A2A Java SDK
onnx-djl Running ONNX Models in Java with ONNX Runtime and DJL a DistilBERT sentiment model: DJL tokenizer, ONNX Runtime directly and through DJL's engine, batching, and a Spring Boot endpoint with latency numbers
jlama Local LLM Inference in Pure Java with Jlama load and run a 4-bit TinyLlama inside the JVM on the Vector API, tokens per second, streaming, and the same measurement against Ollama

Versions (verified on Maven Central, 2026-10-11)

Component Version Note
Embabel embabel-agent-api 1.5.3 brings Spring AI 2.0.1 and Spring Boot 4.1.1 transitively
A2A Java SDK io.github.a2asdk 1.0.0.Alpha3 alpha; targets A2A protocol 1.0. The newest stable release, 0.3.3.Final, targets the 0.3 protocol and its API differs
Jlama jlama-core (com.github.tjake) 0.8.4 needs --add-modules jdk.incubator.vector at compile and run time
Model tjake/TinyLlama-1.1B-Chat-v1.0-Jlama-Q4 1.1 GB, downloaded by Jlama into jlama/models/ on first run (git-ignored)
ONNX Runtime com.microsoft.onnxruntime:onnxruntime 1.31.0 overrides the 1.21.1 that DJL's engine declares
DJL (ai.djl) 0.38.0 api, huggingface:tokenizers, onnxruntime:onnxruntime-engine
Spring Boot 4.1.1 onnx-djl endpoint only
Model distilbert/distilbert-base-uncased-finetuned-sst-2-english onnx/model.onnx (268 MB), onnx/tokenizer.json, config.json in onnx-djl/models/sst2/ (git-ignored)
JDK 25 LTS
JUnit 6.1.3

Run

export JAVA_HOME=/path/to/jdk-25
mvn test                      # everything
mvn test -pl a2a -am          # one module

No API key is needed. The A2A agents contain no model call, because the protocol is the subject. The Embabel tests use ScriptedLlmOperations, the scripted stand-in shipped inside embabel-agent-api, and the Spring AI comparison uses a small scripted ChatModel. Each module writes its transcripts to <module>/output/ when you run the tests.

The jlama module needs a model and takes minutes

mvn test -pl jlama downloads the 1.1 GB model on the first run and then runs a few minutes of inference on the CPU. The Ollama comparison is skipped unless you pass the address of a running Ollama that has a model named tl:

# Modelfile: FROM ./tinyllama-1.1b-chat-v1.0.Q4_0.gguf  (TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF)
# plus the TinyLlama chat TEMPLATE and PARAMETER stop "</s>"
ollama create tl -f Modelfile
mvn test -pl jlama -Dollama.url=http://127.0.0.1:11434

The committed numbers in jlama/output/ come from one 2-vCPU cloud machine, not a laptop. Run it on yours and expect different figures.

The onnx-djl module needs a model on disk

Download four files from https://huggingface.co/distilbert/distilbert-base-uncased-finetuned-sst-2-english/resolve/main/ into onnx-djl/models/sst2/: onnx/model.onnx, onnx/tokenizer.json, onnx/tokenizer_config.json and config.json, then run mvn test -pl onnx-djl. Each test class runs in its own JVM (reuseForks=false) because only one creator of ONNX Runtime's shared environment may exist per JVM.