Versions, and honest limits. Jlamajlama-core0.8.4 (groupcom.github.tjake), Java 25, JUnit 6.1.3, and the modeltjake/TinyLlama-1.1B-Chat-v1.0-Jlama-Q4(1.1 GB). All the code is in thejlamamodule of asmhatre/java-ai-agents, and every console block below is quoted from a file underjlama/output/, written by a test. The numbers are from a cloud sandbox with 2 virtual CPUs (Intel Xeon at 2.1 GHz), not a laptop, and the transcripts report that machine. Your hardware will give different figures and may give a different ratio against Ollama. I did not test a GPU, the optionaljlama-nativelibrary, any model other than TinyLlama, or running several requests at once. Inference ran in the test JVM only; there is no Spring Boot or LangChain4j integration here.
The only setup is a dependency and one JVM flag
The dependency is one artifact,jlama-core. Its transitive dependencies are Jackson, Guava and a template engine for chat prompts, listed in 08-dependencies.txt. The part that trips people up is the Vector API. In JDK 25 it is still an incubator module, so the compiler and the running JVM both have to be told to include it. These are the two relevant settings from pom.xml, the first for the compiler and the second for the JVM that runs the tests:
<compilerArgs>
<arg>--add-modules</arg>
<arg>jdk.incubator.vector</arg>
</compilerArgs>
<argLine>--add-modules jdk.incubator.vector</argLine>
I ran a second JVM without that flag to see what happens. The test starts one with the same classpath and no --add-modules. Its output, from 05-without-vector-module.txt:
Second JVM, same classpath, started WITHOUT --add-modules jdk.incubator.vector.
Exit code: 1
Relevant lines:
Exception in thread "main" java.lang.RuntimeException: java.lang.reflect.InvocationTargetException
Caused by: java.lang.ClassNotFoundException: jdk.incubator.vector.FloatVector
There is no fallback to slower scalar code: the model refuses to load, and the exception message does not mention a flag. If you see ClassNotFoundException: jdk.incubator.vector.FloatVector, add --add-modules jdk.incubator.vector to the JVM options.
Going deeper: the README asks for a second flag
Jlama’s README, which I read on GitHub, tells you to set
JDK_JAVA_OPTIONS to --add-modules jdk.incubator.vector --enable-preview. With version 0.8.4 on JDK 25 I ran the model with --add-modules alone and it loaded and answered, and the test setup above uses only that flag. I do not know whether other models or code paths need --enable-preview; I only exercised a Llama-architecture model. The same README says Jlama needs Java 20 or later; I only ran JDK 25.
Load a model and ask a question
Jlama’s own API has a handful of steps: find or download the model directory, load it with a working-memory type and a quantisation type, build a prompt through the model’s chat template, and callgenerate. LocalModel.java wraps those so the tests read as “load, ask, measure”:
public static LocalModel load(String modelsDir) throws IOException {
File dir = SafeTensorSupport.maybeDownloadModel(modelsDir, MODEL);
return new LocalModel(ModelSupport.loadModel(dir, DType.F32, DType.I8));
}
public Answer ask(String question, int maxTokens, float temperature) {
return stream(question, maxTokens, temperature, (token, ms) -> {
});
}
/** onToken receives each piece of text as the model produces it. */
public Answer stream(String question, int maxTokens, float temperature, BiConsumer<String, Float> onToken) {
PromptContext prompt = model.promptSupport().orElseThrow().builder().addUserMessage(question).build();
var r = model.generate(UUID.randomUUID(), prompt, temperature, maxTokens, onToken);
return new Answer(r.responseText, r.promptTokens, r.generatedTokens, r.promptTimeMs, r.generateTimeMs);
}
maybeDownloadModel fetches the model from Hugging Face into a local folder the first time and finds it there afterwards; downloading the 1.1 GB took 11 seconds in my sandbox. The first test loads the model and asks one question. Its output is 01-load-and-ask.txt:
Machine: 2 CPUs, JDK 25.0.4.1+1-LTS, Vector API preferred float width 512 bits
Model: tjake/TinyLlama-1.1B-Chat-v1.0-Jlama-Q4, 1132 MB on disk
Load time: 665 ms; process resident memory 119 MB before, 175 MB after
Question: In one sentence, what is a Java record?
Answer: A Java record is a data structure that allows for the creation of immutable, named, and indexed collections of objects.
Prompt tokens: 26, generated tokens: 24
Prompt time: 4838 ms, generation time: 2762 ms
The preferred vector width on this CPU was 512 bits, so the Vector API could process sixteen floats per instruction. The load took under a second here; the file was on a local disk and the resident memory figures (119 MB before, 175 MB after) suggest the 1.1 GB of weights are not copied onto the Java heap. The answer is a fluent sentence that is not quite right: a record is not a collection of indexed objects. That is a property of a 1.1-billion-parameter model, not of Jlama, and it is a reminder that these tiny models are for plumbing and experiments, not for answers you rely on.
How fast is it?
Speed is reported in tokens per second, where a token is a word or a piece of one. Jlama returns two timings for each call: how long it took to read the prompt, and how long it took to generate the reply. They behave very differently. The benchmark test asks the same question once as a warm-up (presumably the JIT compiler and the memory pages are cold on the first call) and then five times, with temperature 0. Output, from 02-tokens-per-second.txt:Same question, 5 measured runs after 1 warm-up run, temperature 0, limit 64 tokens.
run 1: prompt 26 tokens in 4902 ms (5.3 tok/s), generated 24 tokens in 3040 ms (7.9 tok/s)
run 2: prompt 26 tokens in 4762 ms (5.5 tok/s), generated 24 tokens in 3061 ms (7.8 tok/s)
run 3: prompt 26 tokens in 4527 ms (5.7 tok/s), generated 24 tokens in 2856 ms (8.4 tok/s)
run 4: prompt 26 tokens in 4892 ms (5.3 tok/s), generated 24 tokens in 3007 ms (8.0 tok/s)
run 5: prompt 26 tokens in 5203 ms (5.0 tok/s), generated 24 tokens in 2894 ms (8.3 tok/s)
Generation tokens/s: min 7.8, median 8.0, max 8.4
Prompt tokens/s: min 5.0, median 5.3, max 5.7
Prompt + generated tokens in the longest run: 50 (limit was 64)
Generation ran at about 8 tokens per second, steady across runs. Reading the prompt was slower than generating: about 5 tokens per second, so a 26-token question took almost 5 seconds before the first word appeared. For a chat interface that is the number you feel. Longer prompts, such as retrieved documents, would take proportionally longer to read.
Going deeper: what the benchmark does and does not measure
The timings are the ones Jlama reports on each response, not a stopwatch around my code, and each run produced the same 24 tokens. Five runs on one machine at one moment is a small sample; the spread (7.8 to 8.4 tokens per second) is the spread I saw, not a confidence interval. A shared cloud CPU may be noisier than a laptop. I did not vary the quantisation setting, the number of threads, or the prompt length, so I cannot tell you which of those would help most.
- JlamaTest.java
- 03-streaming.txt — the same prompt time seen from the streaming side.
Streaming, and a token limit that includes the prompt
generate takes a callback that receives each piece of text as it is produced, which is how you would feed a web page or a server-sent-events stream. The test records when each piece arrives. Output, from 03-streaming.txt:
Callback receives one piece of text per generated token.
Pieces received: 24, generated tokens reported: 24
Time to first piece: 5230 ms (includes processing the 26-token prompt)
Time to last piece: 7997 ms; whole call 8114 ms
First pieces: [[A], [ Java], [ record], [ is], [ a], [ data], [ structure], [ that]]
Joined: A Java record is a data structure that allows for the creation of immutable, named, and indexed collections of objects.
Nothing arrives for the first five seconds while the prompt is read, then pieces come steadily, roughly one every 120 milliseconds. The pieces are tokens, not words: the first ones are “A”, “ Java”, “ record” with their leading spaces. Joined, they equal the final answer.
generate also takes a maximum number of tokens. I assumed it limited the reply. It limits the prompt and the reply together. From 07-token-limit.txt:
The question is 26 prompt tokens long.
Limit 30: generated 3 tokens (prompt + generated = 29)
Limit 20: IllegalArgumentException: Prompt exceeds max tokens
With a limit of 30 and a 26-token prompt, the model produced only 3 tokens and stopped. With a limit of 20, below the prompt length, the call threw an exception instead of returning an empty answer. If you port code from an API where the limit means “new tokens”, set it to the prompt length plus what you want.
Temperature: repeatable at zero
Temperature controls how much randomness the model allows when it picks the next token. At 0 it always picks the most likely one. The test asks the same question at temperature 0 and at 0.9. Output, from 04-temperature.txt:Temperature 0.0, asked twice:
1: A Java record is a data structure that allows for the creation of immutable, named, and indexed collections of
2: A Java record is a data structure that allows for the creation of immutable, named, and indexed collections of
identical: true
Temperature 0.9, asked three times:
1: In Java, a record is a type of data structure that stores and retrieves data in a logical, indexed
2: A Java record is a type of nested class (or interface) that represents a collection of related data objects in
3: A Java record is a data structure that contains multiple values and provides a way to represent multiple instances of the same
distinct answers: 3
Temperature 0.0, limit 64, asked six more times: 1 distinct answer(s)
6x: A Java record is a data structure that allows for the creation of immutable, named, and indexed collections of objects.
At temperature 0 the two asks with a 48-token limit gave identical answers, and so did all six asks with a 64-token limit, which ended at the same 24 tokens. At 0.9 each of three asks gave a different sentence. If you want repeatable output for a test, use 0. I only checked repeatability inside one JVM on one machine; I did not check that a different CPU gives the same tokens.
The same model against Ollama
To put the numbers in context I ran the same kind of 4-bit TinyLlama through Ollama on the same machine, stopping Ollama while the Jlama tests ran so they did not compete for the two CPUs. The weights are not byte-identical: Jlama uses its own converted Q4 files, and Ollama uses the GGUF Q4_0 file from TheBloke. The baseline is a test too, OllamaBaselineTest.java, which only runs when you give it the address of an Ollama server. Output, from 06-ollama-baseline.txt:Ollama 0.40.3, model "tl" created from tinyllama-1.1b-chat-v1.0.Q4_0.gguf (TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF).
Same machine, same question, 1 warm-up run then 5 measured runs, temperature 0, limit 64 tokens.
run 1: prompt 36 tokens (16 served from cache) in 190 ms, generated 40 tokens in 1196 ms (33.4 tok/s)
run 2: prompt 36 tokens (35 served from cache) in 38 ms, generated 40 tokens in 1172 ms (34.1 tok/s)
run 3: prompt 36 tokens (35 served from cache) in 28 ms, generated 40 tokens in 1187 ms (33.7 tok/s)
run 4: prompt 36 tokens (35 served from cache) in 28 ms, generated 40 tokens in 1169 ms (34.2 tok/s)
run 5: prompt 36 tokens (35 served from cache) in 29 ms, generated 40 tokens in 1215 ms (32.9 tok/s)
Generation tokens/s: min 32.9, median 33.7, max 34.2
Answer of the last run: A Java record is a data structure that combines the properties of a class and a struct. It is a type of data structure that allows for the creation of immutable objects with specific properties.
| Measure (this 2-vCPU machine) | Jlama 0.8.4 | Ollama 0.40.3 |
|---|---|---|
| Generation, median of 5 runs | 8.0 tokens/s | 33.7 tokens/s |
| Reading the prompt, first measured run | 26 tokens in about 4.9 s | 36 tokens (16 from cache) in 190 ms |
| Setup | A Maven dependency and one JVM flag | A separate server and a model import |
Going deeper: how I set up the Ollama side
I downloaded the Ollama Linux release from GitHub, unpacked it, imported
tinyllama-1.1b-chat-v1.0.Q4_0.gguf with a Modelfile that sets the TinyLlama chat template and a stop token, and ran the server on a spare port with a single parallel slot. The exact steps are in the module’s README. Ollama’s answer to the same question differs from Jlama’s (“combines the properties of a class and a struct”), and the token counts differ (40 versus 24), because the two use different prompt templates and weights, so compare the rates, not the texts.
Should you even do this?
A fair answer. If you want the fastest local model, use a native runtime such as Ollama; on this machine it was about four times faster, and it handles bigger models. Jlama earns its place when “no native install” matters more than speed: an offline test that exercises your prompt plumbing without a server, a small tool shipped as one jar, a demo you can run with only a JDK, or a way to study how inference works in readable Java. The costs today are an incubator module flag, slow prompt reading, and a model small enough that its answers need checking. What this article did not test: a GPU; the jlama-native add-on; models other than TinyLlama; larger models, where memory and speed will both change; several requests at once; JDK versions other than 25; a laptop; and integration with Spring AI or LangChain4j. The tests show that the model loads and answers in a plain JVM and what the speed was on one small machine.
Further reading
- The code for this article: asmhatre/java-ai-agents, with the transcripts under
jlama/output/and the README. - Jlama on GitHub
- Run LLMs locally with Spring AI and Ollama — the native-runtime route.
- A2A protocol in Java and Embabel: goal-oriented AI agents on the JVM — other modules from the same repository.
No Comments yet!