Versions, and honest limits. ONNX Runtimecom.microsoft.onnxruntime:onnxruntime1.31.0, DJL 0.38.0 (api,huggingface:tokenizersandonnxruntime-engine), Spring Boot 4.1.1, Java 25, JUnit 6.1.3, and the modeldistilbert/distilbert-base-uncased-finetuned-sst-2-englishfrom Hugging Face (its ONNX export, 268 MB). All the code is in theonnx-djlmodule of asmhatre/java-ai-agents, and every console block below is quoted from a file underonnx-djl/output/, written by a test. The latency numbers come from a cloud sandbox with 2 virtual CPUs, not a laptop or a production server. I tested only this one classification model, only on the CPU, and only with English text. I did not test a GPU, embedding models, quantised models, or a long-running load.
What the model expects
An ONNX file describes a computation graph, and ONNX Runtime can tell you what that graph takes and returns. The first test asks. It also tokenizes a sentence with DJL’s Hugging Face tokenizer, because a model like this does not read text. It reads integers. From 01-model-and-tokenizer.txt:Model: distilbert-base-uncased-finetuned-sst-2-english, ONNX export, 267 MB
ONNX Runtime 1.31.0
Inputs of the ONNX graph:
input_ids TensorInfo(javaType=INT64,onnxType=ONNX_TENSOR_ELEMENT_DATA_TYPE_INT64,shape=[-1, -1],dimNames=[batch_size,sequence_length])
attention_mask TensorInfo(javaType=INT64,onnxType=ONNX_TENSOR_ELEMENT_DATA_TYPE_INT64,shape=[-1, -1],dimNames=[batch_size,sequence_length])
Outputs:
logits TensorInfo(javaType=FLOAT,onnxType=ONNX_TENSOR_ELEMENT_DATA_TYPE_FLOAT,shape=[-1, 2],dimNames=[batch_size,""])
Tokenizing "I loved this film.":
tokens: [[CLS], i, loved, this, film, ., [SEP]]
input_ids: [101, 1045, 3866, 2023, 2143, 1012, 102]
attention_mask: [1, 1, 1, 1, 1, 1, 1]
Two texts in one batch are padded to the same length:
[[CLS], i, loved, this, film, ., [SEP]]
mask [1, 1, 1, 1, 1, 1, 1]
[[CLS], bad, ., [SEP], [PAD], [PAD], [PAD]]
mask [1, 1, 1, 1, 0, 0, 0]
The graph takes two integer arrays, input_ids and attention_mask, both of shape [batch_size, sequence_length], and returns logits with two numbers per text, one per class. Tokenizing “I loved this film.” produced seven ids: a start marker [CLS], the lower-cased word pieces, and an end marker [SEP]. The model is “uncased”, so capital letters are dropped before it ever sees them.
The last two lines are the reason for the attention mask. When two texts of different lengths go in one batch, the shorter is padded with [PAD] tokens so the arrays are rectangular, and the mask marks the padding with zeros so the model ignores it. You will see next that this matters for correctness, not just shape.
Going deeper: why the tokenizer is a separate dependency
The ONNX file contains the neural network only. The text-to-ids step lives in
tokenizer.json, which ships next to the model, and something has to run it. DJL’s ai.djl.huggingface:tokenizers loads that file and runs the Hugging Face tokenizer through a native library that it unpacks from its own jar into ~/.djl.ai on first use. You can use that tokenizer without using DJL for the model, which is what the first approach below does. I configured it to pad to the longest text in the batch and to cut texts at 128 tokens; the cut-off is my choice, not the model’s, which accepts up to 512.
- Deep Java Library documentation
- OnnxSentiment.java — the tokenizer settings are in the
tokenizermethod.
Running it with ONNX Runtime directly
ONNX Runtime has its own Java API. You make a session from the model file, wrap each input array in a tensor, run, and read the output. The class below does that for a whole batch at once, and turns the two logits into a label and a probability with a softmax. From OnnxSentiment.java: public List<Sentiment> classify(List<String> texts) {
Encoding[] encodings = tokenizer.batchEncode(texts.toArray(new String[0]));
int rows = encodings.length;
int cols = encodings[0].getIds().length;
long[][] ids = new long[rows][];
long[][] mask = new long[rows][];
for (int i = 0; i < rows; i++) {
ids[i] = encodings[i].getIds();
mask[i] = encodings[i].getAttentionMask();
}
try (OnnxTensor idTensor = OnnxTensor.createTensor(env, ids);
OnnxTensor maskTensor = OnnxTensor.createTensor(env, mask);
OrtSession.Result result = session.run(Map.of("input_ids", idTensor, "attention_mask", maskTensor))) {
float[][] logits = (float[][]) result.get(0).getValue();
return java.util.stream.Stream.of(logits).map(OnnxSentiment::softmaxWinner).toList();
} catch (OrtException e) {
throw new IllegalStateException(e);
}
}
static Sentiment softmaxWinner(float[] logits) {
double max = Math.max(logits[0], logits[1]);
double e0 = Math.exp(logits[0] - max);
double e1 = Math.exp(logits[1] - max);
int best = e1 > e0 ? 1 : 0;
return new Sentiment(LABELS[best], Math.max(e0, e1) / (e0 + e1));
}
Six sentences go through it in one batch. The result is in 02-classification.txt:
ONNX Runtime directly, DJL tokenizer, one batch of 6 texts.
POSITIVE 1.00 I loved this film, the acting was wonderful.
NEGATIVE 1.00 The service was slow and the food arrived cold.
POSITIVE 1.00 It works.
POSITIVE 1.00 Not bad at all.
NEGATIVE 1.00 I wanted to love it, but it was a mess.
NEGATIVE 0.95 The build finished in four minutes.
The clear cases are right and confident. The last line is the useful warning. “The build finished in four minutes.” is a neutral sentence, but this model only knows two labels, so it was pushed to NEGATIVE with 0.95 probability. A softmax always sums to one over the classes it has; a high score means “more like this class than the other”, not “I am sure”. If your texts include neutral content, a two-class model needs a threshold or a third class.
The same model through DJL
DJL offers a different way to run the same file. You describe a translator that converts your input type to arrays and the output arrays back, and ask for the ONNX Runtime engine. DJL then loads the model and gives you a predictor to run it. From DjlSentiment.java: Criteria<String[], Sentiment[]> criteria = Criteria.builder()
.setTypes(String[].class, Sentiment[].class)
.optModelPath(modelDir.resolve("onnx"))
.optModelName("model")
.optEngine("OnnxRuntime")
.optTranslator(new BatchTranslator(tokenizer))
.build();
this.model = criteria.loadModel();
this.predictor = model.newPredictor();
}
The translator is the nested BatchTranslator: it tokenizes in processInput, names the two arrays input_ids and attention_mask, and in processOutput runs the same softmax. A test classifies the same six sentences both ways. Result, from 03-djl-engine.txt:
DJL engine: OnnxRuntime 1.21.1
direct POSITIVE 0.9999 | DJL POSITIVE 0.9999 | difference 0.00e+00
direct NEGATIVE 0.9998 | DJL NEGATIVE 0.9998 | difference 0.00e+00
direct POSITIVE 0.9999 | DJL POSITIVE 0.9999 | difference 0.00e+00
direct POSITIVE 0.9993 | DJL POSITIVE 0.9993 | difference 0.00e+00
direct NEGATIVE 0.9996 | DJL NEGATIVE 0.9996 | difference 0.00e+00
direct NEGATIVE 0.9491 | DJL NEGATIVE 0.9491 | difference 0.00e+00
Same labels: true, largest score difference: 0.00e+00
Same labels and the same scores to four decimals; the largest difference was exactly zero. That is expected, since both end up in the same ONNX Runtime native library, but it is good to see it rather than assume it. The first line shows something to be careful about: DJL prints its engine version as 1.21.1, the ONNX Runtime version it was built against, while the library actually in use is 1.31.0, as the first transcript showed. I only know that from asking ONNX Runtime directly for its version.
Going deeper: the version override, and a startup-order trap
DJL 0.38.0’s engine declares ONNX Runtime 1.21.1. My POM declares 1.31.0 itself, and Maven picks the version declared closest to the project, so 1.31.0 won; the verbose dependency tree in 08-dependencies.txt says “omitted for conflict with 1.31.0”. It worked for this model and these calls. DJL’s authors did not necessarily test against 1.31.0, so treat the override as something to verify for your model, not a given.
The second trap cost me a failing build. ONNX Runtime has one shared environment object per JVM. DJL’s engine creates it with its own thread-pool settings, and if your own code called
OrtEnvironment.getEnvironment() first, DJL fails to start. A test starts both in one JVM in each order, in 05-engine-order.txt:
The same two things started in one JVM, in both orders (a separate JVM for each).
direct-first (ONNX Runtime created by my code, then DJL):
Caused by: java.lang.IllegalStateException: Tried to specify the thread pool when creating an OrtEnvironment, but one already exists.
djl-first (ONNX Runtime created by DJL, then my code):
both started
If you use both in one application, start DJL’s engine first, or use only one of the two. In this module the Spring Boot endpoint uses only the direct API, and the Maven Surefire plugin runs each test class in its own JVM (reuseForks=false) so the two kinds of test do not interfere.
Does batching change the answers, and does it help?
Two questions about batching, both answered by one test. First: does padding change the result? The test classifies 32 texts one at a time and again as a single batch of 32 and compares. Second: does a bigger batch cost less per text? It times the same 32 texts at different batch sizes. Result, from 04-batching.txt:Padding does not change the answers: 32 texts one by one versus one batch of 32.
same labels: true, largest score difference 0.00e+00
Time to classify 32 texts, median of 5 runs after a warm-up, by batch size:
batch size 1: 230 ms in total, 7.2 ms per text
batch size 4: 204 ms in total, 6.4 ms per text
batch size 8: 166 ms in total, 5.2 ms per text
batch size 16: 142 ms in total, 4.4 ms per text
batch size 32: 134 ms in total, 4.2 ms per text
Padding changes nothing: the labels match and the largest score difference is zero, which is the attention mask doing its job. Batching also helps, moderately: classifying 32 texts took 230 ms one at a time and 134 ms as a single batch, which is 7.2 against 4.2 milliseconds per text. The gain is real but not dramatic on two CPU cores, and the middle sizes are not perfectly smooth, since each figure is the median of only five runs.
Going deeper: what the batch timing leaves out
Each figure is the median of five timed runs after a warm-up, on a shared 2-vCPU machine. Texts in a batch are padded to the longest, so a batch that mixes a three-word text with a hundred-word text does the work of the longer one for both; my sentences are all short and similar, so the test does not show that cost. I did not measure how latency changes with longer texts, or how many threads ONNX Runtime used.
Behind a Spring Boot endpoint
The endpoint takes a list of texts and returns a label and score for each. The model is a single bean: an ONNX Runtime session can be called from several threads at once, so one instance serves every request. The batch size is capped at 64 so one request cannot occupy the model for long. From SentimentApplication.java: @Bean(destroyMethod = "close")
OnnxSentiment onnxSentiment(@Value("${sentiment.model-dir}") String modelDir) throws IOException, OrtException {
return new OnnxSentiment(Path.of(modelDir));
}
record Request(List<String> texts) {
}
record Response(List<Sentiment> results) {
}
@RestController
static class SentimentController {
private static final int MAX_BATCH = 64;
private final OnnxSentiment model;
SentimentController(OnnxSentiment model) {
this.model = model;
}
@PostMapping("/sentiment")
Response classify(@RequestBody Request request) {
if (request.texts() == null || request.texts().isEmpty() || request.texts().size() > MAX_BATCH) {
throw new ResponseStatusException(HttpStatus.BAD_REQUEST, "send between 1 and " + MAX_BATCH + " texts");
}
return new Response(model.classify(request.texts()));
}
A test starts the real application on a random port and calls it over HTTP. The transcript is 06-endpoint.txt:
POST /sentiment
{"texts":["I loved this film, the acting was wonderful.","The service was slow and the food arrived cold."]}
HTTP 200
{
"results" : [ {
"label" : "POSITIVE",
"score" : 0.9998835686297906
}, {
"label" : "NEGATIVE",
"score" : 0.9997554818864473
} ]
}
Empty list: HTTP 400
65 texts (limit is 64): HTTP 400
Spring Boot 4 serialises the records with Jackson 3 without any configuration. Empty and oversized batches are rejected with a 400.
Latency over HTTP
Last, the numbers a caller would see. The client and the server share one JVM and one machine, so network time is absent and the figures include JSON handling, Tomcat and the model. The test measures requests of different sizes, then 100 single-text requests one after another and from four clients at once. From 07-http-latency.txt:Round trip over HTTP on the same machine (client and server in one JVM), median of 20 requests.
1 text(s) per request: 17.6 ms per request, 17.6 ms per text
4 text(s) per request: 31.7 ms per request, 7.9 ms per text
16 text(s) per request: 80.0 ms per request, 5.0 ms per text
64 text(s) per request: 297.2 ms per request, 4.6 ms per text
100 requests of 1 text: one client after another 1.7 s (59 requests/s), 4 clients at once 1.4 s (70 requests/s)
A single text costs about 17 milliseconds end to end, and the model call itself was about 7 of those in the direct measurement above, so HTTP and JSON add about as much as the model here. Sending 16 or 64 texts per request brings the cost per text down to about 5 milliseconds, which agrees with the batching test. Four clients in parallel raised throughput only from 59 to 70 requests per second on two cores, which suggests the model call already keeps both cores busy; I did not check CPU use.
Going deeper: what this measures and what it does not
The medians are of 20 requests each, from a client in the same JVM as the server, on a shared 2-vCPU machine, so treat them as a rough order of magnitude. There is no network, no TLS, no request queueing beyond what four threads create, and no sustained load, so I cannot tell you the tail latency or what happens at a hundred concurrent requests. If you need those, run a load test on the hardware you will deploy to.
Should you even do this?
A fair answer. For one classifier or embedder inside a Java service, running ONNX in-process is simple: a few dependencies, no extra server, and single-digit milliseconds per text on a CPU, as measured here. It stops being simple when you need many different models, a GPU, or model updates without a redeploy; then a dedicated model server is worth its operational cost. Calling a hosted API is easier still if sending your text to a third party is acceptable. The real costs here were an operational one (a 268 MB model file to distribute), a version question (DJL’s engine against a newer ONNX Runtime), and a startup-order trap. What this article did not test: a GPU; embedding models; quantised variants, which are smaller and faster but may change the scores; languages other than English; texts longer than 128 tokens; the other tokenizer settings; sustained or concurrent load beyond four clients; and DJL’s own model zoo or its PyTorch engine. The tests show that one model gives identical answers by two routes, that batching does not change them, and what the timing looked like on one small machine.
Further reading
- The code for this article: asmhatre/java-ai-agents, with the transcripts under
onnx-djl/output/and the README. - ONNX Runtime for Java
- Deep Java Library documentation
- Local LLM inference in pure Java with Jlama — the same idea for a language model, with an honest speed comparison.
- Run LLMs locally with Spring AI and Ollama
No Comments yet!