Versions, and an honest limit. Spring Boot 4.1.1, Spring AI 2.0.1, Jackson 3 and Java 25. The code is themultimodalmodule of asmhatre/spring-ai, and every console block below is quoted from a file under itsoutput/directory. I did not run a vision model, and nothing here measures one. There are no API keys in my build environment, so the realOpenAiChatModel,AnthropicChatModelandOllamaChatModelclasses talk to a local server that stands in for all three services. That server is not a language model: it decodes the image it receives, runs thetesseractOCR program on the pixels and parses the text with a few regular expressions. So the requests are the real ones and the answers depend on the image, but every accuracy figure below describes OCR plus a parser, and says nothing about GPT, Claude or Gemini. Token figures are arithmetic from a published formula, not a bill. The receipts are ten I wrote by hand.
An image is just another part of the message
When you talk to a chat model through Spring AI, a user message is text plus, optionally, media: an object that holds some bytes (or a link) and a MIME type saying what kind of bytes they are. You attach it in the same call that carries your question. Spring AI then turns that into whatever shape the chosen provider’s API wants. You write the call once; the provider-specific packaging is not your job.ReceiptExtractor.java. The system prompt tells the model to extract only what is printed and to use null rather than guess; the user part is a sentence plus the image.
public Receipt extractOnce(byte[] image, MimeType type) {
return client.prompt().system(SYSTEM)
.user(u -> u.text("Extract this receipt.").media(type, new ByteArrayResource(image)))
.call().entity(Receipt.class);
}
To see what each provider actually receives I ran this one call through the three real Spring AI model classes against the local server, with the same 24,530-byte PNG (01-wire-formats.txt, from WireFormatTest.java). Long strings are shortened for reading; the images are not.
openai image as sent:
[{"text":"Extract this receipt. Your response shou... (1455 characters)","type":"text"},{"image_url":{"url":"data:image/png;base64,<32708 base64 characters>"},"type":"image_url"}]
mime label: image/png | bytes the server decoded: 24530 | sha256 matches: true
anthropic image as sent:
[{"text":"Extract this receipt. Your response shou... (1455 characters)","type":"text"},{"source":{"data":"<32708 base64 characters>","media_type":"image/png","type":"base64"},"type":"image"}]
mime label: image/png | bytes the server decoded: 24530 | sha256 matches: true
ollama image as sent:
{"role":"user","content":"Extract this receipt. Your response shou... (1455 characters)","images":["<32708 base64 characters>"]}
mime label: (none: Ollama sends bare base64) | bytes the server decoded: 24530 | sha256 matches: true
Three shapes, one image. OpenAI wants a data:image/png;base64,... URI inside an image_url part. Anthropic wants an image block with the base64 in a source. Ollama wants a bare base64 string in an images list on the message, with no MIME type at all. The test decodes what the server received and compares the SHA-256 hash with the original: all three match, so the bytes arrive unchanged.
Going deeper: size limits, formats, and the 33% cost of base64
Base64 writes every three bytes as four characters, so an image sent as bytes is about a third bigger on the wire than the file. In
07-cost.txt a 27,560-byte PNG becomes a 36,748-byte string. That is a request-size cost, separate from the token cost covered below. Each provider also has its own limits. When I read Anthropic’s vision page on 9 October 2026 it listed JPEG, PNG, GIF and WebP, a 10 MB limit for a base64-encoded image on the direct API, and 8000 x 8000 pixels as the maximum dimension. OpenAI’s image guide listed PNG, JPEG, WebP and non-animated GIF. These pages change, so check the current ones before you depend on a number.
- Spring AI reference: Multimodality — lists the models that accept images and the
MediaAPI. - Anthropic: vision and OpenAI: images and vision
Media a URL instead of bytes. What happens then differs by provider, and one of the three does something you would not guess (06-url-vs-bytes.txt, from UrlVsBytesTest.java):
openai request sent; image part carries: URL: https://example.invalid/receipts/2026-03-14.png
anthropic request sent; image part carries: URL: https://example.invalid/receipts/2026-03-14.png
ollama request sent; image part carries: NOT AN IMAGE: https://example.invalid/receipts/2026-03-14.png
OpenAI and Anthropic requests carry the URL itself, so the provider’s servers must be able to fetch it. Ollama’s request puts the URL string into the images list, where its API documents base64 bytes. I did not run a real Ollama, so I cannot say what it does with that string, but a local model server is not going to be downloading images for you on the strength of that field, so the safe habit is to send bytes to Ollama.
A URL is a promise someone else keeps. If you pass a link, the provider fetches it later from its own network: it must be reachable from there, it must still exist, and a private or expiring link will fail in a way that looks like a model problem. Sending the bytes costs bandwidth and removes that whole class of failure.
Ask for a record, not for text
A model that reads a receipt will happily answer in a paragraph. A paragraph is useless to a program. What you want is a Java object you can validate and store, and Spring AI will ask the model for one: you giveentity(...) a class, and it adds instructions to the prompt describing the JSON shape, then parses the reply into the class.
The class is a record, Receipt.java. Money is BigDecimal, not double, and that choice matters later: with doubles, 0.1 + 0.2 is not 0.3, and the arithmetic check in the next section would fail on correct receipts.
public record Receipt(String merchant, LocalDate date, String currency, List<Line> items,
BigDecimal subtotal, BigDecimal tax, BigDecimal total) {
public record Line(String description, int quantity, BigDecimal unitPrice, BigDecimal lineTotal) {
}
}
Here is the whole path on a clean image: what the stand-in backend read from the pixels, and the record Spring AI built (02-extraction.txt, from ExtractionTest.java):
what the stand-in backend read from the pixels (tesseract OCR):
| CORNER CAFE
|
| 2026-03-14
|
| CURRENCY: USD
|
| Flat White 2 x 4.50 9.00
| Croissant 1 x 3.75 3.75
| SUBTOTAL 12.75
|
| TAX 1.12
|
| TOTAL 13.87
the record Spring AI built from the model's JSON:
Receipt[merchant=CORNER CAFE, date=2026-03-14, currency=USD, items=[Line[description=Flat White, quantity=2, unitPrice=4.50, lineTotal=9.00], Line[description=Croissant, quantity=1, unitPrice=3.75, lineTotal=3.75]], subtotal=12.75, tax=1.12, total=13.87]
validation problems: none; attempts used: 1
The first block is the OCR text, including the blank lines OCR leaves between lines of the image. The second is the record. Notice how little the prompt had to say: the instructions that make the model return JSON of exactly this shape were generated from the record’s fields and appended by Spring AI. That appended text is long (the request shown in the retry section runs to 57 lines in all), and it is sent with every call, so it is a small recurring token cost for the convenience.
Going deeper: what entity() actually sends, and what I did not test
entity(Receipt.class) works by appending the JSON Schema of the class and a few instructions (return only JSON, no markdown fences) to the user text, then parsing the reply. That is why the request in 01-wire-formats.txt shows the user text as 1,455 characters for a one-sentence question. It relies on the model following format instructions; some providers also offer a native structured-output mode that constrains the reply on their side. I did not test any native mode here, and the local server returns JSON regardless, so this article says nothing about how often a real model returns malformed JSON.
- Structured output in Spring AI 2.0 — the converters behind
entity(). Receipt.java— the record, with the reason forBigDecimal.
Valid JSON is not a correct receipt
Once the reply parses into a record, it is tempting to stop. Do not. A model can misread a digit and still return perfectly valid JSON; so can an OCR engine. The parse succeeds, the types are right, and the total is wrong. The only way to notice is to use something the receipt itself promises: the numbers add up. Each line is quantity times unit price, the lines sum to the subtotal, and subtotal plus tax is the total.ReceiptValidator.java, with no model involved:
BigDecimal sum = BigDecimal.ZERO;
for (Receipt.Line l : r.items()) {
if (l.unitPrice() == null || l.lineTotal() == null) {
out.add("line '" + l.description() + "' has no price");
continue;
}
BigDecimal expected = l.unitPrice().multiply(BigDecimal.valueOf(l.quantity()));
if (expected.compareTo(l.lineTotal()) != 0) {
out.add("line '" + l.description() + "': " + l.quantity() + " x " + l.unitPrice() + " is "
+ expected + ", not " + l.lineTotal());
}
sum = sum.add(l.lineTotal());
}
if (r.subtotal() == null || sum.compareTo(r.subtotal()) != 0) {
out.add("lines add up to " + sum + ", subtotal says " + r.subtotal());
}
if (r.subtotal() != null && r.tax() != null && r.total() != null
&& r.subtotal().add(r.tax()).compareTo(r.total()) != 0) {
out.add("subtotal " + r.subtotal() + " + tax " + r.tax() + " is " + r.subtotal().add(r.tax())
+ ", total says " + r.total());
}
if (r.tax() == null || r.total() == null) {
I fed it a receipt corrupted in one specific way per rule, and two receipts with a typo (03-validator.txt, from ValidatorTest.java):
correct receipt -> PASSES validation
line total is not qty x unit price -> [line 'Flat White': 2 x 4.50 is 9.00, not 9.50]
a digit misread in one unit price (3.75 -> 3.15) -> [lines add up to 12.15, subtotal says 12.75]
total does not equal subtotal + tax -> [subtotal 12.75 + tax 1.12 is 13.87, total says 13.97]
currency the app does not support -> [currency US$ is not one of [EUR, GBP, INR, USD]]
merchant name typo (text, not arithmetic) -> PASSES validation
item name typo (text, not arithmetic) -> PASSES validation
The last two are wrong and pass: arithmetic cannot check spelling.
The first four are caught, each with a message that says exactly what disagrees. The misread digit (a unit price of 3.15 instead of 3.75) is caught not at that line but one step later, when the lines no longer add up to the subtotal, which is the reason to check the sums rather than only the shape. The last two lines are the limit of the method: a misspelled merchant and a misspelled item both pass, because arithmetic cannot check spelling.
Reconciliation catches numbers, not words. If the merchant name or an item description matters (for matching against a vendor list, say), validate it against something that knows the right answer, such as a table of known merchants. A passing check means “the arithmetic is consistent”, not “the receipt is right”.
Going deeper: choosing the rules
The rules above are the ones a receipt guarantees by construction. Others depend on your data: tax within a plausible percentage of the subtotal, a date not in the future, a total below some limit that routes large amounts to review, a currency from a list your system supports (the code uses four). Rounding is the usual trap: a receipt that rounds tax per line will not equal the sum of line taxes, so compare against what the printed total says, with a one-cent tolerance if your merchants round that way. The code here compares exactly because the test receipts were built to be exact.
ReceiptValidator.java— all the rules in one place.Fixtures.java— the ten ground-truth receipts the tests use.
Retry only helps if the answer can change
The obvious reaction to a failed check is to ask again.ReceiptExtractor.java does that once, and puts the problems into the second request so the model is told what was wrong:
public Result extract(byte[] image) {
Receipt first = extractOnce(image, MimeTypeUtils.IMAGE_PNG);
List<String> problems = ReceiptValidator.problems(first);
if (problems.isEmpty()) {
return new Result(first, problems, 1);
}
Receipt second = client.prompt().system(SYSTEM)
.user(u -> u.text("Extract this receipt. A previous attempt had these problems, so re-read the "
+ "image carefully: " + String.join("; ", problems))
.media(MimeTypeUtils.IMAGE_PNG, new ByteArrayResource(image)))
.call().entity(Receipt.class);
return new Result(second, ReceiptValidator.problems(second), 2);
}
I tested two cases. In the first, I told the stand-in to simulate a misread: on the first attempt only, it adds 10 cents to the total. This is invented, to exercise the loop; real models make errors of their own shapes. In the second I gave it an image it cannot read, with no simulation (04-repair-loop.txt and 04b-unreadable.txt, from RepairLoopTest.java):
SIMULATED misread (the stand-in adds 0.10 to the total on the first attempt only):
attempts: 2, problems after: none, total: 13.87
text sent with the second attempt:
| Extract this receipt. A previous attempt had these problems, so re-read the image carefully: subtotal 12.75 + tax 1.12 is 13.87, total says 13.97
| (+ 56 more lines: the system prompt and the JSON-schema format instructions Spring AI appends for entity())
noisy image, no simulation: attempts 2, requests made 2
problems left: [date missing, currency null is not one of [EUR, GBP, INR, USD], no line items]
The second attempt saw the same pixels and made the same mistake: a retry only helps when the error is not deterministic.
In the first case the second attempt succeeds, and you can see exactly what was added to the request: the sentence naming the problem. In the second case the retry changes nothing, because the second attempt sees the same pixels and makes the same mistake. That is the point worth taking away: a retry helps with errors that are random, and does nothing for errors caused by the input. It also costs a second request that carries the whole image again (the test checks that both requests hold the same image hash), so a failing image is paid for twice.
Decide what happens after the last retry before you ship. A receipt that still fails validation should go to a person, or be stored flagged as unverified, never quietly accepted. Money values that nobody checked are the worst kind of wrong, because they look exactly like right ones.
Going deeper: why I retry once, and what else could change between attempts
With a real model at a temperature above zero, a second attempt can differ from the first, which is what makes the retry useful. At temperature zero most providers are close to deterministic but not guaranteed identical, so a retry may or may not change anything. Other things that can change between attempts: a larger or more capable model for the second try, a higher-resolution copy of the image, or a prompt that asks for the lines first and the totals second. I did not try any of these. One retry keeps the worst-case cost at twice the image; unlimited retries on a bad photo is a good way to make a large bill.
ReceiptExtractor.java— the loop, with the system prompt.- Testing LLM apps in Java — measuring how a real model’s answers vary across runs.
How the photo changes the result
Real receipts are not clean scans. They are photographed at an angle, shrunk to save space, or smudged. To see what that does to the pipeline I drew each of the ten receipts four ways (clean, tilted four degrees, shrunk to a third of the size, and speckled with noise) and ran every image through the whole path, then compared the record to the receipt I started from (05-accuracy.txt, from AccuracyTest.java):
condition fields right exact receipts pass validation passed but WRONG
clean 156/156 (100%) 10/10 10/10 0/10
tilted 4 degrees 149/156 (96%) 5/10 5/10 0/10
small (1/3 size) 105/156 (67%) 0/10 0/10 0/10
noisy 2/156 (1%) 0/10 0/10 0/10
Going deeper: building a test set, and why synthetic images are only a start
The images come from
ReceiptImages.java, which draws a receipt as text on a white page and then tilts, shrinks or speckles it with a seeded random generator, so the same input always gives the same bytes. That makes the test repeatable, and also makes it easy: a rendered page has none of the creases, glare, shadows, curved paper or handwriting of a real photo. For a real project, collect a few dozen genuine receipts (with permission, and see the privacy note at the end), write down the correct values by hand once, and keep that set as your regression suite.
AccuracyTest.java— the field-by-field scoring, including how a missing or extra line item counts.FakeVisionServer.java— the stand-in backend, including the OCR call and the parser.
What an image costs, in tokens
Images are billed as input tokens, and the rule for how many varies by provider and by model, and has changed. I am only going to compute one, because it is the only one I could read an exact rule for. Anthropic’s vision documentation (as I read it on 9 October 2026) says an image is cut into 28 x 28 pixel patches and costs one token per patch, so width / 28 rounded up, times height / 28 rounded up, tokens, after the image is scaled down if it exceeds a size limit that depends on the model (a long edge of 1,568 pixels and 1,568 tokens for most models, 2,576 pixels and 4,784 tokens for the newest tier). The code applies that rule inImageTokens.java; its downscaling is a simple search, so it can differ from the service by a token or two. It is an estimate from a documented formula, not a measurement.
look pixels png bytes base64 bytes tokens std tokens hi-res
clean 640x522 27560 36748 437 437
small (1/3 size) 211x172 6892 9192 56 56
phone photo (given) 4032x3024 - - 1564 4740
1000x1000 (docs) 1000x1000 - - 1296 1296
(07-cost.txt, from CostTest.java.) The first two rows are images I generated and the last two are given sizes, not images I have. A 640 x 522 receipt is about 437 tokens. Shrinking it to a third drops that to 56, an eighth, which sounds like a bargain until the previous section: the same shrink took the OCR backend from a perfect read to no fully correct receipt. A 4032 x 3024 phone photo is scaled down by the service before it is counted, so on the standard tier it lands near the 1,568-token ceiling no matter how many megapixels the camera has.
The practical consequence is that you control cost mainly by controlling the image you send, and the trade is against how well the text can be read. A common middle path is to downscale a phone photo yourself to a size that keeps small text legible, rather than uploading the original and letting the provider decide. Whether a given size is legible to your model is something to measure on your own receipts, the way the last section measured it.
Going deeper: other providers, and pricing
I did not compute OpenAI’s cost. Its documentation lists several detail levels (
low, high, original and auto) with sizing rules that differ by model and have changed, and the page points to its own cost calculator; use that rather than a formula from a blog. I also did not run or price Google’s Gemini models through Spring AI, or any Ollama model, where the cost is your own hardware. To turn tokens into money, multiply by your provider’s current per-token input price; the article on observability shows how to read real token usage out of each response and compute cost per endpoint, which replaces any estimate once you have a real key.
ImageTokens.java— the patch rule and the two tiers.- Anthropic: vision — the rule and the current limits.
- OpenAI: images and vision
Should you even do this?
For receipts at volume, probably yes, but with a person behind the checks. A vision model needs no template per merchant, which is what usually makes template-based OCR pipelines costly to maintain. But a receipt is money, so keep the arithmetic validation, route anything that fails (or any total above a limit) to a human, and keep the original image so someone can look. If your documents are all one fixed form, a plain OCR engine with regular expressions may be cheaper, faster and fully offline, which is exactly what the stand-in server in this article is. Two things to think about before you send images to a hosted model: receipts contain names, card fragments and purchase history, so check your agreement with the provider and your own privacy rules first; and text inside an image is still text the model reads, so a receipt that says “ignore your instructions” is the same prompt-injection risk as a poisoned document (see the guardrails article, which tests text only, not images). What this article did not test: any real vision model, Gemini, provider-native structured output, PDFs, multi-page documents, handwriting, non-Latin scripts, streaming, and anything about latency.
Further reading
- The code for this article: the
multimodalmodule of asmhatre/spring-ai, with all the transcripts underoutput/and its README. - Structured output in Spring AI 2.0 — what
entity()does under the hood. - Run LLMs locally with Spring AI and Ollama — the local option for sending images without a hosted provider.
- Spring AI 2.0 ChatClient on Spring Boot 4.1 — the client used throughout.
- Observability for Spring AI — real token usage and cost per endpoint.
- Testing LLM apps in Java — scoring a noisy model against a golden set.
- Prompt injection defence in Spring AI
- Spring AI reference: Multimodality, Anthropic: vision, OpenAI: images and vision and Ollama API
No Comments yet!