Skip to main content

Multimodal Spring AI: Extract Structured Data from Images (Receipts to Java Records)

You have a shoebox of receipts, or a phone full of photos of them, and what you want is rows: merchant, date, line items, total. Typing them in is the job nobody wants. A model that can look at an image can do it, and Spring AI lets you hand it the image almost as easily as you hand it a string. The easy part is sending the picture. The part that decides whether you can trust the result is everything after: getting a typed Java record back instead of a paragraph, noticing when the numbers do not add up, and knowing what happens when the photo is tilted, small or grainy. This article builds that, step by step, and measures it. The depth is in expandable sections, so you can read straight through or open only what you need.
Versions, and an honest limit. Spring Boot 4.1.1, Spring AI 2.0.1, Jackson 3 and Java 25. The code is the multimodal module of asmhatre/spring-ai, and every console block below is quoted from a file under its output/ directory. I did not run a vision model, and nothing here measures one. There are no API keys in my build environment, so the real OpenAiChatModel, AnthropicChatModel and OllamaChatModel classes talk to a local server that stands in for all three services. That server is not a language model: it decodes the image it receives, runs the tesseract OCR program on the pixels and parses the text with a few regular expressions. So the requests are the real ones and the answers depend on the image, but every accuracy figure below describes OCR plus a parser, and says nothing about GPT, Claude or Gemini. Token figures are arithmetic from a published formula, not a bill. The receipts are ten I wrote by hand.

An image is just another part of the message

When you talk to a chat model through Spring AI, a user message is text plus, optionally, media: an object that holds some bytes (or a link) and a MIME type saying what kind of bytes they are. You attach it in the same call that carries your question. Spring AI then turns that into whatever shape the chosen provider’s API wants. You write the call once; the provider-specific packaging is not your job.
Receipt imagePNG bytesMediaMIME type + bytesOpenAIcontent part: image_url, data:image/png;base64,…Anthropiccontent block: image, source type base64Ollamamessage field: images, bare base64 strings
One image goes in on the left and three different requests come out on the right. They carry the same pixels but package them differently, which is why you never build these requests by hand. The code that does the attaching is short. It is the whole of the “send an image” side of this article: This is ReceiptExtractor.java. The system prompt tells the model to extract only what is printed and to use null rather than guess; the user part is a sentence plus the image.
    public Receipt extractOnce(byte[] image, MimeType type) {
        return client.prompt().system(SYSTEM)
                .user(u -> u.text("Extract this receipt.").media(type, new ByteArrayResource(image)))
                .call().entity(Receipt.class);
    }
To see what each provider actually receives I ran this one call through the three real Spring AI model classes against the local server, with the same 24,530-byte PNG (01-wire-formats.txt, from WireFormatTest.java). Long strings are shortened for reading; the images are not.
openai    image as sent:
          [{"text":"Extract this receipt. Your response shou... (1455 characters)","type":"text"},{"image_url":{"url":"data:image/png;base64,<32708 base64 characters>"},"type":"image_url"}]
          mime label: image/png | bytes the server decoded: 24530 | sha256 matches: true
anthropic image as sent:
          [{"text":"Extract this receipt. Your response shou... (1455 characters)","type":"text"},{"source":{"data":"<32708 base64 characters>","media_type":"image/png","type":"base64"},"type":"image"}]
          mime label: image/png | bytes the server decoded: 24530 | sha256 matches: true
ollama    image as sent:
          {"role":"user","content":"Extract this receipt. Your response shou... (1455 characters)","images":["<32708 base64 characters>"]}
          mime label: (none: Ollama sends bare base64) | bytes the server decoded: 24530 | sha256 matches: true
Three shapes, one image. OpenAI wants a data:image/png;base64,... URI inside an image_url part. Anthropic wants an image block with the base64 in a source. Ollama wants a bare base64 string in an images list on the message, with no MIME type at all. The test decodes what the server received and compares the SHA-256 hash with the original: all three match, so the bytes arrive unchanged.
Going deeper: size limits, formats, and the 33% cost of base64
Base64 writes every three bytes as four characters, so an image sent as bytes is about a third bigger on the wire than the file. In 07-cost.txt a 27,560-byte PNG becomes a 36,748-byte string. That is a request-size cost, separate from the token cost covered below. Each provider also has its own limits. When I read Anthropic’s vision page on 9 October 2026 it listed JPEG, PNG, GIF and WebP, a 10 MB limit for a base64-encoded image on the direct API, and 8000 x 8000 pixels as the maximum dimension. OpenAI’s image guide listed PNG, JPEG, WebP and non-animated GIF. These pages change, so check the current ones before you depend on a number.
You can also give Media a URL instead of bytes. What happens then differs by provider, and one of the three does something you would not guess (06-url-vs-bytes.txt, from UrlVsBytesTest.java):
openai    request sent; image part carries: URL: https://example.invalid/receipts/2026-03-14.png
anthropic request sent; image part carries: URL: https://example.invalid/receipts/2026-03-14.png
ollama    request sent; image part carries: NOT AN IMAGE: https://example.invalid/receipts/2026-03-14.png
OpenAI and Anthropic requests carry the URL itself, so the provider’s servers must be able to fetch it. Ollama’s request puts the URL string into the images list, where its API documents base64 bytes. I did not run a real Ollama, so I cannot say what it does with that string, but a local model server is not going to be downloading images for you on the strength of that field, so the safe habit is to send bytes to Ollama.
A URL is a promise someone else keeps. If you pass a link, the provider fetches it later from its own network: it must be reachable from there, it must still exist, and a private or expiring link will fail in a way that looks like a model problem. Sending the bytes costs bandwidth and removes that whole class of failure.

Ask for a record, not for text

A model that reads a receipt will happily answer in a paragraph. A paragraph is useless to a program. What you want is a Java object you can validate and store, and Spring AI will ask the model for one: you give entity(...) a class, and it adds instructions to the prompt describing the JSON shape, then parses the reply into the class. The class is a record, Receipt.java. Money is BigDecimal, not double, and that choice matters later: with doubles, 0.1 + 0.2 is not 0.3, and the arithmetic check in the next section would fail on correct receipts.
public record Receipt(String merchant, LocalDate date, String currency, List<Line> items,
        BigDecimal subtotal, BigDecimal tax, BigDecimal total) {

    public record Line(String description, int quantity, BigDecimal unitPrice, BigDecimal lineTotal) {
    }
}
Here is the whole path on a clean image: what the stand-in backend read from the pixels, and the record Spring AI built (02-extraction.txt, from ExtractionTest.java):
what the stand-in backend read from the pixels (tesseract OCR):
    | CORNER CAFE
    | 
    | 2026-03-14
    | 
    | CURRENCY: USD
    | 
    | Flat White 2 x 4.50 9.00
    | Croissant 1 x 3.75 3.75
    | SUBTOTAL 12.75
    | 
    | TAX 1.12
    | 
    | TOTAL 13.87

the record Spring AI built from the model's JSON:
    Receipt[merchant=CORNER CAFE, date=2026-03-14, currency=USD, items=[Line[description=Flat White, quantity=2, unitPrice=4.50, lineTotal=9.00], Line[description=Croissant, quantity=1, unitPrice=3.75, lineTotal=3.75]], subtotal=12.75, tax=1.12, total=13.87]
validation problems: none; attempts used: 1
The first block is the OCR text, including the blank lines OCR leaves between lines of the image. The second is the record. Notice how little the prompt had to say: the instructions that make the model return JSON of exactly this shape were generated from the record’s fields and appended by Spring AI. That appended text is long (the request shown in the retry section runs to 57 lines in all), and it is sent with every call, so it is a small recurring token cost for the convenience.
Going deeper: what entity() actually sends, and what I did not test
entity(Receipt.class) works by appending the JSON Schema of the class and a few instructions (return only JSON, no markdown fences) to the user text, then parsing the reply. That is why the request in 01-wire-formats.txt shows the user text as 1,455 characters for a one-sentence question. It relies on the model following format instructions; some providers also offer a native structured-output mode that constrains the reply on their side. I did not test any native mode here, and the local server returns JSON regardless, so this article says nothing about how often a real model returns malformed JSON.

Valid JSON is not a correct receipt

Once the reply parses into a record, it is tempting to stop. Do not. A model can misread a digit and still return perfectly valid JSON; so can an OCR engine. The parse succeeds, the types are right, and the total is wrong. The only way to notice is to use something the receipt itself promises: the numbers add up. Each line is quantity times unit price, the lines sum to the subtotal, and subtotal plus tax is the total.
Imagephoto or scanModelreturns a recordValidatordo the sums agree?Acceptstore the recordRetry once, then a humanquote the problems back
The figure shows where the check sits: between the model and the database, and it decides which of two paths a record takes. The check is plain Java in ReceiptValidator.java, with no model involved:
        BigDecimal sum = BigDecimal.ZERO;
        for (Receipt.Line l : r.items()) {
            if (l.unitPrice() == null || l.lineTotal() == null) {
                out.add("line '" + l.description() + "' has no price");
                continue;
            }
            BigDecimal expected = l.unitPrice().multiply(BigDecimal.valueOf(l.quantity()));
            if (expected.compareTo(l.lineTotal()) != 0) {
                out.add("line '" + l.description() + "': " + l.quantity() + " x " + l.unitPrice() + " is "
                        + expected + ", not " + l.lineTotal());
            }
            sum = sum.add(l.lineTotal());
        }
        if (r.subtotal() == null || sum.compareTo(r.subtotal()) != 0) {
            out.add("lines add up to " + sum + ", subtotal says " + r.subtotal());
        }
        if (r.subtotal() != null && r.tax() != null && r.total() != null
                && r.subtotal().add(r.tax()).compareTo(r.total()) != 0) {
            out.add("subtotal " + r.subtotal() + " + tax " + r.tax() + " is " + r.subtotal().add(r.tax())
                    + ", total says " + r.total());
        }
        if (r.tax() == null || r.total() == null) {
I fed it a receipt corrupted in one specific way per rule, and two receipts with a typo (03-validator.txt, from ValidatorTest.java):
correct receipt                                      -> PASSES validation
line total is not qty x unit price                   -> [line 'Flat White': 2 x 4.50 is 9.00, not 9.50]
a digit misread in one unit price (3.75 -> 3.15)     -> [lines add up to 12.15, subtotal says 12.75]
total does not equal subtotal + tax                  -> [subtotal 12.75 + tax 1.12 is 13.87, total says 13.97]
currency the app does not support                    -> [currency US$ is not one of [EUR, GBP, INR, USD]]
merchant name typo (text, not arithmetic)            -> PASSES validation
item name typo (text, not arithmetic)                -> PASSES validation

The last two are wrong and pass: arithmetic cannot check spelling.
The first four are caught, each with a message that says exactly what disagrees. The misread digit (a unit price of 3.15 instead of 3.75) is caught not at that line but one step later, when the lines no longer add up to the subtotal, which is the reason to check the sums rather than only the shape. The last two lines are the limit of the method: a misspelled merchant and a misspelled item both pass, because arithmetic cannot check spelling.
Reconciliation catches numbers, not words. If the merchant name or an item description matters (for matching against a vendor list, say), validate it against something that knows the right answer, such as a table of known merchants. A passing check means “the arithmetic is consistent”, not “the receipt is right”.
Going deeper: choosing the rules
The rules above are the ones a receipt guarantees by construction. Others depend on your data: tax within a plausible percentage of the subtotal, a date not in the future, a total below some limit that routes large amounts to review, a currency from a list your system supports (the code uses four). Rounding is the usual trap: a receipt that rounds tax per line will not equal the sum of line taxes, so compare against what the printed total says, with a one-cent tolerance if your merchants round that way. The code here compares exactly because the test receipts were built to be exact.

Retry only helps if the answer can change

The obvious reaction to a failed check is to ask again. ReceiptExtractor.java does that once, and puts the problems into the second request so the model is told what was wrong:
    public Result extract(byte[] image) {
        Receipt first = extractOnce(image, MimeTypeUtils.IMAGE_PNG);
        List<String> problems = ReceiptValidator.problems(first);
        if (problems.isEmpty()) {
            return new Result(first, problems, 1);
        }
        Receipt second = client.prompt().system(SYSTEM)
                .user(u -> u.text("Extract this receipt. A previous attempt had these problems, so re-read the "
                        + "image carefully: " + String.join("; ", problems))
                        .media(MimeTypeUtils.IMAGE_PNG, new ByteArrayResource(image)))
                .call().entity(Receipt.class);
        return new Result(second, ReceiptValidator.problems(second), 2);
    }
I tested two cases. In the first, I told the stand-in to simulate a misread: on the first attempt only, it adds 10 cents to the total. This is invented, to exercise the loop; real models make errors of their own shapes. In the second I gave it an image it cannot read, with no simulation (04-repair-loop.txt and 04b-unreadable.txt, from RepairLoopTest.java):
SIMULATED misread (the stand-in adds 0.10 to the total on the first attempt only):
  attempts: 2, problems after: none, total: 13.87
  text sent with the second attempt:
    | Extract this receipt. A previous attempt had these problems, so re-read the image carefully: subtotal 12.75 + tax 1.12 is 13.87, total says 13.97
    | (+ 56 more lines: the system prompt and the JSON-schema format instructions Spring AI appends for entity())
noisy image, no simulation: attempts 2, requests made 2
problems left: [date missing, currency null is not one of [EUR, GBP, INR, USD], no line items]
The second attempt saw the same pixels and made the same mistake: a retry only helps when the error is not deterministic.
In the first case the second attempt succeeds, and you can see exactly what was added to the request: the sentence naming the problem. In the second case the retry changes nothing, because the second attempt sees the same pixels and makes the same mistake. That is the point worth taking away: a retry helps with errors that are random, and does nothing for errors caused by the input. It also costs a second request that carries the whole image again (the test checks that both requests hold the same image hash), so a failing image is paid for twice.
Decide what happens after the last retry before you ship. A receipt that still fails validation should go to a person, or be stored flagged as unverified, never quietly accepted. Money values that nobody checked are the worst kind of wrong, because they look exactly like right ones.
Going deeper: why I retry once, and what else could change between attempts
With a real model at a temperature above zero, a second attempt can differ from the first, which is what makes the retry useful. At temperature zero most providers are close to deterministic but not guaranteed identical, so a retry may or may not change anything. Other things that can change between attempts: a larger or more capable model for the second try, a higher-resolution copy of the image, or a prompt that asks for the lines first and the totals second. I did not try any of these. One retry keeps the worst-case cost at twice the image; unlimited retries on a bad photo is a good way to make a large bill.

How the photo changes the result

Real receipts are not clean scans. They are photographed at an angle, shrunk to save space, or smudged. To see what that does to the pipeline I drew each of the ten receipts four ways (clean, tilted four degrees, shrunk to a third of the size, and speckled with noise) and ran every image through the whole path, then compared the record to the receipt I started from (05-accuracy.txt, from AccuracyTest.java):
condition              fields right exact receipts  pass validation     passed but WRONG
clean                156/156 (100%)          10/10            10/10                 0/10
tilted 4 degrees      149/156 (96%)           5/10             5/10                 0/10
small (1/3 size)      105/156 (67%)           0/10             0/10                 0/10
noisy                    2/156 (1%)           0/10             0/10                 0/10
clean100%tilted 4 deg96%small (1/3)67%noisy1%fields read correctly, 156 fields per condition (OCR backend, not an LLM)
Read the table first. A “field” is one value: the merchant, the date, a quantity, a unit price. The ten receipts hold 156 of them. Clean images are read perfectly. A four-degree tilt already costs half the receipts (5 of 10 come out exact), shrinking to a third of the size leaves two thirds of the fields right and no receipt fully right, and heavy noise leaves almost nothing. The bars show the same thing as a share of fields. Two cautions about how to read this. First, the numbers belong to OCR and a regular-expression parser, not to a vision model; a vision model will behave differently from tesseract under every one of these conditions, and I have no measurement of how. What carries over is the method: build a set of receipts with known contents, degrade them in the ways your users will, and score field by field. Second, the last column, “passed but WRONG”, is zero in every row. That is a result about this backend, not reassurance about the method: every receipt it got wrong also failed validation. The validation section showed receipts that are wrong and pass. Measure that column on your own model and your own data.
Going deeper: building a test set, and why synthetic images are only a start
The images come from ReceiptImages.java, which draws a receipt as text on a white page and then tilts, shrinks or speckles it with a seeded random generator, so the same input always gives the same bytes. That makes the test repeatable, and also makes it easy: a rendered page has none of the creases, glare, shadows, curved paper or handwriting of a real photo. For a real project, collect a few dozen genuine receipts (with permission, and see the privacy note at the end), write down the correct values by hand once, and keep that set as your regression suite.

What an image costs, in tokens

Images are billed as input tokens, and the rule for how many varies by provider and by model, and has changed. I am only going to compute one, because it is the only one I could read an exact rule for. Anthropic’s vision documentation (as I read it on 9 October 2026) says an image is cut into 28 x 28 pixel patches and costs one token per patch, so width / 28 rounded up, times height / 28 rounded up, tokens, after the image is scaled down if it exceeds a size limit that depends on the model (a long edge of 1,568 pixels and 1,568 tokens for most models, 2,576 pixels and 4,784 tokens for the newest tier). The code applies that rule in ImageTokens.java; its downscaling is a simple search, so it can differ from the service by a token or two. It is an estimate from a documented formula, not a measurement.
look                    pixels  png bytes base64 bytes tokens std tokens hi-res
clean                  640x522      27560        36748        437          437
small (1/3 size)       211x172       6892         9192         56           56

phone photo (given)  4032x3024          -            -       1564         4740
1000x1000 (docs)     1000x1000          -            -       1296         1296
(07-cost.txt, from CostTest.java.) The first two rows are images I generated and the last two are given sizes, not images I have. A 640 x 522 receipt is about 437 tokens. Shrinking it to a third drops that to 56, an eighth, which sounds like a bargain until the previous section: the same shrink took the OCR backend from a perfect read to no fully correct receipt. A 4032 x 3024 phone photo is scaled down by the service before it is counted, so on the standard tier it lands near the 1,568-token ceiling no matter how many megapixels the camera has. The practical consequence is that you control cost mainly by controlling the image you send, and the trade is against how well the text can be read. A common middle path is to downscale a phone photo yourself to a size that keeps small text legible, rather than uploading the original and letting the provider decide. Whether a given size is legible to your model is something to measure on your own receipts, the way the last section measured it.
Going deeper: other providers, and pricing
I did not compute OpenAI’s cost. Its documentation lists several detail levels (low, high, original and auto) with sizing rules that differ by model and have changed, and the page points to its own cost calculator; use that rather than a formula from a blog. I also did not run or price Google’s Gemini models through Spring AI, or any Ollama model, where the cost is your own hardware. To turn tokens into money, multiply by your provider’s current per-token input price; the article on observability shows how to read real token usage out of each response and compute cost per endpoint, which replaces any estimate once you have a real key.

Should you even do this?

For receipts at volume, probably yes, but with a person behind the checks. A vision model needs no template per merchant, which is what usually makes template-based OCR pipelines costly to maintain. But a receipt is money, so keep the arithmetic validation, route anything that fails (or any total above a limit) to a human, and keep the original image so someone can look. If your documents are all one fixed form, a plain OCR engine with regular expressions may be cheaper, faster and fully offline, which is exactly what the stand-in server in this article is. Two things to think about before you send images to a hosted model: receipts contain names, card fragments and purchase history, so check your agreement with the provider and your own privacy rules first; and text inside an image is still text the model reads, so a receipt that says “ignore your instructions” is the same prompt-injection risk as a poisoned document (see the guardrails article, which tests text only, not images). What this article did not test: any real vision model, Gemini, provider-native structured output, PDFs, multi-page documents, handwriting, non-Latin scripts, streaming, and anything about latency.

Further reading

No Comments yet!

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.