Files

3.4 KiB

multimodal

Companion code for Multimodal Spring AI: Extract Structured Data from Images (Receipts to Java Records), part of the Spring AI series on ankurm.com.

A receipt photo goes in through Media, a Java record comes out through ChatClient.entity(...), and arithmetic checks decide whether to believe it.

No vision model was used, and no claim is made about vision models. The three real Spring AI chat models (OpenAiChatModel, AnthropicChatModel, OllamaChatModel) talk to a local server, FakeVisionServer. That server decodes the image it receives, runs the tesseract OCR program on the pixels and parses the text with a few regular expressions. So the request each model builds is the real one, and the answer depends on what is in the image, but every accuracy figure here describes OCR plus a parser, not GPT, Claude or Gemini. Token figures are arithmetic from a documented formula, not a bill. The receipts, the fixtures and the corruption in the retry test are all written for this module.

Versions

Component Version
Spring Boot 4.1.1 (parent)
Spring AI 2.0.1 (spring-ai-openai, spring-ai-anthropic, spring-ai-ollama)
Jackson 3 (tools.jackson)
tesseract 5.3.4 (test only)
Java 25 (LTS)

Quickstart

scripts/run-all.sh     # runs the suite and regenerates output/01 .. 07

Needs tesseract on the PATH. Two consecutive runs produce byte-identical files.

What's here

File What it is
Receipt.java The record the model must fill; money is BigDecimal
ReceiptExtractor.java Image to record, with one repair attempt that quotes the problems
ReceiptValidator.java Lines times quantity, lines to subtotal, subtotal plus tax to total, currency
ReceiptImages.java Draws a receipt as a deterministic PNG, clean, tilted, shrunk or noisy
ImageTokens.java Token estimate from Anthropic's documented 28-pixel patch rule
FakeVisionServer.java OpenAI, Anthropic and Ollama endpoints on one port, OCR behind them

Output files

File Written by
01-wire-formats.txt WireFormatTest: one image, three wire formats, bytes intact
02-extraction.txt ExtractionTest: what the backend read and the record that came out
03-validator.txt ValidatorTest: each rule, and two errors arithmetic cannot see
04-repair-loop.txt, 04b-unreadable.txt RepairLoopTest: a retry that fixes a simulated misread, and one that cannot
05-accuracy.txt AccuracyTest: ten receipts under four photo conditions
06-url-vs-bytes.txt UrlVsBytesTest: Media built from a URI
07-cost.txt CostTest: bytes on the wire and estimated tokens