Add multimodal module: receipt images to Java records on the real OpenAI, Anthropic and Ollama models against an OCR-backed local server, validation, repair retry, accuracy by photo condition

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
Claude
2026-10-09 07:07:57 +00:00
parent b0bba995e6
commit d67ac0630b
31 changed files with 1339 additions and 0 deletions
+48
View File
@@ -0,0 +1,48 @@
# multimodal
Companion code for [Multimodal Spring AI: Extract Structured Data from Images (Receipts to Java Records)](https://ankurm.com/multimodal-spring-ai-extract-structured-data-from-images-receipts-java-records/), part of the [Spring AI series](../README.md) on ankurm.com.
A receipt photo goes in through `Media`, a Java record comes out through `ChatClient.entity(...)`, and arithmetic checks decide whether to believe it.
**No vision model was used, and no claim is made about vision models.** The three real Spring AI chat models (`OpenAiChatModel`, `AnthropicChatModel`, `OllamaChatModel`) talk to a local server, `FakeVisionServer`. That server decodes the image it receives, runs the `tesseract` OCR program on the pixels and parses the text with a few regular expressions. So the request each model builds is the real one, and the answer depends on what is in the image, but every accuracy figure here describes OCR plus a parser, not GPT, Claude or Gemini. Token figures are arithmetic from a documented formula, not a bill. The receipts, the fixtures and the corruption in the retry test are all written for this module.
## Versions
| Component | Version |
|---|---|
| Spring Boot | 4.1.1 (parent) |
| Spring AI | 2.0.1 (`spring-ai-openai`, `spring-ai-anthropic`, `spring-ai-ollama`) |
| Jackson | 3 (`tools.jackson`) |
| tesseract | 5.3.4 (test only) |
| Java | 25 (LTS) |
## Quickstart
```bash
scripts/run-all.sh # runs the suite and regenerates output/01 .. 07
```
Needs `tesseract` on the PATH. Two consecutive runs produce byte-identical files.
## What's here
| File | What it is |
|---|---|
| [`Receipt.java`](src/main/java/com/ankurm/multimodal/Receipt.java) | The record the model must fill; money is `BigDecimal` |
| [`ReceiptExtractor.java`](src/main/java/com/ankurm/multimodal/ReceiptExtractor.java) | Image to record, with one repair attempt that quotes the problems |
| [`ReceiptValidator.java`](src/main/java/com/ankurm/multimodal/ReceiptValidator.java) | Lines times quantity, lines to subtotal, subtotal plus tax to total, currency |
| [`ReceiptImages.java`](src/main/java/com/ankurm/multimodal/ReceiptImages.java) | Draws a receipt as a deterministic PNG, clean, tilted, shrunk or noisy |
| [`ImageTokens.java`](src/main/java/com/ankurm/multimodal/ImageTokens.java) | Token estimate from Anthropic's documented 28-pixel patch rule |
| [`FakeVisionServer.java`](src/test/java/com/ankurm/multimodal/support/FakeVisionServer.java) | OpenAI, Anthropic and Ollama endpoints on one port, OCR behind them |
## Output files
| File | Written by |
|---|---|
| [`01-wire-formats.txt`](output/01-wire-formats.txt) | `WireFormatTest`: one image, three wire formats, bytes intact |
| [`02-extraction.txt`](output/02-extraction.txt) | `ExtractionTest`: what the backend read and the record that came out |
| [`03-validator.txt`](output/03-validator.txt) | `ValidatorTest`: each rule, and two errors arithmetic cannot see |
| [`04-repair-loop.txt`](output/04-repair-loop.txt), [`04b-unreadable.txt`](output/04b-unreadable.txt) | `RepairLoopTest`: a retry that fixes a simulated misread, and one that cannot |
| [`05-accuracy.txt`](output/05-accuracy.txt) | `AccuracyTest`: ten receipts under four photo conditions |
| [`06-url-vs-bytes.txt`](output/06-url-vs-bytes.txt) | `UrlVsBytesTest`: `Media` built from a URI |
| [`07-cost.txt`](output/07-cost.txt) | `CostTest`: bytes on the wire and estimated tokens |