Add multimodal module: receipt images to Java records on the real OpenAI, Anthropic and Ollama models against an OCR-backed local server, validation, repair retry, accuracy by photo condition
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,13 @@
|
||||
# The same receipt image through three real Spring AI chat models (24530 bytes, PNG)
|
||||
|
||||
openai image as sent:
|
||||
[{"text":"Extract this receipt. Your response shou... (1455 characters)","type":"text"},{"image_url":{"url":"data:image/png;base64,<32708 base64 characters>"},"type":"image_url"}]
|
||||
mime label: image/png | bytes the server decoded: 24530 | sha256 matches: true
|
||||
anthropic image as sent:
|
||||
[{"text":"Extract this receipt. Your response shou... (1455 characters)","type":"text"},{"source":{"data":"<32708 base64 characters>","media_type":"image/png","type":"base64"},"type":"image"}]
|
||||
mime label: image/png | bytes the server decoded: 24530 | sha256 matches: true
|
||||
ollama image as sent:
|
||||
{"role":"user","content":"Extract this receipt. Your response shou... (1455 characters)","images":["<32708 base64 characters>"]}
|
||||
mime label: (none: Ollama sends bare base64) | bytes the server decoded: 24530 | sha256 matches: true
|
||||
|
||||
distinct image hashes seen by the server across the three providers: 1
|
||||
@@ -0,0 +1,20 @@
|
||||
# A clean receipt through ChatClient.entity(Receipt.class)
|
||||
|
||||
what the stand-in backend read from the pixels (tesseract OCR):
|
||||
| CORNER CAFE
|
||||
|
|
||||
| 2026-03-14
|
||||
|
|
||||
| CURRENCY: USD
|
||||
|
|
||||
| Flat White 2 x 4.50 9.00
|
||||
| Croissant 1 x 3.75 3.75
|
||||
| SUBTOTAL 12.75
|
||||
|
|
||||
| TAX 1.12
|
||||
|
|
||||
| TOTAL 13.87
|
||||
|
||||
the record Spring AI built from the model's JSON:
|
||||
Receipt[merchant=CORNER CAFE, date=2026-03-14, currency=USD, items=[Line[description=Flat White, quantity=2, unitPrice=4.50, lineTotal=9.00], Line[description=Croissant, quantity=1, unitPrice=3.75, lineTotal=3.75]], subtotal=12.75, tax=1.12, total=13.87]
|
||||
validation problems: none; attempts used: 1
|
||||
@@ -0,0 +1,11 @@
|
||||
# What the arithmetic checks catch, and what they cannot
|
||||
|
||||
correct receipt -> PASSES validation
|
||||
line total is not qty x unit price -> [line 'Flat White': 2 x 4.50 is 9.00, not 9.50]
|
||||
a digit misread in one unit price (3.75 -> 3.15) -> [lines add up to 12.15, subtotal says 12.75]
|
||||
total does not equal subtotal + tax -> [subtotal 12.75 + tax 1.12 is 13.87, total says 13.97]
|
||||
currency the app does not support -> [currency US$ is not one of [EUR, GBP, INR, USD]]
|
||||
merchant name typo (text, not arithmetic) -> PASSES validation
|
||||
item name typo (text, not arithmetic) -> PASSES validation
|
||||
|
||||
The last two are wrong and pass: arithmetic cannot check spelling.
|
||||
@@ -0,0 +1,7 @@
|
||||
# Validate, then ask once more with the problems quoted
|
||||
|
||||
SIMULATED misread (the stand-in adds 0.10 to the total on the first attempt only):
|
||||
attempts: 2, problems after: none, total: 13.87
|
||||
text sent with the second attempt:
|
||||
| Extract this receipt. A previous attempt had these problems, so re-read the image carefully: subtotal 12.75 + tax 1.12 is 13.87, total says 13.97
|
||||
| (+ 56 more lines: the system prompt and the JSON-schema format instructions Spring AI appends for entity())
|
||||
@@ -0,0 +1,5 @@
|
||||
# Retrying an image the backend cannot read
|
||||
|
||||
noisy image, no simulation: attempts 2, requests made 2
|
||||
problems left: [date missing, currency null is not one of [EUR, GBP, INR, USD], no line items]
|
||||
The second attempt saw the same pixels and made the same mistake: a retry only helps when the error is not deterministic.
|
||||
@@ -0,0 +1,9 @@
|
||||
# Ten receipts x four photo conditions through the real pipeline (OCR backend, NOT an LLM)
|
||||
|
||||
condition fields right exact receipts pass validation passed but WRONG
|
||||
clean 156/156 (100%) 10/10 10/10 0/10
|
||||
tilted 4 degrees 149/156 (96%) 5/10 5/10 0/10
|
||||
small (1/3 size) 105/156 (67%) 0/10 0/10 0/10
|
||||
noisy 2/156 (1%) 0/10 0/10 0/10
|
||||
|
||||
'passed but WRONG' = the arithmetic checks were satisfied and the record still differs from the receipt.
|
||||
@@ -0,0 +1,7 @@
|
||||
# Media given as a URI instead of bytes
|
||||
|
||||
openai request sent; image part carries: URL: https://example.invalid/receipts/2026-03-14.png
|
||||
anthropic request sent; image part carries: URL: https://example.invalid/receipts/2026-03-14.png
|
||||
ollama request sent; image part carries: NOT AN IMAGE: https://example.invalid/receipts/2026-03-14.png
|
||||
|
||||
(example.invalid cannot resolve, so a provider that fetches the URL itself would fail here; the fake server never fetches.)
|
||||
@@ -0,0 +1,8 @@
|
||||
# What an image costs: bytes now, tokens by documented formula
|
||||
|
||||
look pixels png bytes base64 bytes tokens std tokens hi-res
|
||||
clean 640x522 27560 36748 437 437
|
||||
small (1/3 size) 211x172 6892 9192 56 56
|
||||
|
||||
phone photo (given) 4032x3024 - - 1564 4740
|
||||
1000x1000 (docs) 1000x1000 - - 1296 1296
|
||||
Reference in New Issue
Block a user