Add rag module: Spring AI 2.0 RAG with pgvector, chunking, reranking and a faithfulness check

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01B38FGKKam5SCGgwgduVAh3
This commit is contained in:
Claude
2026-09-21 19:09:05 +00:00
commit 1d4625a1c2
62 changed files with 3715 additions and 0 deletions
+29
View File
@@ -0,0 +1,29 @@
# LLM reranking: calls, order, latency and failure
--- one model call per candidate; only topN survive (topN = 2) ---
candidates in: 4, model calls made: 4, chunks out: 2
order from the vector search, best first:
similarity 0.5891 page 2
similarity 0.4077 page 3
similarity 0.2535 page 1
similarity 0.2023 page 4
order after reranking, best first:
rerank_score 8 page 2
rerank_score 6 page 3
the rating prompt for the page 2 candidate (calls run concurrently, so pick it by content):
Rate how well the PASSAGE helps answer the QUESTION, from 0 (irrelevant) to 10 (answers it).
Reply with a single integer and nothing else.
QUESTION: How many days of annual leave do employees get?
PASSAGE: 4.1 Annual Leave Entitlement. Full-time employees are entitled to 20 working days
--- latency: 20 candidates, each rating call takes 200 ms (a Thread.sleep in the fake model) ---
the 20 calls one after another take 4000 ms or more: true
LlmReranker, one virtual thread per candidate, takes under 1000 ms: true
a real API adds its own rate limits, which this test cannot show
--- failure: the model does not reply with a bare integer ---
reply "Score: 8" for every candidate -> failures counted: 4 of 4
scores assigned: [0, 0]
pages kept, in order: [2, 3] (the vector-search order, because every score is 0)
reply " 9\n" (padded) -> score 9