Add rag module: Spring AI 2.0 RAG with pgvector, chunking, reranking and a faithfulness check
Co-Authored-By: Claude Sonnet 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01B38FGKKam5SCGgwgduVAh3
This commit is contained in:
@@ -0,0 +1,29 @@
|
||||
# LLM reranking: calls, order, latency and failure
|
||||
|
||||
|
||||
--- one model call per candidate; only topN survive (topN = 2) ---
|
||||
candidates in: 4, model calls made: 4, chunks out: 2
|
||||
order from the vector search, best first:
|
||||
similarity 0.5891 page 2
|
||||
similarity 0.4077 page 3
|
||||
similarity 0.2535 page 1
|
||||
similarity 0.2023 page 4
|
||||
order after reranking, best first:
|
||||
rerank_score 8 page 2
|
||||
rerank_score 6 page 3
|
||||
the rating prompt for the page 2 candidate (calls run concurrently, so pick it by content):
|
||||
Rate how well the PASSAGE helps answer the QUESTION, from 0 (irrelevant) to 10 (answers it).
|
||||
Reply with a single integer and nothing else.
|
||||
QUESTION: How many days of annual leave do employees get?
|
||||
PASSAGE: 4.1 Annual Leave Entitlement. Full-time employees are entitled to 20 working days
|
||||
|
||||
--- latency: 20 candidates, each rating call takes 200 ms (a Thread.sleep in the fake model) ---
|
||||
the 20 calls one after another take 4000 ms or more: true
|
||||
LlmReranker, one virtual thread per candidate, takes under 1000 ms: true
|
||||
a real API adds its own rate limits, which this test cannot show
|
||||
|
||||
--- failure: the model does not reply with a bare integer ---
|
||||
reply "Score: 8" for every candidate -> failures counted: 4 of 4
|
||||
scores assigned: [0, 0]
|
||||
pages kept, in order: [2, 3] (the vector-search order, because every score is 0)
|
||||
reply " 9\n" (padded) -> score 9
|
||||
Reference in New Issue
Block a user