Files
spring-ai/rag/docs/output/08-reranking.txt
T

30 lines
1.3 KiB
Plaintext

# LLM reranking: calls, order, latency and failure
--- one model call per candidate; only topN survive (topN = 2) ---
candidates in: 4, model calls made: 4, chunks out: 2
order from the vector search, best first:
similarity 0.5891 page 2
similarity 0.4077 page 3
similarity 0.2535 page 1
similarity 0.2023 page 4
order after reranking, best first:
rerank_score 8 page 2
rerank_score 6 page 3
the rating prompt for the page 2 candidate (calls run concurrently, so pick it by content):
Rate how well the PASSAGE helps answer the QUESTION, from 0 (irrelevant) to 10 (answers it).
Reply with a single integer and nothing else.
QUESTION: How many days of annual leave do employees get?
PASSAGE: 4.1 Annual Leave Entitlement. Full-time employees are entitled to 20 working days
--- latency: 20 candidates, each rating call takes 200 ms (a Thread.sleep in the fake model) ---
the 20 calls one after another take 4000 ms or more: true
LlmReranker, one virtual thread per candidate, takes under 1000 ms: true
a real API adds its own rate limits, which this test cannot show
--- failure: the model does not reply with a bare integer ---
reply "Score: 8" for every candidate -> failures counted: 4 of 4
scores assigned: [0, 0]
pages kept, in order: [2, 3] (the vector-search order, because every score is 0)
reply " 9\n" (padded) -> score 9