Add evaluation module: RelevancyEvaluator and FactCheckingEvaluator, golden dataset with a pass-rate gate, deterministic CI judge and simulated judge noise

Co-Authored-By: Claude Sonnet 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
Claude
2026-10-09 06:15:41 +00:00
parent 80db4eb8ac
commit 65f580c5bc
37 changed files with 1353 additions and 0 deletions
@@ -0,0 +1,42 @@
# What RelevancyEvaluator and FactCheckingEvaluator send to the judge
--- RelevancyEvaluator: prompt sent to the judge (1 message(s)) ---
Your task is to evaluate if the response for the query
is in line with the context information provided.
You have two options to answer. Either YES or NO.
Answer YES, if the response for the query
is in line with context information otherwise NO.
Query:
How long do I have to request a refund on an annual plan?
Response:
You can request a refund on an annual plan within 30 days of purchase.
Context:
Annual plans can be refunded within 30 days of purchase.
Monthly plans are not refundable.
Answer:
--- RelevancyEvaluator: judge said "yes" ---
pass=true score=1.0 feedback='' metadata={}
--- FactCheckingEvaluator: prompt sent to the judge ---
Evaluate whether or not the following claim is supported by the provided document.
Respond with "yes" if the claim is supported, or "no" if it is not.
Document:
Annual plans can be refunded within 30 days of purchase.
Monthly plans are not refundable.
Claim:
You can request a refund on an annual plan within 30 days of purchase.
--- FactCheckingEvaluator: judge said "yes" ---
pass=true score=0.0 feedback='' metadata={}
--- RelevancyEvaluator: judge said "no" ---
pass=false score=0.0 feedback=''
+13
View File
@@ -0,0 +1,13 @@
# Which judge replies the built-in evaluators accept
judge reply relevancy fact-chk wrapped
'yes' true true true
'YES' true true true
' yes\n' true true true
'Yes.' false false true
'yes!' false false true
'Yes, the response is in line with the context.' false false true
'**Yes**' false false true
'no' false false false
'No.' false false false
'' false false false
+37
View File
@@ -0,0 +1,37 @@
# Golden dataset: 12 cases, deterministic judge, pass-rate gate at 0.90
== healthy build ==
case relevant grounded has facts passed
refund-annual true true true true
refund-monthly true true true true
storage-pro true true true true
support-hours true true true true
sso-plan true true true true
data-region true true true true
api-rate-limit true true true true
backup-retention true true true true
macos-agent true true true true
reset-link true true true true
extra-seat true true true true
trial-card true true true true
pass rate: 1.00 failing: []
== regressed build ==
case relevant grounded has facts passed
refund-annual false false false false
refund-monthly true true true true
storage-pro true true true true
support-hours true true true true
sso-plan false false false false
data-region true true true true
api-rate-limit true true true true
backup-retention true true true true
macos-agent true true true true
reset-link true true true true
extra-seat true true false false
trial-card true true true true
pass rate: 0.75 failing: [refund-annual, sso-plan, extra-seat]
== the gate ==
healthy build : requirePassRate(0.90) returned normally
regressed build : java.lang.AssertionError: pass rate 0.75 is below the 0.90 threshold; failing cases: [refund-annual, sso-plan, extra-seat]
+17
View File
@@ -0,0 +1,17 @@
# Simulated judge noise: 200 runs of a 12-case suite per build
noise-free pass rate: healthy build 1.00, regressed build 0.75
== 3% of judge verdicts flipped (mean pass rate: healthy 0.947, regressed 0.706) ==
gate healthy build fails it regressed build fails it
every case must pass 96 of 200 runs 200 of 200 runs
pass rate >= 0.90 29 of 200 runs 200 of 200 runs
pass rate >= 0.85 29 of 200 runs 200 of 200 runs
pass rate >= 0.80 2 of 200 runs 200 of 200 runs
== 8% of judge verdicts flipped (mean pass rate: healthy 0.859, regressed 0.640) ==
gate healthy build fails it regressed build fails it
every case must pass 166 of 200 runs 200 of 200 runs
pass rate >= 0.90 102 of 200 runs 200 of 200 runs
pass rate >= 0.85 102 of 200 runs 200 of 200 runs
pass rate >= 0.80 47 of 200 runs 200 of 200 runs
@@ -0,0 +1,16 @@
# GradedEvaluator (minimum grade 4) and CompositeEvaluator
judge reply pass score feedback
'5' true 1.0 grade 5 of 5
'4/5' true 0.8 grade 4 of 5
'Score: 4 - mostly right, misses the date' true 0.8 grade 4 of 5
'3' false 0.6 grade 3 of 5
'I would say excellent' false 0.0 judge did not return a grade: I would say excellent
'10' false 0.0 judge did not return a grade: 10
== CompositeEvaluator on the regressed 'extra-seat' answer ("12 USD" instead of "8 USD") ==
answer : Extra seats cost 12 USD per month.
pass : false
score : 0.5
feedback : contains-facts failed (missing: [8 USD]);
verdicts : {relevancy=true, contains-facts=false}
+8
View File
@@ -0,0 +1,8 @@
# RuleBasedJudge(0.6) on hand-labelled answers to: Can I get a refund on a monthly plan?
answer type correct? judge says verdict
correct, same words true pass right
WRONG: negation dropped false pass WRONG
WRONG: roles swapped false pass WRONG
correct, paraphrased true fail WRONG
correct, but a refusal true fail WRONG
+7
View File
@@ -0,0 +1,7 @@
# Model calls per testing layer (stub app model, 12 golden cases)
layer app calls judge calls network
1 unit: prompt and wiring, no evaluator 1 0 none
2 golden run, RuleBasedJudge (pull request) 12 24 none
3 golden run, real judge (nightly) n/a n/a needs OPENAI_API_KEY; skipped here
pass rate at layer 2: 1.00
@@ -0,0 +1,4 @@
# LiveJudgeTest as recorded by surefire when OPENAI_API_KEY is not set
testcase: "realJudgeAgreesWithTheKnownGoodAndKnownBadBuilds"
skipped : "Environment variable [OPENAI_API_KEY] does not exist"