Add evaluation module: RelevancyEvaluator and FactCheckingEvaluator, golden dataset with a pass-rate gate, deterministic CI judge and simulated judge noise
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,42 @@
|
||||
# What RelevancyEvaluator and FactCheckingEvaluator send to the judge
|
||||
|
||||
--- RelevancyEvaluator: prompt sent to the judge (1 message(s)) ---
|
||||
Your task is to evaluate if the response for the query
|
||||
is in line with the context information provided.
|
||||
|
||||
You have two options to answer. Either YES or NO.
|
||||
|
||||
Answer YES, if the response for the query
|
||||
is in line with context information otherwise NO.
|
||||
|
||||
Query:
|
||||
How long do I have to request a refund on an annual plan?
|
||||
|
||||
Response:
|
||||
You can request a refund on an annual plan within 30 days of purchase.
|
||||
|
||||
Context:
|
||||
Annual plans can be refunded within 30 days of purchase.
|
||||
Monthly plans are not refundable.
|
||||
|
||||
Answer:
|
||||
|
||||
--- RelevancyEvaluator: judge said "yes" ---
|
||||
pass=true score=1.0 feedback='' metadata={}
|
||||
|
||||
--- FactCheckingEvaluator: prompt sent to the judge ---
|
||||
Evaluate whether or not the following claim is supported by the provided document.
|
||||
Respond with "yes" if the claim is supported, or "no" if it is not.
|
||||
|
||||
Document:
|
||||
Annual plans can be refunded within 30 days of purchase.
|
||||
Monthly plans are not refundable.
|
||||
|
||||
Claim:
|
||||
You can request a refund on an annual plan within 30 days of purchase.
|
||||
|
||||
--- FactCheckingEvaluator: judge said "yes" ---
|
||||
pass=true score=0.0 feedback='' metadata={}
|
||||
|
||||
--- RelevancyEvaluator: judge said "no" ---
|
||||
pass=false score=0.0 feedback=''
|
||||
@@ -0,0 +1,13 @@
|
||||
# Which judge replies the built-in evaluators accept
|
||||
|
||||
judge reply relevancy fact-chk wrapped
|
||||
'yes' true true true
|
||||
'YES' true true true
|
||||
' yes\n' true true true
|
||||
'Yes.' false false true
|
||||
'yes!' false false true
|
||||
'Yes, the response is in line with the context.' false false true
|
||||
'**Yes**' false false true
|
||||
'no' false false false
|
||||
'No.' false false false
|
||||
'' false false false
|
||||
@@ -0,0 +1,37 @@
|
||||
# Golden dataset: 12 cases, deterministic judge, pass-rate gate at 0.90
|
||||
|
||||
== healthy build ==
|
||||
case relevant grounded has facts passed
|
||||
refund-annual true true true true
|
||||
refund-monthly true true true true
|
||||
storage-pro true true true true
|
||||
support-hours true true true true
|
||||
sso-plan true true true true
|
||||
data-region true true true true
|
||||
api-rate-limit true true true true
|
||||
backup-retention true true true true
|
||||
macos-agent true true true true
|
||||
reset-link true true true true
|
||||
extra-seat true true true true
|
||||
trial-card true true true true
|
||||
pass rate: 1.00 failing: []
|
||||
|
||||
== regressed build ==
|
||||
case relevant grounded has facts passed
|
||||
refund-annual false false false false
|
||||
refund-monthly true true true true
|
||||
storage-pro true true true true
|
||||
support-hours true true true true
|
||||
sso-plan false false false false
|
||||
data-region true true true true
|
||||
api-rate-limit true true true true
|
||||
backup-retention true true true true
|
||||
macos-agent true true true true
|
||||
reset-link true true true true
|
||||
extra-seat true true false false
|
||||
trial-card true true true true
|
||||
pass rate: 0.75 failing: [refund-annual, sso-plan, extra-seat]
|
||||
|
||||
== the gate ==
|
||||
healthy build : requirePassRate(0.90) returned normally
|
||||
regressed build : java.lang.AssertionError: pass rate 0.75 is below the 0.90 threshold; failing cases: [refund-annual, sso-plan, extra-seat]
|
||||
@@ -0,0 +1,17 @@
|
||||
# Simulated judge noise: 200 runs of a 12-case suite per build
|
||||
|
||||
noise-free pass rate: healthy build 1.00, regressed build 0.75
|
||||
|
||||
== 3% of judge verdicts flipped (mean pass rate: healthy 0.947, regressed 0.706) ==
|
||||
gate healthy build fails it regressed build fails it
|
||||
every case must pass 96 of 200 runs 200 of 200 runs
|
||||
pass rate >= 0.90 29 of 200 runs 200 of 200 runs
|
||||
pass rate >= 0.85 29 of 200 runs 200 of 200 runs
|
||||
pass rate >= 0.80 2 of 200 runs 200 of 200 runs
|
||||
|
||||
== 8% of judge verdicts flipped (mean pass rate: healthy 0.859, regressed 0.640) ==
|
||||
gate healthy build fails it regressed build fails it
|
||||
every case must pass 166 of 200 runs 200 of 200 runs
|
||||
pass rate >= 0.90 102 of 200 runs 200 of 200 runs
|
||||
pass rate >= 0.85 102 of 200 runs 200 of 200 runs
|
||||
pass rate >= 0.80 47 of 200 runs 200 of 200 runs
|
||||
@@ -0,0 +1,16 @@
|
||||
# GradedEvaluator (minimum grade 4) and CompositeEvaluator
|
||||
|
||||
judge reply pass score feedback
|
||||
'5' true 1.0 grade 5 of 5
|
||||
'4/5' true 0.8 grade 4 of 5
|
||||
'Score: 4 - mostly right, misses the date' true 0.8 grade 4 of 5
|
||||
'3' false 0.6 grade 3 of 5
|
||||
'I would say excellent' false 0.0 judge did not return a grade: I would say excellent
|
||||
'10' false 0.0 judge did not return a grade: 10
|
||||
|
||||
== CompositeEvaluator on the regressed 'extra-seat' answer ("12 USD" instead of "8 USD") ==
|
||||
answer : Extra seats cost 12 USD per month.
|
||||
pass : false
|
||||
score : 0.5
|
||||
feedback : contains-facts failed (missing: [8 USD]);
|
||||
verdicts : {relevancy=true, contains-facts=false}
|
||||
@@ -0,0 +1,8 @@
|
||||
# RuleBasedJudge(0.6) on hand-labelled answers to: Can I get a refund on a monthly plan?
|
||||
|
||||
answer type correct? judge says verdict
|
||||
correct, same words true pass right
|
||||
WRONG: negation dropped false pass WRONG
|
||||
WRONG: roles swapped false pass WRONG
|
||||
correct, paraphrased true fail WRONG
|
||||
correct, but a refusal true fail WRONG
|
||||
@@ -0,0 +1,7 @@
|
||||
# Model calls per testing layer (stub app model, 12 golden cases)
|
||||
|
||||
layer app calls judge calls network
|
||||
1 unit: prompt and wiring, no evaluator 1 0 none
|
||||
2 golden run, RuleBasedJudge (pull request) 12 24 none
|
||||
3 golden run, real judge (nightly) n/a n/a needs OPENAI_API_KEY; skipped here
|
||||
pass rate at layer 2: 1.00
|
||||
@@ -0,0 +1,4 @@
|
||||
# LiveJudgeTest as recorded by surefire when OPENAI_API_KEY is not set
|
||||
|
||||
testcase: "realJudgeAgreesWithTheKnownGoodAndKnownBadBuilds"
|
||||
skipped : "Environment variable [OPENAI_API_KEY] does not exist"
|
||||
Reference in New Issue
Block a user