Add evaluation module: RelevancyEvaluator and FactCheckingEvaluator, golden dataset with a pass-rate gate, deterministic CI judge and simulated judge noise
Co-Authored-By: Claude Sonnet 5.5 <[email protected]> Claude-Session: https://claude.ai/code/session_01JXVi2GMQ7bR5EmbUFdDj7N
This commit is contained in:
@@ -0,0 +1,7 @@
|
||||
# Model calls per testing layer (stub app model, 12 golden cases)
|
||||
|
||||
layer app calls judge calls network
|
||||
1 unit: prompt and wiring, no evaluator 1 0 none
|
||||
2 golden run, RuleBasedJudge (pull request) 12 24 none
|
||||
3 golden run, real judge (nightly) n/a n/a needs OPENAI_API_KEY; skipped here
|
||||
pass rate at layer 2: 1.00
|
||||
Reference in New Issue
Block a user