Files
spring-ai/evaluation/output/04-noisy-judge.txt
T

18 lines
995 B
Plaintext

# Simulated judge noise: 200 runs of a 12-case suite per build
noise-free pass rate: healthy build 1.00, regressed build 0.75
== 3% of judge verdicts flipped (mean pass rate: healthy 0.947, regressed 0.706) ==
gate healthy build fails it regressed build fails it
every case must pass 96 of 200 runs 200 of 200 runs
pass rate >= 0.90 29 of 200 runs 200 of 200 runs
pass rate >= 0.85 29 of 200 runs 200 of 200 runs
pass rate >= 0.80 2 of 200 runs 200 of 200 runs
== 8% of judge verdicts flipped (mean pass rate: healthy 0.859, regressed 0.640) ==
gate healthy build fails it regressed build fails it
every case must pass 166 of 200 runs 200 of 200 runs
pass rate >= 0.90 102 of 200 runs 200 of 200 runs
pass rate >= 0.85 102 of 200 runs 200 of 200 runs
pass rate >= 0.80 47 of 200 runs 200 of 200 runs