Class: Raif::CLI::EvalsCompare
- Defined in:
- lib/raif/cli/evals_compare.rb
Constant Summary collapse
- FORMATS =
["text", "json", "html"].freeze
Instance Attribute Summary
Attributes inherited from Base
Instance Method Summary collapse
-
#run ⇒ Object
Rails is deliberately not loaded: this reads two JSON files and does arithmetic, so booting an app would make the one part of the eval tooling that needs no database or API key the slowest to run.
Methods inherited from Base
Constructor Details
This class inherits a constructor from Raif::CLI::Base
Instance Method Details
#run ⇒ Object
Rails is deliberately not loaded: this reads two JSON files and does arithmetic, so booting an app would make the one part of the eval tooling that needs no database or API key the slowest to run. Required up front rather than after parsing so the --significance help text can name the real default instead of a copy that could drift.
16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 |
# File 'lib/raif/cli/evals_compare.rb', line 16 def run require_relative "../evals/statistics" require_relative "../evals/comparison" require_relative "../evals/comparison_report" require_relative "../utils/colors" format = "text" threshold = nil allow_judge_mismatch = false alpha = nil max_error_rate = Raif::Evals::Comparison::MAX_GATE_ERROR_RATE parser = OptionParser.new do |opts| opts. = "Usage: raif evals:compare BASELINE_RESULTS.json CANDIDATE_RESULTS.json [options]" opts.on("--fail-on-regression N", Float, "Exit non-zero when a pass rate or gated score gets more than N worse than baseline, as a " \ "fraction of it (0.25 = 25% worse)") do |n| threshold = n end opts.on("--significance ALPHA", Float, "Family-wise significance level a regression must clear as well as the threshold " \ "(default: #{Raif::Evals::Comparison::FAMILY_WISE_ALPHA}). 1 waives it and gates on effect size alone.") do |value| alpha = value end opts.on("--format FORMAT", FORMATS, "Output format: #{FORMATS.join(", ")} (default: text)") do |value| format = value end opts.on("--allow-judge-mismatch", "Compare runs that used different judge models anyway") do allow_judge_mismatch = true end opts.on("--max-error-rate N", Float, "Fraction of runs either side may lose to errors before --fail-on-regression declines to " \ "decide (default: #{Raif::Evals::Comparison::MAX_GATE_ERROR_RATE}). 1 gates regardless.") do |value| max_error_rate = value end opts.on("-h", "--help", "Show this help message") do puts opts exit end end (parser) # A level outside (0, 1] is a typo with a silent failure mode at each end: 0 can never be # cleared, so the gate would never fire, and a negative one does the same. if alpha && (alpha <= 0 || alpha > 1) puts "Error: --significance must be greater than 0 and at most 1 (got #{alpha})" exit 1 end if max_error_rate.negative? || max_error_rate > 1 puts "Error: --max-error-rate must be between 0 and 1 (got #{max_error_rate})" exit 1 end baseline_path, candidate_path = args if baseline_path.nil? || candidate_path.nil? puts parser exit 1 end comparison = Raif::Evals::Comparison.new( baseline: load_payload(baseline_path), candidate: load_payload(candidate_path), baseline_label: File.basename(baseline_path), candidate_label: File.basename(candidate_path) ) if comparison.judge_mismatch? && !allow_judge_mismatch puts Raif::Utils::Colors.red(<<~MSG) Refusing to compare runs judged by different models: baseline: #{comparison.baseline_judge.inspect} candidate: #{comparison.candidate_judge.inspect} Scores from two judges measure two different things. Re-run one arm with the other's judge, or pass --allow-judge-mismatch if you know the difference does not matter here. A run with no Raif.config.evals_default_llm_judge_model_key set is graded by the model under test, so comparing two models changes the judge unless one is configured. MSG exit 2 end alpha ||= Raif::Evals::Comparison::FAMILY_WISE_ALPHA report = Raif::Evals::ComparisonReport.new( comparison, threshold: threshold, color: format == "text", alpha: alpha, max_error_rate: max_error_rate ) if format == "html" output_path = html_output_path(candidate_path, baseline_path) File.write(output_path, report.render("html")) # The report went to a file rather than the screen, and a dataset that changed under the # comparison has to be seen before the rest. text and json carry it themselves. dataset_warning = report.dataset_warning puts Raif::Utils::Colors.yellow(dataset_warning) if dataset_warning puts "Comparison report written to: #{output_path}" else puts report.render(format) end # Ahead of the gate rather than after it: a run that lost this much to errors cannot support # a verdict either way. 2 rather than 1, matching the refusals around it. if threshold && comparison.error_rate_unreliable?(max_error_rate: max_error_rate) puts Raif::Utils::Colors.red(<<~MSG) Refusing to pass or fail: too many runs errored to gate on this comparison. baseline: #{format("%.1f%%", comparison.baseline_error_rate * 100)} of runs errored candidate: #{format("%.1f%%", comparison.candidate_error_rate * 100)} of runs errored ceiling: #{format("%g%%", max_error_rate * 100)} (--max-error-rate) Errored runs are excluded from the pass rates rather than counted as failures, so the rates above are sound - but the runs that survived may not be a fair sample of the ones that did not, and a gate cannot tell the difference. Re-run the failing arm, or pass --max-error-rate 1 to gate on the surviving runs anyway. MSG exit 2 end exit 1 if comparison.regressed?(threshold, alpha: alpha) # Something moved past the threshold and none of it could be tested, so the gate has no # answer to give. Exiting 0 would report a possibly-regressed run as clean; 2 matches the # judge mismatch above - refusing to decide, rather than deciding against. if comparison.insufficient_evidence?(threshold, alpha: alpha) puts Raif::Utils::Colors.red(<<~MSG) Refusing to pass or fail: #{comparison.unverifiable_regressions(threshold).count} regression(s) cleared the --fail-on-regression threshold, but none of them can be told apart from run-to-run variation. The gate tests matched dataset cases. An eval with no dataset has none, and --repeat sharpens each case's estimate rather than creating pairs to compare. Give the evals a dataset, or pass --significance 1 to gate on effect size alone. MSG exit 2 end end |