Class: Raif::Evals::ComparisonReport
- Inherits:
-
Object
- Object
- Raif::Evals::ComparisonReport
- Defined in:
- lib/raif/evals/comparison_report.rb
Overview
Renders a Comparison as console text, JSON, or a self-contained HTML file.
Constant Summary collapse
- TEMPLATE_PATH =
File.("comparison_report.html.erb", __dir__)
- STYLESHEET_PATH =
File.("report.css", __dir__)
Instance Attribute Summary collapse
-
#alpha ⇒ Object
readonly
Returns the value of attribute alpha.
-
#comparison ⇒ Object
readonly
Returns the value of attribute comparison.
-
#max_error_rate ⇒ Object
readonly
Returns the value of attribute max_error_rate.
-
#threshold ⇒ Object
readonly
Returns the value of attribute threshold.
Instance Method Summary collapse
-
#cost_split? ⇒ Boolean
Only when a judge actually spent something on one side or the other: a run that used no judge would otherwise gain two rows saying so twice.
-
#dataset_warning ⇒ Object
Loud and not a refusal - see Comparison#dataset_differences.
-
#describe_code(code) ⇒ Object
"unknown" rather than blank for a run recorded before the git sha was: what code produced it was not recorded, which is different from having been produced by no code.
-
#error_fraction(side) ⇒ Object
The number the SUMMARY prints beside "evals errored".
- #evidence_note ⇒ Object
-
#evidence_summary(row) ⇒ Object
How sure the gate is that a row is not noise, in the terms the reader has to act on: what was paired, how it split, and the p-value that came out.
-
#format_cost(cost) ⇒ Object
"-" rather than $0.00 for a run recorded before judge spend was tagged: it is not known to be zero, and a zero would read as "this run used no judge".
- #format_delta(delta) ⇒ Object
-
#format_p(p_value) ⇒ Object
A p-value rounded to 6 places can land on 0.0, which reads as a bug rather than as "smaller than this report shows".
- #format_rate(rate) ⇒ Object
- #gated_rows ⇒ Object
- #h(value) ⇒ Object
- #html ⇒ Object
-
#initialize(comparison, threshold: nil, color: true, alpha: Comparison::FAMILY_WISE_ALPHA, max_error_rate: Comparison::MAX_GATE_ERROR_RATE) ⇒ ComparisonReport
constructor
A new instance of ComparisonReport.
- #render(format) ⇒ Object
- #run_at(value) ⇒ Object
-
#score_headline(row) ⇒ Object
Public so the HTML template shares it with #text and the two cannot drift.
-
#score_relative(row) ⇒ Object
Printed alongside the absolute delta because --fail-on-regression is relative to the baseline: whether -1.0 clears a 0.25 threshold depends on what it is -1.0 of.
-
#score_spread(row) ⇒ Object
A single observation has no standard deviation, so one side can be absent while the other is real.
-
#stylesheet ⇒ Object
The palette and base styles every eval page shares, inlined so the page stays one file.
- #text ⇒ Object
- #unverifiable_advice ⇒ Object
-
#verdict ⇒ Object
The threshold is restated as a percentage: a bare "0.25" next to a list of absolute deltas invites reading it as those deltas' units rather than as a fraction of baseline.
Constructor Details
#initialize(comparison, threshold: nil, color: true, alpha: Comparison::FAMILY_WISE_ALPHA, max_error_rate: Comparison::MAX_GATE_ERROR_RATE) ⇒ ComparisonReport
Returns a new instance of ComparisonReport.
19 20 21 22 23 24 25 26 |
# File 'lib/raif/evals/comparison_report.rb', line 19 def initialize(comparison, threshold: nil, color: true, alpha: Comparison::FAMILY_WISE_ALPHA, max_error_rate: Comparison::MAX_GATE_ERROR_RATE) @comparison = comparison @threshold = threshold @color = color @alpha = alpha @max_error_rate = max_error_rate end |
Instance Attribute Details
#alpha ⇒ Object (readonly)
Returns the value of attribute alpha.
17 18 19 |
# File 'lib/raif/evals/comparison_report.rb', line 17 def alpha @alpha end |
#comparison ⇒ Object (readonly)
Returns the value of attribute comparison.
17 18 19 |
# File 'lib/raif/evals/comparison_report.rb', line 17 def comparison @comparison end |
#max_error_rate ⇒ Object (readonly)
Returns the value of attribute max_error_rate.
17 18 19 |
# File 'lib/raif/evals/comparison_report.rb', line 17 def max_error_rate @max_error_rate end |
#threshold ⇒ Object (readonly)
Returns the value of attribute threshold.
17 18 19 |
# File 'lib/raif/evals/comparison_report.rb', line 17 def threshold @threshold end |
Instance Method Details
#cost_split? ⇒ Boolean
Only when a judge actually spent something on one side or the other: a run that used no judge would otherwise gain two rows saying so twice.
135 136 137 |
# File 'lib/raif/evals/comparison_report.rb', line 135 def cost_split? [comparison.to_h[:baseline], comparison.to_h[:candidate]].any? { |side| side[:judge_cost].to_f.positive? } end |
#dataset_warning ⇒ Object
Loud and not a refusal - see Comparison#dataset_differences. Public so evals:compare can print it when --format html sent the report to a file instead of the screen.
103 104 105 106 107 108 109 110 111 112 113 114 115 |
# File 'lib/raif/evals/comparison_report.rb', line 103 def dataset_warning return if comparison.dataset_differences.empty? lines = ["Warning: these two runs did not measure the same datasets:"] comparison.dataset_differences.each do |row| lines << " #{row[:name]} (#{row[:eval_set]}): #{row[:baseline]} -> #{row[:candidate]}" end lines << " Cases are joined by id, so a difference the dataset caused reads as a difference the model caused." lines << " Re-run the baseline against the current dataset to compare the two models alone." lines.join("\n") end |
#describe_code(code) ⇒ Object
"unknown" rather than blank for a run recorded before the git sha was: what code produced it was not recorded, which is different from having been produced by no code.
119 120 121 122 123 |
# File 'lib/raif/evals/comparison_report.rb', line 119 def describe_code(code) return "unknown" if code.nil? "#{code["git_sha"].to_s[0, 12]}#{" (dirty)" if code["dirty"]}" end |
#error_fraction(side) ⇒ Object
The number the SUMMARY prints beside "evals errored". Not a pass rate's denominator: this one is every run, since the question is how much of the run was lost.
162 163 164 |
# File 'lib/raif/evals/comparison_report.rb', line 162 def error_fraction(side) "#{side[:errored_evals]}/#{side[:total_evals]} (#{format("%.1f%%", side[:error_rate].to_f * 100)})" end |
#evidence_note ⇒ Object
201 202 203 204 205 206 |
# File 'lib/raif/evals/comparison_report.rb', line 201 def evidence_note return "evidence not required (--significance #{alpha})" if alpha.nil? || alpha.to_f >= 1.0 count = comparison.candidate_regressions(threshold).count "at a family-wise #{alpha} over #{count} candidate row#{"s" if count != 1}" end |
#evidence_summary(row) ⇒ Object
How sure the gate is that a row is not noise, in the terms the reader has to act on: what was paired, how it split, and the p-value that came out. Nothing to say for a row that is not a regression candidate, which is most of them.
93 94 95 96 97 98 99 |
# File 'lib/raif/evals/comparison_report.rb', line 93 def evidence_summary(row) case row[:evidence] when :paired_cases then "#{row[:worsened]}/#{row[:pairs]} cases worse, #{format_p(row[:p_value])}" when :repeats then "repeats only, #{format_p(row[:p_value])}" else "no matched cases to test - add dataset cases, or --significance 1 to gate on size alone" end end |
#format_cost(cost) ⇒ Object
"-" rather than $0.00 for a run recorded before judge spend was tagged: it is not known to be zero, and a zero would read as "this run used no judge".
127 128 129 130 131 |
# File 'lib/raif/evals/comparison_report.rb', line 127 def format_cost(cost) return "-" if cost.nil? "$#{format("%.2f", cost.to_f)}" end |
#format_delta(delta) ⇒ Object
152 153 154 |
# File 'lib/raif/evals/comparison_report.rb', line 152 def format_delta(delta) delta.to_f.positive? ? "+#{delta}" : delta.to_s end |
#format_p(p_value) ⇒ Object
A p-value rounded to 6 places can land on 0.0, which reads as a bug rather than as "smaller than this report shows".
145 146 147 148 149 150 |
# File 'lib/raif/evals/comparison_report.rb', line 145 def format_p(p_value) return "p n/a" if p_value.nil? return "p<0.000001" if p_value.zero? "p=#{p_value}" end |
#format_rate(rate) ⇒ Object
156 157 158 |
# File 'lib/raif/evals/comparison_report.rb', line 156 def format_rate(rate) format("%.2f", rate.to_f) end |
#gated_rows ⇒ Object
139 140 141 |
# File 'lib/raif/evals/comparison_report.rb', line 139 def gated_rows @gated_rows ||= comparison.significant_regressions(threshold, alpha: alpha) end |
#h(value) ⇒ Object
213 214 215 |
# File 'lib/raif/evals/comparison_report.rb', line 213 def h(value) ERB::Util.html_escape(value.to_s) end |
#html ⇒ Object
36 37 38 |
# File 'lib/raif/evals/comparison_report.rb', line 36 def html ERB.new(File.read(TEMPLATE_PATH), trim_mode: "-").result(binding) end |
#render(format) ⇒ Object
28 29 30 31 32 33 34 |
# File 'lib/raif/evals/comparison_report.rb', line 28 def render(format) case format.to_s when "json" then JSON.pretty_generate(comparison.to_h) when "html" then html else text end end |
#run_at(value) ⇒ Object
217 218 219 |
# File 'lib/raif/evals/comparison_report.rb', line 217 def run_at(value) value.to_s.sub("T", " ")[0, 16] end |
#score_headline(row) ⇒ Object
Public so the HTML template shares it with #text and the two cannot drift.
59 60 61 62 63 64 |
# File 'lib/raif/evals/comparison_report.rb', line 59 def score_headline(row) parts = [score_relative(row), score_spread(row)].compact parts << "not gated" unless row[:gated] "#{row[:name]} #{row[:baseline_mean]} -> #{row[:candidate_mean]} #{format_delta(row[:delta])} (#{parts.join(", ")})" end |
#score_relative(row) ⇒ Object
Printed alongside the absolute delta because --fail-on-regression is relative to the baseline: whether -1.0 clears a 0.25 threshold depends on what it is -1.0 of. nil when the baseline mean is zero and there is no fraction to take.
69 70 71 72 73 74 |
# File 'lib/raif/evals/comparison_report.rb', line 69 def score_relative(row) baseline = row[:baseline_mean].to_f.abs return if baseline.zero? format("%+.1f%%", (row[:delta].to_f / baseline) * 100) end |
#score_spread(row) ⇒ Object
A single observation has no standard deviation, so one side can be absent while the other is real. The sd fragment is dropped only when neither side has one. spread_n is named apart from n because they differ on a dataset score: n counts every observation, sd is over the per-case means, and printing only n would overstate what the spread was measured on.
80 81 82 83 84 85 86 87 88 |
# File 'lib/raif/evals/comparison_report.rb', line 80 def score_spread(row) baseline = row[:baseline_stddev] candidate = row[:candidate_stddev] parts = ["n=#{row[:candidate_n]}"] parts << "over #{row[:spread_n]} cases" if row[:spread_n] && row[:spread_n] != row[:candidate_n] parts << "sd #{baseline || "-"} -> #{candidate || "-"}" unless baseline.nil? && candidate.nil? parts.join(", ") end |
#stylesheet ⇒ Object
The palette and base styles every eval page shares, inlined so the page stays one file.
41 42 43 |
# File 'lib/raif/evals/comparison_report.rb', line 41 def stylesheet File.read(STYLESHEET_PATH) end |
#text ⇒ Object
45 46 47 48 49 50 51 52 53 54 55 56 |
# File 'lib/raif/evals/comparison_report.rb', line 45 def text lines = ["", "Comparing eval runs", ""] lines.concat(header_lines) lines.concat(section("NEW FAILURES", comparison.new_failures) { |row| rate_lines(row, :red) }) lines.concat(section("FIXED", comparison.fixed) { |row| rate_lines(row, :green) }) lines.concat(section("SCORE MOVES", comparison.score_moves) { |row| score_lines(row) }) lines.concat(section("ERROR RATES", comparison.error_moves) { |row| error_lines(row) }) lines.concat(section("NOT COMPARABLE", comparison.not_comparable) { |row| not_comparable_lines(row) }) lines.concat(section("REGRESSION GATE", comparison.candidate_regressions(threshold)) { |row| regression_lines(row) }) lines.concat(summary_lines) lines.join("\n") end |
#unverifiable_advice ⇒ Object
208 209 210 211 |
# File 'lib/raif/evals/comparison_report.rb', line 208 def unverifiable_advice "a dataset gives the gate matched cases to test; --repeat sharpens each case rather than " \ "creating pairs. Pass --significance 1 to gate on effect size alone." end |
#verdict ⇒ Object
The threshold is restated as a percentage: a bare "0.25" next to a list of absolute deltas invites reading it as those deltas' units rather than as a fraction of baseline.
Five outcomes, not two. A row can clear the size bar and still not be distinguishable from run-to-run variation, and saying so keeps the reader from concluding either that nothing moved or that something definitely did. The error ceiling is checked ahead of all of it, since a run that lost too much to errors cannot support any of those readings.
173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 |
# File 'lib/raif/evals/comparison_report.rb', line 173 def verdict return "no regression threshold set (--fail-on-regression)" if threshold.nil? if comparison.error_rate_unreliable?(max_error_rate: max_error_rate) return "gate declined: #{format("%.1f%%", [comparison.baseline_error_rate, comparison.candidate_error_rate].max * 100)} " \ "of runs errored, above the #{format("%g", max_error_rate.to_f * 100)}% ceiling (exit 2)" end gate = "--fail-on-regression #{threshold} (#{format("%g", threshold.to_f * 100)}% worse than baseline)" candidates = comparison.candidate_regressions(threshold) significant = comparison.significant_regressions(threshold, alpha: alpha) return "no regression beyond #{gate}" if candidates.empty? if significant.any? return "#{significant.count} regression#{"s" if significant.count != 1} beyond #{gate}, " \ "#{evidence_note} (exit 1)" end if comparison.insufficient_evidence?(threshold, alpha: alpha) return "#{candidates.count} regression#{"s" if candidates.count != 1} beyond #{gate}, none of them testable - " \ "#{unverifiable_advice} (exit 2)" end "#{candidates.count} regression#{"s" if candidates.count != 1} beyond #{gate}, none distinguishable from " \ "run-to-run variation #{evidence_note}" end |