Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,13 @@ kept as history. This entry is the one to read if you are coming from **1.15.1**
is offline CLIP via the `informers` gem; custom backends (local VLM,
decision APIs) register by name via `SnapDiff::AI.register`. Zero new hard
dependencies. See `docs/ai.md`.
- **Report contributions registry.** Core reporters and the assertion failure
message no longer reference the AI module at all: `SnapDiff::Contributions`
is the single extension point (Minitest/SimpleCov style — modules
self-register, core renders plain `{source:, text:, data:}` payloads).
Loading `snap_diff/ai` opts into annotations; `fail_on:` claims the one
failure-suppression slot. Any module can contribute annotations the same
way. See `docs/ai.md`.

### Upgrading from 1.15.1: change the version, run your suite

Expand Down
42 changes: 31 additions & 11 deletions docs/ai.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,20 +18,39 @@ One line per diff in the test output, and — automatically, via the
shared store — a verdict badge plus one-line summary per failure in
`snap_diff_report.html`. No files, no duplicate storage.

## How the HTML report gets AI annotations
## How reports get AI annotations

Both reporters read the same shared store: `AISimple` writes each result
into `SnapDiff::AI` (keyed by screenshot name), and the HTML reporter
attaches `SnapDiff::AI[name]` to the failure entry — behind a `defined?`
guard, so `html.rb` never requires the AI module and nothing changes
when AI triage isn't loaded. No wiring between the two reporters:
Core never names the AI module. It exposes one registry,
`SnapDiff::Contributions`, and loading `snap_diff/ai` self-registers
into it — the same shape as Minitest plugins appending to the
`CompositeReporter`, or SimpleCov formatters receiving a plain payload:

- `AISimple` writes each result into the shared `SnapDiff::AI` store
(keyed by screenshot name).
- `SnapDiff::AI.annotate(name)` adapts a stored result to the generic
`{source:, text:, data:}` contribution shape.
- The HTML reporter and the assertion failure message ask
`Contributions.annotations_for(name)` and render whatever comes back —
empty when AI was never loaded, so the no-AI report is byte-clean.

```ruby
require "snap_diff/reporters/html" # already auto-registers
require "snap_diff/reporters/ai_simple"
require "snap_diff/reporters/ai_simple" # self-registers into Contributions
SnapDiff::Reporting.register(SnapDiff::Reporters::AISimple.new)
```

Your own module can contribute the same way — no edits to core:

```ruby
module TicketLinker
def self.annotate(name)
ticket = JIRA_FOR[name]
ticket && {source: "jira", text: ticket}
end
end
SnapDiff::Contributions.register(TicketLinker)
```

The report then shows, per failure: an `AI: real bug` / `AI: flaky` badge
in the sidebar and top strip, plus `similarity · confidence · summary`
when the backend provides them. Under fork-parallel, results merge in the
Expand All @@ -47,11 +66,12 @@ parent before the report renders.
## Architecture

```
lib/snap_diff/ai.rb # backend registry, verdict thresholds, shared result store
lib/snap_diff/contributions.rb # core registry: annotations + the one failure gate
lib/snap_diff/ai.rb # backend registry, verdict thresholds, shared store; self-registers
lib/snap_diff/ai/backends/clip.rb # built-in offline backend (informers)
lib/snap_diff/reporters/ai_simple.rb # record/finalize/summary; writes the store
lib/snap_diff/reporters/html.rb # annotates failures from the store (defined? guard)
lib/snap_diff/screenshot_assertion.rb# validate: fail-gate + AI line in the failure message
lib/snap_diff/reporters/ai_simple.rb # record/finalize/summary; suppression via Contributions
lib/snap_diff/reporters/html.rb # renders Contributions.annotations_for — no AI reference
lib/snap_diff/screenshot_assertion.rb# validate: Contributions gate + annotation lines
```

A backend is any object with `#call(name:, base:, current:, meta:) -> Hash`.
Expand Down
1 change: 1 addition & 0 deletions lib/snap_diff.rb
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ def self.assert_single_gem!(loaded_specs = Gem.loaded_specs)
require "capybara/dsl"
require "snap_diff/config"
require "snap_diff/comparison"
require "snap_diff/contributions"
require "snap_diff/legacy_shims"
require "snap_diff/version"
# SnapDiff.session/.reset/.pending_screenshots_message are part of the
Expand Down
28 changes: 17 additions & 11 deletions lib/snap_diff/ai.rb
Original file line number Diff line number Diff line change
@@ -1,14 +1,17 @@
# frozen_string_literal: true

require "snap_diff/contributions"

# Optional AI triage. A backend is any object responding to
# #call(name:, base:, current:, meta:) -> Hash; register one by name or
# pass an instance. Builders run lazily, so optional gems load only when
# their backend is used. Advisory by default: pixel diff stays the verdict
# unless a reporter is configured with fail_on: (see AISimple).
#
# Results live in a process-wide store so ANY consumer can read them --
# the AISimple reporter writes, the HTML reporter annotates from it, the
# fail-gate consults it. Keyed by screenshot name; later writes win.
# the AISimple reporter writes, and reports pick them up through the
# SnapDiff::Contributions registry (no consumer names this module).
# Keyed by screenshot name; later writes win.
module SnapDiff
module AI
# The only verdicts a backend may return; anything else maps to
Expand All @@ -20,16 +23,14 @@ module AI
@mutex = Mutex.new

class << self
# Optional fail-gate, set by AISimple when configured with fail_on:.
# ScreenshotAssertion#validate consults it on a pixel diff: verdicts
# the gate accepts suppress the failure, the rest still fail.
attr_accessor :gate

def gated_result(name, difference) = gate&.gated_result(name, difference)

# No gate -> everything fails, exactly as without AI.
def fails?(verdict) = gate ? gate.fails?(verdict) : true
# Report contribution (SnapDiff::Contributions): whatever the store
# holds for this screenshot, rendered one line plus the raw payload.
def annotate(name)
result = self[name]
result && {source: "ai", text: format(result), data: result}
end

# Register a lazy backend factory under a symbolic name, replacing any prior factory.
def register(name, &build)
@mutex.synchronize { @backends[name.to_sym] = build }
end
Expand Down Expand Up @@ -113,3 +114,8 @@ def merge_state!(state)
end

require "snap_diff/ai/backends/clip"

# Minitest-style self-registration: loading this file opts into
# annotating reports. Core (HTML reporter, assertion message) only ever
# talks to SnapDiff::Contributions and never references this module.
SnapDiff::Contributions.register(SnapDiff::AI)
61 changes: 61 additions & 0 deletions lib/snap_diff/contributions.rb
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# frozen_string_literal: true

# Contribution points: how OPTIONAL modules (AI triage today, anything
# else tomorrow) feed the reports and the failure decision without the
# reporters or the assertion knowing those modules exist.
#
# The contract follows the two proven shapes in the ecosystem:
# Minitest's CompositeReporter (plugins append themselves to a
# core-owned list; core calls a uniform interface and never names a
# plugin) and SimpleCov's formatter pipeline (consumers receive a plain
# data payload, never a plugin's class). A contributor is any object
# responding to #annotate(name) -> {source:, text:, data:} or nil;
# reports render whatever comes back.
#
# Failure suppression is deliberately a SINGLE slot (it was
# SnapDiff::AI.gate before): two gates with different accept-lists
# would silently suppress each other's real bugs, so registration
# replaces the previous gate rather than stacking.
module SnapDiff
module Contributions
@providers = []
@suppression = nil
@mutex = Mutex.new

class << self
# Register a report contributor. The provider must respond to
# #annotate(name), returning {source:, text:, data: (optional)}
# or nil. Registering the same object twice is a no-op -- identity,
# not ==: two distinct providers that happen to compare equal
# (e.g. Structs with equal fields) must BOTH contribute.
def register(provider)
@mutex.synchronize { @providers << provider unless @providers.any? { |p| p.equal?(provider) } }
end

# All contributions for one screenshot, in registration order.
# Empty when nothing is registered -- the no-AI default.
def annotations_for(name)
@mutex.synchronize { @providers.dup }.filter_map { |provider| provider.annotate(name) }
end

# The one failure gate: #suppress(name, difference) ->
# {source:, text:} (failure waived) or nil (failure stands).
# nil clears the slot (test teardown, reconfiguration).
def register_suppression(provider)
@mutex.synchronize { @suppression = provider }
end

# Cheap probe so callers can skip building `difference` entirely
# when no gate is registered.
def any_suppressor?
@mutex.synchronize { !@suppression.nil? }
end

# Ask the current suppressor to evaluate a screenshot difference.
# Return its {source:, text:} waiver, or nil when no gate waives the failure.
def suppression_for(name, difference)
@mutex.synchronize { @suppression }&.suppress(name, difference)
end
end
end
end
27 changes: 20 additions & 7 deletions lib/snap_diff/reporters/ai_simple.rb
Original file line number Diff line number Diff line change
Expand Up @@ -33,9 +33,13 @@ def initialize(backend: nil, flaky: nil, intentional: nil, fail_on: nil)
# fresh regression.
@memo = {}
@memo_mutex = Mutex.new
AI.gate = self if @fail_on
# The one failure-gate slot in SnapDiff::Contributions -- core
# consults it without knowing AI exists.
Contributions.register_suppression(self) if @fail_on
end

# Analyze differing assertions and store their results, reusing gate-time analysis.
# Do nothing when the backend is unavailable.
def record(assertions)
return unless @backend

Expand All @@ -47,15 +51,24 @@ def record(assertions)
end
end

# Fail-gate entry point, called by ScreenshotAssertion#validate on a
# pixel diff. Analysis runs there (before the error message is
# built) and is memoized, so #record never re-analyzes.
def gated_result(name, difference)
return unless @fail_on && @backend
# Failure-gate contract (SnapDiff::Contributions), consulted by
# ScreenshotAssertion#validate on a pixel diff, before the error
# message is built. Analysis is memoized, so #record never
# re-analyzes. Returns {source:, text:} to waive the failure,
# nil to let it stand. "unknown" ALWAYS fails -- AI can downgrade
# a diff, never vouch for one it could not classify.
def suppress(name, difference)
return unless @backend

# nil analysis (the backend raised) -> the pixel failure stands.
result = analyze_once(name, difference)
return unless result

analyze_once(name, difference)
{source: "ai", text: AI.format(result)} unless fails?(result[:verdict])
Comment thread
sourcery-ai[bot] marked this conversation as resolved.
Comment thread
coderabbitai[bot] marked this conversation as resolved.
end

# Return whether a verdict must fail under the configured fail_on policy.
# Unknown verdicts always fail; requires a reporter configured with fail_on.
def fails?(verdict) = verdict == "unknown" || @fail_on.include?(verdict)

# Results are already in the shared store -- nothing to write out.
Expand Down
27 changes: 14 additions & 13 deletions lib/snap_diff/reporters/html.rb
Original file line number Diff line number Diff line change
Expand Up @@ -96,8 +96,9 @@ def summary
"[snap_diff] Report: #{output_path}" if @finalized
end

# Attach available contributions and return the rendered HTML report.
def render
attach_ai_annotations
attach_annotations
ERB.new(File.read(self.class.template_path)).result(binding)
end

Expand All @@ -112,6 +113,7 @@ def self.default_output_path

private

# Build a report entry from a comparison, omitting unavailable images and metrics.
def failure_entry_for(name, compare)
difference = compare.difference
{
Expand All @@ -127,19 +129,18 @@ def failure_entry_for(name, compare)
}.compact
end

# Advisory AI triage annotations, attached at RENDER time: HTML
# records before the AI reporter (auto-registration runs first) and
# fork-parallel merges land after record, so only the final render
# can see every verdict. HTML never requires the AI module -- the
# annotation appears iff the user opted into AI triage.
def attach_ai_annotations
return unless defined?(SnapDiff::AI)

# Contributed annotations (AI triage, ...), attached at RENDER
# time: HTML records before contributing reporters
# (auto-registration runs first) and fork-parallel merges land
# after record, so only the final render can see everything.
# HTML never names a contributor -- it renders whatever
# SnapDiff::Contributions returns, empty by default.
def attach_annotations
failures.each do |entry|
# || would create a nil :ai key on misses; keep the entry clean.
if (annotation = SnapDiff::AI[entry[:name]])
entry[:ai] ||= annotation
end
# ||= would create a nil :annotations key on misses; keep the
# entry clean.
annotations = SnapDiff::Contributions.annotations_for(entry[:name])
entry[:annotations] ||= annotations unless annotations.empty?
end
end

Expand Down
38 changes: 29 additions & 9 deletions lib/snap_diff/reporters/templates/report.html.erb
Original file line number Diff line number Diff line change
Expand Up @@ -274,7 +274,15 @@
'<div class="thumb-img-wrap"><img src="' + esc(thumb) + '" alt="" loading="lazy"/></div>' +
'<div class="thumb-footer">' +
'<div class="thumb-name">' + esc(item.name) + '</div>' +
(item.ai && item.ai.verdict ? '<span class="ai-badge' + aiClass(item.ai.verdict) + '">' + esc(item.ai.verdict.replace('_', ' ')) + '</span>' : '') +
(item.annotations || []).map(function(note) {
var v = note.data && note.data.verdict;
/* verdict annotations badge the verdict; text-only ones
(TicketLinker, ...) must show their text or they render
as a bare source label */
var label = note.source.toUpperCase() +
(v ? ': ' + v.replace('_', ' ') : (note.text ? ': ' + note.text : ''));
return '<span class="ai-badge' + (v ? aiClass(v) : '') + '">' + esc(label) + '</span>';
}).join(' ') +
'<span class="thumb-badge ' + (hasDiff ? 'thumb-badge-fail' : 'thumb-badge-pass') + '">' + badgeText + '</span>' +
'</div>';
btn.addEventListener('click', function() { selectItem(+this.dataset.idx); });
Expand Down Expand Up @@ -311,20 +319,32 @@
}
topBadge.className = 'diff-badge ' + (hasDiff ? 'diff-badge-fail' : 'diff-badge-pass');

/* AI triage strip: shown only when an advisory verdict exists */
/* Contribution strip: the verdict layout when a contributor supplied
one (AI triage), source + text for text-only annotations */
var aiBar = $('ai-bar');
if (item.ai && item.ai.verdict) {
var note = (item.annotations || []).find(function(a) { return a.data && a.data.verdict; });
Comment thread
coderabbitai[bot] marked this conversation as resolved.
if (note) {
var d = note.data;
var aiV = $('ai-verdict');
aiV.textContent = 'AI: ' + item.ai.verdict.replace('_', ' ');
aiV.className = 'ai-badge' + aiClass(item.ai.verdict);
aiV.textContent = note.source.toUpperCase() + ': ' + d.verdict.replace('_', ' ');
aiV.className = 'ai-badge' + aiClass(d.verdict);
var bits = [];
if (item.ai.similarity != null) bits.push('similarity ' + item.ai.similarity);
if (item.ai.confidence != null) bits.push('confidence ' + item.ai.confidence);
if (item.ai.summary) bits.push(item.ai.summary);
if (d.similarity != null) bits.push('similarity ' + d.similarity);
if (d.confidence != null) bits.push('confidence ' + d.confidence);
if (d.summary) bits.push(d.summary);
$('ai-text').textContent = bits.join(' · ');
aiBar.className = 'visible';
} else {
aiBar.className = '';
var plain = (item.annotations || [])[0];
if (plain) {
var plainV = $('ai-verdict');
plainV.textContent = plain.source.toUpperCase();
plainV.className = 'ai-badge';
$('ai-text').textContent = plain.text || '';
aiBar.className = 'visible';
} else {
aiBar.className = '';
}
}

/* Resolve image sources based on view + annotated toggle */
Expand Down
20 changes: 12 additions & 8 deletions lib/snap_diff/screenshot_assertion.rb
Original file line number Diff line number Diff line change
Expand Up @@ -60,22 +60,26 @@ def inspect
"#<#{self.class.name} #{name.inspect} #{state} new=#{compare.image_path} base=#{compare.base_image_path}>"
end

# Return an annotated failure message for an unsuppressed screenshot mismatch.
# Return nil for absent, matching, or suppressed comparisons; archive matching baselines.
def validate
return unless compare

if compare.different?
# Optional AI gate (SnapDiff::Reporters::AISimple with fail_on:):
# the verdict is computed here, before the message is built, so
# the failure text can quote it and accepted verdicts can skip it.
ai = SnapDiff::AI.gated_result(name, compare.difference) if defined?(SnapDiff::AI) && SnapDiff::AI.gate
if ai && !SnapDiff::AI.fails?(ai[:verdict])
$stdout.puts "[snap_diff:ai] #{name}: failure suppressed -- #{SnapDiff::AI.format(ai)}"
# Optional contribution gate (e.g. AI triage with fail_on:): a
# registered suppressor runs BEFORE the message is built, so a
# waived diff never fails and a standing failure can quote what
# the contributors know. any_suppressor? keeps the no-gate path
# from even touching compare.difference.
if Contributions.any_suppressor? && (suppressed = Contributions.suppression_for(name, compare.difference))
$stdout.puts "[snap_diff:#{suppressed[:source]}] #{name}: failure suppressed -- #{suppressed[:text]}"
return nil
end

message = "Screenshot does not match for '#{name}': #{compare.error_message}\n#{caller.join("\n")}"
ai ||= SnapDiff::AI[name] if defined?(SnapDiff::AI)
message += "\n AI triage: #{SnapDiff::AI.format(ai)}" if ai
Contributions.annotations_for(name).each do |note|
message += "\n #{note[:source].upcase} triage: #{note[:text]}"
end
message
else
archive_baseline!
Expand Down
Loading
Loading