Compare
Re-run a fixture scenario against two OpenRouter models and show both ProofResults side-by-side.
Implements spec §8.3 AC #13 and fixture skill.model_comparison. The UI never flattens divergence
into a composite verdict (spec §8.4): you see both proofs and any disagreement.
Why two proofs and never one number
Averaging two models into a single score destroys the only interesting information: where they
disagreed. A composite verdict that reads 0.5 cannot be distinguished from two runs
that each half-worked, and those are different findings. Both ProofResults are shown
whole.
- What differing traces tell you
- If two models reach the same outcome by different tool sequences, the skill is closer to
portable. If only one reaches it, the skill is bound to that model and the manifest'sbindingshould say so rather than claim otherwise. - What differing gates tell you
- A gate that passes for one model and fails for the other localizes the divergence to a named claim — an undeclared capability, a state hash that did not reproduce — instead of a difference in overall vibe.
- What this page needs
- Two live model calls, so it is one of the two surfaces that asks for an OpenRouter key. Recording a session you already ran needs no key at all.
No OpenRouter key in this session. Set one on the Teach page before running a comparison,
or record a session you already ran — that needs no key.
Model A
—
Model B
—