v0.4.0-alpha

Compare

Re-run a fixture scenario against two OpenRouter models and show both ProofResults side-by-side. Implements spec §8.3 AC #13 and fixture skill.model_comparison. The UI never flattens divergence into a composite verdict (spec §8.4): you see both proofs and any disagreement.

Why two proofs and never one number

Averaging two models into a single score destroys the only interesting information: where they disagreed. A composite verdict that reads 0.5 cannot be distinguished from two runs that each half-worked, and those are different findings. Both ProofResults are shown whole.

What differing traces tell you
If two models reach the same outcome by different tool sequences, the skill is closer to portable. If only one reaches it, the skill is bound to that model and the manifest's binding should say so rather than claim otherwise.
What differing gates tell you
A gate that passes for one model and fails for the other localizes the divergence to a named claim — an undeclared capability, a state hash that did not reproduce — instead of a difference in overall vibe.
What this page needs
Two live model calls, so it is one of the two surfaces that asks for an OpenRouter key. Recording a session you already ran needs no key at all.
No OpenRouter key in this session. Set one on the Teach page before running a comparison, or record a session you already ran — that needs no key.

Model A

Model B