The outcome moat
Proof of lift, not promises.
Every experiment North Signal runs records a controlled before/after: a change, a held-out control, and a significance test. Aggregated across customers, that becomes the one thing a monitoring tool can never build, evidence of which changes actually move AI answers, by how much, and how reliably. No brand names, no per-customer data; only the lever → lift signal.
Which levers actually move answers
Ranked by verified wins. Win rate = share of experiments where the change beat its control significantly; average lift is measured against that control. Directional until sample sizes are large.
Why this is honest:AI answers vary run to run, so a single before/after is near-noise. We only count a win when the treated group moved significantly beyond a held-out control that absorbs model drift. The uncomfortable cases, "that change did nothing", are recorded too. variance-honest
See where you stand, then ship a fix and prove it, the way this data was built.