What you get
During the sprint
- Consequence tier and threat model for the agent, agreed in week one
- The pre-frozen assurance coverage denominator, signed by your owner and our lead methodologist before anything runs
- Invariant suite: automated synthesis for the standard relational families plus authored fixtures for your proprietary logic
- Challenge-suite hash committed and nonce issued before execution
- A shared channel and a 30-minute call each week
At the end
- A dual-signed audit bundle in your S3 bucket or Snowflake stage, with a notary record you can check publicly
- The three-tier readout: board scorecard, engineering triage, scientific appendix
- Every finding in one of four states, with reproducible trace and counterfactual proof for each
DEFECT_CONFIRMED - ICFY against your native evaluation, with a 95% Wilson interval
- The invariant suite itself, versioned, in your repo — it is yours whether or not you continue
How it runs
| Week | What happens | What we need from you |
|---|---|---|
| 0 | 30-minute readiness review. Which agent, which schema surface, which consequence tier. We say whether a sprint is worth your money. | 30 minutes, and someone who knows the agent |
| 1 | Kickoff. Metadata extraction (dbt manifest.json, INFORMATION_SCHEMA). Threat model. Denominator frozen and signed. Automated invariant generation runs; we report how many rules it produced and how long it took. | Read-only catalog access or one run of the local extractor; the exec sponsor, a technical lead and the Business Definition Owner in one room |
| 2 | Targeted human authoring of Type B/C fixtures for your proprietary logic. Challenge suite committed by hash. First in-perimeter execution via CLI or stored procedure. | A staging branch or sandbox the suite can run against |
| 3 | Sequential trials complete. Adjudication: Tier-3 review on exceptions, Tier-4 blinded pair on any Sev-1. INDETERMINATE items go to your Business Definition Owner (48-hour SLA), then re-run. | Definitions, within 48 hours |
| 4 | Bundle signed. CDAO readout. ICFY reported. A binary conversation about the platform, on the date we fixed in week one. | A decision |
What it costs, and why
One fixed fee. Billed by corporate card or Stripe so procurement is not on the critical path. Kickoff is gated on 100% upfront receipt. Fixed fees, agreed before we start. No hourly billing, no success fee, no scope surprises.
The fee is anchored to the work, and the work is mostly authoring. A 100-case suite is about 50 hours of senior assurance engineering in Sprint 1 — adapting archetypes, verifying custom cases against your dbt models, and writing the bespoke fixtures for your proprietary logic — plus intake validation, execution, adjudication and the readout. The number that has to fall for this to be a software business rather than a consultancy is that authoring figure: targeted at under 5 hours by Sprint 10 as manifest-driven synthesis takes over. We report the actual hours in every readout, whether or not they flatter us.
| Authoring type | What it covers | Minutes per case | Share of a suite |
|---|---|---|---|
| Type A — adapted archetype | Standard failure patterns from the reusable catalogue (join fan-out, basic SCD-2 logic) | 15–20 | 50% |
| Type B — validated custom case | Test queries verified against your dbt models and schema relationships | 30–45 | 30% |
| Type C — net-new bespoke fixture | Edge fixtures for proprietary stored procedures, multi-table rollups, custom UDFs | 60–90 | 20% |
Success criteria, agreed before we start
Numeric, written into the agreement, reported whether or not we hit them:
- ICFY ≥ 40% — at least two in five of the critical and high defects we confirm were not flagged by your native evaluation.
- Decision Change Rate ≥ 30% — findings that actually altered a go/no-go, a prompt or a semantic-layer definition.
- Finding Overturn Rate < 5% — if you overturn more than one in twenty of our findings on adjudication, our method has a problem and we say so.
- Authoring hours reported — the actual human hours, against the automation target.
A sprint that ends “the native evaluation already catches this” is a real outcome. We would rather write that than pad a number; our own kill rule depends on it being true.
What this is not
- Not a certification. The bundle is evidence you can hand to an auditor. It is not a mark of approval, and your CDAO keeps ownership of the deployment decision. See the standards mapping.
- Not a data-quality monitor. Observability tools watch pipelines and tables. We test what the agent’s reasoning does with them.
- Not remediation. We never author your production prompts or dbt models; that would break the Independence Charter.
- Not a public benchmark. Your results are yours. We publish methodology, not customer scores.
We are running the first sprints under a 90-day falsification programme with a cohort of 3 design partners. If you engage at this stage you get founder-led delivery, the design-partner terms, and direct influence on the invariant catalogue. What you do not get is a list of logos to check. There is none, and we will not borrow one.