Company

What we are, and what we stopped being

EvalQA is the independent technical verification, validation and continuous assurance layer for enterprise analytics agents that write SQL against cloud warehouses — beginning with Snowflake Cortex and extending across Databricks and custom Text-to-SQL stacks. It used to claim more. This page records the change.

Last updated 13 September 2026 · Protocol v9.3

Operating stance

Empirical technical defect discovery, consequence-weighted invariants and continuous software assurance — not premature statutory certification. We believe an independent, automated, consequence-weighted assurance engine is necessary before an agent is granted unattended production execution against a warehouse holding a company’s ground truth. We also believe that belief is unproven until paying customers say so, and we have written the rule that ends the company’s platform ambition if they do not.

The 12 September 2026 truth audit

Before accepting commercial payment, we audited our own public footprint against a sixty-point adversarial review of the plan. Four changes were blocking and were made the same day:

  1. Every claim of “statutory compliance certification” was removed. We produce evidence; we do not certify. See standards.
  2. “The evaluation layer for everything AI” was retired. Scope is warehouse SQL agents. Anything else on the site is education, not a product claim.
  3. Public model comparison leaderboards (alt.qa) were suspended and archived. A benchmark publisher cannot also be an independent auditor of the benchmarked.
  4. The Independence Charter and the technical exclusions were published.

Other corrections from that review are visible across the site where they apply: the NIST identifiers on the standards page, the retracted 10.8x ROI figure on the ROI page, the Snowflake vocabulary in the failure domains, the “qualified” rather than “certified” partner programme, and the relabelling of benchmark families that had been marked validated before they were tested.

The truth hierarchy

Every quantitative figure on this site carries one of four tags. If a number has no tag, that is a bug; tell us.

VERIFIED FACT
An empirically validated historical metric or an enacted statutory provision, with a source. Example: Snowflake’s 14,554 customers as of 31 July 2026 (Form 10-Q).
PLANNING ASSUMPTION
A modelled operating parameter based on industry benchmarks, unvalidated, with a decision gate and a date. Example: 70% automated invariant synthesis.
TARGET
A prescribed internal milestone, acceptance gate or quality standard. Example: the 149-trial zero-miss Sev-1 gate.
ILLUSTRATIVE
A hypothetical case for conceptual modelling. Every sample report, terminal panel and SQL snippet on this site.

Two more tags appear where a claim depends on something outside us: CLAIM: EXTERNAL VERIFICATION REQUIRED (for instance the EU AI Act dates after Regulation 2026/1744) and PENDING ACCREDITATION (IEEE 1012 IV&V).

Team and hiring, gated

Two founders today: one leading methodology and statistical metrology, one leading the modern data stack and delivery. Hiring is staggered on milestones, not on a fundraising story:

DisciplineCurrent ownerHire trigger
Modern data stack (Snowflake / Databricks)Founder, lead data engineerSenior Analytics Engineer at Month 4, on completing two design-partner sprints and closing Pilot #2
Enterprise security & complianceContract CISO advisorFull-time Head of Security at Month 7, on Customer #4 and SOC 2 Type 2 start
Evaluation science & metrologyCo-founder, lead methodologistSenior Evaluation Scientist at $250K contracted ARR and >50 monthly active cases
Reviewer operationsFoundersOperations Manager when the active roster exceeds 50
Enterprise commercialFounder / CEOFirst enterprise AE at 5 deliveries, 3 renewals and $250K ARR

Brands

Contact

Engagements
[email protected] or the readiness review form
Security
[email protected] · security.txt
Privacy
[email protected]
Reviewers
[email protected]
Partners
[email protected]

Methodology → FAQ