What transfers
- Your test cases, as inputs. Anything expressible as a question plus an expected relation can seed an invariant.
- Your golden answers, as a definition source — subject to the 30% independent-case floor, because a suite built only from your own expectations tests your imagination.
- Confirmed defect history, as regression cases.
What does not
- Scores and thresholds. Different rubric, different severity anchors, different adjudication. A score from another tool is not comparable to ours and we will not pretend otherwise by mapping it.
- LLM-judge prompts. Our judge runs against our rubric; importing someone else’s prompt imports their assumptions.
- Trace history. We are not an observability platform and cannot be your system of record for traces. Keep that tool, or keep its export.
A parallel period is the right shape
Run both for one cycle. The comparison you want is not which tool scores higher — scores are not comparable — but which defects one found that the other did not. That is the same Incremental Consequential Finding Yield measurement we hold ourselves to against native warehouse evaluation, and it is the only honest way to compare two assurance methods.
Before you sign with anyone, including us
Current buyer guidance after the OpenAI and Humanloop closures is to attempt an export-and-recreate before committing. Ours is documented: case-level records with input, configuration, output, tool trajectory, error states and cost, in a non-proprietary schema, with a written guide to leaving. Portability & exit →
If a vendor cannot show you this in a 30-minute call, that is your answer about what leaving will look like later.
If we are not the right destination
The comparison page names the situations where a platform or open-source tracing is the better answer, and where native warehouse evaluation is enough on its own.