Compare

EvalQA vs Langfuse

Langfuse solves a problem we do not touch, and we solve one it structurally cannot. If you are already self-hosting it, you are in a good position — this page is about what is still missing.

Last updated 20 September 2026 · Protocol v9.3

Architectural comparison · last reviewed 20 September 2026 · tell us if this is wrong

What Langfuse is genuinely good at

Open-source tracing and scoring you can self-host, keeping every prompt and completion inside your own infrastructure. For teams with data-residency constraints or an aversion to per-seat SaaS, that is a real structural advantage, and it publishes its unit economics openly — a standard we think more of this category should meet.

What self-hosting does not change

Ownership of the tooling is not independence of the judgement. Whoever operates Langfuse still writes the scorers, picks the thresholds and decides what a pass means. Self-hosting removes the vendor from your data path; it does not put a second pair of eyes on the agent.

For most systems that is fine. It stops being fine at the point where somebody outside your team has to rely on the result — and that is a governance requirement, not a tooling one.

LangfuseEvalQA
DeploymentSelf-hosted or cloudRuns inside your perimeter; only hashes and status leave
Data residencyFully yoursFully yours in the default evidence mode (modes)
Who authors the scoringYour teamUs, plus invariants compiled from your dbt manifest
Licence costOpen sourceFixed annual fee
Operational burdenYoursOurs
Third-party verifiable outputNoYes — signed bundle + notary

Where we agree with Langfuse’s philosophy

On lock-in. OpenAI is shutting its Evals platform on 30 November 2026 and Humanloop was absorbed into Anthropic; current buyer guidance is to run an export-and-recreate drill before signing anything. We publish a documented non-proprietary export schema and a written exit guide for the same reason open source is attractive: the ability to leave is what makes staying a choice.

Pick Langfuse if

Add EvalQA if

All comparisons How to leave us →