What is a detector error model?
If you have run a Stim simulation, you have almost certainly produced a detector error model. This post is a tour of that object: what a DEM is, why decoders consume it, and why it deserves more scrutiny in QEC pipelines.
A detector error model (DEM) is best understood as a compressed description of the way physical errors manifest themselves in the syndrome data of a fault-tolerant quantum circuit.
Instead of describing which physical Pauli error happened on which qubit, a DEM describes which detectors would flip, with what probability, and whether a logical observable would flip.
1. What is a detector?
In a QEC circuit, a detector is a parity check on measurement results that is guaranteed to be 0 in the absence of errors.
For example, if a stabilizer is measured repeatedly,
$$ s_t = \text{measurement of stabilizer at round } t, $$then one can define a detection event
$$ d_t = s_t \oplus s_{t-1}. $$Normally $d_t = 0$. An error between the two measurements can make $d_t = 1$ [3].
Each detector is a node $D_i$ in the model. The boolean $d_t$ is the bit a detector produces in a given shot (0 or 1), while $D_i$ is the detector itself. A set of flipped bits $\{d_t = 1\}$ is a set of triggered nodes $\{D_i\}$.
So instead of looking at a huge stream of raw measurement bits, we transform it into a sparse set of detection events (often called defects).
2. What is in a detector error model?
A DEM essentially contains entries of the form
$$ p : D_i, D_j, \ldots, L_k $$each of which is called an error mechanism (Stim calls these error instructions; in the graph picture below, error hyperedges), meaning:
With probability $p$, this error mechanism produces detection events at $D_i, D_j, \ldots$, and potentially flips logical observable $L_k$.
For example:
error(0.001) D5 D6 # fires two detectors — a graphlike "edge"
error(0.002) D6 D7
error(0.0001) D10 D11 L0 # fires two detectors AND flips logical qubit 0
The first mechanism fires detectors D5 and D6. The last fires two detectors and flips a logical observable — a logical error.
Stim's definition is exactly this [1]: a DEM is a collection of error mechanisms, each with a probability, the detectors it flips ("symptoms"), and possible logical-frame changes. Note that an error mechanism is not an arbitrary Pauli error but its downstream effect: it records the symptom signature (which detectors and observables flip) and forgets which qubit, gate, or Pauli type produced it.
This gives you something very close to a weighted graph/hypergraph:
- vertices → detectors
- edges → error mechanisms
- edge weights → error probabilities
- observables → markers on edges saying which error mechanisms flip a logical qubit (often drawn as edges to a boundary node)
For graphlike errors, this becomes particularly intuitive: an ordinary single-qubit error often produces two detection events, so it corresponds to an edge connecting two detectors. More complicated errors can produce >2 detectors and therefore correspond to hyperedges. Stim explicitly supports this distinction and can suggest decompositions of hypererrors into graphlike errors.
The idea of decoding as a matching problem on detection events appears in the QEC literature [3], and the DEM formalism itself is now used beyond Stim, e.g. for circuit design [5]. The explicit .dem file format was introduced by Stim [1] and has become the de facto interchange format, since most QEC tools integrate Stim.
3. Decoding with detector error models
Its primary purpose is decoding.
Suppose your hardware gives you:
001000010000101...
of syndrome measurements.
Using the parity checks from §1, you convert that raw record into detection events:
D17 D18 D43 D44 ...
Given these detection events and the model of the error mechanisms, what is the most likely underlying error / logical correction? A matching decoder such as minimum-weight perfect matching (MWPM) essentially tries to find the most likely collection of edges in the detector graph that explains the observed defects [2].
That's why DEMs are so important to tools such as Stim and PyMatching. Stim can automatically transform a noisy stabilizer circuit into a DEM suitable for configuring a decoder [1].
Essentially, the DEM is an interface between the physical implementation and the decoder.
4. DEMs as an intermediate representation
You can formulate QEC without ever talking about DEMs. However, for large-scale numerical QEC, it is enormously useful. Being a lossy compression format, a DEM throws away a large amount of irrelevant information.
Consider a surface-code memory with thousands of physical qubits and thousands of gates. A physical noise model might specify every gate, qubit, Pauli channel, measurement error, reset error, and correlations between operations, etc. The decoder doesn't necessarily need all of that.
Concretely, a DEM forgets:
- which qubit experienced the error,
- when, and through which operation (timing and gate type),
- the Pauli type of the error (bitflip? phaseflip? a measurement error?).
This is by design: the decoder only cares about which defects were triggered and whether the logical frame flips, so all information about how and where the error happened is irrelevant to it.
The DEM describes which patterns of syndrome defects errors can produce, and with which probability. This makes the DEM a key intermediate representation for decoding — and, increasingly, a design formalism in its own right: the same framework can be used to derive robust syndrome-extraction circuits, measurement schedules, and fault-tolerant logical operations [5].
However, this lossy compression format introduces limitations: two physically very different errors can be equivalent from the decoder's perspective if they produce the same syndrome and logical frame effect.
5. The catch: a wrong DEM looks exactly like a right one
Everything above assumes the DEM is correct, i.e. that it faithfully describes what the circuit's noise actually does. But where does a DEM come from? Stim derives it from two inputs: the circuit and the noise model. Both are written by humans, or generated by compilers, and both can be wrong.
When the inputs are wrong, the DEM is still perfectly valid syntax. It just silently describes the wrong noise. For example:
- A detector that no error mechanism can ever fire. Usually a sign that a noise channel was forgotten somewhere. The decoder receives a permanently-silent check and loses information for free.
- An observable that no mechanism flips. The decoder has no way to protect that logical qubit — every error that affects it is invisible to the model.
- Two mechanisms with identical detector and observable signatures. Often a missed decomposition or a duplicated noise channel.
- Probabilities outside $[0,1]$, or wildly larger than the surrounding mechanisms. A typo in a noise parameter propagates straight through.
None of these are syntax errors. None will crash your simulation. The decoder simply underperforms, and you find out the expensive way: a logical error rate curve that is subtly too high, visible only after $10^6$ shots.
This is exactly the situation classical software was in before static analysis: bugs that compile cleanly and fail silently in production. The fix: linters, type checkers, and dataflow analysis that read the code and find structural problems in milliseconds, without executing it.
emlint applies that idea to detector error models. It parses a DEM and verifies structural properties: every detector is reachable by some error mechanism, every observable is covered, no duplicate mechanisms, probabilities in range, and more. Each failure comes with a counter-example identifying the offending mechanism, so you fix the circuit instead of guessing. It runs in milliseconds, so it fits in CI next to your unit tests — a bug that would take 45 minutes of Monte Carlo sampling to surface is caught before the simulation starts. This is part of a broader program of treating DEMs as a formal object with static guarantees — for example, quasilinear-time equivalence checking of DEM terms [6].
The DEM sits at the interface between the physical implementation and the decoder. It is worth checking what crosses that interface.
stim simulates. sinter samples. emlint verifies.
References
- Craig Gidney. Stim: a fast stabilizer circuit simulator. Quantum 5, 497 (2021). arXiv:2103.02202 — introduces the detector error model formalism and the circuit-to-DEM transformation. The Stim documentation details the
.demfile format (format specification) andDETECTOR/OBSERVABLE_INCLUDEannotations. - Oscar Higgott, Craig Gidney. Sparse Blossom: correcting a million errors per core second with minimum-weight matching. Quantum 9, 1600 (2025). arXiv:2303.15933 — the matching decoder (PyMatching) that consumes DEMs as its input model.
- Eric Dennis, Alexei Kitaev, Andrew Landahl, John Preskill. Topological quantum memory. Journal of Mathematical Physics 43, 4452 (2002). arXiv:quant-ph/0110143 — the original formulation of error correction as a matching problem on detection events, the conceptual ancestor of the detector graph.
- emlint. Static linter for Stim detector error models. PyPI — verifies the structural properties discussed in §5.
- Peter-Jan H.S. Derks, Alex Townsend-Teague, Ansgar G. Burchards, Jens Eisert. Designing fault-tolerant circuits using detector error models. Quantum 9, 1905 (2025). arXiv:2407.13826 — a pedagogical introduction to the DEM formalism, and a demonstration that DEMs can drive circuit design, not just decoding.
- Mathys Rennela. Quasilinear Equivalence Checking for Detector Error Models. arXiv:2606.14677 (2026) — a sound, terminating, confluent rewriting system for DEMs with a quasilinear-time normal form, giving the first static decision procedure for DEM equivalence.