Checks and triage
Write known failure patterns as deterministic checks, understand match, no_match and unknown, and the strict rules for when triage may conclude without a model.
A check is a falsifiable known pattern
A check describes a failure you already know how to recognise, written as a hypothesis: what it explains, which facts it needs, what those facts should look like if it is happening, and what would prove it is not.
checks:
- id: pod-oom-killed
terminal: false
hypothesis:
id: oom-killed
statement: A container was OOM-killed during the incident.
causal_path: [service:shop:checkout]
evidence_needed: [checkout-oom-kills]
predictions:
- {entity_id: "service:shop:checkout", key: oom_kills, operator: gt, value: 0}
falsifiers:
- {entity_id: "service:shop:checkout", key: oom_kills, operator: eq, value: 0}Operators are eq, ne, gt, ge, lt and le. Ordered comparisons need numbers.
Findings
| Finding | Meaning |
|---|---|
match | Every prediction is supported by a usable fact and no falsifier holds. |
no_match | A prediction fails, or a falsifier holds. |
unknown | A needed fact is missing, degraded or conflicting. |
When triage may conclude
A matched check ends triage only if all five conditions hold:
- It is marked
terminal: trueand declaresexplains_entities. - It predicts at least two distinct entity/key facts, supplied by independent queries.
- Its predictions are supported and no falsifier holds.
- It covers every affected entity, and every other check in scope is contradicted.
- An optional caller-supplied
TriageGuardalso accepts it.
Writing good checks
- Start nonterminal. A matched nonterminal check is a lead handed to the investigator or to a person.
- Make a check terminal only when two independent signals both point the same way, and both report a real value when the failure happens (see missing data).
- Write falsifiers. A check that cannot be contradicted cannot be tested.
- Keep one check per mechanism. Overlapping checks that both match will block a terminal conclusion, by design.
In the GridCast evaluation, ten checks covered the ten original scenarios. Only one was terminal: “planning API scaled to zero”, which concluded in about 60 milliseconds and was right every time. The other nine were leads.
Source: Incident investigation ↗ in the SDK repository.