How Lumis works
The end-to-end incident flow, who decides what, and the vocabulary used across the SDK, explained in plain language.
The flow, end to end
- Prepare. Lumis builds an operational graph from what you declare and what it can discover (Kubernetes, Prometheus service graphs, Tempo, Prefect), then keeps only the part around the affected services.
- Triage. Your checks are tested against facts from your registered queries. Each check either matches, does not match, or is unknown. A sufficient terminal check ends here, without a model.
- Investigate (optional). If triage is not sufficient and you enabled it, one tool-using model explores: it reads the graph, asks for evidence by query ID, reads recent changes and allowlisted code, and proposes explanations.
- Assess. Every explanation, from a check or from the model, states what it predicts and what would prove it wrong. Lumis checks those statements against the facts it collected itself.
- Report. The result is one structured report for a person. Nothing is executed.
Who decides what
| Role | Decides |
|---|---|
| You, the operator | Which sources Lumis may read, which queries exist, which checks to run, which files the investigator may see, and every budget. |
| The model (optional) | Which registered queries to ask for, and which explanations to propose. Nothing else. |
| Lumis | Whether each explanation is supported, contradicted or unresolved, and whether the evidence is enough to conclude. |
| A person | What actually happened, and what to do about it. |
The working rule is “the model proposes; Lumis tests”. A model cannot add evidence, mark something as confirmed, or take an action.
Vocabulary
| Term | Meaning |
|---|---|
| Incident | The affected entities and a time window. The input to every investigation. |
| Entity | Something in your estate with a stable ID, for example service:shop:checkout or k8s:shop:deployment:checkout. |
| Operational graph | Entities and their relationships. It scopes what Lumis looks at; it is not proof of cause. |
| Registered query | An operator-written query (PromQL, LogQL, SQL…) with an ID, an entity and a key. The only way facts are collected. |
| Observation (fact) | One value for one entity and key at one time, with its source. For example, up = 0 for service:api. |
| Check | A known failure pattern written as a hypothesis. Tested deterministically during triage. |
| Finding | A check's result: match, no_match or unknown. |
| Hypothesis | A falsifiable explanation: a statement, a causal path through the graph, predictions and falsifiers. |
| Prediction / falsifier | Facts the explanation expects to see, and facts that would contradict it. |
| Assessment | Lumis' mechanical verdict on a hypothesis: supported, contradicted or unresolved. |
| Receipt | A redacted record of every tool call and query, kept in the report. |
| Conclusion | supported_diagnosis, insufficient_evidence or requires_human_expert. |
Design principles
- Deterministic first. Known failures are handled by checks. A model is used only for what checks cannot explain.
- Evidence first. Facts come only from operator-registered queries, with provenance and time-window checks.
- Uncertainty stays visible. Missing, conflicting or degraded facts leave an explanation unresolved; they never become support.
- One cause or no conclusion. If supported explanations disagree on the root cause, the report says so instead of picking one.
- Bounded. Graph size, queries, model requests, tool calls, tokens and time all have limits.
- Read-only. There is no executor. A person keeps every decision.
Source: Architecture ↗ in the SDK repository.