Concepts

How Lumis works

The end-to-end incident flow, who decides what, and the vocabulary used across the SDK, explained in plain language.

v0.1.0 · experimentalPython 3.11+Updated 2026-10-05

The flow, end to end

  • Prepare. Lumis builds an operational graph from what you declare and what it can discover (Kubernetes, Prometheus service graphs, Tempo, Prefect), then keeps only the part around the affected services.
  • Triage. Your checks are tested against facts from your registered queries. Each check either matches, does not match, or is unknown. A sufficient terminal check ends here, without a model.
  • Investigate (optional). If triage is not sufficient and you enabled it, one tool-using model explores: it reads the graph, asks for evidence by query ID, reads recent changes and allowlisted code, and proposes explanations.
  • Assess. Every explanation, from a check or from the model, states what it predicts and what would prove it wrong. Lumis checks those statements against the facts it collected itself.
  • Report. The result is one structured report for a person. Nothing is executed.

Who decides what

RoleDecides
You, the operatorWhich sources Lumis may read, which queries exist, which checks to run, which files the investigator may see, and every budget.
The model (optional)Which registered queries to ask for, and which explanations to propose. Nothing else.
LumisWhether each explanation is supported, contradicted or unresolved, and whether the evidence is enough to conclude.
A personWhat actually happened, and what to do about it.

The working rule is “the model proposes; Lumis tests”. A model cannot add evidence, mark something as confirmed, or take an action.

Vocabulary

TermMeaning
IncidentThe affected entities and a time window. The input to every investigation.
EntitySomething in your estate with a stable ID, for example service:shop:checkout or k8s:shop:deployment:checkout.
Operational graphEntities and their relationships. It scopes what Lumis looks at; it is not proof of cause.
Registered queryAn operator-written query (PromQL, LogQL, SQL…) with an ID, an entity and a key. The only way facts are collected.
Observation (fact)One value for one entity and key at one time, with its source. For example, up = 0 for service:api.
CheckA known failure pattern written as a hypothesis. Tested deterministically during triage.
FindingA check's result: match, no_match or unknown.
HypothesisA falsifiable explanation: a statement, a causal path through the graph, predictions and falsifiers.
Prediction / falsifierFacts the explanation expects to see, and facts that would contradict it.
AssessmentLumis' mechanical verdict on a hypothesis: supported, contradicted or unresolved.
ReceiptA redacted record of every tool call and query, kept in the report.
Conclusionsupported_diagnosis, insufficient_evidence or requires_human_expert.

Design principles

  • Deterministic first. Known failures are handled by checks. A model is used only for what checks cannot explain.
  • Evidence first. Facts come only from operator-registered queries, with provenance and time-window checks.
  • Uncertainty stays visible. Missing, conflicting or degraded facts leave an explanation unresolved; they never become support.
  • One cause or no conclusion. If supported explanations disagree on the root cause, the report says so instead of picking one.
  • Bounded. Graph size, queries, model requests, tool calls, tokens and time all have limits.
  • Read-only. There is no executor. A person keeps every decision.

Source: Architecture ↗ in the SDK repository.