Concepts

Evidence and queries

How Lumis collects facts: operator-registered queries, typed observations with provenance, the incident time window, and why missing data is never a zero.

v0.1.0 · experimentalPython 3.11+Updated 2026-10-05

Registered queries

Every fact Lumis uses comes from a query you registered in lumis.yaml. A query has an ID, a provider, the entity it describes, the key of the fact it produces, a description, and provider parameters such as PromQL or SQL. Neither a check nor the model can run anything else.

yaml
queries:
  - id: checkout-error-logs
    provider: loki
    entity_id: service:shop:checkout
    key: error_entries
    description: Checkout error log lines during the incident
    parameters:
      logql: '{namespace="shop", app="checkout"} |= "error"'
      output: count
ProviderTypical use
prometheusMetrics as one instant scalar at the end of the incident.
lokiLog entries or counts in the incident window.
tempoTrace searches, durations and span reads.
prefectFlow and task run states and durations.
sqlOne read-only scalar from PostgreSQL.
changesRecent Git commits and Kubernetes rollouts for an entity.
snapshotFacts replayed from a file (tests, offline runs).
probeResults of an opt-in sandbox experiment (degraded quality).

Details for each provider are in connectors.

What an observation contains

FieldMeaning
query_id, entity_id, keyWhich registered query produced it, about which entity, for which fact. Must match the registration.
valueOne scalar: number, boolean or short text. Booleans are never treated as 1 or 0.
observed_atWhen it was observed. Must fall inside the incident window.
source, retrieval_methodProvenance: where it came from and how.
qualityobserved, or degraded when a result was capped, partial or synthetic.

Missing data is unknown, never zero

If a query fails, times out or returns nothing, no fact is recorded and any check that needed it stays unknown. A failed query is not a false measurement. This sounds obvious, but it is the most common way investigation tools go wrong.

Degraded facts (a capped log result, a sandbox probe) can guide a person or the investigator, but cannot satisfy a terminal check.

Time window and budgets

Queries run against the incident window, started_at to ended_at: Prometheus instant queries are evaluated at ended_at, and log, trace, workflow and SQL queries are bounded by the window. List queries in initial_query_ids to collect them before anything else runs.

Triage and the investigator share one query budget (budget.max_queries). Repeated requests for the same query reuse the first result; failed queries are not retried invisibly.

Source: YAML reference ↗ in the SDK repository.