Concepts

The operational graph

How Lumis models your estate as entities and relationships, how identities work, how the graph is discovered and how it is scoped to each incident.

v0.1.0 · experimentalPython 3.11+Updated 2026-10-05

Why a graph

An incident rarely starts where the alert fires. A slow pipeline may be caused by a feature service, its database, or a vendor feeding it. The operational graph records what depends on what, so an investigation can look upstream of the symptom without reading the whole estate. It is context for reasoning; a connected path is never treated as proof of cause.

Under the hood it is a Pydantic-validated snapshot loaded into a NetworkX MultiDiGraph, so parallel relationships of different kinds between the same two entities are kept. No graph database is required.

Where the graph comes from

Declared entities carry what discovery cannot know: owners, criticality, external vendors, idle dependencies. Enabled sources add what exists right now. If an enabled source fails, preparation stops with a sanitized report rather than continuing with a partial graph.

Identities and directions

Identity or relationshipConvention
Logical serviceservice:<namespace>:<name>, from app labels, OpenTelemetry or service-graph metrics.
Kubernetes resourcek8s:<namespace>:<kind>:<name>. Kept distinct from the logical service.
hostsResource → logical service. Says where a service runs, not that it was called.
servesServer or dependency → client. So “upstream of the frontend” includes its database.
Declared lineageDataset → job → dataset (feeds, produces), with explicit provenance.
AliasesExplicit raw ID → canonical ID. Bare names are never guessed or joined.

Namespaces keep two estates from merging by accident. Conflicting kinds or metadata fail validation instead of being silently resolved.

Scoping to an incident

For each incident Lumis keeps the neighbourhood of the affected entities, bounded by budget.graph_hops and budget.max_entities. If the neighbourhood would exceed the entity budget, preparation fails rather than truncating silently. Both triage and the investigator work only inside this scope.

python
from lumis_sdk.runtime import YamlProject

prepared = await YamlProject.from_file("lumis.yaml").prepare()
graph = prepared.graph
upstream = graph.upstream_of("service:shop:checkout", hops=3, max_entities=100)
local = graph.dependencies_within("service:shop:checkout", hops=2, max_entities=100)
scoped = graph.scope(["service:shop:checkout"], hops=2, max_entities=100)
nx_copy = graph.to_networkx()
dot = graph.to_dot()

Use top-level await in a notebook, or wrap the code in asyncio.run(main()) in a script. Exports are deep copies, so changing them cannot change an investigation.

Limits

  • Live Kubernetes discovery reads current resources, not historical cluster state. Archive topology snapshots yourself if you need exact replay.
  • Recent Git commits and rollouts are evidence about entities, not graph nodes. See evidence and queries.
  • There is no live OpenTelemetry receiver and no OpenLineage ingestion yet.

Source: Graph and lineage ↗ in the SDK repository.