Build

YAML configuration

Reference for the lumis.dev/operational-v1alpha1 project file: top-level fields, budgets, the incident and observation formats, and validation rules.

v0.1.0 · experimentalPython 3.11+Updated 2026-10-05

A strict project file

The project format is lumis.dev/operational-v1alpha1. Unknown fields, duplicate keys, YAML aliases, excessive nesting and files over 1 MiB are rejected. Run lumis doctor --project lumis.yaml after every edit; it validates locally and makes no network calls.

Top-level fields

FieldPurpose
api_versionAlways lumis.dev/operational-v1alpha1.
projectname and environment.
sourcesRead-only connectors. Each is disabled unless enabled: true.
identity.aliasesMap discovered IDs to canonical IDs.
discoveryLimits for building the graph.
graphDeclared entities and relationships.
queriesThe registered query catalog.
initial_query_idsQueries collected before triage.
checksKnown failure patterns for triage.
modelsOptional model provider and ID.
investigatorOptional investigator budgets, repositories and sandbox.
budgetInvestigation limits: hops, entities, queries, hypotheses, context and time.
policiesdefault_action_mode: read_only, the only accepted mode.
observations_fileFacts to replay for snapshot queries.

rule_hypotheses also exists for the lower-level candidate-only baseline (lumis investigate); it does not take part in incident triage.

A complete example

The small-project file is a complete, valid project. Add models and investigator to enable the investigator.

yaml
api_version: lumis.dev/operational-v1alpha1
project:
  name: my-api
  environment: local

sources:
  prometheus:
    enabled: true
    endpoint: http://localhost:9090

policies:
  default_action_mode: read_only

graph:
  entities:
    - id: service:api
      kind: service
      name: API
  relationships: []

queries:
  - id: api-up
    provider: prometheus
    entity_id: service:api
    key: up
    description: Was the API scrape target up at the end of the incident?
    parameters:
      promql: 'min(up{job="api"})'
  - id: api-probe
    provider: prometheus
    entity_id: service:api
    key: probe_success
    description: Did the external HTTP health probe succeed?
    parameters:
      promql: 'min(probe_success{job="blackbox-api"})'

checks:
  - id: api-down
    terminal: true
    explains_entities: [service:api]
    hypothesis:
      id: api-unavailable
      statement: The API is down; both its scrape target and an external HTTP probe fail.
      causal_path: [service:api]
      evidence_needed: [api-up, api-probe]
      predictions:
        - {entity_id: "service:api", key: up, operator: eq, value: 0}
        - {entity_id: "service:api", key: probe_success, operator: eq, value: 0}
      falsifiers:
        - {entity_id: "service:api", key: up, operator: eq, value: 1}
        - {entity_id: "service:api", key: probe_success, operator: eq, value: 1}

Budgets

yaml
discovery:
  timeout_seconds: 30
  max_entities: 5000
  max_relationships: 10000
  max_service_graph_series: 1000
  max_response_bytes: 2000000
budget:
  graph_hops: 3
  max_entities: 100
  max_queries: 8
  max_hypotheses: 5
  max_model_output_tokens: 3000
  max_context_characters: 20000
  query_timeout_seconds: 10
  source_timeout_seconds: 30
  total_timeout_seconds: 120

Example values, not recommendations. When a limit would be exceeded, Lumis refuses rather than silently truncating. Preparation and investigation have separate deadlines.

Incident and observation files

json
{
  "id": "checkout-001",
  "affected_entities": ["service:shop:checkout"],
  "symptoms": ["Checkout p95 latency above 2s"],
  "started_at": "2026-10-05T10:00:00Z",
  "ended_at": "2026-10-05T10:20:00Z"
}
json
[
  {
    "id": "obs-1",
    "query_id": "checkout-up",
    "entity_id": "service:shop:checkout",
    "key": "up",
    "value": 0,
    "observed_at": "2026-10-05T10:20:00Z",
    "source": "replay",
    "retrieval_method": "snapshot-replay"
  }
]

Timestamps must include a timezone. Each observation must match a registered query's entity and key and fall inside the incident window.

Source: YAML reference ↗ in the SDK repository.