YAML configuration
Reference for the lumis.dev/operational-v1alpha1 project file: top-level fields, budgets, the incident and observation formats, and validation rules.
A strict project file
The project format is lumis.dev/operational-v1alpha1. Unknown fields, duplicate keys, YAML aliases, excessive nesting and files over 1 MiB are rejected. Run lumis doctor --project lumis.yaml after every edit; it validates locally and makes no network calls.
Top-level fields
| Field | Purpose |
|---|---|
api_version | Always lumis.dev/operational-v1alpha1. |
project | name and environment. |
sources | Read-only connectors. Each is disabled unless enabled: true. |
identity.aliases | Map discovered IDs to canonical IDs. |
discovery | Limits for building the graph. |
graph | Declared entities and relationships. |
queries | The registered query catalog. |
initial_query_ids | Queries collected before triage. |
checks | Known failure patterns for triage. |
models | Optional model provider and ID. |
investigator | Optional investigator budgets, repositories and sandbox. |
budget | Investigation limits: hops, entities, queries, hypotheses, context and time. |
policies | default_action_mode: read_only, the only accepted mode. |
observations_file | Facts to replay for snapshot queries. |
rule_hypotheses also exists for the lower-level candidate-only baseline (lumis investigate); it does not take part in incident triage.
A complete example
The small-project file is a complete, valid project. Add models and investigator to enable the investigator.
api_version: lumis.dev/operational-v1alpha1
project:
name: my-api
environment: local
sources:
prometheus:
enabled: true
endpoint: http://localhost:9090
policies:
default_action_mode: read_only
graph:
entities:
- id: service:api
kind: service
name: API
relationships: []
queries:
- id: api-up
provider: prometheus
entity_id: service:api
key: up
description: Was the API scrape target up at the end of the incident?
parameters:
promql: 'min(up{job="api"})'
- id: api-probe
provider: prometheus
entity_id: service:api
key: probe_success
description: Did the external HTTP health probe succeed?
parameters:
promql: 'min(probe_success{job="blackbox-api"})'
checks:
- id: api-down
terminal: true
explains_entities: [service:api]
hypothesis:
id: api-unavailable
statement: The API is down; both its scrape target and an external HTTP probe fail.
causal_path: [service:api]
evidence_needed: [api-up, api-probe]
predictions:
- {entity_id: "service:api", key: up, operator: eq, value: 0}
- {entity_id: "service:api", key: probe_success, operator: eq, value: 0}
falsifiers:
- {entity_id: "service:api", key: up, operator: eq, value: 1}
- {entity_id: "service:api", key: probe_success, operator: eq, value: 1}Budgets
discovery:
timeout_seconds: 30
max_entities: 5000
max_relationships: 10000
max_service_graph_series: 1000
max_response_bytes: 2000000
budget:
graph_hops: 3
max_entities: 100
max_queries: 8
max_hypotheses: 5
max_model_output_tokens: 3000
max_context_characters: 20000
query_timeout_seconds: 10
source_timeout_seconds: 30
total_timeout_seconds: 120Example values, not recommendations. When a limit would be exceeded, Lumis refuses rather than silently truncating. Preparation and investigation have separate deadlines.
Incident and observation files
{
"id": "checkout-001",
"affected_entities": ["service:shop:checkout"],
"symptoms": ["Checkout p95 latency above 2s"],
"started_at": "2026-10-05T10:00:00Z",
"ended_at": "2026-10-05T10:20:00Z"
}[
{
"id": "obs-1",
"query_id": "checkout-up",
"entity_id": "service:shop:checkout",
"key": "up",
"value": 0,
"observed_at": "2026-10-05T10:20:00Z",
"source": "replay",
"retrieval_method": "snapshot-replay"
}
]Timestamps must include a timezone. Each observation must match a registered query's entity and key and fall inside the incident window.
Source: YAML reference ↗ in the SDK repository.