Your first real project
Connect Lumis to one service and a Prometheus server, write a deterministic check with two independent signals, and add the optional investigator.
What you will build
One service (service:api), one Prometheus server and one check that answers a single question: is the API down? The example is the SDK's own small-project example ↗, which is exercised by its test suite.
pip install "lumis-sdk[http]"The project file
api_version: lumis.dev/operational-v1alpha1
project:
name: my-api
environment: local
sources:
prometheus:
enabled: true
endpoint: http://localhost:9090
policies:
default_action_mode: read_only
graph:
entities:
- id: service:api
kind: service
name: API
relationships: []
queries:
- id: api-up
provider: prometheus
entity_id: service:api
key: up
description: Was the API scrape target up at the end of the incident?
parameters:
promql: 'min(up{job="api"})'
- id: api-probe
provider: prometheus
entity_id: service:api
key: probe_success
description: Did the external HTTP health probe succeed?
parameters:
promql: 'min(probe_success{job="blackbox-api"})'
checks:
- id: api-down
terminal: true
explains_entities: [service:api]
hypothesis:
id: api-unavailable
statement: The API is down; both its scrape target and an external HTTP probe fail.
causal_path: [service:api]
evidence_needed: [api-up, api-probe]
predictions:
- {entity_id: "service:api", key: up, operator: eq, value: 0}
- {entity_id: "service:api", key: probe_success, operator: eq, value: 0}
falsifiers:
- {entity_id: "service:api", key: up, operator: eq, value: 1}
- {entity_id: "service:api", key: probe_success, operator: eq, value: 1}| Section | What it says |
|---|---|
sources.prometheus | Where to read metrics. Lumis runs read-only instant queries. |
graph.entities | The one service this project knows about. |
queries | Two operator-owned PromQL queries. Each produces one named fact about the service: up and probe_success. |
checks | One signature: “the API is down” predicts both facts are 0, and is falsified if either is 1. |
The model never writes PromQL. If you later enable the investigator, it can only ask for these queries by ID.
Run it
{
"id": "api-001",
"affected_entities": ["service:api"],
"symptoms": ["Health check failing"],
"started_at": "2026-10-05T05:00:00Z",
"ended_at": "2026-10-05T05:10:00Z"
}lumis doctor --project lumis.yaml
lumis incident --project lumis.yaml --incident incident.jsonQueries are evaluated at the end of the incident window. The check produces one of three findings:
| Finding | When | What happens |
|---|---|---|
match | Both facts are 0 | The check is terminal, so triage concludes supported_diagnosis with no model call. |
no_match | A falsifier holds: the API is up | The incident goes to a person, or to the investigator if enabled. |
unknown | A query returned no data | The same. Missing data is never read as “down”. |
Why the check looks like this
Lumis will not end triage on weak evidence, and the YAML enforces it. A terminal check must name the entities it explains (explains_entities) and predict at least two distinct facts from independent queries. With only one signal, set terminal: false: the check becomes a lead instead of a conclusion.
Both signals here report a real 0 when the API is down: Prometheus writes up = 0 for a failed scrape, and a blackbox HTTP probe writes probe_success = 0. A request rate would not work. When the process is down it stops producing samples, the query returns nothing, and Lumis records unknown.
Replayed facts (--observations) apply only to provider: snapshot queries, as in the lumis init scaffold. They do not override live Prometheus queries.
Add the investigator
When no check matches, a model can look further, still read-only and bounded. Add a model and budget to the project:
models:
provider: openrouter # or openai, anthropic, gemini
model: deepseek/deepseek-v4-pro
api_key_env: OPENROUTER_API_KEY
reasoning: high
investigator:
budget:
request_limit: 12pip install "lumis-sdk[http,agent]"
export OPENROUTER_API_KEY=...
lumis incident --project lumis.yaml --incident incident.json --use-agentConfiguration alone never makes a paid call; only --use-agent does. See the investigator for what it can and cannot do, and model providers for choosing a model.
Grow it
- More signals: logs from Loki, traces from Tempo, workflow runs from Prefect, read-only SQL, and recent Git commits and Kubernetes rollouts. See connectors.
- More services: add entities and relationships, or let Kubernetes and the Prometheus service graph discover them. See the operational graph.
- Code context: allowlist the files and Git history the investigator may read. See the investigator.
For a large worked example with seven sources, 40 queries and ten checks, see the GridCast cookbook ↗. Smaller cookbooks are planned.
Source: Small-project guide ↗ in the SDK repository.