Evaluation on the GridCast estate
What we measured when we ran Lumis against fifteen injected failures on a live reference estate, how it compares with simpler approaches, and what went wrong.
The estate and the protocol
GridCast is an open reference estate we built for this purpose: a synthetic electricity-demand forecaster with ten services on Kubernetes (kind), Prometheus, Loki, Tempo, Prefect and PostgreSQL, external weather and telemetry vendors, and changes made through GitOps. We injected fifteen different failures through ordinary channels (releases, configuration, rotated secrets, resource limits, a model promotion, vendor outages and silent data problems). Five of them were added later, specifically to be hard. Everything is public in the GridCast cookbook ↗.
For each failure the estate's own alerts opened one incident, which was then frozen: every system answered the same incident with the same evidence window. The ground truth was read only after all reports existed. A run counts as correct only if its top answer names the right component and the right mechanism.
What was compared
| System | What it gets |
|---|---|
| Rules only | Lumis' deterministic checks, no model. |
| Model, alert only | The alert, the affected services and the time window. Like pasting an alert into a chat assistant. |
| Model + graph | The same, plus the service graph. |
| One model call with evidence | The graph plus the facts Lumis collected for its checks, in one completion. |
| Lumis | Triage, then the investigator, then mechanical assessment. |
Every model-based system used the same model: DeepSeek v4 pro through OpenRouter, reasoning effort high.
Results
| Rules only | Model, alert only | Model + graph | One call with evidence | Lumis | |
|---|---|---|---|---|---|
| Correct (component and mechanism) | 0.57 | 0.04 | 0.18 | 0.54 | 0.89 |
| Hard set (silent, partial, decoy faults) | 0 / 20 | 0 / 8 | 2 / 8 | 2 / 8 | 7 / 8 |
| Model cost per run | $0 | $0.005 | $0.015 | $0.035 | $0.14 |
| Median time per run | 0.13 s | — | — | 224 s | 225 s |
Figures exclude one scenario (N) after we found a ground-truth leak; with it, Lumis scored 0.90. Lumis concluded 22 times on the remaining scenarios and was right 21 times. When evidence did not settle a question, it said so instead of guessing. The rules score reflects leads: rules concluded only on the one scenario with a sufficient terminal check, and were right every time.
- Without evidence, the model guesses. Given only the alert, it named the right cause and mechanism once in 28 runs. It usually found the symptomatic service and invented a plausible mechanism.
- Evidence does most of the work. One evidence-fed call reached 0.54. The largest single jump is from no evidence to curated evidence.
- Reach matters on hard faults. The facts that decided the hard scenarios (a timeout commit, a CPU-limit commit, a missing zone) were not in any pre-collected bundle. An investigator that can ask for more found them.
- A capable model with raw tools matched Lumis on the one clean scenario we could compare (2 / 2 each), but needed about twice the tool calls and model requests, and its answer was unchecked free text.
What went wrong
- Our estate produced false alerts in the first run. Two processes shared one telemetry identity and corrupted rate calculations. We fixed the estate and re-ran everything.
- False zeros. All four wrong conclusions in the first run rested on telemetry that reported a zero that was really missing data (see evidence).
- A ground-truth leak, found after the runs: a docstring in an allowlisted file named one scenario, and the unguided tool agent could read our own development history. That scenario is excluded, and a test now guards against labels in readable files.
- Rubric revisions. The regex rubric that scores mechanisms was revised twice after we inspected outputs; each revision applied to every system.
- Thirteen SDK defects, from redaction masking decimals to a routing parameter that broke one provider, were found and fixed during the integration. They are the bulk of what changed before 0.1.0.
Limits
Every number, transcript and correction is in the research notes ↗. The estate, harness and evaluation were built with AI assistance (Claude Opus 5.5); the research design and most scenarios came from the author. Claude was not one of the systems evaluated.