The investigator
The optional tool-using model: when it runs, the tools it can use, its budgets, how its output is validated, and what it can never do.
When it runs
Only when triage is not sufficient and you opted in, with --use-agent on the CLI or use_agent=True in Python. It is one Pydantic AI agent with typed output, not a multi-agent system. It starts from the triage findings and the evidence already collected.
What it can do
The investigator has two tools, inspect and probe. inspect has a fixed set of operations:
| Operation | What it allows |
|---|---|
catalog | List the registered query IDs, approved files and enabled capabilities. |
graph | Read a one-hop neighbourhood inside the incident scope. |
evidence | Run a registered query by ID. It cannot write PromQL, LogQL, TraceQL or SQL. |
changes | List recent commits and rollouts that touch scoped entities, newest first. |
code.read, code.search | Read or literally search an explicit allowlist of text files. |
git.log, git.diff | Fixed, read-only Git commands on approved paths. |
hypothesis.register | Register a falsifiable explanation before testing it. |
probe runs a generated Python experiment in a disabled-by-default, network-less container. Its results are marked degraded. See safety and limits.
Giving it code context
investigator:
repositories:
- id: application
root: ./approved-source
entity_ids: ["service:shop:checkout"]
files: [src/checkout/handler.py, deploy/releases.yaml]
include_commit_subjects: falseEach repository is mapped to the entities it belongs to, and only the listed files are readable. Files are snapshotted once per investigation, redacted and hashed. Paths outside the list, symlinks, binary or oversized files, and secret or dot directories are refused.
Budgets
investigator:
budget:
request_limit: 8
tool_calls_limit: 10
max_probes: 2
output_tokens_limit: 12000
max_tool_characters: 8000
max_total_tool_characters: 32000
validation_retries: 2These are example values. Every tool attempt, including a denied or malformed one, counts. When a budget runs out, the evidence and hypotheses collected so far are still assessed and reported.
How its output is checked
The model's final answer is checked against the same rules Lumis applies afterwards: graph IDs must exist, evidence_needed must name registered queries, a revised hypothesis needs a new ID, and suggestions may cite only evidence Lumis actually issued. Problems are sent back to the model to repair, up to validation_retries times. Anything still invalid is dropped, and the reason is listed in the report.
| Stop reason | Meaning |
|---|---|
agent_completed | The investigator returned a valid answer. |
agent_budget_exhausted | A request, tool or token limit was reached. |
agent_output_invalid | The answer could not be repaired within the retries. |
deadline_exceeded | The total time budget ran out. |
investigator_rejected_or_unavailable | Provider or other failure; the report names the exception type and HTTP status. |
What it can never do
- Write its own queries, read files outside the allowlist, or run shell commands.
- Add facts. It can only ask for registered queries; Lumis records what they return.
- Mark anything as confirmed, or decide the conclusion. Lumis assesses its hypotheses mechanically.
- Change your systems. Suggestions, including patch text, are never applied.
Source: Incident investigation ↗ in the SDK repository.