Concepts

The investigator

The optional tool-using model: when it runs, the tools it can use, its budgets, how its output is validated, and what it can never do.

v0.1.0 · experimentalPython 3.11+Updated 2026-10-05

When it runs

Only when triage is not sufficient and you opted in, with --use-agent on the CLI or use_agent=True in Python. It is one Pydantic AI agent with typed output, not a multi-agent system. It starts from the triage findings and the evidence already collected.

What it can do

The investigator has two tools, inspect and probe. inspect has a fixed set of operations:

OperationWhat it allows
catalogList the registered query IDs, approved files and enabled capabilities.
graphRead a one-hop neighbourhood inside the incident scope.
evidenceRun a registered query by ID. It cannot write PromQL, LogQL, TraceQL or SQL.
changesList recent commits and rollouts that touch scoped entities, newest first.
code.read, code.searchRead or literally search an explicit allowlist of text files.
git.log, git.diffFixed, read-only Git commands on approved paths.
hypothesis.registerRegister a falsifiable explanation before testing it.

probe runs a generated Python experiment in a disabled-by-default, network-less container. Its results are marked degraded. See safety and limits.

Giving it code context

yaml
investigator:
  repositories:
    - id: application
      root: ./approved-source
      entity_ids: ["service:shop:checkout"]
      files: [src/checkout/handler.py, deploy/releases.yaml]
      include_commit_subjects: false

Each repository is mapped to the entities it belongs to, and only the listed files are readable. Files are snapshotted once per investigation, redacted and hashed. Paths outside the list, symlinks, binary or oversized files, and secret or dot directories are refused.

Budgets

yaml
investigator:
  budget:
    request_limit: 8
    tool_calls_limit: 10
    max_probes: 2
    output_tokens_limit: 12000
    max_tool_characters: 8000
    max_total_tool_characters: 32000
    validation_retries: 2

These are example values. Every tool attempt, including a denied or malformed one, counts. When a budget runs out, the evidence and hypotheses collected so far are still assessed and reported.

How its output is checked

The model's final answer is checked against the same rules Lumis applies afterwards: graph IDs must exist, evidence_needed must name registered queries, a revised hypothesis needs a new ID, and suggestions may cite only evidence Lumis actually issued. Problems are sent back to the model to repair, up to validation_retries times. Anything still invalid is dropped, and the reason is listed in the report.

Stop reasonMeaning
agent_completedThe investigator returned a valid answer.
agent_budget_exhaustedA request, tool or token limit was reached.
agent_output_invalidThe answer could not be repaired within the retries.
deadline_exceededThe total time budget ran out.
investigator_rejected_or_unavailableProvider or other failure; the report names the exception type and HTTP status.

What it can never do

  • Write its own queries, read files outside the allowlist, or run shell commands.
  • Add facts. It can only ask for registered queries; Lumis records what they return.
  • Mark anything as confirmed, or decide the conclusion. Lumis assesses its hypotheses mechanically.
  • Change your systems. Suggestions, including patch text, are never applied.

Source: Incident investigation ↗ in the SDK repository.