Skip to main content

Agentic RCA

Agentic RCA means the AI runs the investigation: it decides what to query next from what it just found, and it stops at findings you can check yourself. That is a different thing from answering one prompt with whatever text you pasted into it.

Your work is Windows endpoints and vendor sprawl (backup, EDR, identity, line of business), not one cloud service with clean telemetry. Senior time is the scarce resource. The trade is minutes of gather-and-structure against a 45-minute RDP tour. You still own what happens next.

Chat or a written report

You wantUse
An answer now, with follow-upsChat with the plugin or MCP
A written, cited report to attach to the ticket/sparklogs-investigate
Candidate causes once findings exist/sparklogs-analyze-cause (explicit; labeled as analysis)
Handoff text, or "why did you say that?"/sparklogs-summary, /sparklogs-explain

Chat is not the shallow option. It queries the same data and follows the issue as far as it goes. /sparklogs-investigate costs more turns because it gathers, cites, and structures everything it reports, so run it when you want that document.

What a good investigation looks like

/sparklogs-investigate returns a system condition summary: what was observed in a window, each claim linked to the SparkLogs result behind it (open any of them in Explore to verify).

It also states what was not checked because it sits outside agent visibility: a backup target with no agent, an EDR cloud console, a network path. That list is what separates evidence from a fluent guess.

A backup-failure pass looks past the error line: VSS writer state at job time, free space on the target volume, what installed that morning, whether the same failure shape hit other machines at the client. The log named the symptom. State and fleet context are where the cause usually lives.

You will not click every citation. Being able to is what makes the report safe to put in front of a client.

Cause analysis is a second step

Facts and hypotheses live in separate commands, so speculation never lands in the "what we saw" section.

/sparklogs-analyze-cause labels its output as analysis, anchors each hypothesis to findings from the report, and gives the check that would confirm or refute it. Where the evidence runs out, it says so instead of filling the gap.

Trust, in one list

  • Factual claims cite data the agent actually queried.
  • Confidence is banded to the evidence, not to how fluent the answer reads.
  • What was not checked is listed, not omitted.
  • You remain accountable for operational decisions.

Overview: IT Fleet Intelligence.