Skip to main content
For service managers

Cut MTTR. Reduce repeat tickets. Find root-cause for critical client issues in minutes.

Improve service by standardizing ticket work through SparkLogs data + AI + your tailored process, operating in your current systems and workflow. Get root-cause writeups for any ticket in minutes. Analyze your IT fleet every day for unknown issues, unexpected system changes, and recurring problems.

#7623 · Harborview
Backups failed on fs-02 three nights running
Signals
  • VSS writer failed ×3, SQL Server writer, 01:58 nightly
  • shadow storage near cap D, since Sat
  • ntfs delayed write lost ×2
State
  • data volume space low D, open 3d 4h, Serious
Changes
  • none on hv-fs-02 in 14 days
Scope
  • 1 host, 1 client
Could not see
  • Management channels, not yet collected
CauseD filled, shadow storage capped, and the VSS writer failed before each job.
#7598 · Northridge
dc-01 rebooted at 2 am
Signals
  • unexpected shutdown Mon 02:04
  • storage controller reset ×4, Sun to Mon
  • bugcheck Mon 02:03
State
  • disk unresponsive opened 02:01, cleared on reboot
Changes
  • driver updated storage, Mon 01:40
  • reboot pending set Mon 01:41
Scope
  • 1 host; same driver on 11 more Northridge hosts
Could not see
  • crash dump, not configured on this host
CauseStorage driver updated at 01:40, the disk stalled, and the host bugchecked.
What changes for the IT service manager

Better outcomes on every ticket through standardized root-cause analysis.

Head off misdiagnosis and dead-ends early.

For IT service managers

multiply your capacity + level-up quality

Conduct process-driven fleet-wide hunts for priority-defects, new security risks, process violations, and to surface unknown-unknowns.

It's like having your L3 team hunt for issues fleet-wide every day, analyzing 1000s of signals on each device comprehensively. Find client issues earlier. Create more time for L3 to focus on higher-impact work.

Scheduled · daily 06:00 · every client
service crashed, service start failed, service hang
Found today3 hosts · 2 clients
lv-app-03 · Spoolernr-db-01 · SQLAgenthv-web-02 · W3SVC

Every investigation is rigorous, process-driven, and tailored for you.

The /sparklogs-investigate skill analyzes and reports signals, state, changes and scope. Open-source playbooks and skills can be customized to fit your process and team. Achieve better-informed and more consistent ticket resolution.

Queue · evidence shape
PR #7623Harborview
DM #7611Lakeview
JK #7598Northridge
AS #7584Pinecrest
signals · state · changes · scope could not see, named

Level up your engineering bench.

Give junior engineers a senior-engineer-like AI companion that follows your process and rigorously analyzes all available signals. Reduce misidentification of issues and head off dead-ends in investigations early.

Post-mortem · #7598 · dc-01 rebooted 02:00
Checked
  • System log · dirty shutdown
  • Storage · disk unresponsive
  • Drivers · storage driver updated, Mon
Could not see
  • Firewall logs, not yet collected
  • Crash dump, not configured
Cause: storage driver update Monday; disk stalled, host bugchecked.
Change analysis

Ask what changed on the box before it failed.

Drivers, system identity and Windows servicing record every change on every endpoint. Put those changes on one clock with the failure and the cause is usually the last change before it.

What changed on dc-01 before the 2 am reboot?
  1. Sun 22:10win servicing session finalized, cumulative update committedWindows servicing
  2. Mon 01:40driver updated, storage, 2.31 to 2.40Drivers
  3. Mon 01:41reboot pending, setSystem identity
  4. Mon 01:52storage controller reset, ×4System log
  5. Mon 02:01disk unresponsive, openedDevice IO
  6. Mon 02:03bugcheckSystem log
  7. Mon 02:04unexpected shutdown, recorded on bootSystem log
CauseStorage driver updated at 01:40. The disk stalled 12 minutes later and the host bugchecked. 11 more Northridge hosts took the same driver, and none has stalled yet.
See for yourself

Walk the signal catalog, then watch an end-to-end investigation with citations.

Browse feeds and graded conditions on the data-feeds page. The demo shows the same cited report shape your team can run on demand.

Make IT problems die young

IT Fleet Intelligence
Ask your fleet anything
- what new failures are happening this week we haven't seen before?
- strange IO slowdowns after patch tuesday. why? who's affected?
- we just fixed VPN flakiness on the lakeview ts. anyone else affected?
- bob's account got hacked at 3:43pm. audit everything. impact?
Any question, any scope
One host, one client, or the whole fleet
Logs plus deep system state
CPU, RAM, IO and disk trends, projected forward. Snapshots plus instant deltas: services and VSS, BitLocker, leaks and blue screens, patch and servicing state, installs, event logs
Fleet-wide patterns before they become tickets
Spot issues emerging across clients from curated endpoint data
Agentic RCA · cited evidence
From question to root-cause and ticket-ready report
- ticket 7623 is off track. dig in and write it up
- backups failed on fs-02 three nights running. why?
- why is lakeview file server super slow every afternoon?
- why did dc-01 at northridge reboot at 2am last night?
/sparklogs-investigate writes the report
Your agent queries fleet evidence, follows what it finds, and returns a cited root-cause report in minutes
Works where you work
Claude, Microsoft Copilot, Cursor, Codex, or your own MCP agents
Open-source skills, cited reports
/sparklogs-investigate writes the ticket-ready report; click any claim to verify the evidence; then dive deep into likely root causes and fixes