Correlate alarms and telemetry into one incident, rank the hypotheses that could explain it with the evidence behind each, and leave the declaration of cause to the named duty engineer.
Ingesting alarms, counters, topology and change records from supported assurance or ticketing sources.
02
Normalising element names, timestamps and severities, and carrying each signal forward with its source.
Reason
03
Grouping related alarms into one incident, and stating what the grouping is based on.
04
Applying the correlation, topology and maintenance-window rules configured for the network.
05
Assembling the evidence pack behind each hypothesis — topology, dependency and timing.
Decide
06
Ranking hypotheses with what would falsify each one, rather than naming a single cause.
07
Surfacing the reportability question, planned maintenance included, to the duty engineer.
Out
08
Retaining the correlation, the signals it set aside, the evidence and the engineer's decision.
09
Executing write actions only inside the approval boundaries agreed during implementation.
→Product statement
The notification clock runs from discovery, not from diagnosis; the copilot proposes and a duty engineer decides.
Example workflow
One incident, alarm to declaration
AgentHuman
1Alarm burst receivedFault management, performance telemetry, change record or trouble ticket
2Evidence assembledTopology, dependency, timing and recent changes, each with its named source
3Hypotheses rankedCauses, evidence, falsifiers and confidence
4Controls appliedGrouping limits, 911 and 988 path exclusions, maintenance-window checks and confidence threshold
No human action required
Stages 1 to 4 run unaided, and nothing is declared or closed at any of them — the copilot is assembling, and the engineer's lane opens at the confidence gate.
5DecisionBranches at the confidence threshold
High confidence
Goes to the duty engineer to accept.
Low confidence
Adds a senior NOC read first.
Engineer review
The incident is held with its evidence, its ranked hypotheses and the confidence.
Accept · Revise · Send to senior review
Accepted — cause declared▼
6Incident record updatedOnly where write access and approval policy allow it
7Outcome evaluatedRank at acceptance, revisions, set-aside-signal outcomes and post-incident corrections
Revisions
Every engineer revision is counted.
What should not run autonomously
Human approval stays in control
Outside the boundary — human approval required8 items
Declaring the root cause of an incident as a finding.
Closing an incident or declaring service restored.
Deciding whether an outage or event is reportable.
Notifying a PSAP, the FCC or any other regulator.
Automation boundaryAgent acts unaided
✓Correlate alarms and telemetry into a single incident view.
✓Assemble the evidence pack behind each candidate hypothesis.
✓Rank the hypotheses, and state what would falsify each one.
✓Surface the reportability question and its clock to an engineer.
Writes stay inside the agreed boundaries and never reach a network element — that belongs to remediation.
Changing any network element or its configuration.
Suppressing any signal on a 911 or 988 call path.
Writing the best-known cause into a notification.
Changes to correlation, suppression or escalation rules.
Example output
One incident, annotated
Everything the copilot proposes is attached to the signals it was drawn from.
Root-cause output · single incidentIllustrative example
Incident
Leading hypothesis
Alarm group
Evidence class
Confidence
Escalation
Regional transport
Protection switch on a shared fibre span, not the access nodes that alarmed
One transport ring
Topology and timing
71%
Duty engineer, clock shown
As receivedTaken from fault management and the change record — nothing on this side is inferred by the copilot.
Do not read the revision rate whole — break it out by incident class. No audited production figure exists for automated root-cause analysis in live networks, so Nestack reports what your engineers revise, by slice.
Slice performance — reported separately, not only in aggregateIllustrative example
Slice
Failure rate
Lift
Lift vs. threshold
Status
Alarm-storm incidents
9.9%
3.9×
Review
Shared-infrastructure faults
6.9%
2.7×
Review
Post-change incidents
4.6%
1.8×
Watch
Single-element faults
1.5%
0.6×
Normal
Bar: engineer-revision-rate lift vs. single-element baseline · scale 0–4.0× · tick marks the 2.0× review threshold2 of 4 slices over threshold
Evidence-linked improvement
A post-mortem is not the end of the loop
The cycle ends in a regression case, not in a meeting about what happened. That suite is what the next incident into the NOC is measured against.
Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect
Engineer-revision rate rises in an incident class.
02Diagnose
The evidence pack behind each disputed hypothesis is pulled and read until the miss narrows to one.
03Improve
Every change is versioned against the incidents that exposed it.
04Verify
A failing case holds the release back.
05Learn
The case joins the suite for good, and the correlation rules are revisited.
Learn → DetectThe return edge. The next detection runs against a suite one case longer.
Typical build scope
Twelve workstreams across six weeks
The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, triage workflow, evaluation, integration, then production validation and handover.
WorkstreamWeek 1Week 2Week 3Week 4Week 5Week 6
01NOC workflow discovery and automation-boundary definition.
02Assurance and telemetry source assessment.
03Topology, dependency and correlation-rule mapping.
04Alarm ingestion and normalisation.
05Correlation and hypothesis ranking.
06Confidence scoring and escalation routing.
07Duty-engineer review workflow.
08Assurance and ticketing integration.
09Causal-claim regression cases.
10Guardrails and escalation controls.
11Incident-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallelFinal scope and sequence confirmed in discovery
Engagement tiers
What each tier includes
Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.
Capability✓ in scope · — not at this tierPilotOne domain, one NOCProductionProduction assurance stackAdvancedMultiple domains / regions
Introduced at Pilot
Correlation to your topology✓✓✓
Engineer sign-off✓✓✓
Hypothesis-quality baseline✓✓✓
Introduced at Production
Reporting by incident class—✓✓
Review workflow in your systems—✓✓
Approved incident write-back—✓✓
Assurance-stack integration—✓✓
Introduced at Advanced
Multi-vendor correlation rules——✓
Multi-stage NOC escalation——✓
High alarm volume——✓
Multi-domain incident controls——✓
Build priceFrom $5,000From $8,000Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.
What we need from you
What you bring, and what we build with it
Each input maps to a piece of build scope and a week in the delivery timeline.
You bringWe build with it
01Your alarm sources and topology model→Alarm ingestion and topology mappingWeek 1
02Representative past incidents→Correlation baseline and hypothesis rankingWeek 2
03Your correlation and suppression rules→Correlation-rule mapping and automation-boundary definitionWeek 1
04Access to relevant APIs, feeds or exports→Assurance and telemetry assessment, then integration setupWeek 2
05Causes you would not want acted on→Causal cases and the evaluation suiteWeek 4
06What must reach an engineer before anything is touched→Confidence scoring, escalation routing, guardrails and review controlsWeek 3
07Named duty engineers to review incidents→Duty-engineer review workflow, then pilot and production validationWeeks 5–6
Nothing else is requiredDeployment, documentation and Agent Care handover are ours.
Delivery timeline
Four phases across six weeks
The bands follow the real work, which is why evaluation and pilot share the fifth week.
PhaseW1W2W3W4W5W6
DiscoveryW1
BuildW2 – W3
EvaluateW4 – W5
Pilot & LaunchW5 – W6
Week focusW1NOC workflow discovery, rule mapping and the automation boundaryW2Assurance integration and the correlation baselineW3Triage workflow, confidence logic and escalation controlsW4Evaluation suite, exclusion checks and failure-mode testingW5Ticketing integration, pilot incidents and targeted correctionsW6One incident cycle triaged under the NOC, then handover
Reading the bandA bar covers the weeks its work is named in, and nothing else. The week 5 overlap is real, not padding.
At the end of W6The cycle closes validation and Agent Care assumes monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.
Next step · Telecom AI agent
Build a NOC copilot around the engineers who declare cause.
Show us your alarm sources, your topology and who declares cause. Being wrong costs twice: the fault stays live and the clock keeps running.