Nestack Agent Care
Industries / Telecom / Troubleshooting agent

Telecom AI agent · Device support

Device & Connectivity Troubleshooting AI Agent

Run the configured line and device tests, read the result back with what it does not cover, and give documented non-destructive steps — holding anything that would interrupt the line for an engineer.

4–6 weeksTypical delivery
Your stackDeployment
Non-destructiveEngineer review
Agent CareAfter launch

What this agent does

Tests what it can see, and names what it cannot

In
01

When a fault is reported, ingest the line record and service state from supported assurance sources.

02

Where the symptom arrives in the customer's own words, normalise it against the documented fault classes.

Reason
03

Run the configured read-only line, radio and CPE tests, and record what each one returned.

04

When an open outage or planned work already covers the address, match the fault to it first.

05

Where the line carries a monitored alarm, a medical alert or telecare equipment, mark it before any step is offered.

Decide
06

Where a step would interrupt the line, and with it the ability to reach 911, hold it rather than issue it.

07

When the evidence runs out at the demarcation point, say so and route the fault to a named engineer.

Out
08

Retain the symptom, the tests run, the steps given and the engineer's decision against the fault.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent tests and proposes steps; a qualified engineer decides what is done to the line, and the carrier stays responsible for the service.

Example workflow

One fault, report to restoration

AgentHuman
1Fault reportedCare contact, self-service app, assurance alarm or CPE telemetry
2Evidence gatheredLine state, sync and error counters, provisioning status and any open outage, each timestamped
3Tests run and read backLine and device results, gaps and confidence
4Controls appliedContinuity check, life-safety screening, assistive-configuration check, destructive-step bar and confidence threshold
No human action required

Stages 1 to 4 run unaided, and nothing on the line is changed at any of them — the agent is testing, and the engineer's lane opens at the confidence gate.

5DecisionBranches at the confidence threshold
High confidence

Goes to the duty engineer to authorise.

Low confidence

Adds a service-continuity read first.

Engineer review

The fault is held with its test results, its flagged risks and the confidence.

Approve · Amend · Escalate to assurance
Approved — released to the customer
6Care and assurance systems updatedOnly where write access and approval policy allow it
7Outcome evaluatedStep outcomes, repeat faults, engineer amendments and corrections after restoration
Amendments

Every engineer amendment is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Any step that would take the line out of service.
Factory resets and anything that destroys data.
Disabling a firewall, encryption or DNS filtering.
Restoring defaults on an assistive configuration.
Automation boundaryAgent acts unaided
Run the configured line, radio and CPE tests, read-only.
Read the results back with what those tests do not cover for the named owner.
Match the fault against open outages, planned work and history.
Give documented non-destructive steps, and hold the rest.
Any write happens inside the boundaries agreed at implementation, never ahead of the engineer.
Steps on a monitored alarm or medical alert line.
Wiring, an ONT or anything past the demarcation.
Dispatching an engineer or booking a field visit.
Changes to test, threshold or escalation rules.

Example output

One fault, annotated

Everything the agent proposes is attached to the tests it was drawn from.

Diagnostic output · single faultIllustrative example
Fault
Reported symptom
Line state
Test source
Confidence
Continuity
Broadband fault
Drop-outs each evening, wired and wireless alike
In service, sync held
Configured line test
84%
Line held through the tests
As receivedTaken from the line record and the configured tests — nothing on this side is written by the agent.
Evidence used Line test result CPE telemetry Open-outage check
Why this stepIt changes nothing on the line — the reason it can be given without an engineer.
ActionApproveAmendEscalate to assurance
What the score decidesBelow the configured threshold the fault picks up a continuity read before it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every faultFrom the line record
03Testing

Test what the network shows

Draw on line state, CPE telemetry, provisioning status and the open-outage picture.

01Approved path

Test the line, keep the call

Routine line and device checks are run before anyone is asked to try anything.

02Human review

Send the risk to an engineer

Continuity risk, life-safety lines and faults past the demarcation point are marked, so an engineer's time starts where the agent cannot see.

04Build an evidence trail

The symptom, the tests run against it and the engineer who took it over stay on the fault.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Assurance and line testTR-069 · TR-369 / USP
Nokia · Ericsson assurance
Care platformSalesforce · Zendesk
ServiceNow · Genesys
Subscriber and CPE recordsAmdocs · Netcracker
Device and SIM inventory

Agent

Device & connectivity troubleshooting

Reads the line
Runs the tests
Holds for an engineer

Workforce and outageField service · scheduling
Outage and incident systems
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the line

Every layer wraps the next. What none of them catches is in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modePull the agent back to read-only testing when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, test-rule and escalation-configuration changes.Track
L4Fault traceabilityRecord the symptom, the tests, the steps given and the engineer's decision.Record
L3Engineer approvalHold the fault for a named engineer; it governs release, not whether the step works.Gate
L2Continuity guardrailsTest each step against the continuity, life-safety and assistive rules; a failure returns the step.Restrict
L1Confidence thresholdsRoute low-confidence faults to a continuity read before an engineer sees them.Require review
Model coreSteps proposed — line results, what is untested, held steps and confidence
L1 – L2Test whether a step may be given
L3Puts the release in an engineer's hands
L4 – L5Hold the tests the diagnosis rested on
L6Falls back to line status when signals degrade

How Nestack evaluates it

Evaluate the whole fault path — not only the step that was given.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the step the customer is given
Depth of coverage ▼
E1Final-output evaluationDid the step match what the tests actually returned?
E2Step-level evaluationDid the agent use the right line record, test set and escalation rules?
E3Tool evaluationDid it read the correct line and run the correct test?
E4Confidence calibrationDo low-confidence faults actually attract more engineer amendments?
E5Slice evaluationHow does performance change across specific fault classes?
E6Business outcomeHow many faults came back after the customer was told the line was fine?
Floor — whether the line still carries a call

Failure modes

Where each failure originates in the agent

Seven ways a diagnosis goes wrong, by stage.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
XC-03

Stale line telemetry

Line state read from a cached or superseded test.

Stage gathersLine record, CPE telemetry, outage feed and test rules
02 · Reasoning2 modes
XC-04

Overconfident cause

A partial test becomes a definite cause for the fault.

XC-06

Continuity rule bypassed

A step is offered that would take the line out of service.

Stage proposesLine and device results, gaps and confidence
03 · Tool / write2 modes
XC-02

Life-safety line missed

A monitored alarm or medical alert line is not screened.

XC-05

Assistive setting reset

A telecoil, TTY or captions setting is cleared by a step.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
XC-01

Beyond-demarcation fault

A fault past the demarcation point is answered anyway.

Stage returnsThe steps the engineer approves and
05 · Change / Version1 mode
XC-07

Silent step-list regression

A model or rule change widens what the agent will instruct.

Stage tracksModel, prompt, test rules and escalation config
Sev-1 · a step runs outside the boundary Sev-2 · service and 911 interrupted unwarned Sev-3 · telemetry degrades, fault routes to review

Affected slices

Overall health can hide a concentrated harm

Break the engineer-takeover rate down by fault class before reading it as healthy. A few cohorts absorb most of the takeovers The cohorts that carry it are named, not averaged away..

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Alarm and medical alert lines6.8%3.7× Review
Faults past the demarcation point5.2%2.8× Review
Assistive and relayed contacts3.5%1.9× Watch
Single-device, in-service faults1.5%0.8× Normal
Bar: engineer-takeover-rate lift vs. in-service-fault baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

The loop closes on a case, not a cause

The loop shuts when the miss is a case in the suite Each cycle leaves the next one a harder test to pass..

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Engineer-takeover rate rises in a fault class.

02Diagnose

Pull the faults and the tests behind them until the cause narrows to one.

03Improve

Whatever changes ships against a version, with the faults that prompted it attached.

04Verify

Nothing ships until the affected cases pass a second time.

05Learn

The suite grows by one case; so does the safe-step list.

Learn → DetectThe return edge. The next pass is measured against a suite that grew.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, test workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Fault workflow discovery and automation-boundary definition.
02Assurance and CPE source assessment.
03Continuity, life-safety and escalation rule mapping.
04Line-record ingestion and normalisation.
05Test orchestration and result binding.
06Confidence scoring and held-step routing.
07Engineer review workflow.
08Care and assurance system integration.
09Service-continuity cases.
10Guardrails and escalation controls.
11Fault-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne network, one queue ProductionProduction assurance systems AdvancedMultiple networks / markets
Introduced at Pilot
Testing to your line and rules
Engineer approval
Diagnosis-consistency baseline
Introduced at Production
Reporting by fault class
Review workflow in your systems
Approved write-back
Assurance-system integration
Introduced at Advanced
Multi-network and multi-vendor rules
Multi-stage engineer approvals
High fault volume
Multi-network diagnostic controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your line records and CPE inventory Line-record ingestion and fault mappingWeek 1
02Representative closed fault tickets Test orchestration baseline and result bindingWeek 2
03Your documented safe-step list and escalation rules Continuity, life-safety and escalation rule mappingWeek 1
04Access to relevant APIs, feeds or exports Assurance and CPE assessment, then integration setupWeek 2
05Steps you would not want a customer to take Continuity cases and failure-mode testingWeek 4
06Where a step must stop and wait for an engineer Confidence scoring, held-step routing, guardrails and approval controlsWeek 3
07Named engineers to review held faults Engineer review workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

The bands follow real work rather than a plan, so evaluation and pilot genuinely share the fifth week.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Fault workflow discovery, safe-step mapping and the automation boundary W2Assurance and CPE integration, and the test baseline W3Test workflow, confidence logic and escalation controls W4Evaluation suite, continuity guardrails and failure-mode testing W5Care integration, pilot faults and targeted corrections W6A live fault queue worked under supervision, then Agent Care handover
Reading the bandEach bar spans the weeks its work is named in and no others. The fifth week genuinely carries two.
At the end of W6Once the queue validates, Agent Care owns the running agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Telecom AI agent

Build a troubleshooting agent around your service-continuity rules.

Show us your line tests, your safe-step list and who takes a fault over. If you name the lines carrying alarm or medical equipment, we'll draw the boundary around those first.

Nestack Agents · Device & connectivity troubleshootingAGT-TL-03 · Agent Care available after launch