Nestack Agent Care
Industries / Operations / KPI-anomaly agent

Operations AI agent · KPI-anomaly detection

KPI-Anomaly Detection AI Agent

Suppress what is already explained — month end, a bank holiday, a planned outage — raise the rest to a named owner with what was seen, and record what was suppressed.

4–6 weeksTypical delivery
Your stackDeployment
Suppression-ledNamed owner
Agent CareAfter launch

What this agent does

Raises the alert, never decides what it means

In
01

A series moves, and the movement is measured against the shape that series has held before.

02

A threshold trips at month end, and the calendar behind the trip is attached before anyone is woken.

Reason
03

A series stops arriving, and the drop is read as a possible broken feed rather than a fall in throughput.

04

A movement is already explained — a planned outage, a warehouse reopening — and it is suppressed and logged.

05

A series is under investigation, and further movement on it joins that investigation rather than opening another.

Decide
06

An alert has no owner who could act on it, and it is not sent — that is a design rule, not an omission.

07

A movement survives the known explanations, and it goes to a named owner with the window, the shape and the source.

Out
08

A series definition changes, and its own history is marked as no longer comparable rather than compared.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Watching, separating and routing belong to the agent. Deciding what a movement means, closing an investigation and acting on a flagged series belong to a named owner.

Example workflow

One movement, series to owner

AgentHuman
1Series observedMetric stores, warehouse tables, operational extracts or the monitoring stack
2Known events attachedThe calendar, the planned outages, the onboardings and the investigations already open
3Movement separatedThe window, the shape, what explains it and confidence
4Controls appliedSuppression checks, feed-health checks, ownership checks and detection confidence
No human action required

Stages 1 to 4 run unaided, and nobody is alerted at any of them — the agent is separating, and the owner lane opens at the routing gate.

5DecisionSplits at the routing gate
Movement unexplained

Goes to the named series owner.

Anything feed-shaped

Adds a data-quality read first.

Owner review

The alert is held with the window, the shape, what was ruled out and the confidence.

Acknowledge · Mark explained · Send to data-quality review
Acknowledged — by the named owner
6Alerting and investigation records updatedOnly where write access and records policy allow it
7Outcome evaluatedAlerts acted on, suppressions held, owner corrections and what review found
Corrections

Each series-owner correction is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Deciding what a movement means.
Closing an investigation on a series.
Acting on a series the agent flagged.
Setting the threshold a series is watched at.
Automation boundaryAgent acts unaided
Watch each series against the shape it has held.
Separate the movement a known event explains from the rest of it.
Route what is left to a named owner with what was seen.
Hold a record of what was suppressed and the rule that suppressed it.
No alert leaves except to a named owner, inside the boundaries agreed at implementation.
Judging whether an alert was worth sending.
Standing a suppression rule up or down.
Telling a board that a trend has turned.
Changes to thresholds, routing or suppression rules.

Example output

One flagged movement, annotated

Campaign analytics and spend-pacing triage watch spend; this one watches operational movement, and its subject is the alert nobody acted on.

Anomaly output · single seriesIllustrative example
Series
Recorded as
Window
Evidence of record
Confidence
Held for
Units despatched, daily
Moved beyond its usual shape
Unexplained
Metric store, 3 August 2026
Held unsent
The series owner, by name
As receivedTaken from the metric store and the event calendar on file — nothing on this side is inferred.
What the record holds Metric-store series Event calendar Feed health record
Why no verdict hereSaying what a movement means is a judgement the series owner makes.
ActionAcknowledgeMark explainedSend to data-quality review
What the score decidesBelow the configured threshold a movement gets a quality read before the owner sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every movementFrom the series that carries it
03Explanation

Where the explanation comes from

The agent does not decide what a movement means, only what the series did, what the calendar already explains, and what is left over for a person.

01Approved path

Most movement means nothing

Statistical unusualness is trivial to detect and almost always uninteresting: a series moves at quarter end, on a bank holiday, when a big customer onboards, and nothing is wrong. A detector that fires on all of it gets muted in the third week, and the real signal goes down with the noise.

02Human review

What was checked, and not found

No statute or standard was located that sets which operational series must be watched, what counts as an anomaly, or how quickly a movement has to be looked at, and none was found for the series in scope here. What holds this page up is the alerting policy the company writes for itself: the threshold is a product decision that gets written down, the suppression rules are reviewed on a stated cycle, and an alert nobody could act on is not sent.

04Build an evidence trail

The alert, the movement that raised it and the person it was routed to stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Metric and series storesPrometheus · InfluxDB · Datadog
The operational series and their history
Warehouse of recordSnowflake · BigQuery
Modelled series and the metric layer
Operational systemsERP · WMS · service desk
The events a series is built out of

Agent

KPI-anomaly detection

Reads the series
Separates the movement
Holds for the owner

Alerting and on-callPagerDuty · Opsgenie · Slack
Where an alert lands and who saw it
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six meshes between the model and the channel

Six meshes over one drain, the last the finest. What gets through is drawn in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeStop raising alerts and hold the movement for a person when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, threshold and suppression rules, and note the version each alert was raised under.Track
L4TraceabilityRecord each alert, the movement under it, the rule that suppressed it and who it was routed to.Record
L3Owner routingRoute the alert to a named owner; the routing governs who sees it, not whether it was worth seeing.Gate
L2Suppression rulesTest each movement against the known-event calendar, and return one whose explanation is already on file.Restrict
L1Confidence thresholdsRoute a movement that looks like a feed failure to a data-quality read before an owner is alerted.Require review
Model coreMovement detected — the window, the shape, what explains it and confidence
L1 – L2Test whether an alert may stand
L3Leaves the meaning to a named series owner
L4 – L5Keep the alert and the movement behind it
L6Raises nothing at all when signals degrade

How Nestack evaluates it

Evaluate the whole watch — not only the alert that comes out of it.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the alert an owner receives
Depth of coverage ▼
E1Final-output evaluationDid the alert name something the owner could act on?
E2Step-level evaluationDid the agent read the right series, the right window and the live calendar?
E3Tool evaluationDid it read and write the correct series and the correct alert?
E4Confidence calibrationDo low-confidence alerts actually attract more owner corrections?
E5Slice evaluationHow does performance change across specific series?
E6Business outcomeHow many alerts led to an action, and how many were closed unread?
Floor — the action an alert asks for

Failure modes

Where each failure originates in the agent

Seven failure modes, each placed at the stage the alert goes wrong.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
MY-03

Series read incomplete

The window read was missing part of the series.

Stage gathersThe series, the windows, the events and the owners
02 · Reasoning2 modes
MY-04

Alert raised, nothing to do

An alert goes out that names no action.

MY-06

Known event read as anomaly

A seasonal shape is raised as new.

Stage proposesThe window, the shape and what explains it
03 · Tool / write2 modes
MY-02

Thin movement passed forward

A movement is raised without the data-quality read.

MY-05

Alert routed to nobody

The named owner has left and nobody replaced them.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
MY-01

Real signal suppressed

A stale rule held back a movement that mattered.

Stage returnsThe alert an owner receives and has to act on
05 · Change / Version1 mode
MY-07

Silent threshold drift

A rule change widens what the agent will raise.

Stage tracksModel, prompt, threshold rules and event dates
Sev-1 · an alert sent with no owner Sev-2 · a real movement suppressed Sev-3 · feed degrades, alerts held back

Affected slices

The noise concentrates in the new series

A series-level actionability figure can read clean while newly instrumented series carry most of the alerts nobody acted on. Nestack reports the unactioned-alert rate by series, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Newly instrumented series8.0%3.7× Review
Strongly seasonal series5.7%2.6× Review
Low-volume series3.6%1.7× Watch
Established daily series1.4%0.6× Normal
Bar: unactioned-alert-rate lift vs. established-series baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What a muted channel costs

A cycle ends when the alert nobody could act on is a regression case. That suite is what the next rule shipped is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Unactioned alerts rise on newly instrumented series.

02Diagnose

The alert channel everyone muted in the third week is opened again and read alert by alert until the rule that filled it is found.

03Improve

Number the change; the alerts that drove it ride along with it.

04Verify

One alert case still failing holds the release where it is.

05Learn

The case stays on, and the threshold rules are amended in the same commit.

Learn → DetectThe return edge. The next rule shipped is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, detection workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Series-ownership and automation-boundary mapping.
02Metric-store and warehouse sources.
03Known-event calendar and suppression-rule mapping.
04Series and event-calendar ingestion.
05Series, window and owner binding.
06Actionability scoring and review routing.
07Owner routing workflow.
08Alerting-system integration.
09Threshold and routing cases.
10Guardrails and suppression controls.
11Alert-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne series set, one cycle ProductionProduction alerting workflow AdvancedMultiple sites / series sets
Introduced at Pilot
Detection tuned to your series
Named owner routing
Series-history baseline
Introduced at Production
Reporting by series
Data-quality review workflow in your systems
Approved write-back
Metric-store integration
Introduced at Advanced
Multi-source series reconciliation
Cross-site alert packs
Large series catalogues
Multi-series suppression controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, series volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live series and the owner each one answers to Series ingestion and ownership captureWeek 1
02Representative history for the series in scope Series binding, detection logic and the alert baselineWeek 2
03Your event calendar and your planned outages Known-event mapping, suppression rules and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Metric-store and warehouse source assessment, then integration setupWeek 2
05Alerts you would not want explained False-positive cases and the evaluation runWeek 4
06What no alert may decide Actionability scoring, review routing, guardrails and suppression controlsWeek 3
07A named owner for each series in scope Routing to the named owner, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

The fifth week shows two bands. That is not a drawing error; those two phases genuinely coincide.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Series discovery, ownership mapping and the automation boundary W2Source integration and the series-history baseline W3Detection workflow, suppression logic and routing controls W4Evaluation suite, false-positive cases and failure-mode testing W5Alerting-system integration, pilot series and targeted corrections W6One tuning cycle run under the analytics lead, then Agent Care handover
Reading the bandA band sits on the weeks its own work is named for, and the fifth week honestly carries two.
At the end of W6When the alert record validates, Agent Care picks the agent up.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Operations AI agent

Build an anomaly agent around the alert channel your team stopped reading.

Show us the series you watch and the channel the alerts land in. If nobody can say what they would have done differently on receiving the last one, then it cost attention and bought nothing. We tune for silence first, and name an owner for what is left.

Nestack Agents · KPI-anomaly detectionAGT-OP-09 · Agent Care available after launch