Nestack Agent Care
Industries / Manufacturing / Downtime agent

Manufacturing AI agent · Downtime

Downtime Root-Cause Analysis AI Agent

Reconcile stop signals with operator reason codes, group recurring stops, assemble the evidence around an event and propose candidate factors — an engineer establishes the cause and owns the countermeasure.

4–6 weeksTypical delivery
Your stackDeployment
Never the agentCause decision
Agent CareAfter launch

What this agent does

Assembles the evidence, does not name the cause

In
01

Take stop signals from the historian, PLC tags or MES alongside the reason code and notes the shift entered.

02

Pull what surrounded the stop — alarms, process values, changeover, shift, material lot and the last intervention.

Reason
03

Reconcile the recorded stop against the signal — start, end, duration, asset — and mark where the two disagree.

04

Trace a stop that cascaded down the line back to where it started, and mark blocked and starved time as such.

05

Group stops that behave like the same failure, and hold apart the ones that only share a reason code.

Decide
06

Propose candidate contributing factors, each with the evidence under it and what would rule it out.

07

Flag events too thin to investigate, and estimate the stopped time that never reached the record at all.

Out
08

Present the event, its group and the candidate factors for an engineer to accept, reject or send back.

09

Track whether a countermeasure held at 30 and 90 days, and reopen the group when the stops return.

Product statement

The agent assembles and proposes. An engineer or a team establishes the cause, decides the countermeasure and owns whether it worked.

Example workflow

One stop, end to end

AgentHuman
1Stop detectedHistorian, PLC tag, MES downtime record or the line's own counter
2Record reconciledSignal start, end, duration and asset against the reason code and notes the shift entered
3Event groupedMatched to an existing recurring-stop group, or opened as one of its own
4Evidence assembledAlarms, process values, changeover, shift, material lot and the intervention before it
No human action required

Stages 1 to 4 run without a person in the loop — reconciliation, grouping and evidence assembly finish before anyone is asked to read anything. An event that will not reconcile ends that stretch early.

5DecisionSplits on whether the event reconciles and the evidence is sufficient
Reconciled, evidence sufficient

Reaches the engineer as a proposed factor set.

Unreconciled or thin record

Goes to a person with the gaps named, and no factors.

Reliability or CI engineer

Reads the event, its group and the evidence under each proposed factor, then decides what the cause was and what to do about it.

Accept factor · Reject · Request more evidence
Cause agreed — handed back
6Findings filedWritten to the downtime record only where write access and policy allow; the cause field stays empty
7Outcome evaluatedReason-code accuracy, engineer agreement, cascade attribution and what the countermeasure did at 90 days
Rejections

Factors the engineer rejects are counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Declaring the cause of a downtime event.
Approving or scheduling a countermeasure.
Deciding that a countermeasure worked.
Naming an operator, crew or shift as the cause.
Automation boundaryAgent acts unaided
Reconcile stop signals against reason codes and mark the disagreements.
Group recurring stops, and hold apart the ones that only look alike.
Assemble the alarms, process values, changeover, lot and shift context around an event.
Propose candidate factors with the evidence for each, and flag records too thin to work.
Write actions run only inside the approval boundaries agreed during implementation. The cause field is not one of them.
Overwriting a reason code the shift entered.
Re-attributing downtime between assets in the record.
Changing the reason-code taxonomy or capture thresholds.
Raising, changing or closing a maintenance work order.

Example output

One stop, annotated

Everything the agent proposes stays attached to the event and the signals it was assembled from.

Downtime output · single stop eventIllustrative example
Stop signal
Reason code entered
Next asset
Candidate factor
Confidence
Cause
Filler 2 stopped for 14 min 20 s
Changeover, keyed at shift end
Starved 11 min
Cap-feed jam, not changeover
78%
Not established by the agent
As receivedThe signal, the code the shift keyed and what the next asset did — as recorded, contradictions included.
Evidence used Alarm three minutes before No recipe change logged Nine like stops in 30 days
Why this factor is proposedNo recipe change was logged; one alarm precedes nine stops. Correlation, not cause.
ActionAccept factorRejectRequest more evidence
What the score decidesConfidence decides how much evidence an engineer should demand, not who is at fault.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every recorded stopHistorian, PLC, MES or the line counter
03Reconciliation

Apply the plant's own taxonomy

Use the client's reason-code list, asset hierarchy, line topology, shift pattern and changeover records.

01Approved path

Hand the engineer an assembled event

Alarms, process values, lot and changeover context arrive gathered against the stop, so an investigation starts from a record rather than from a query.

02Human review

Show which stops share a pattern

Recurring stops are grouped and the thin or contradictory records are separated out, so the meeting argues about evidence instead of about the data.

04Build an evidence trail

Retain the signal, the entered code, the reconciliation, the sources read, the proposed factors, the engineer's decision and the 30- and 90-day check.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Machine & process dataPLC and SCADA tags · historian · OPC UA
Alarm and event logs · edge collectors
MES & downtime captureSiemens Opcenter · Rockwell FactoryTalk
In-house MES · operator terminals
Maintenance & work ordersSAP PM · IBM Maximo
Fiix and Limble · work-order history

Agent

Downtime root-cause analysis

Reconciles the stop
Assembles the evidence
Proposes candidate factors

Production contextShift calendars · changeover records
Lot genealogy · material receipts
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the downtime record

Each control wraps the one inside it. An event clears every layer before an engineer reads it, and the cause sits outside all six.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeReturn grouping and factor proposal to manual if evaluations or production signals degrade.Roll back
L5TraceabilityRecord the signal, the entered code, the sources read, the factors and every rejection.Record
L4Engineer gateCause, countermeasure and taxonomy changes stay with a named engineer under change control.Gate
L3No-attribution ruleFactors name conditions, assets and steps. A person, crew or shift is never proposed as one.Block
L2Evidence sufficiencyA factor is proposed only where the evidence under it is named and reachable.Require
L1Reconciliation gateAn event whose signal and reason code will not reconcile is marked unusable, not explained.Withhold
Model coreCandidate factors proposed — the reconciled event, its group, the evidence under each factor and a confidence
L1 – L2Decide whether an event can be investigated
L3Keeps people out of the factor list
L4 – L5Keep the cause with a person and the trail intact
L6Pulls automation back when signals degrade

How Nestack evaluates it

Evaluate the whole investigation — not only the factor that was accepted.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the event record the engineer opens
Depth of coverage ▼
E1Final-output evaluationWhich proposed factors did the engineer accept, and which were rejected?
E2Reconciliation accuracyDid duration, asset and reason code match a sample a person reviewed?
E3Grouping evaluationWere separate failures merged, or one failure split across groups?
E4Unrecorded-stop estimationHow much stopped time never reached the record at all?
E5Slice evaluationHow does accuracy change across lines, shifts and stop lengths?
E6Business outcomeDid the countermeasure still hold at 30 and at 90 days?
Floor — the countermeasure that had to hold

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle.

Agent lifecycleDirection of processing →
01 · Capture2 modes
DT-01

Reason code keyed by default

One code is batched across a shift's stops at the end of it.

DT-02

Micro-stops never recorded

Stops under the capture threshold leave no event to investigate.

Stage gathersStop signals, reason codes, alarms and shift notes
02 · Reconciliation1 mode
DT-03

Cascade hits the wrong asset

A starved downstream machine is logged as the failure.

Stage matchesSignal against entered code, duration and asset
03 · Grouping1 mode
DT-04

Different failures merged

One group hides two unrelated stop mechanisms.

Stage clustersRecurring stops, and the ones that only look alike
04 · Proposal2 modes
DT-05

Correlation read as cause

A factor lands firmly enough that the team stops looking.

DT-06

A factor points at people

Evidence resolves to a crew or a shift rather than a condition.

Stage proposesCandidate factors and the evidence under each one
05 · Follow-up1 mode
DT-07

Countermeasure credited wrongly

The line ran well for unrelated reasons and the group closed.

Stage tracksCountermeasure durability and version changes
Sev-1 · cause or blame is overstepped Sev-2 · the record misleads the work Sev-3 · real loss, absent from the data

Affected slices

The record is worst where the stops are shortest

The stops that get recorded well are the long ones, and they dominate any plant-wide figure. Short stops, night shifts and cascades are where the entered code and the signal disagree. Reported by slice, not in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Stops under two minutes7.2%3.8× Review
Night and weekend shifts5.5%2.9× Review
Cascaded line stops3.7%1.9× Watch
Long stops, single asset1.5%0.8× Normal
Bar: mis-recorded-event rate vs. long-stop baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

A rejected factor is worth as much as an accepted one

What an engineer throws out says something the evidence did not. So does a countermeasure that held for a month and then quietly stopped holding.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Stops return in a closed group, or reason-code accuracy slips on a reviewed sample.

02Diagnose

Traced to the signal, the entered code, the asset map, the grouping rule or evidence that was never pulled.

03Improve

The reconciliation rule, grouping threshold or evidence set changes under change control, with a named approver.

04Verify

Re-run over stored events from that cohort, including the factors the engineer rejected.

05Learn

The rejected factor becomes a negative case, and the group keeps its 30- and 90-day check.

Learn → DetectThe return edge. A group is closed only after the 90-day check, and when a conclusion is withdrawn the people who acted on it are told.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, signals and taxonomy, reconciliation and grouping, evaluation, review workflow, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and automation boundary.
02Historian, PLC and MES signal assessment.
03Asset hierarchy and line-topology mapping.
04Reason-code taxonomy review and mapping.
05Stop-signal reconciliation logic.
06Cascade and blocked/starved attribution.
07Recurring-stop grouping and thresholds.
08Evidence assembly from alarms and context.
09Candidate factors, confidence and rule-outs.
10Unrecorded-stop study on the pilot line.
11Engineer review and rejection workflow.
12Countermeasure tracking and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne line, one asset group ProductionPlant-wide downtime record AdvancedMulti-line / multi-plant
Introduced at Pilot
Stop-signal and reason-code reconciliation
Recurring-stop grouping
Evidence assembly around an event
Baseline evaluation
Introduced at Production
Client reason-code taxonomy mapping
Candidate factors with evidence and rule-outs
Cascade and blocked/starved attribution
Engineer review and rejection workflow
Observability and evaluation
Introduced at Advanced
Countermeasure durability tracking
Multi-plant and enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on signal sources, asset and line topology, reason-code taxonomy, number of lines and plants, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your reason-code list and how operators actually key it Reason-code taxonomy review and mappingWeek 2
02Asset hierarchy and how the line is physically connected Asset hierarchy, line topology and cascade attributionWeek 1
03Access to the historian, PLC tags and MES downtime records Signal assessment, then stop-signal reconciliation logicWeek 2
04A sample of events an engineer has already worked through Reconciliation baseline and the reviewed-sample comparisonWeek 3
05Investigations that reached the wrong conclusion Evaluation suite, regression cases and failure-mode testingWeek 4
06Where your capture threshold sits and what falls under it Unrecorded-stop study on the pilot lineWeek 5
07Named reliability or CI engineers to review events Engineer review workflow, then pilot events and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. Week 5 carries both the unrecorded-stop study and the first supervised review sessions.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Workflow discovery, asset hierarchy and line topology W2Signal access, reason-code taxonomy and reconciliation logic W3Cascade attribution and recurring-stop grouping W4Evidence assembly, candidate factors and failure-mode testing W5Unrecorded-stop study, supervised review sessions and corrections W6Engineers work live events, then Agent Care starts
Reading the bandThe unrecorded-stop study runs in week 5, before anyone builds a Pareto from this data. A chart drawn from an incomplete record is worse than no chart.
At the end of W6Engineers have worked real events from the assembled record alongside their existing process, and the first 30- and 90-day countermeasure checks are already scheduled.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Manufacturing AI agent

Build a downtime agent around your own stop records.

Show us how a stop reaches your system today, your reason-code list, and one recurring stop nobody has settled. We'll reconcile a month of your own events against a sample your engineers have already worked, and show you what the record cannot tell you.

Nestack Agents · Downtime root-cause analysisAGT-MFG-09 · Agent Care available after launch