Nestack Agent Care
Industries / Quality Assurance / Defect triage agent

Quality assurance AI agent · Defect triage

Defect Triage AI Agent

Date the hour someone first knew, link the duplicates without burying the one that names an injury, and leave the reportability call to the qualified person who signs it.

4–6 weeksTypical delivery
Your stackDeployment
Clock firstNamed reviewer
Agent CareAfter launch

What this agent does

Prepares the judgement, never signs it

In
01

A report arrives, and the hour an employee first held it is stamped beside it.

02

A clock is already running, because 21 CFR 803.3 dates awareness to any employee.

Reason
03

A duplicate is linked and not merged, so the one line naming harm survives the cluster.

04

A cluster is a trend analysis, and 21 CFR 803.53 runs five work days from one.

05

A severity band is set, and it is kept apart from the reportability question.

Decide
06

A consumer hazard surfaces, and 16 CFR 1115.14 allows ten days to investigate.

07

A not-reportable view is drafted, and 21 CFR 803.20(c)(2) leaves that call to a qualified person.

Out
08

One defect meets two regimes, and CRA Article 14 adds a third from 11 September 2026.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Intake, dating, linking and routing belong to the agent. The reportability decision, and the name that stands under it, belong to a qualified person the rule itself describes.

Example workflow

One report, arrival to decision

AgentHuman
1Report receivedA complaint, a field report, a support thread, a returned unit or a failed test
2Awareness dated and recordedWho first held the information, the hour it reached them, the product and what it describes
3Duplicates linked, not mergedThe cluster, each member left whole, the outlier lines and confidence
4Controls appliedHarm-language checks, clock checks, regime checks and triage confidence
No human action required

Stages 1 to 4 run unaided, and no report is closed at any of them — the agent is preparing, and the reviewer lane opens at the reportability gate.

5DecisionSplits at the reportability gate
No harm language present

Goes to the qualified reviewer to decide.

Anything naming injury

Adds a regulatory-affairs read first.

Reviewer decision

The report is held with its arrival hour, the cluster it joined and the clocks it sits under.

Decide · Attach evidence · Send to regulatory review
Decided — by the qualified reviewer
6Complaint and event-file records updatedOnly where write access and event-file policy allow it
7Outcome evaluatedDecision accuracy, reopened duplicates, reviewer corrections and what the later audit turned up
Corrections

Each correction the reviewer makes counts in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Concluding that an event is not reportable.
Signing the written report that goes to the Commission.
Closing a report whose text names an injury.
Deciding a quarterly summary form is enough.
Automation boundaryAgent acts unaided
Date the hour the first employee became aware.
Link duplicate reports and preserve each reportability determination.
Record the deliberation a reviewer relied on, dated and sourced.
Hold the report for the reviewer who decides it.
No report closes as not reportable except by a qualified person, inside agreed boundaries.
Judging whether a device may have contributed.
Telling a regulator the review is complete.
Setting which day a statutory clock is counted from.
Changes to the triage rules or the reportability gate.

Example output

One defect report, annotated

An engineering agent keeps the cause and a warranty agent keeps the unit; this one keeps whether a report must be told to a regulator. Below is one report exactly as the agent leaves it.

Triage record · single reportIllustrative example
Report
Recorded as
Awareness
Evidence of record
Confidence
Held for
Customer report, one product line
Unexpected speed change, injury named
Linked, not merged
Support thread, 3 August 2026
Held undecided
The qualified reviewer, by name
As receivedDrawn from the support thread and the complaint file, and it concludes nothing those two do not carry.
What the record holds Support thread lines Complaint file entry Linked duplicate set
Why no decision hereCalling a report not reportable is an act a qualified reviewer owns.
ActionDecideAttach evidenceSend to regulatory review
What the score decidesBelow the configured threshold a report gets a regulatory read before the reviewer sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Each reportFrom the hour it was first held
03Report

Where the report is used

The best published duplicate detectors put the true match in a top-ten shortlist about three times in five, so a linked cluster is a shortlist for a reader and not a closure.

01Approved path

The clock started already

A treadmill importer settled with the Commission this August over a hazard known from March 2018 and reported in October 2022; those reports were triaged, not missed.

02Human review

What was checked, and not found

No live standard defines defect-report severity: IEEE 1044-2009 has been Inactive-Reserved since 5 March 2020, and the classes in IEC 62304 and ISO 26262 attach to a software item or a hazard, never to a report. Corrective action moved to ISO 13485 clause 8.5 via § 820.10(c) on 2 February 2026.

04Build an evidence trail

The report, the hour it arrived and the person who judged it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Defect and complaint trackingJira · Bugzilla · complaint logs
Field and test defect records
Support and field intakeHelpdesk · CRM · field reports
First-contact report lines
Regulatory event filesMDR event files · CPSC submissions
Reportability decisions and dates

Agent

Defect triage and reportability

Reads the reports
Dates and links them
Holds the decision

Trend and analysis recordsComplaint trending · dashboards
Cluster reviews and escalation logs
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six screens between the model and the file

Six screens down one chute, the last the tightest. What lands is drawn in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeNarrow the agent to intake and dating when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and triage rules, and note the version each report was prepared under.Track
L4TraceabilityRecord each report, the hour it arrived, the cluster it joined and each read of the file.Record
L3Reviewer releaseHold the report for a qualified reviewer; the hold governs release, not whether the call is right.Gate
L2Triage guardrailsTest each draft against the harm language in the source, and return one that closes on silence.Restrict
L1Confidence thresholdsRoute a thin or injury-adjacent report to a regulatory read before the reviewer sees it.Require review
Model coreReport prepared — the arrival hour, the linked cluster, the clocks and confidence
L1 – L2Test whether a report may close
L3Leaves the decision to a qualified person
L4 – L5Keep the report and the judgement behind it
L6Routes to a qualified reader when signals degrade

How Nestack evaluates it

Evaluate the whole triage run — not only the decision that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the decision an event file carries
Depth of coverage ▼
E1Final-output evaluationDid the record hold the deliberation the decision rested on?
E2Step-level evaluationDid the agent read the right reports, the right hours and the live triage rules?
E3Tool evaluationDid it read and write the correct report and the correct event file?
E4Confidence calibrationDo low-confidence triage drafts actually attract more reviewer corrections?
E5Slice evaluationHow does performance change from one defect class to another?
E6Business outcomeHow many reports needed a correction before the reviewer decided?
Floor — the file a regulator opens first

Failure modes

Where each failure originates in the agent

Seven failure modes, each placed where it first becomes visible.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
VE-03

The report that never reached the tracker

A clock ran inside a thread nobody triaged.

Stage gathersThe reports, the hours, the units and the notes
02 · Reasoning2 modes
VE-04

Merged away with the injury line

The one report naming harm joins a cluster.

VE-06

Two regimes, one due date

The tracker counts down the wrong clock.

Stage proposesThe arrival hour, the cluster and the clocks
03 · Tool / write2 modes
VE-02

Severity gates the reportability read

A low-banded report skips the review.

VE-05

Cluster seen, nobody told

A trend surfaces on a screen, not to a name.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
VE-01

Closed, the deliberation unrecorded

The file shows a call, not the reasoning under it.

Stage returnsThe call a file carries and a regulator reads
05 · Change / Version1 mode
VE-07

Quiet accuracy drift

The wording moves on while the linking model does not.

Stage tracksModel, prompt, triage rules and clock dates
Sev-1 · a reportable report closed Sev-2 · a due date counted from the wrong hour Sev-3 · signals thin, report held back

Affected slices

Injury-adjacent reports absorb the corrections

A class-level triage figure can read clean while injury-adjacent complaints carry most of the reviewer corrections. Nestack reports the correction rate by defect class, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Injury-adjacent complaints13.6%3.7× Review
Cross-regime field reports9.6%2.6× Review
Duplicate-cluster outliers6.0%1.6× Watch
Routine cosmetic defects2.7%0.7× Normal
Bar: correction-rate lift vs. routine-cosmetic baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What a low-filed report costs

A cycle ends when the report held below the bar nobody set is a standing case. That suite is what the next intake window is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on injury-adjacent reports.

02Diagnose

The corrective action opened on a low-filed report, which stopped no clock and discharged no reporting duty, is worked backwards until one cause is left standing.

03Improve

The queue goes out numbered, and the reports beneath it ride with it.

04Verify

One reportability case still failing is enough to hold the queue back.

05Learn

It is kept for good, and the triage rules are amended in the same commit.

Learn → DetectThe return edge. The next intake meets a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, triage workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Reportability triage and automation-boundary scope.
02Complaint, tracker and event-file sources.
03Reportability-clock and qualified-reviewer mappings.
04Complaint and report intake.
05Report dating and cluster binding.
06Confidence scoring and regulatory routing.
07Qualified-reviewer decision workflow.
08Event-file system integration.
09Reportability and severity cases.
10Guardrails and routing controls.
11Report-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne defect class, one cycle ProductionProduction triage workflow AdvancedMultiple regimes / product lines
Introduced at Pilot
Triage to your reporting rules
Qualified-reviewer decision
Complaint-intake baseline
Introduced at Production
Reporting by defect class
Regulatory review workflow in your systems
Approved event-file write-back
Tracker-and-file integration
Introduced at Advanced
Multi-regime triage rules
Cross-regime report packs
High report volume
Multi-regime deadline controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, report volume, routing controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live defect classes and the duty each one can trigger Triage-rule capture and clock versioningWeek 1
02Representative reports, complaint sets and past decisions Source binding, linking logic and the intake baselineWeek 2
03Your escalation path and the reviewers it names Triage-rule mapping, clock binding and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Complaint, tracker and event-file source assessment, then integration setupWeek 2
05Reports you would not want dated Reportability cases and the failure roundWeek 4
06What no triage may conclude Confidence scoring, review routing, guardrails and hold controlsWeek 3
07A qualified person who decides reportability Handover to the qualified reviewer, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Week five carries two bands because those two phases coincide, not because the column reads better.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Triage-rule discovery, clock versioning and the automation boundary W2Complaint and tracker integration and the intake baseline W3Report dating, linking logic and hold controls W4Evaluation suite, reportability cases and failure-mode testing W5Event-file integration, pilot reports and targeted corrections W6One intake cycle run under the quality lead, then Agent Care handover
Reading the bandEach bar covers the weeks its own work is named for, and week five holds a pair.
At the end of W6When the intake record validates, Agent Care takes the agent on.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Quality assurance AI agent

Build a defect triage agent around the report your last cycle filed below the line.

Show us one report you closed as not reportable. Not how quickly it was triaged. The hour an employee first knew, what the file holds of the reasoning, and whose name stands under the call. A clock that ran while nobody counted comes back as a case.

Nestack Agents · Defect triageAGT-QA-01 · Agent Care available after launch