Nestack Agent Care
Industries / Quality Assurance / Release gate agent

Quality AI agent · Release decisions

Release Gate Copilot AI Agent

Green means the checks ran, not that the product works, so the agent assembles the evidence, lists the anomalies you are shipping knowingly, and holds the release for the person who owns it.

4–6 weeksTypical delivery
Your stackDeployment
Evidence firstNamed owner
Agent CareAfter launch

What this agent does

Assembles the pack, never releases it

In
01

A build is put up for release, and for most products no law requires that any check pass first.

02

A gate turns green, and that says the checks ran, which is not the same as the product working.

Reason
03

The product is a device or an aircraft system, and the same pack becomes a controlled record.

04

An anomaly ships knowingly, and IEC 62304 clause 5.8 has it documented and each one evaluated.

05

A coverage figure is met, and stronger forms give no more insight; MC/DC is a Level A duty.

Decide
06

A suite is quarantined, and the pack says so rather than reporting a pass over what still ran.

07

An exit rule is waived, and the waiver becomes a record rather than a quiet edit to the criteria.

Out
08

A release cannot be evidenced against its own criteria, and that is recorded as the finding.

09

Execute write actions only inside the approval boundaries fixed during implementation.

Product statement

Assembly and the pack record belong to the agent; the release belongs to a named person. FDA proposed carrying the record signature into 820.35, then deleted it in the final rule.

Example workflow

One release, evidence to decision

AgentHuman
1Evidence receivedTest results, defect records, coverage reports, waivers and the criteria in force
2Release class settledWhether nothing binds this release at all, or a quality system and a declaration of conformity do
3Pack assembled against that classThe criteria, the results against them, the residual anomalies and the gaps
4Checks appliedCriteria coverage, citation currency, waiver completeness and assembly confidence
No human action required

Stages 1 to 4 run unaided, and nothing ships at any of them — the agent is assembling, and the owner lane opens at the release gate.

5DecisionSplits at the release gate
Complete against the fixed class

Goes to the named quality owner to release.

Anything with a gap

Adds a regulatory-affairs read before that.

Owner review

The pack is held with its criteria, its residual list and the evidence each result came from.

Release · Append evidence · Send to regulatory review
Released — by the named quality owner
6Quality and release records updatedOnly where write access and the release-records policy allow it
7Release scoredCriteria accuracy, evidence ties, owner corrections and what regulatory review found
Quality-owner corrections

Each correction the owner makes is counted in the release evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Releasing a build to a market.
Deciding a residual anomaly is acceptable.
Judging which regime reaches this release.
Signing a declaration of conformity.
Automation boundaryAgent acts unaided
Fix the exit criteria before the evidence is read.
Assemble the residual-anomaly list and mark the rationales missing.
Hold the pack, its results and its gaps in one record.
Record the waiver, its justification and the approval it carries.
Nothing ships to a market except by a named quality owner, inside the agreed boundaries.
Deciding a known defect will harm nobody.
Telling a regulator the checks were adequate.
Choosing the exit criteria a release is held to.
Changes to the criteria, the pack or the record.

Example output

One release pack, annotated

This serves a quality team whose submission must give a rationale for each anomaly left unfixed; below is one release exactly as the agent leaves it.

Release pack · one candidate buildIllustrative example
Release
Recorded as
Class
Evidence of record
Confidence
Held for
Candidate build, one regulated product
Assembled against the criteria in force
Regulated, quality system
Test records, 7 August 2026
Held unreleased
The quality owner, by name
As receivedBuilt from the test records and the defect file behind them, and it claims nothing past either one.
What the pack record holds Test records Defect file Waiver register
Why no release hereDeciding that a known defect may ship is a call a named person makes.
ActionReleaseAppend evidenceSend to regulatory review
What the score decidesBelow the configured threshold a pack gets a regulatory read before the owner sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Each release assembledFrom the criterion that gates it
03Evidence

Where that evidence lands

The agent does not vouch for a check, only for what it reported and which criterion it answered; the standards behind those criteria are paywalled, so the pack cites clause numbers and quotes none.

01Approved path

Green means the checks ran

A warning letter of 23 January 2026 cited finished sensors released without accuracy testing on the finished device, under two sections reserved ten days later.

02Human review

Looked for, and not found

No section of the current Code of Federal Regulations requires a test to pass before ordinary software ships, and residual anomalies is nowhere in it; Regulation (EU) 2026/1744 defers the AI Act declaration to 2 December 2027 for Annex III and 2 August 2028 for Annex I.

04Build an evidence trail

The build, the evidence assembled for it and the person who released it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Test and verification systemsTest runners · result archive
Planned suites and executed runs
Defect and anomaly recordsIssue tracker · anomaly file
Open anomalies, by severity
Quality management recordsQMS · procedures and criteria
Exit criteria and the waivers on file

Agent

Release gate assembly

Reads the evidence
Assembles the pack
Holds for the owner

Standards and rule libraryClause registers · rule library
Clause numbers and where they stand
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six locks between the model and the release

Six locks along one channel, the last the heaviest. What passes is set out in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeNarrow the agent to listing results when evaluation or production signals degrade.Roll back
L5Version trackingTrack model, prompt and criteria logic, and note the version each pack was assembled under.Track
L4Pack trailRecord each pack, the results beneath it, the criteria it was built on and each read of the file.Record
L3Owner releaseHold the pack for a named owner; MDR Article 15(3)(a) puts the check before release on one registered person.Gate
L2Exit-rule guardrailsTest each criterion against the rule it cites, and refuse 21 CFR 820.30 or 820.80, reserved since February 2026.Restrict
L1Confidence routingRoute a low-confidence pack to a regulatory read before the release reaches the owner.Require review
Model corePack assembled — the criteria, the results, the residual list and its evidence
L1 – L2Test whether a pack may stand
L3Leaves the release to a named owner
L4 – L5Keep the release and the residual list behind it
L6Presents the pack unsigned when signals degrade

How Nestack evaluates it

Evaluate the whole assembly run — not only the pack that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the pack a release decision rests on
Depth of coverage ▼
E1Final-output evaluationDid the pack record the criteria it was actually assembled against?
E2Step-level evaluationDid the agent read the right build, the right suite and the live criteria?
E3Tool evaluationDid it read and write the correct build and the correct result?
E4Confidence calibrationDo low-confidence packs actually attract more owner corrections?
E5Slice evaluationHow does performance change across classes, from a patch to a Level A build?
E6Business outcomeHow many packs needed a correction before the owner released?
Floor — the evidence a release rests on

Failure modes

Where each failure originates in the agent

Seven failure modes, set at the stage each first shows itself.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
VF-02

Skipped suites read as green

The pass rate covers what ran, not what was planned.

Stage gathersThe results, the defects, the waivers and the rules
02 · Reasoning2 modes
VF-04

Coverage read as correctness

A met number is read as proof the product works.

VF-05

Anomaly listed, never evaluated

The list exists; no rationale sits beside it.

Stage proposesThe criteria, the results, the gaps and confidence
03 · Tool / write2 modes
VF-03

Override left off the record

A waiver is taken under deadline and not logged.

VF-07

Approval logged without attribution

The record shows approval but not who held what.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
VF-06

Criteria aimed at the wrong object

Each criterion passes; none touches what ships.

Stage returnsThe pack an owner signs and an auditor reads
05 · Change / Version1 mode
VF-01

Superseded rule read as live

A reserved section is cited as though in force.

Stage tracksModel, prompt, exit rules and clause dates
Sev-1 · a release out on no evaluation Sev-2 · a superseded rule reaches the pack Sev-3 · signals degrade, pack held unsigned

Affected slices

Regulated first releases absorb the corrections

A programme-level correction figure can read clean while regulated first releases carry most of the rework. Nestack reports the correction rate by release class, not only across the programme.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Regulated first releases12.1%3.7× Review
Cross-market submissions8.6%2.6× Review
Waived exit criteria5.4%1.7× Watch
Routine patch releases2.5%0.8× Normal
Bar: correction-rate lift vs. routine-release baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What an unevaluated anomaly costs

The loop shuts when the anomaly shipped without a written evaluation is a standing case. That suite is what the next release pack is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on regulated first releases.

02Diagnose

The green gate that tested a part rather than the finished device, with eight hundred and sixty serious injuries and seven deaths behind it, is worked backwards until one cause is left standing.

03Improve

Every pack goes out numbered, with the evidence that filled it attached.

04Verify

Nothing ships while one touched acceptance case is still red.

05Learn

The case stays on, and the exit rules are rewritten alongside it.

Learn → DetectThe return edge. The next release pack meets a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, gate assembly, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Exit-criteria capture and automation-boundary work.
02Test, defect and criteria sources.
03Result-to-criterion and residual-anomaly list mapping.
04Test-result ingestion.
05Criteria, build and result binding.
06Evidence-gap scoring and review routing.
07Owner release workflow.
08Tracker and QMS integration.
09Acceptance and residual cases.
10Guardrails and release controls.
11Evidence-pack instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne product, one release ProductionProduction release workflow AdvancedMultiple products / markets
Introduced at Pilot
Pack assembly to your exit criteria
Named quality-owner release
Exit-criteria baseline
Introduced at Production
Reporting by build class
Regulatory review workflow in your systems
Approved records write-back
Tracker-and-pipeline integration
Introduced at Advanced
Multi-standard criteria sets
Cross-release evidence packs
Dense release calendars
Multi-market conformity controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, evidence volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live release classes and the criteria each one is held to Criteria capture and exit-rule versioningWeek 1
02Representative test, defect and waiver material Evidence binding, assembly logic and the release baselineWeek 2
03Your release schedule and the owners it names Criteria mapping, evidence binding and the automation boundaryWeek 1
04Access to the relevant runners, trackers or exports Test, tracker and QMS assessment, then integration setupWeek 2
05Releases you would not want reconstructed Acceptance cases and the exit roundWeek 4
06What no gate may certify Evidence-gap scoring, review routing, guardrails and release controlsWeek 3
07A named quality owner who releases Quality-owner sign-off, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Widths follow what a phase costs rather than the grid, so week five draws one band beneath another.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Criteria discovery, exit-rule versioning and the automation boundary W2Test and tracker integration, and the exit-criteria baseline W3Assembly logic, residual-anomaly handling and release controls W4Evaluation suite, acceptance cases and failure-mode testing W5QMS and records integration, pilot packs and targeted corrections W6One release cycle run under the quality owner, then Agent Care handover
Reading the bandEach bar runs across the weeks its own work is named for, and week five carries two.
At the end of W6Once the release record validates, Agent Care takes the agent up.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Release gate AI agent

Build a release gate agent around the anomalies your last release shipped without a written evaluation.

Show us one release you ship on a schedule and the known defects that went out with it. If nothing beside them records why each was acceptable, that list is an inventory rather than a decision. A release nobody can evidence comes back as a finding.

Nestack Agents · Release decisionsAGT-QA-02 · Agent Care available after launch