Nestack Agent Care
Industries / Engineering / R&D / Autonomous coding agent

Engineering AI agent · Autonomous coding

Autonomous Coding AI Agent

Bound each task before the agent starts — the repositories it may touch, the systems it may reach, the changes it may merge — and name the person whose grant carries what follows.

4–6 weeksTypical delivery
Your stackDeployment
Scope grantedNamed owner
Agent CareAfter launch

What this agent does

Works the task, never grants the scope

In
01

A task is scoped, and the grant that opened it is written down with the name that signed it.

02

A change is merged, and no statute required a person to read the diff before it went out.

Reason
03

A task runs to an outcome nobody expected, and UNCITRAL MLAC Art. 7(3) removes surprise as a defence.

04

A change moves a customer-facing term, and UETA § 14 forms the contract unreviewed.

05

A release record is signed, and 21 CFR 11.200(a)(2) reserves that signature to its genuine owner.

Decide
06

A task reaches past the estate, and 18 U.S.C. § 1030(a)(5)(A) needs no unauthorised access.

07

A change is machine-written throughout, and a ticket is a prompt, which is not human control.

Out
08

A change is blamed for an outage, and no company has published a root-cause analysis naming an agent.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Scoping, execution and the record belong to the agent. The grant belongs to a named owner, who authorises the task and answers for the change once it is running.

Example workflow

One task, grant to merge

AgentHuman
1Task receivedA ticket, a failing build, a backlog item or a scheduled maintenance job
2Scope and grant assembledThe repositories, the environments, the credentials, the blast radius and the owner who granted them
3Change producedThe diff, the tests it ran, the systems it reached and confidence
4Controls appliedScope checks, blast-radius checks, third-party reach checks and confidence
No human action required

Stages 1 to 4 run unaided, and nothing merges at any of them — the agent is working, and the owner lane opens at the confidence gate.

5DecisionSplits at the confidence gate
Inside the grant

Goes to the named owner to authorise.

Anything outside it

Adds a delivery-lead read first.

Owner authorisation

The change is held with its scope grant, the systems it reached and the confidence.

Authorise · Narrow scope · Send to lead review
Authorised — by the named owner
6Repository and runtime records updatedOnly where write access and merge policy allow it
7Outcome evaluatedScope adherence, blast radius, owner narrowings and post-merge defects
Narrowings

Each owner narrowing is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Deciding what an agent may touch at all.
Widening a scope after the work has started.
Signing a release record as an individual.
Deciding an unexpected change was acceptable.
Automation boundaryAgent acts unaided
Work the task inside the scope it was granted.
Record the grant, each system reached and the change produced.
Stop the task at the edge of the grant and hold the change there.
Flag a change that reaches outside the estate.
Nothing merges except under a grant a named owner made, inside the agreed boundaries.
Reaching a system the company does not control.
Accepting a contract term on the company.
Saying that a merged change was reviewed.
Changes to scope rules or automation thresholds.

Example output

One scoped task, annotated

The code-generation sibling suggests and an engineer commits; here nobody writes a line, and the only human act left is the grant of scope.

Task output · single changeIllustrative example
Task
Recorded as
Repository
Evidence of record
Confidence
Held for
Retire the legacy billing adapter
Stamped with the grant it ran under
Billing service, main
Scope grant, 3 August 2026
Held unmerged
The named owner, by name
As receivedTaken from the ticket and the scope grant recorded against it, and it claims nothing beyond them.
What the record holds The grant and its date Systems reached Merge target
Why no merge hereTaking a change into production is a judgement a named owner makes.
ActionAuthoriseNarrow scopeSend to lead review
What the score decidesBelow the configured threshold a change gets a delivery-lead read before the owner sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Each taskFrom the grant it runs under
03Authorisation

Which rules reach an agent

The AI Act Art. 14 oversight duty binds providers and deployers of high-risk systems and is deferred to 2 December 2027; the vendor terms already ask the customer to review the output.

01Approved path

Someone authorised this

Restatement (Third) of Agency § 1.01 needs two persons, so the software is not an agent, and the tribunal in Moffatt refused that framing outright.

02Human review

What was checked, and not found

No statute found requires a person to read a code change before it merges, an agent to explain itself, a kill switch or an audit log. The AI exclusions insurers filed name generated content, not autonomous action, and the operative wording could not be obtained.

04Build an evidence trail

The change, the scope it was given and the person who granted it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Source forgesGitHub · GitLab · Bitbucket
Branch, merge and status-check APIs
Task and ticket sourcesJira · Linear · GitHub Issues
Scoped work items and grants
Runtimes and environmentsKubernetes · serverless runtimes
Deployment and rollback boundaries

Agent

Autonomous coding

Reads the task
Works the change
Holds for the owner

Identity and secretsSecret vaults · service accounts
Credential grants and audit trails
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six turnstiles between the model and the merge

Six turnstiles down one passage, the last the stiffest. What enters is set out in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeStop inside the granted scope when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, scope-rule and credential changes, and note the grant each change ran under.Track
L4TraceabilityRecord the task, the grant it ran under, the systems it reached and the change it produced.Record
L3Owner authorisationHold the change for a named owner; the hold governs release, not whether the change is right.Gate
L2Scope guardrailsTest each action against the granted scope in code; a breach stops the task, and scope written only in prose is not what this enforces.Restrict
L1Confidence thresholdsRoute a wide-blast-radius or low-confidence change to a delivery-lead read before the owner sees it.Require review
Model coreChange produced — the diff, the grant it ran under, the systems reached and confidence
L1 – L2Test whether a change may stand
L3Leaves the authorisation to a named owner
L4 – L5Keep the change and the scope behind it
L6Stops inside its scope when signals degrade

How Nestack evaluates it

Evaluate the whole task run — not only the change that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the change that reaches production
Depth of coverage ▼
E1Final-output evaluationDid the change record the grant it was actually carried out under?
E2Step-level evaluationDid the agent read the right repository, the right ticket and the live scope rules?
E3Tool evaluationDid it read and write the correct repository and the correct environment?
E4Confidence calibrationDo low-confidence changes actually attract more owner narrowings?
E5Slice evaluationHow does performance change across specific task classes?
E6Business outcomeHow many changes needed a narrowing before the owner authorised?
Floor — the estate the company must answer for

Failure modes

Where each failure originates in the agent

Seven failure modes, set at the stage where each one first shows.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
QY-03

Stale scope read

The grant read is not the one now in force.

Stage gathersThe ticket, the grant, the systems and the rules
02 · Reasoning2 modes
QY-04

Task widened beyond its grant

The work extends past what was granted.

QY-06

Retired grant read as live

A superseded scope grant is worked as current.

Stage proposesThe diff, the systems reached and the grant
03 · Tool / write2 modes
QY-02

Thin change passed forward

A change moves on without the lead read.

QY-05

Change merged to the wrong target

The work lands on another branch or service.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
QY-01

Merged, nobody read it

The change is running with no read behind it.

Stage returnsThe change that merges and the service that runs it
05 · Change / Version1 mode
QY-07

Silent scope drift

A grant widens while the stored record keeps the old one.

Stage tracksModel, prompt, scope rules and credential grants
Sev-1 · a change merged outside its grant Sev-2 · a wrong change reaches production Sev-3 · signals degrade, the task stops

Affected slices

Cross-service refactors absorb the narrowings

A class-level scope figure reads clean while cross-service refactors carry most of the narrowings and the reach past scope. Nestack reports the narrowing rate by task class, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Cross-service refactors10.6%3.7× Review
Schema and data migrations7.5%2.6× Review
Third-party integrations4.7%1.6× Watch
Single-service bug fixes2.3%0.8× Normal
Bar: narrowing-rate lift vs. single-service baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What an unread merge costs

A cycle closes when the change made outside its granted scope is a case. That suite is what the next task accepted is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Narrowing rate rises on cross-service refactors.

02Diagnose

The change nobody wrote, merged under a grant nobody re-read, is worked backwards until one decision is left standing.

03Improve

Numbered changes leave, with the tasks that prompted them attached beneath.

04Verify

One red scope case is enough to hold the whole release.

05Learn

It is kept permanently, and the scope rules change in that same commit.

Learn → DetectThe return edge. The next task is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, task workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Merge workflow discovery and automation-boundary work.
02Forge, runtime and identity assessment.
03Scope-rule mapping and credential-boundary capture.
04Task and grant capture.
05Scope binding and change execution.
06Blast-radius scoring and review routing.
07Owner authorisation workflow.
08Forge and runtime integration.
09Scope and authorisation cases.
10Guardrails and merge controls.
11Task-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne service, one team ProductionProduction engineering workflow AdvancedMultiple services / teams
Introduced at Pilot
Task execution to your scope rules
Named owner authorisation
Scope-grant baseline
Introduced at Production
Reporting by task class
Delivery-lead review workflow in your forge
Approved merge and deploy
Forge-and-runtime integration
Introduced at Advanced
Multi-service scope rules
Multi-stage owner approvals
Large task volumes
Multi-repository scope controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, task volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live services and the scope an agent may be granted Scope-grant capture and boundary versioningWeek 1
02Representative tickets, branches and merge history Task binding, scope rules and the grant baselineWeek 2
03Your delivery calendar and the owners it names Scope-rule mapping, grant capture and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Forge, runtime and identity assessment, then integration setupWeek 2
05Changes you would not want unattributed Authorisation cases and the roundWeek 4
06What no task may authorise Blast-radius scoring, review routing, guardrails and release controlsWeek 3
07A named owner who authorises the change Owner sign-off workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Every band is drawn to its phase cost, so two of them share week five and nothing was stretched.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Delivery workflow discovery, scope-rule mapping and the automation boundary W2Forge and runtime integration and the scope-grant baseline W3Task workflow, confidence logic and merge controls W4Evaluation suite, authorisation cases and failure-mode testing W5Runtime integration, pilot tasks and targeted corrections W6One delivery cycle run under the engineering owner, then Agent Care handover
Reading the bandEach bar spans the weeks its own work is named in, and the week five overlap is real, not padding.
At the end of W6Once the scope record validates, Agent Care picks the agent up.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Engineering AI agent

Build an autonomous coding agent around a scope grant somebody actually signed.

Show us one service and the last change merged into it without a person reading the diff. Somebody authorised that: the engineer who granted the agent its scope, whose name is probably written nowhere on the change. Name them before an incident review picks one for you.

Nestack Agents · Autonomous codingAGT-ENG-06 · Agent Care available after launch