Nestack Agent Care
Industries / Accounting / Evidence-testing agent

Accounting AI agent · Audit

Audit Evidence-Testing Agent (Auditor-Side)

Run the tests of details you scope — population checks, documented selections, three-way matching and journal-entry work — and write up every exception, with evaluation, audit trails and the sufficiency judgement left with the auditor.

4–6 weeksTypical delivery
Your stackDeployment
Before relianceHuman review
Agent CareAfter launch

What this agent does

Runs the procedure and documents what it found

In
01

Ingest the population extract, the audit programme step and the documents the test calls for.

02

Normalise each item into entity, period, account, assertion and the test it belongs to.

Reason
03

Agree the extract to the ledger control total before a single item is drawn from it.

04

Draw the selection the auditor specified, and record the method, parameters and seed used.

05

Perform the defined test on every item selected — match, recalculate, agree the date, reperform.

Decide
06

Flag exceptions, items the test could not be applied to, and evidence that misses the assertion.

07

Return every exception and every untested item to the auditor instead of clearing it.

Out
08

Retain the extract, the selection, each item tested, the evidence seen and the result.

09

Write to the workpaper file only inside the approval boundaries agreed during implementation.

Product statement

The agent executes the procedures the auditor scoped and documents what it found; it gives no assurance, and the auditor judges sufficiency and forms every conclusion.

Example workflow

One procedure, end to end

AgentHuman
1Procedure receivedAudit programme step, workpaper request or the engagement team's own scope
2Scope resolvedEntity, period, account, assertion, population definition and the selection the auditor set
3Test performedExtract agreed to the control total, the specified selection drawn, then each item tested
4Controls appliedPopulation-completeness check, evidence-to-assertion check, source-data checks and confidence threshold
No human action required

Stages 1 to 4 run without a person in the loop — nothing has been relied on yet, so the lane stays empty until the gate.

5DecisionSplits on the confidence and exception gate
High confidence

Recorded as a worked step with its evidence.

Low confidence

Held with the exception it could not resolve.

Auditor review

The step is held with its population, selection, evidence and the exception that stopped it.

Accept · Amend · Extend testing
Accepted — handed back
6Step written to the fileRecorded as work performed, never as a conclusion; the auditor accepts the step
7Outcome evaluatedException precision, population-completeness checks, auditor overturns and business outcome
Overturns

Amendments and extended testing at review are counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Whether the evidence is sufficient and appropriate.
Sample size, method and materiality decisions.
Any conclusion on an assertion.
Whether an exception is a misstatement.
Automation boundaryAgent acts unaided
Agree the population extract to the ledger control total.
Draw the selection the auditor specified and record how.
Perform the defined test on every item selected.
Document each exception and hand it to the auditor.
Workpapers are written only inside the approval boundaries agreed during implementation.
Treating an exception as an isolated anomaly.
Accepting entity-produced data as reliable.
Extending, reducing or stopping a test.
Sign-off of any workpaper or audit step.

Example output

One tested item, annotated

Everything the agent records is attached to the procedure and the population it came from.

Worked step · single procedureIllustrative example
Procedure
Population
Period
Result
Confidence
Population check
Three-way match
Posted supplier invoices
FY25
Two open exceptions
89%
Agrees to the control total
As receivedThe procedure, population and period as the audit programme defines them — the agent redefines neither.
Evidence used Purchase orders Goods-received notes Posted supplier invoices
Why it holdsBoth exceptions are receipts with no matching order — whether that is a misstatement is the auditor's call.
ActionAcceptAmendExtend testing
What the score decidesBelow the configured threshold the step is held for the auditor instead of being recorded as done.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
All scoped proceduresFrom the audit programme
03Execution

Apply the engagement's own scope

Use the population definition, materiality and selection method the engagement team set, never a default.

01Approved path

Cut the tick-and-tie hours

Routine, in-scope tests come back performed, documented and repeatable instead of being worked item by item.

02Human review

Put auditors on the exceptions

Exceptions, unmatched items and population gaps reach the auditor instead of every item doing so.

04Build an evidence trail

Retain the extract, the selection and its parameters, every item tested, the evidence seen, evaluator result and auditor decision — on both paths.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Audit managementCaseware · CCH Axcess Engagement
TeamMate+ · Inflo
ERP, GL & journalsSAP · Oracle · NetSuite
Dynamics 365 · GL and journal extracts
Documents & extractionNetDocuments · iManage · SharePoint
DataSnipper · IDEA

Agent

Audit evidence testing

Reads the population
Performs the test
Routes exceptions

ConfirmationsConfirmation.com · Bank portals
Legal confirmations · AR confirmations
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the workpaper

Each control wraps the one inside it. A test clears every layer before it is recorded as work performed.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeRestrict automation if evaluations or production signals degrade.Roll back
L5Version monitoringTrack ledger re-extracts, exception rules, model and prompt changes.Track
L4Reproducible recordRecord scope, extract, selection parameters, items tested and evidence seen.Record
L3Auditor acceptanceDefine which tests may be recorded without a named auditor accepting them.Gate
L2Assurance boundaryThe agent performs and documents; it never concludes.Restrict
L1Population checksThe extract must agree to the ledger control total.Reconcile
Model coreTest result proposed — items tested, evidence seen, exceptions and confidence
L1 – L2Decide whether the test may stand
L3Decides whether an auditor accepts it
L4 – L5Keep the run repeatable and the rules visible
L6Pulls automation back when signals degrade

How Nestack evaluates it

Evaluate the full workflow — not only the exception list.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the step the auditor reviews
Depth of coverage ▼
E1Final-output evaluationDid the test address the assertion it was scoped to?
E2Population evaluationDid the extract agree to the ledger it was drawn from?
E3Selection evaluationWas the draw made as specified, and can it be repeated?
E4Exception calibrationAre flagged exceptions real, and are real ones flagged?
E5Slice evaluationHow does performance change across specific test cohorts?
E6Business outcomeHow many steps were overturned or re-performed at review?
Floor — the outcome the client pays for

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle.

Agent lifecycleDirection of processing →
01 · Retrieval2 modes
AE-01

Selection drifts from the plan

Stratification or interval no longer matches the agreed sampling plan.

AE-02

Reperformance not reproducible

Re-running the step against the same source returns a different result.

Stage gathersPopulation extract, source documents and prior evidence
02 · Reasoning2 modes
AE-03

Untested entity report used

A company-produced report is relied on without being tested.

AE-04

Assertion not evidenced

A document shows existence but not the assertion tested.

Stage proposesThe selection, the test result, exceptions and confidence
03 · Tool / write1 mode
AE-05

Recorded before review

A step enters the file as though an auditor had accepted it.

Stage writesOnly where file access and approval boundaries allow it
04 · Output1 mode
AE-06

Exception closed by rule

An item is cleared on a threshold instead of investigated.

Stage returnsThe worked step the auditor and the file see
05 · Change / Version1 mode
AE-07

Narrowed exception rules

A rule change quietly shrinks what counts as an exception.

Stage tracksLedger re-extracts, rule, model and prompt changes
Sev-1 · untested work recorded as tested Sev-2 · wrong evidence behind the item Sev-3 · testing degrades, step goes to review

Affected slices

Exception risk concentrates by procedure type

Journal-entry work and scanned evidence behave nothing like a three-way match, and they carry most of the missed and mis-flagged exceptions. Nestack reports performance by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Journal-entry testing runs6.3%3.5× Review
Scanned-document evidence5.4%3.0× Review
Confirmation exception runs3.4%1.9× Watch
Three-way match runs0.9%0.5× Normal
Bar: lift vs. three-way-match baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

Scoring comes from the auditor, not the agent

What the reviewer changed — a missed exception, or one raised that was not real — is what the next release is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Overturns rise in a test cohort, or exception precision drops.

02Diagnose

Failure isolated to the extract, the draw, the evidence match or the rule.

03Improve

The rule or the match logic is corrected — approved, version-linked and dated.

04Verify

Archived engagement data is re-run and the two results are compared.

05Learn

The overturned item becomes a case in the regression suite.

Learn → DetectThe return edge. A release that cannot reproduce last cycle's overturns does not ship.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, populations and extracts, test execution, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Procedure scoping and assurance-boundary definition.
02Source-system and workpaper assessment.
03Population definition and control-total mapping.
04Extract reconciliation and selection logic.
05Test execution, matching and recalculation.
06Exception rules and confidence scoring.
07Auditor review and acceptance workflow.
08Workpaper and document-system integration.
09Evaluation suite and regression cases.
10Guardrails and documentation controls.
11Observability and trace instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them. No tier adds assurance or the sufficiency judgement.

Capability✓ in scope · — not at this tier PilotOne procedure ProductionProduction integration AdvancedGroup engagements
Introduced at Pilot
Defined-procedure execution
Population completeness checks
Auditor acceptance
Baseline evaluation
Introduced at Production
Reproducible selection records
Exception write-ups
Approved workpaper writes
Observability and evaluation
Introduced at Advanced
Group and multi-entity testing
Multi-stage review
Enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on the procedures in scope, population volumes, source systems, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01The procedures in scope and the programme steps they sit under Procedure scoping and the population definition for eachWeek 1
02The assurance boundary and who accepts each step Assurance-boundary definition and acceptance controlsWeek 1
03Access to the ledger, the extracts and the document store Source-system and workpaper assessmentWeek 2
04Population definitions, control totals and materiality Extract reconciliation and control-total mappingWeek 2
05Selection methods and the exception thresholds you use Selection logic, exception rules and confidence scoringWeek 3
06Exceptions your reviewers overturned last year Evaluation suite, regression cases and failure-mode testingWeek 4
07Named auditors or test users Auditor acceptance workflow, then pilot testing and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. Weeks 5 and 6 run a real audit step, which is the only place the exception rules meet live data.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Procedure scoping, assurance boundary and population definitions W2Ledger and document access, extract reconciliation and the selection baseline W3Test execution, exception rules and auditor acceptance controls W4Evaluation suite, documentation guardrails and failure-mode testing W5Workpaper integration, a pilot procedure and targeted corrections W6One audit step run end to end, verification and Agent Care handover
Reading the bandA bar covers only the weeks its work is named in. Week 5 carries two because the pilot procedure runs against the evaluation suite, not after it.
At the end of W6One procedure has been run end to end and accepted by a named auditor, then Agent Care monitors population checks, exception precision and overturns.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Accounting AI agent

Build an evidence-testing agent around your audit programme.

Show us the procedures you want run, the populations they draw from and who accepts each step. We'll take one procedure end to end, agree what the agent may never conclude and scope the rest from there.

Nestack Agents · Audit evidence-testing agentAGT-ACC-09 · Agent Care available after launch