Nestack Agent Care
Industries / Finance & Accounting / Journal-entry anomaly detection

Finance AI agent · Entry scoring

Journal-Entry Anomaly Detection AI Agent

Score what has already posted, keep the reason beside the score, and hand a reading order — not a verdict — to the controller whose name goes on the certification.

4–6 weeksTypical delivery
Your stackDeployment
Queue firstThe controller
Agent CareAfter launch

What this agent does

Scores the entry, never condemns it

In
01

An entry is scored, and the signals that raised the score are recorded beside the entry itself.

02

Odd is not wrong: what comes out is a queue in reading order, not a finding about an entry.

Reason
03

Fraudulent entries are rare enough that most of what a good detector flags is sound accounting.

04

AS 2401.61 names round numbers and seldom-used accounts; odd hours are practice, not authority.

05

Certification paragraph 5(b) reaches fraud whether or not material, where an ICFR role is held.

Decide
06

Rule 13a-14(c) bars signing that certification by power of attorney, so two named people sign.

07

Handed over, the queue is company-produced information AS 1105.10 makes the auditor test.

Out
08

No rule requires a company to score its own entries, and that absence is recorded, not filled.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Extraction, scoring and holding belong to the agent. Disposition belongs to a named controller, who clears the entry and answers for the log at the close.

Example workflow

One entry, extract to disposition

AgentHuman
1Entry population extractedGeneral ledger and sub-ledger postings, manual and automated, by entity and by period
2Population reconciled and datedThe entity, the period, the line totals, the debits and credits agreed to trial balance and the day of extract
3Signals raised and scoredThe account pairing, the poster, the timing, the round-number pattern and the score
4Controls appliedPopulation checks, poster checks, threshold checks and scoring confidence
No human action required

Stages 1 to 4 run unaided, and no entry is called improper at any of them — the agent is scoring, and the controller lane opens at the disposition gate.

5DecisionSplits at the disposition gate
A score inside tolerance

Goes to the named controller to clear.

Anything top-side

Adds an assistant-controller read first.

Controller review

The entry is held with its score, its signals and the support the agent could not locate.

Clear · Request support · Send to controller review
Cleared — by the named controller
6Close and control records updatedOnly where write access and records policy allow it
7Outcome evaluatedClearance outcomes, queue age, controller corrections and what the audit committee was told
Corrections

Each disposition the controller changes counts in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Concluding that an entry is improper.
Signing the Exhibit 31 certification.
Disclosing fraud to the auditors.
Clearing an entry off the review queue.
Automation boundaryAgent acts unaided
Score each posted entry against its own history.
Keep the population reconciliation beside each scored ledger period.
Separate the signals AS 2401 names from the practitioner ones.
Hold each flagged entry for the named controller.
No entry leaves the queue except by a named controller, inside the agreed boundaries.
Judging whether an entry is supported at all.
Telling the auditor the population is complete.
Choosing the threshold a queue is drawn at.
Changes to the scoring rules or the queue gate.

Example output

One flagged entry, annotated

This serves the controller who owns the ledger, and whose work the auditor considers under AS 2605 where internal audit runs the queue; below is one flagged entry exactly as the agent leaves it.

Flagged entry · single postingIllustrative example
Entry
Recorded as
Scored
Evidence of record
Confidence
Held for
Top-side adjustment, one entity
Round number, post-closing, unsupported
Signals only
Ledger extract, 12 August 2026
Held uncleared
The named controller, by name
As receivedDrawn from the posting record, the approval trail and the account history, and it asserts nothing further.
What the record holds Posting record Approval trail Account history
Why no clearance hereClearing a flagged entry is a judgement the named controller makes.
ActionClearRequest supportSend to controller review
What the score decidesBelow the configured threshold an entry gets an assistant-controller read first.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Each flagged entryFrom the ledger it sits in
03Score

Where the score is read

Rule 13b2-1 binds any person, not only officers: in AAER-4562 of 4 February 2025 it was a controller who took a one-year bar over topside adjustments.

01Approved path

Odd is not the same as wrong

A round number is an accrual, an allocation or a rent charge far more often than a fraud, and the same holds for the late entry and the unfamiliar poster.

02Human review

What was checked, and not found

No rule located, in the United States or in Europe, obliges a company to score its own journal entries for anomalies. SAF-T and GoBD force ledger data into a machine-readable export; neither makes detection a duty, and that absence is reported rather than filled.

04Build an evidence trail

The entry, the score it drew and the reviewer who cleared it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

General ledger and sub-ledgersERP ledgers · sub-ledger feeds
Posted entries and their lines
Approval and workflow recordsWorkflow · approval histories
Who posted and who approved
Close and consolidationClose checklists · consolidation tools
Top-side and post-close items

Agent

Journal-entry scoring and review

Reads the ledger
Scores the entries
Holds the queue

Control and evidence recordsSOX testing evidence · deficiency logs
Control results and outcomes
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six sieves between the score and the close

Six sieves down one column, the last the closest. What settles is set out in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeNarrow the agent to population extract when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and threshold versions, and note the version each entry was scored under.Track
L4TraceabilityRecord each score, the signals under it, the population it came from and each read of the log.Record
L3Controller releaseHold the entry for a named controller; the hold governs disposition, not whether the entry is wrong.Gate
L2Population guardrailsTest each score against its reconciled population, and refuse one drawn from a partial extract.Restrict
L1Confidence thresholdsRoute a top-side or thinly supported entry to an assistant-controller read before the controller sees it.Require review
Model coreEntry scored — the signals, the population it came from, the threshold and what stays unexplained
L1 – L2Test whether an entry may be cleared
L3Leaves the disposition to a named controller
L4 – L5Keep the entry and the score behind it
L6Queues the entry for a human read when signals degrade

How Nestack evaluates it

Evaluate the whole scoring run — not only the entry that reaches the queue.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the queue a controller works
Depth of coverage ▼
E1Final-output evaluationDid the score stay inside what the signals actually establish?
E2Step-level evaluationDid the agent read the right ledger, the right period and the live threshold?
E3Tool evaluationDid it read the correct entry and write the correct disposition record?
E4Confidence calibrationDo low-confidence scores actually attract more controller corrections?
E5Slice evaluationHow does performance change across entry classes and entities?
E6Business outcomeHow many entries needed a second read before the controller cleared one?
Floor — the log read back at the audit

Failure modes

Where each failure originates in the agent

Seven failure modes, each at the stage where it first surfaces.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
VG-03

Population never reconciled

A ledger is missing and the score reads clean.

Stage gathersThe entries, the posters, the hours and the accounts
02 · Reasoning2 modes
VG-04

Pilot precision does not hold

A tuned sample flatters the live rate.

VG-06

Re-baselined after a change

An ERP move resets what counts as normal.

Stage proposesThe signals, the score, the threshold and the version
03 · Tool / write2 modes
VG-02

Boundary learned by the poster

Flags fall away and it reads as improvement.

VG-05

Top-side layer left unscored

The consolidation entries never reach it.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
VG-01

Queue stops being cleared

Flags pile up with no disposition recorded.

Stage returnsThe queue a controller works and then clears
05 · Change / Version1 mode
VG-07

Threshold lowered quietly

Sensitivity drops with no record of why.

Stage tracksModel, prompt, scoring rules and thresholds
Sev-1 · a queue nobody cleared Sev-2 · a partial ledger scored clean Sev-3 · signals thin, entry held back

Affected slices

Top-side entries absorb the corrections

An entity-level clearance figure can read clean while top-side adjustments carry most of the controller corrections. Nestack reports the correction rate by entry class, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Top-side and consolidation entries11.6%3.7× Review
Post-close manual entries8.2%2.6× Review
Unfamiliar poster entries5.1%1.6× Watch
Routine automated entries2.3%0.7× Normal
Bar: correction-rate lift vs. automated-entry baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What an unread queue costs

The loop shuts when the flagged entry nobody ever opened is a standing case. That suite is what the next ledger period scored is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on top-side and consolidation entries.

02Diagnose

The large, round number, unauthorized and unsupported topside entries that ran for two years before anyone stopped them are worked backwards until one cause is left standing.

03Improve

Every review goes out numbered, with the entries that raised it attached.

04Verify

Nothing closes while one touched review case is still red.

05Learn

The case stays on, and the scoring rules are redrawn beside it.

Learn → DetectThe return edge. The next ledger period is scored against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, scoring logic, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Entry-scoring scope and the automation-boundary work.
02Ledger, workflow and close sources.
03Certification-duty and population-boundary mapping.
04Ledger entry extraction.
05Signal binding and scoring logic.
06Score confidence and review routing.
07Controller disposition workflow.
08Close-system integration.
09Scoring and disposition cases.
10Guardrails and review controls.
11Entry-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne ledger, one close cycle ProductionProduction review workflow AdvancedMultiple ledgers / entities
Introduced at Pilot
Scoring to your close calendar
Named controller disposition
Ledger-population baseline
Introduced at Production
Reporting by entry class
Controller review workflow in your systems
Approved close write-back
Ledger-and-workflow integration
Introduced at Advanced
Multi-ledger scoring rules
Cross-entity review packs
High entry volume
Multi-ledger scoring controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, entry volume, disposition controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live ledgers and the close calendar each one runs on Threshold capture and model versioningWeek 1
02Representative entries, extracts and past dispositions Signal binding, scoring logic and the ledger-population baselineWeek 2
03Your escalation path and the controller it names Threshold mapping, signal binding and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Ledger, workflow and close source assessment, then integration setupWeek 2
05Entries you would not want produced Disposition cases and the evaluation roundWeek 4
06What no score may conclude Score confidence, review routing, guardrails and release controlsWeek 3
07A named controller who clears the entry Handover to the named controller, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Widths follow the work and not the grid, which is why one band sits beneath another in the fifth week.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Threshold discovery, model versioning and the automation boundary W2Ledger and workflow integration and the population baseline W3Signal binding, scoring logic and release controls W4Evaluation suite, disposition cases and failure-mode testing W5Close-system integration, pilot scoring runs and targeted corrections W6One close cycle run under the controller, then Agent Care handover
Reading the bandEach bar stops at the weeks its own work is named for, and week five carries the one overlap.
At the end of W6Once the review record validates, Agent Care assumes the agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Finance AI agent

Build a journal-entry anomaly agent around the queue your last close never cleared.

Show us one close and the entries nobody could explain afterwards. Not how fast the model flagged them. Which ledgers the extract actually covered, what the score established and did not, and whose name goes under the certification. A queue nobody worked comes back as a case.

Nestack Agents · Entry scoringAGT-FIN-02 · Agent Care available after launch