Nestack Agent Care
Industries / Accounting / Tax copilot

Accounting AI agent · Tax

Tax-Question Copilot (Jurisdiction-Aware)

Answer internal tax questions with a drafted response, the jurisdiction and tax year it applies to and every authority cited — held for a qualified reviewer, with evaluation, audit trails and guardrails built in.

4–6 weeksTypical delivery
Your stackDeployment
Before releaseHuman review
Agent CareAfter launch

What this agent does

Drafts the answer and shows its authority

In
01

Ingest tax questions from supported chat, email, ticketing or research-tool channels.

02

Normalise each question into jurisdiction, tax year, entity type and the facts supplied.

Reason
03

Retrieve statute, regulations, rulings and published guidance in force for that period.

04

Apply the firm's precedent library, agreed house positions and the engagement file.

05

Draft an answer that separates the rule, its application and what it does not cover.

Decide
06

Check every citation exists, still stands and supports the sentence it sits under.

07

Route advice-class, unsettled and low-confidence questions to a qualified reviewer.

Out
08

Retain the question, retrieved authorities, draft, confidence and reviewer decision.

09

Release or write answers only inside the approval boundaries agreed during implementation.

Product statement

The agent drafts and cites inside the approval boundaries agreed during implementation; a qualified person signs the position.

Example workflow

One question, end to end

AgentHuman
1Question receivedChat, email, ticket or the firm's research tool
2Scope resolvedJurisdiction, tax year, entity type and the facts already on the engagement file
3Authority retrievedStatute, regulations, rulings, published guidance and firm precedent in force for the period
4Controls appliedCitation-existence checks, currency checks, scope guardrails and confidence threshold
No human action required

Stages 1 to 4 run without a person in the loop — nothing has left the agent yet, so the lane stays empty until the gate.

5DecisionSplits on the confidence and question-class gate
High confidence

Returns to the requester as cited draft research.

Low confidence

Held for a qualified reviewer.

Reviewer sign-off

The draft is held with its citations, jurisdiction, tax year and confidence.

Approve · Amend · Escalate
Signed — handed back
6Answer issued and filedReturned to the requester as draft; issued outside the firm only after the reviewer signs
7Outcome evaluatedCitation accuracy, authority currency, reviewer amendments, escalations and business outcome
Amendments

Corrections and escalations made in review are counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Any answer issued outside the firm.
Filing positions and disclosure decisions.
Questions where the authorities genuinely conflict.
Questions turning on facts not on file.
Automation boundaryAgent acts unaided
Retrieve authority in force for the jurisdiction and period.
Draft the answer with its citations and its reasoning.
Check that cited authority exists, still stands and applies.
Route advice-class and low-confidence questions to review.
Answers are released and written only inside the approval boundaries agreed during implementation.
Positions resting on non-precedential written determinations.
Answers on legislation still inside transition.
Anything framed as planning or structuring advice.
Changes to the authority corpus or confidence thresholds.

Example output

One drafted answer, annotated

Everything the agent proposes is attached to the question and the authority it came from.

Draft answer · single questionIllustrative example
Question
Entity
Tax year
Draft answer
Confidence
Authority cited
Availability of a new relief
Multi-state operating company
2025
Available federally; states differ
88%
Statute, regulation, notices
As receivedThe question, entity and tax year exactly as the requester gave them — the agent does not assume a jurisdiction that was not stated.
Evidence used Statute in force Current regulation text State conformity notices
Why it splitsThe states in scope do not adopt the federal provision for this period, so the two answers are kept apart.
ActionApproveAmendEscalate
What the score decidesBelow the configured threshold the draft is held for a qualified reviewer instead of returning to the requester.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
All tax questionsFrom the firm's channels
03Research

Apply the firm's own positions

Use the precedent library, agreed house positions and the facts already on the engagement file.

01Approved path

Cut the routine research time

Settled, in-scope questions come back drafted and cited instead of being researched from scratch.

02Human review

Put reviewers on the hard questions

Advice-class, unsettled and low-confidence questions reach a qualified reviewer instead of every question doing so.

04Build an evidence trail

Retain the question, retrieved authorities, draft, confidence, evaluator result and reviewer decision — on both paths.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Tax research platformsCheckpoint · CCH AnswerConnect
Bloomberg Tax · IBFD
Primary sourcesLegislation · regulations · rulings
Published guidance · authority feeds
Firm precedentPrior memos · house positions
Document management · engagement files

Agent

Tax-question copilot

Reads authority
Drafts and cites
Routes to review

Practice workflowEmail · chat · ticketing
Practice management · approval workflow
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and a signed answer

Each control wraps the one inside it. A draft clears every layer before anyone is asked to rely on it.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeRestrict automation if evaluations or production signals degrade.Roll back
L5Version monitoringTrack authority updates, model, prompt and configuration changes.Track
L4TraceabilityRecord question, authorities, draft, evaluation, decision and amendment.Record
L3Reviewer sign-offDefine which question classes may return without a qualified signer.Gate
L2Scope guardrailsRestrict answers to the jurisdictions, taxes and question classes in scope.Restrict
L1Citation checksCited authority must exist, still stand and support the claim.Verify
Model coreDraft answer proposed — position, citations, jurisdiction, tax year and confidence
L1 – L2Decide whether the answer may stand
L3Decides whether a qualified person signs it
L4 – L5Keep the record and the authority current
L6Pulls automation back when signals degrade

How Nestack evaluates it

Evaluate the full workflow — not only the final answer.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the answer the reviewer sees
Depth of coverage ▼
E1Final-output evaluationWas the answer correct for that jurisdiction, year and entity type?
E2Citation evaluationDoes every cited authority exist, apply and still stand?
E3Step-level evaluationDid it retrieve and weigh the right authority for the period?
E4Escalation calibrationAre unsettled and advice-class questions actually held back?
E5Slice evaluationHow does performance change across specific question cohorts?
E6Business outcomeHow many answers were amended or escalated in review?
Floor — the outcome the client pays for

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle.

Agent lifecycleDirection of processing →
01 · Retrieval2 modes
TQ-01

Superseded authority cited

A revoked, obsoleted or modified ruling is treated as current.

TQ-02

Authority coverage gap

The corpus holds nothing for the jurisdiction or period asked.

Stage gathersStatute, regulations, rulings, guidance and firm precedent
02 · Reasoning2 modes
TQ-03

Fabricated citation

A section, ruling or case that does not exist is quoted.

TQ-04

Jurisdiction or period mismatch

The answer is right for another state, country or tax year.

Stage proposesThe answer, the authority behind it and confidence
03 · Tool / write1 mode
TQ-05

Advice-boundary crossing

A draft is released as though a qualified person had signed it.

Stage writesOnly where release policy and approval boundaries allow it
04 · Output1 mode
TQ-06

Unsettled point answered flat

A genuinely contested position is returned without escalation.

Stage returnsThe answer the requester and the engagement file see
05 · Change / Version1 mode
TQ-07

Silent corpus drift

An authority update is not reflected in the retrieved set.

Stage tracksAuthority updates, model and prompt changes
Sev-1 · unsupported answer could be relied on Sev-2 · wrong answer for that taxpayer Sev-3 · retrieval degrades, routes to review

Affected slices

Overall health can hide concentrated risk

Aggregate answer quality can look acceptable while a small number of question cohorts carry most of the citation and currency failures. Nestack reports performance by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Cross-border questions5.8%3.4× Review
Recent-legislation questions3.9%2.3× Review
Multi-state nexus questions2.9%1.7× Watch
Settled federal questions1.0%0.6× Normal
Bar: lift vs. settled-question baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

The loop does not end at Learn

Each cycle leaves the corpus re-tagged, a position re-verified and an escalation rule in force — that is what the next detection is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Citation or currency failures rise in a question cohort.

02Diagnose

Failure isolated to retrieval, authority weight, prompt or scope.

03Improve

Corpus re-tagged or an escalation rule added — approved and version-linked.

04Verify

Affected questions are re-asked against the updated corpus.

05Learn

The question becomes a standing test; the signed position joins firm precedent.

Learn → DetectThe return edge. The next cycle starts against a re-tagged corpus and one more escalation rule.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, authority corpus, agent workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and advice-boundary definition.
02Source-system and research-platform assessment.
03Jurisdiction, tax and question-class scoping.
04Authority corpus and effective-date tagging.
05Retrieval, citation checking and drafting.
06Confidence scoring and escalation routing.
07Reviewer sign-off workflow.
08Research and document-system integration.
09Evaluation suite and regression cases.
10Guardrails and advice-boundary controls.
11Observability and trace instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne jurisdiction ProductionProduction integration AdvancedMultiple jurisdictions
Introduced at Pilot
Drafted answers with citations
Reviewer sign-off
Baseline evaluation
Introduced at Production
Firm precedent library
Escalation workflow
Approved release actions
Observability and evaluation
Introduced at Advanced
Multi-jurisdiction coverage
Multi-stage review
High query volume
Enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on jurisdictions in scope, research sources, question volume, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01The jurisdictions, taxes and question types in scope Jurisdiction, tax and question-class scopingWeek 1
02The advice boundary and who may sign what Advice-boundary definition and approval controlsWeek 1
03Access to your research subscriptions and sources Source-system and research-platform assessmentWeek 2
04Prior memos, house positions and precedent files Authority corpus build, effective-date tagging and precedent indexingWeek 2
05Confidence and escalation thresholds Confidence scoring, escalation routing and guardrailsWeek 3
06Questions your team has found hard to answer Evaluation suite, regression cases and failure-mode testingWeek 4
07Named reviewers or test users Reviewer sign-off workflow, then pilot workflow and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. Week 5 carries both evaluation and launch work.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Workflow discovery, advice boundary and question-class scoping W2Authority corpus, source access and the research baseline W3Agent workflow, citation checking and reviewer controls W4Evaluation suite, advice-boundary guardrails and failure modes W5Research-platform integration, pilot and targeted corrections W6Sign-off on live questions, verification and Agent Care handover
Reading the bandBars cover only the weeks their work is named in. Build starts once the advice boundary is agreed in week 1.
At the end of W6Live questions are answered under sign-off, then Agent Care monitors the corpus and the escalations.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Accounting AI agent

Build a tax-question copilot around your research workflow.

Show us the questions your team asks, the sources you already licence and who signs what. We'll map the workflow, set the advice boundary and recommend the safest path to production.

Nestack Agents · Tax-question copilot (jurisdiction-aware)AGT-ACC-02 · Agent Care available after launch