Nestack Agent Care
Industries / Human Resources / Review drafting copilot

HR AI agent · Performance reviews

Performance Review Drafting AI Agent

Draft the review with each assertion tied to the record it came from, keep what the manager changed, and hold the document for the named manager who signs it.

4–6 weeksTypical delivery
Your stackDeployment
Sourced linesManager signs
Agent CareAfter launch

What this agent does

Drafts the review, never signs it

In
01

A review is drafted, and the record behind each assertion is named beside it.

02

Ames, No. 23-1039, 5 June 2025: pretext asks whether a stated reason is the true one.

Reason
03

Dominguez-Cruz, 1st Cir., 2 February 2000: shifting explanations let a jury infer pretext.

04

EO 14281, 90 FR 17537: agencies deprioritise disparate impact. Treatment claims are untouched.

05

29 CFR 1602.14: once a charge lands, relevant records are held until final disposition.

Decide
06

Reg (EU) 2026/1744, in force 27 July 2026: Annex III duties wait until 2 December 2027.

07

Article 5(1)(f) was not deferred: emotion inference at work is prohibited, live 2 February 2025.

Out
08

A record is too thin to evaluate, and that thinness is recorded as the finding, not written over.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Drafting, sourcing and holding belong to the agent. The signature belongs to a named manager, who adopts each sentence and answers for it later.

Example workflow

One review, notes to signature

AgentHuman
1Source records receivedGoal records, one-to-one notes, peer input, delivery telemetry and the prior review
2Records ranked and boundWhich note supports which claim, who wrote that note, and the day it was written
3Draft review builtThe narrative, the record behind each assertion, the lines no record reaches and the diff against last cycle
4Controls appliedSource checks, prior-review diffing, protected-material checks and release confidence
No human action required

Stages 1 to 4 run unaided, and nothing reaches an employee at any of them — the agent is drafting, and the manager lane opens at the confidence gate.

5DecisionForks at the release threshold
High confidence

Goes to the named manager to sign.

Low confidence

Adds a business-partner read first.

Manager review

The draft is held with its source records, its unsourced lines and the confidence.

Sign · Attach record · Send to HR
Signed — by the named manager
6Review and case records updatedOnly where write access and records policy allow it
7Outcome evaluatedSource accuracy, edit depth, manager corrections and what review found
Corrections

Each manager amendment is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Signing a review of another adult.
Setting the rating a narrative supports.
Deciding a performance concern exists.
Countersigning a plan or a discharge.
Automation boundaryAgent acts unaided
Cite the source record behind each drafted line.
Surface any sentence whose evidence the record cannot substantiate.
Keep protected-activity material out of the drafting context window.
Diff each draft against the review it follows.
Nothing goes into a personnel file except by a named manager, inside the agreed boundaries.
Judging whether the record supports a rating.
Telling anyone a review is defensible.
Choosing what a calibration outcome means.
Changes to the rating, the plan or the file.

Example output

One drafted review, annotated

This serves HR and line managers who may have to explain, years later, who wrote a given sentence in a personnel file; below is one drafted review exactly as the agent leaves it.

Drafted review · single cycleIllustrative example
Report
Draft
Status
Record cited
Confidence
Held for
Mid-year review, one direct report
Two goals met, one carried over from the prior cycle
Draft, unsigned
Goal record, 4 August 2026
Two lines unsourced
The named manager, by name
As receivedDrawn from the goal records and the notes on file, and it asserts nothing that those records do not.
What the record holds Goal record One-to-one note Peer input
Why no signature hereOwning a sentence as your own judgement is an act the manager makes.
ActionSignAttach recordSend to HR
What the score decidesBelow the configured threshold a draft gets an HR read before the manager sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Each drafted lineFrom the record that carries it
03Drafting

Where the review is used

The agent vouches for what a record says and the day it was written, not for what a manager believed. C-203/22, 27 February 2025: which data were used, and how, has to be answerable.

01Approved path

Read back in a deposition

SCHUFA, C-634/21: a preparatory score becomes the decision where the human draws strongly on it — and how strongly is a question about manager edits, which a product can measure.

02Human review

What was checked, and not found

No rule checked obliges an employer to keep the notes, the prompt or the drafts a manager rejected. But 29 CFR 1602.14 preserves relevant personnel records until final disposition once a charge is filed, and a pipeline cannot hold what it deleted at the close of the cycle.

04Build an evidence trail

The review, the evidence under it and the manager whose name it carries stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Performance managementWorkday · Lattice · Culture Amp
Goals, ratings and past reviews
Manager notesOne-to-one notes · check-ins
What was observed, and when
Delivery telemetryJira · Salesforce · GitHub · Zendesk
Tickets closed, never judgement

Agent

Performance review drafting

Reads the records
Drafts the review
Holds for the manager

Peer and calibration inputPeer feedback · calibration records
What colleagues wrote, and when
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six passes between the record and the file

Six passes over one page, the last the closest. What survives is drawn in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeNarrow the agent to record listing when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and drafting rules, and note the version each review was drafted under.Track
L4TraceabilityRecord each draft, the sources under it, the manager edits and every read of the file.Record
L3Manager signatureHold the draft for a named manager; the hold governs signature, not whether the review is fair.Gate
L2Source guardrailsTest each assertion against the record cited, and refuse a line that no source reaches.Restrict
L1Confidence thresholdsRoute a low-confidence draft to an HR read before the review reaches an employee.Require review
Model coreReview drafted — the narrative, the records under it, the unsourced lines and the diff
L1 – L2Test whether a draft may be signed
L3Leaves the signature to a named manager
L4 – L5Keep the review and the evidence behind it
L6Hands the draft back unsigned when signals degrade

How Nestack evaluates it

Evaluate the whole drafting run — not only the review that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the review an employee reads
Depth of coverage ▼
E1Final-output evaluationDid the review record the source each assertion was drafted from?
E2Step-level evaluationDid the agent read the right period, the right notes and the right report?
E3Tool evaluationDid it read and write the correct employee and the correct cycle?
E4Confidence calibrationDo low-confidence drafts actually attract deeper manager edits?
E5Slice evaluationHow does performance change across specific report classes?
E6Business outcomeHow many drafts needed a correction before the manager signed?
Floor — the records a review rests on

Failure modes

Where each failure originates in the agent

Seven failure modes, each set at the stage it first becomes visible.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
UL-03

Protected material retrieved

A leave, a complaint or an accommodation enters the context.

Stage gathersThe notes, the goals, the peer input and the metrics
02 · Reasoning2 modes
UL-04

Invented specific detail

A project, a date or an incident that no record holds.

UL-06

Prior review regenerated

Last year's narrative returns as this year's judgement.

Stage proposesThe narrative, its sources, the diff and confidence
03 · Tool / write2 modes
UL-02

Narrative against the rating

The prose and the rating tell different stories.

UL-05

Affect signal reaches the draft

Tone or sentiment inference enters an evaluative line.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
UL-01

Signed with no drafting record

The file keeps the text and the name, not who wrote it.

Stage returnsThe review a worker signs and a lawyer reads back
05 · Change / Version1 mode
UL-07

Silent wording drift by class

The register moves by group while each review reads fine.

Stage tracksModel, prompt, wording rules and review dates
Sev-1 · a sentence the manager never meant Sev-2 · protected activity reaches the file Sev-3 · source degrades, draft held back

Affected slices

Thinly recorded reports absorb the corrections

A cycle-level accuracy figure can read clean while thinly recorded reports carry the bulk of the manager corrections. Nestack reports the correction rate by report class, and not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Thinly recorded reports9.9%3.7× Review
Newly transferred reports7.0%2.6× Review
Rating-boundary cases4.4%1.6× Watch
Long-tenured steady reports2.0%0.7× Normal
Bar: correction-rate lift vs. the long-tenured baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What an unmeant sentence costs

A loop closes when the sentence the manager never meant is a standing case. That suite is what the next review cycle is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on thinly recorded reports.

02Diagnose

The line that entered the file because the deadline was nearer than the evidence is worked backwards until one cause is left standing.

03Improve

Reviews go out numbered, and the evidence beneath them rides with it.

04Verify

One wording case still failing is enough to hold the cycle back.

05Learn

It is kept for good, and the drafting rules change in that same commit.

Learn → DetectThe return edge. The next review is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, review drafting, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Review narratives and automation-boundary definition.
02Goal, note and telemetry sources.
03Sentence-to-source and provenance-coverage tracking.
04Review-record ingestion.
05Record, goal and employee binding.
06Source checks and review routing.
07Manager signature workflow.
08Review-record integration.
09Wording and evidence cases.
10Guardrails and drafting controls.
11Review-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne team, one review cycle ProductionProduction review workflow AdvancedMultiple teams / cycles
Introduced at Pilot
Review drafting to your competencies
Named manager signature
Prior-review baseline
Introduced at Production
Reporting by rating class
Manager review workflow in your systems
Approved write-back
Goal-and-record integration
Introduced at Advanced
Conflicting source records
Cross-team calibration packs
Large review populations
Multi-manager wording controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, review volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live review cycle and the competencies each report is rated on Record capture and source versioningWeek 1
02Representative goal, note and telemetry sources Source binding, drafting logic and the prior-review baselineWeek 2
03Your review calendar and the managers it names Competency mapping, source binding and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Goal, note and telemetry source assessment, then integration setupWeek 2
05Sentences you would not want exhibited Wording cases and the failure roundWeek 4
06What no review may assert Source checks, review routing, guardrails and release controlsWeek 3
07A named manager who signs the review Manager sign-off workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Two of these phases do share the fifth week, and no band was widened to make the column look neat.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Cycle discovery, competency mapping and the automation boundary W2Source integration and the prior-review baseline W3Review drafting, source logic and release controls W4Evaluation suite, wording cases and failure-mode testing W5Review-record integration, pilot drafts and targeted corrections W6One review cycle run under the named manager, then Agent Care handover
Reading the bandEach bar spans only the weeks its own work is named for, and week five carries two.
At the end of W6Once the review record validates, Agent Care picks the agent up.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · HR AI agent

Build a review drafting agent around the sentence your last cycle cannot account for.

Show us one review cycle your managers run and the notes behind the last review written. If nobody can say which sentences the manager composed, the file holds a document with a name on it and no author. A review no one can account for comes back as a finding.

Nestack Agents · Performance reviewsAGT-HR-09 · Agent Care available after launch