Nestack Agent Care
Industries / Education / Proctoring agent

Education AI agent · Proctoring

Proctoring & Academic-Integrity AI Agent

Assemble the evidence behind an exam sitting, check it against the accommodations on file, and hold each flag for the integrity officer, who decides whether anything is a charge.

4–6 weeksTypical delivery
Your stackDeployment
Pre-chargeOfficer review
Agent CareAfter launch

What this agent does

Assembles the evidence, not the charge

In
01

Ingest sitting artefacts, roster context and the accommodations record from supported proctoring.

02

Normalise session events and timestamps, and carry each observation forward with the artefact it came from.

Reason
03

Suppress the signal classes an approved accommodation already covers, before anything is scored.

04

Apply the integrity policy and the stated sitting conditions configured for the institution.

05

Bind each observation to a timestamped window in the artefact, and mark what the artefact leaves ambiguous.

Decide
06

Describe the behaviour a reviewer should watch in plain words rather than a bare score.

07

Route each flagged sitting to the named academic integrity officer for review.

Out
08

Retain the artefact, the observation, the reviewer's finding and the outcome against the sitting.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent proposes observations; the integrity officer decides whether a charge issues, and the institution stays the decision-maker.

Example workflow

One sitting, artefact to review

AgentHuman
1Sitting artefacts receivedProctoring session, assessment log, submission or roster record
2Artefacts assembledSession events, sitting conditions and the accommodations on file, each with its named source
3Observations draftedTimestamped windows, behaviour in words, suppressed classes and confidence
4Accommodation checks appliedAccommodation suppression, integrity-policy checks, prohibited-signal checks and confidence threshold
No human action required

The first four stages run unaided, and no charge exists at any of them — the agent is describing a sitting, and the officer's lane opens at the confidence gate.

5DecisionBranches at the flag threshold
Weak signal

Goes to the integrity officer to review.

Strong signal

Adds a disability-services check first.

Officer review

The sitting is held with its timestamped windows and the behaviour described in words.

Dismiss · Annotate · Send to disability services
Reviewed — recorded against the sitting
6Review record updatedOnly where write access and review policy allow it
7Sitting evaluatedDismissal rate, cohort flag rates, accommodation misses and post-review corrections
Dismissals

Each reviewer dismissal is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Charging a student with academic misconduct.
Assigning a grade, a penalty or a transcript notation.
Requiring, viewing or storing a room scan.
Granting, denying or changing a testing accommodation.
Automation boundaryAgent acts unaided
Assemble the sitting's artefacts and index them by timestamp.
Suppress the signal classes an approved accommodation.
Describe observed behaviour in words.
Hold the flagged sitting for the integrity officer, record intact.
Any write happens inside the boundaries agreed at implementation, never ahead of review.
Ending a sitting, or refusing entry to one.
Identifying a person from a face, voice or biometric template.
Deciding an appeal, or the outcome of a hearing.
Changing a detection threshold inside a live exam window.

Example output

One flagged sitting, annotated

Everything the agent reports is attached to the artefact it was drawn from.

Integrity output · single sittingIllustrative example
Sitting
Observed behaviour
Evidence window
Source artefact
Confidence
Accommodation check
Proctored final
Looked away from the screen repeatedly through the second half
00:18:42
Recorded session
88%
No suppression on file
As receivedTaken from the recorded session and the accommodations record.
Evidence used Recorded session window Stated sitting conditions Accommodations record
Why this flagIt describes what the artefact shows, not what the student intended.
ActionDismissAnnotateSend to disability services
What the score decidesBelow the configured threshold the sitting picks up a disability-services check before.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every sittingFrom the assessment platform
03Observation

Describe from the artefact

Work from the recorded session, the stated sitting conditions and the accommodations on file.

01Approved path

Flag the sitting, not the student

Routine sittings clear without anyone watching a recording end to end.

02Human review

Send review where risk concentrates

Flagged windows and low-confidence sittings are marked, so the officer's read starts at the evidence.

04Build an evidence trail

The flag keeps its evidence window, its confidence and its reviewer.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Proctoring platformsProctorio · Respondus
Honorlock · Examity
Assessment platformsCanvas · Blackboard
Institutional exam systems
Similarity and AI detectionTurnitin · Copyleaks
Institution-configured checks

Agent

Proctoring and academic integrity

Reads the artefacts
Describes the sitting
Holds for review

Roster and identityClever · PowerSchool
SIS and roster feeds
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and a charge

The stack nests. What survives a layer is named in the map beneath it.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modePull the agent back to indexing only when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, threshold and integrity-policy configuration changes.Track
L4TraceabilityRecord the artefact, the observation, the suppressions, the review and its time.Record
L3Officer reviewHold flags for the named integrity officer; it governs who reads a flag, not whether the observation was right.Gate
L2Policy guardrailsTest flags against the accommodations record and the configured integrity policy; a failure returns the flag.Restrict
L1Confidence thresholdsRoute low-confidence sittings to a disability-services check before the officer sees them.Require review
Model coreFlag produced — evidence window, behaviour in words, suppressed classes and confidence
L1 – L2Test whether a flag may stand
L3Puts the review in an officer's hands
L4 – L5Keep the evidence window with the flag
L6Stops flagging and records raw sittings instead

How Nestack evaluates it

Evaluate the whole review path — not only the flag that surfaces.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the flag the officer sees
Depth of coverage ▼
E1Final-output evaluationDid the described behaviour match what the artefact actually shows?
E2Step-level evaluationDid the agent use the right sitting, integrity policy and accommodations record?
E3Tool evaluationDid it read the correct session and write to the correct review record?
E4Confidence calibrationDo low-confidence flags actually attract more dismissals?
E5Slice evaluationHow far apart do flag rates sit across specific student cohorts?
E6Business outcomeHow many flags were dismissed, and how many needed correcting afterwards?
Floor — the charge a student has to answer

Failure modes

Where each failure originates in the agent

Seven modes, from artefact to charge.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
PI-03

Stale accommodations record

A plan approved mid-term is missed, and allowed movement is scored.

Stage gathersSession artefacts, with the source each came from
02 · Reasoning2 modes
PI-04

Style read as authorship

A narrow written register is treated as evidence a machine wrote it.

PI-06

Detection gap read as evasion

Repeated loss of the face in frame is scored as deliberate.

Stage proposesThe evidence window, the behaviour and confidence
03 · Tool / write2 modes
PI-02

Record written before review

A suspicion reaches a gradebook or conduct file first.

PI-05

Artefact leaves the school chain

A session clip is copied outside the institution's control.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
PI-01

Score presented as finding

A confidence figure heads the review card and carries the read.

Stage returnsThe flag the officer reads and acts on
05 · Change / Version1 mode
PI-07

Silent threshold regression

A model or threshold change moves flag volume inside a live term.

Stage tracksModel, prompt, threshold and integrity-policy config
Sev-1 · a charge or grade written by Sev-2 · a wrong flag reaches the officer Sev-3 · artefact degrades

Affected slices

Overall flag rates can hide one cohort

An aggregate flag rate can look defensible while a few cohorts carry most of the flags Nestack reports it by slice rather than in aggregate The cohorts that carry it are named, not averaged away..

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Students using assistive technology9.2%2.8× Review
Students the face model reads worst6.9%2.1× Review
Non-native English writers5.9%1.8× Watch
Students in none of these cohorts2.3%0.7× Normal
Bar: flag-rate lift vs. the all-sitting baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

Every cycle leaves a case behind

An explanation closes nothing here. The cycle ends with a case the next release has to clear.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Flag rate rises in one student cohort.

02Diagnose

The recording, the suppression log and the policy version are read side by side until one of them explains it.

03Improve

The change is stamped to a version, with the flagged sittings attached.

04Verify

A failing case stops the release, not a reviewer's judgement.

05Learn

The case sticks, and the flag-review standard moves with it.

Learn → DetectThe return edge. Every later detection is measured against the larger suite.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, artefacts, flagging workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Integrity workflow discovery and boundary definition.
02Assessment and proctoring source review.
03Integrity policy and accommodation-suppression mapping.
04Sitting-artefact ingestion and indexing.
05Observation logic and evidence windows.
06Confidence scoring and flag routing.
07Integrity officer review workflow.
08Assessment-platform integration.
09Flag regression cases.
10Guardrails and charge controls.
11Evidence-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne exam window, one school ProductionProduction assessment systems AdvancedMultiple campuses / systems
Introduced at Pilot
Flagging to your policy and record
Officer review
Flag-quality baseline
Introduced at Production
Reporting by sitting type
Review workflow in your systems
Approved write-back
Assessment-platform integration
Introduced at Advanced
Multi-campus and multi-state rules
Multi-stage conduct review
High sitting volume
Enterprise integrity controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, sitting volume, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your integrity policy and its sanction ladder Integrity policy and suppression mappingWeek 1
02Representative proctored sittings Flagging baseline, signal extraction and evidence-window bindingWeek 2
03Your accommodations record and its update path Integrity policy and accommodation-suppression mappingWeek 1
04Access to relevant APIs, feeds or exports Proctoring, assessment and roster assessment, then integration setupWeek 2
05Flags that were not misconduct Paired-sitting cases and the evaluation suiteWeek 4
06What a flag must reach before it becomes a charge Confidence scoring, flag routing, guardrails and review controlsWeek 3
07Named integrity officers to review flags Officer review workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Bands track the weeks work actually runs in, so the fifth carries two phases rather than one padded one.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Integrity workflow discovery, policy mapping and the automation boundary W2Artefact integration and the flagging baseline W3Flagging workflow, confidence logic and review controls W4Evaluation suite, cohort flag-rate checks and failure-mode testing W5Platform integration, a pilot exam window and targeted corrections W6One exam window reviewed by the integrity officer, then handover
Reading the bandEach band covers only the weeks its work is named in. The doubled fifth week is real, not padding.
At the end of W6Validation complete; Agent Care holds the monitoring from week seven.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Education AI agent

Build a proctoring agent around your institution's conduct process.

Show us your assessment platform, your integrity policy and who reads a flag. Nothing reaches a conduct file until we have mapped the review path, set the automation boundary and named what stays with the officer.

Nestack Agents · Proctoring & academic integrityAGT-ED-06 · Agent Care available after launch