Nestack Agent Care
Industries / Biotechnology / Literature copilot

Biotechnology AI agent · Literature

Research-Literature Copilot (Evidence-Traced)

Run a recorded search, screen records against the stated criteria and extract outcomes into an evidence table — each line bound to the source and page behind it, for a named reviewer to sign.

4–6 weeksTypical delivery
Your stackDeployment
Named reviewerSign-off
Agent CareAfter launch

What this agent does

Assembles the evidence, does not conclude

In
01

Take the review question, the inclusion criteria and the sources in scope from the protocol.

02

Run the agreed search across the sources in scope for this build, and record it as run.

Reason
03

Screen records against the criteria as written, keeping the reason each decision turns on.

04

Extract design, population, comparator, outcome and result into the evidence table.

05

Mark preprints and translated text, and pull retraction and correction status with the record.

Decide
06

Hold any line whose source check fails — no page, no matching text, or a status that moved.

07

Route safety-relevant reports and out-of-scope questions to the team that owns them.

Out
08

Present the evidence table, the search record and the screening log for a reviewer to work from.

09

Retain the strategy, the records seen, every decision reason and each reviewer correction.

Product statement

The agent searches, screens, extracts and cites inside the approval boundaries agreed during implementation; a named reviewer draws the conclusion.

Example workflow

One evidence pack, end to end

AgentHuman
1Question and criteria receivedReview question, inclusion criteria, sources in scope and the period the search covers
2Search run and recordedThe databases, registries and preprint servers in scope, with the strategy, filters and counts kept as run
3Records screenedEach record judged against the criteria as written, with the reason for the decision kept
4Data extracted and checkedCharacteristics and outcomes into the table, each line checked back to its source and page
No human action required

Stages 1 to 4 run without a person in the loop — the reviewer is not asked to read anything until the source checks have run. A safety-relevant report leaves for the safety team the moment it is seen.

5DecisionSplits on the source check and the confidence gate
Sources check out

Goes to the reviewer as a drafted pack.

Check fails or evidence is thin

Held with the failing lines marked.

Scientist or information specialist

Reads the pack against the search record, the flagged lines and what could not be sourced.

Sign off · Amend · Return to search
Signed off — handed back
6Pack assembled and filedWritten to the review file only where access and policy allow; the conclusion field stays empty
7Outcome evaluatedSource-check failures, screening disagreement, reviewer corrections and what went back to search
Corrections

Lines the reviewer rewrites or removes are counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Concluding that the evidence supports a use.
Any comparative, off-label or promotional statement.
Signing off a review, summary or evidence pack.
Grading certainty of evidence or risk of bias.
Automation boundaryAgent acts unaided
Run the agreed strategy and record it as executed.
Screen records against the criteria as written.
Extract characteristics and outcomes with the source and page.
Hold failing lines, mark preprints and route safety reports out.
Write actions run only inside the approval boundaries agreed during implementation. The conclusion is not one of them.
Deciding a retraction or concern can be ignored.
Screening the literature for safety reporting.
Releasing a pack outside the organisation.
Changing the criteria, strategy or source list.

Example output

One extracted line, annotated

Everything the agent puts in the table is attached to the record and the page it came from.

Evidence table · single extracted studyIllustrative example
Record
Retrieved from
Located at
Extracted
Confidence
Record status
Randomised trial report
Indexed database, full text
p. 6, table 2
Primary outcome, as reported
89%
Checked at the search date
As receivedThe record as the database returned it, beside the strategy that found it — nothing on this side is inferred.
Evidence used Full text, not the abstract Retraction status pulled Page and table located
Why the status carries a dateRetraction notices reach the databases late, so the status carries the date it was checked.
ActionSign offAmendReturn to search
What the score decidesBelow the configured threshold the line is held for the reviewer instead of entering the table.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every record retrievedFrom the sources in scope
03Search & screening

Apply the protocol as written

Use the stated criteria, the agreed source list and the review's own extraction form.

01Approved path

Make the first pass checkable

Screening decisions and extracted lines arrive with their reason and their source, so a reviewer starts from something they can check rather than from the raw record set.

02Human review

Put reviewers on the flagged lines

Failed source checks, moved statuses and split judgements reach a named reviewer instead of the whole record set doing so.

04Build an evidence trail

Retain the strategy as run, the records seen, each screening reason, the extracted line with its source and page, the confidence and the reviewer's correction — on both paths.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Bibliographic databasesPubMed / MEDLINE · Embase · Web of Science
Scopus · Cochrane Library
Registries & preprintsTrial registries · preprint servers
WHO ICTRP · EU CTIS
Full text & statusPublisher full text · library holdings
Crossref · Retraction Watch data

Agent

Research-literature copilot

Runs the search
Screens and extracts
Routes to review

Review & document systemsCovidence · DistillerSR · EndNote / Zotero
Veeva Vault · SharePoint
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the evidence pack

Each control wraps the one inside it. A line clears every layer before it reaches the table a reviewer reads.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeRestrict automation if evaluations or production signals degrade.Roll back
L5Trail and versioningStrategy, screening reasons, corrections and configuration changes are recorded.Record
L4Reviewer sign-offDefine what may reach a requester without a named reviewer signing.Gate
L3Claim guardrailsDraft wording is checked for conclusions, comparisons and off-label claims.Restrict
L2Status checksStatus is pulled where a source carries it, or the line reads unchecked.Flag
L1Source bindingEach line is checked against the record and page it names.Verify
Model coreExtracted line proposed — study, outcome, the source and page behind it, and confidence
L1 – L2Decide whether a line may stand
L3Holds the wording inside the evidence
L4 – L5Put the conclusion with a person, keep the trail
L6Pulls automation back when signals degrade

How Nestack evaluates it

Evaluate the whole search — not only the table at the end.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the pack the reviewer opens
Depth of coverage ▼
E1Final-output evaluationDid the extracted line match the study, and say no more than the source?
E2Source-check evaluationDoes each cited record exist, contain the sentence and still stand?
E3Screening evaluationWere include and exclude decisions made against the stated criteria?
E4Search reproducibilityDoes the recorded strategy return the same set when it is rerun?
E5Slice evaluationHow does performance change across specific record cohorts?
E6Business outcomeHow much of the pack was corrected or sent back to search?
Floor — the review a named person signs

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle.

Agent lifecycleDirection of processing →
01 · Search2 modes
LR-01

Retraction arrives after the copy

A record pulled before the notice still reads as current.

LR-02

Strategy will not rerun

A filter or interface difference is missing from the record.

Stage gathersDatabases, registries, preprints and the strategy as run
02 · Screening1 mode
LR-03

Preprint status lost in the pack

An unrefereed version is read as the published one.

Stage screensRecords against the criteria the protocol states
03 · Extraction2 modes
LR-04

Source does not carry the claim

A sentence is bound to a paper that does not say it.

LR-05

Wrong arm or timepoint extracted

A result is lifted from the neighbouring column.

Stage extractsCharacteristics and outcomes, with source and page
04 · Pack / output1 mode
LR-06

Safety report left in the pack

A case report is summarised instead of routed, and the clock runs.

Stage returnsThe evidence table and log a reviewer works from
05 · Change / Version1 mode
LR-07

Corpus narrows between runs

A licence or index change drops sources unannounced.

Stage tracksCorpus updates, model, prompt and criteria changes
Sev-1 · an unsupported claim could stand Sev-2 · the table misstates the study Sev-3 · retrieval thins, more goes back

Affected slices

Source-check failures cluster by record type

Four cohorts, one pack. The failure rate counts lines whose source check failed or whose extracted value a reviewer had to change, because a scanned or translated report behaves nothing like a well-indexed trial.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Scanned pre-2000 full text4.7%3.0× Review
Non-English source articles3.4%2.2× Review
Preprints and conference abstracts2.7%1.7× Watch
Well-indexed randomised trials1.2%0.8× Normal
Bar: source-check failure lift vs. indexed-trial baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

A failed check does not close with the pack it came from

A line the reviewer had to correct changes the strategy, the extraction form or the status check that let it through — and the next pack is measured against that.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Source-check or extraction failures rise in a record cohort.

02Diagnose

The search record is what makes this reproducible; that is what it is for.

03Improve

The strategy, form or status rule is changed under version control.

04Verify

Affected records are searched, screened and extracted again.

05Learn

The failed line stays as a test case and the form carries the correction.

Learn → DetectThe return edge. A change to the criteria or the strategy is a protocol change — agreed and recorded before the next search runs, not after.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, search and screening, extraction and status checks, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and claim-boundary definition.
02Source, database and licence assessment.
03Review question, criteria and protocol scoping.
04Search strategy build and strategy recording.
05Screening logic and decision-reason capture.
06Extraction form and source-and-page binding.
07Retraction, correction and preprint status checks.
08Reviewer sign-off workflow.
09Claim-scope guardrails and safety routing.
10Evaluation suite and regression records.
11Review-platform and document integration.
12Observability, deployment and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne review question ProductionProduction integration AdvancedSeveral programmes / sources
Introduced at Pilot
Recorded search strategy
Screening against stated criteria
Extraction with source and page
Retraction and preprint status checks
Claim-scope guardrails and safety routing
Reviewer sign-off
Baseline evaluation
Introduced at Production
Review-platform and document integration
Observability and evaluation
Introduced at Advanced
Several databases and licensed sources
Multi-programme and enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on the databases and licences in scope, review types, question volume, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01The review question, criteria and protocol Review question, criteria and protocol scopingWeek 1
02Where your claim boundary sits and who signs Claim-boundary definition and sign-off rulesWeek 1
03Access to the databases and licences you hold Source, database and licence assessmentWeek 2
04A search your information specialist has already run Search strategy build and strategy recordingWeek 2
05Your extraction form and how outcomes are recorded Extraction form and source-and-page bindingWeek 3
06Packs that went wrong — a bad citation, a retracted paper Evaluation suite, regression records and failure-mode testingWeek 4
07Named scientists, reviewers or information specialists Reviewer sign-off workflow, then pilot packs and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. Week 5 carries both the cohort slices and the first packs a reviewer reads.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Review question, criteria and the claim boundary agreed W2Databases, licences and the search strategy recorded as run W3Screening reasons, extraction form and source-and-page binding W4Source-check testing, status checks and cohort slices W5Review-platform integration, first reviewed packs and corrections W6Production validation, sign-off on live questions and Agent Care handover
Reading the bandBars cover only the weeks their work is named in. Week 2 builds the search record; a strategy nobody can rerun later is where a review quietly comes apart.
At the end of W6Your reviewers sign, and nothing the agent produced leaves unsigned. Agent Care then watches the corpus and the status checks.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Biotechnology AI agent

Build a literature copilot around the review you already run.

Show us a review question, the sources you licence and the pack a reviewer signs today. If your reviewers cannot reproduce a search from its own record, extraction is the wrong place to start, and we will say so in the first week.

Nestack Agents · Research-literature copilotAGT-BT-01 · Agent Care available after launch