Nestack Agent Care
Industries / Transportation / WISMO chatbot agent

Transportation AI agent · Delivery WISMO

Delivery and WISMO Chatbot AI Agent

Answer 'where is my order' from the order record itself, disclose that it is AI before anything else, and leave promising a date and granting a remedy to a person.

4–6 weeksTypical delivery
Your stackDeployment
Discloses firstPerson promises
Agent CareAfter launch

What this agent does

Answers from the record, and promises nothing

In
01

Ingest the order, the shipment record and the merchant's own policy text from supported order, desk or carrier sources.

02

Disclose at the first interaction, proactively and in plain language, because Maine's rule carries no intent element.

Reason
03

Meet the EU disclosure duty where an EU person is on the other end, enforceable since August.

04

Quote policy from the merchant's controlled corpus rather than composing it, and cite the record each answer rests on.

05

State a carrier estimate as an estimate, and a committed date only where the record actually holds one.

Decide
06

Enforce remedy caps, entitlement and eligibility in the tool layer, server-side, never by prompt instruction.

07

Decline and escalate anything it cannot substantiate, in preference to a plausible answer.

Out
08

Retain the conversation verbatim, because the transcript is the representation your firm made.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent answers; a named person promises a date or grants a remedy, and whatever the bot said, your firm said.

Example workflow

One enquiry, question to answer

AgentHuman
1Enquiry receivedChat, email, portal message or a post-purchase tracking page
2Records gatheredThe order, the shipment record, the carrier's latest status and the merchant's policy text, each with its source
3Answer preparedAnswer, sources, remedy state, confidence
4Controls appliedDisclosure checks, source-binding checks, remedy caps and the confidence threshold
No human action required

Stages 1 to 4 run unaided and promise nothing — the answer is drafted, and the person's lane opens at the confidence gate.

5DecisionBranches at the confidence threshold
High confidence

Answers from the record.

Low confidence

Hands to an agent first.

A person takes it

The answer is held with its source records and the confidence.

Send · Amend · Hand to an agent
Answered — or handed to a person
6Order records updatedOnly where write access and approval policy allow it
7Outcome evaluatedDisclosures delivered, sources cited, hand-offs taken and what each enquiry actually needed
Hand-offs

Every hand-off to a person is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Promising a consumer a delivery or shipping date.
Re-promising a date once the first one has slipped.
Granting a refund, a credit or any other remedy.
Stating a policy that is not in the policy corpus.
Automation boundaryAgent acts unaided
Disclose that it is an AI system at the first interaction.
Answer from the order and shipment records, with the source.
Quote the merchant's own policy text rather than writing it for the named owner.
Hand over to a person the moment it cannot substantiate.
Any write happens inside the boundaries agreed at implementation, never ahead of approval.
Deciding whether an order may still be cancelled.
Answering on liability for loss or for damage.
Turning the AI disclosure off for a given market.
Changes to the policy corpus or the remedy caps.

Example output

One answer, annotated

No US appellate court has ruled on a chatbot's words, so everything here is tied to its record.

Answer output · single enquiryIllustrative example
Order
What was answered
Carrier estimate
Record it came from
Confidence
Disclosure given
Two-item parcel
In transit, with the carrier's own estimate quoted as an estimate
Est. Thursday
Carrier tracking record
92%
AI disclosed at first message
As receivedTaken from the order and shipment records — nothing on this side is recalled by the model or written by it.
What it was read from The order record itself The carrier's tracking status The merchant's policy text
Why it is not a promiseApparent authority, negligent misrepresentation and FTC Act §5 are the live exposure.
ActionSendAmendHand to an agent
What the score decidesBelow the configured threshold the answer picks up a person before the consumer ever sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every enquiryFrom chat, email or the portal
03Answering

Answer from the record

Draw on the order, the shipment record and the merchant's own controlled policy text.

01Approved path

Tell them what is known

Routine status questions come back answered and sourced.

02Human review

Send the rest to a person

Date requests, remedy asks and low-confidence answers are marked, so the team's read starts where risk concentrates.

04Build an evidence trail

The answer, the tracking record behind it and the agent who took it over stay on the enquiry.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Order and commerceShopify · Salesforce
SAP · Oracle Commerce
Service deskZendesk · Gorgias
Intercom · Freshdesk
Carrier trackingCarrier APIs · EDI
Visibility platforms

Agent

Delivery and WISMO answers

Reads the record
Discloses and answers
Hands over on ask

Policy and transcriptsPolicy corpus · knowledge base
Transcript archives
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the customer

Each control wraps the one inside it. What a layer does not catch is named in the map below it.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeDrop the agent to status-only answers when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, policy-corpus and disclosure-text changes.Track
L4TraceabilityRecord the transcript, the sources cited and the disclosure delivered.Record
L3Person on requestHand the enquiry to a named person on request; it governs who answers next, not what was already said.Gate
L2Remedy capsTest each remedy against caps and entitlement in the tool layer; a failure stops the answer.Restrict
L1Confidence thresholdsRoute low-confidence answers to a person before the consumer sees them.Require review
Model coreAnswer prepared — the text, its sources, the remedy state and confidence
L1 – L2Test whether an answer may go out
L3Puts the reply in a person's hands
L4 – L5Keep the answer and the tracking record behind it
L6Drops the agent to status-only when signals degrade

How Nestack evaluates it

Evaluate the whole conversation — not only the sentence that answered.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the reply the customer reads
Depth of coverage ▼
E1Final-output evaluationDid the disclosure land at the first interaction, before anything else?
E2Step-level evaluationDid the agent read the right order, shipment and policy text?
E3Tool evaluationDid it read the correct order and write the correct case?
E4Confidence calibrationDo low-confidence answers actually attract more hand-offs?
E5Slice evaluationHow does performance change across specific enquiry types?
E6Business outcomeHow many answers were handed over, and how many should have been?
Floor — the promise your firm has to honour

Failure modes

Where each failure originates in the agent

Seven ways an answer goes wrong, placed by stage.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
BA-03

Stale shipment record

The answer rests on a status the carrier has since moved.

Stage gathersOrder, shipment record, carrier status and policy
02 · Reasoning2 modes
BA-04

Estimate stated as a date

A carrier prediction is repeated as a firm commitment.

BA-06

Policy written, not quoted

A rule is composed rather than taken from the corpus.

Stage proposesThe answer, its sources and confidence
03 · Tool / write2 modes
BA-02

Remedy granted in chat

A refund is offered before entitlement was checked.

BA-05

Duplicate case raised

One enquiry opens the same case twice over.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
BA-01

Disclosure arrives late

The AI notice lands after the first substantive reply.

Stage returnsThe reply the customer reads and relies on
05 · Change / Version1 mode
BA-07

Silent policy regression

A model or corpus change widens what the agent will state.

Stage tracksModel, prompt, policy corpus and disclosure text
Sev-1 · remedy with no person Sev-2 · a date the record cannot support Sev-3 · record degrades, answer routes on

Affected slices

Take the total apart by enquiry type first

Do not read the aggregate until it is split: two enquiry types carry the hand-offs, and averaging them into the routine ones makes both invisible. Nestack reports the hand-off rate per enquiry type.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Delayed or missed delivery dates3.5%3.3× Review
Refund and remedy requests2.7%2.6× Review
Lost or damaged parcels1.9%1.8× Watch
Routine in-transit status0.6%0.6× Normal
Bar: hand-off-rate lift vs. routine-status baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

Unless it becomes a test, it is only a note

A cycle closes when the failure is a regression case the next release has to pass. That suite is what the next enquiry answered is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Hand-off rate rises in an enquiry type.

02Diagnose

The transcripts and the records behind them are read until one cause holds.

03Improve

The change ships against a version, with the enquiries that exposed it attached.

04Verify

Release is blocked until the affected regression cases pass again.

05Learn

The case joins the permanent suite and the service playbook.

Learn → DetectThe return edge. The next detection runs against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, answering workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Enquiry workflow discovery and boundary definition.
02Order and carrier source assessment.
03Disclosure, policy-corpus and remedy-cap rule mapping.
04Order and shipment-record ingestion.
05Answer logic and source binding.
06Confidence scoring and hand-off routing.
07Person-on-request workflow.
08Order-system and service-desk integration.
09Promise and remedy cases.
10Guardrails and escalation controls.
11Enquiry-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne market, one channel ProductionProduction order and desk AdvancedMultiple markets / brands
Introduced at Pilot
Answering from your own records
Person on request
Answer-accuracy baseline
Introduced at Production
Reporting by service level
Hand-off workflow in your systems
Approved write-back
Order-system integration
Introduced at Advanced
Multi-market policy corpora
Multi-stage remedy approvals
High enquiry volume
Multi-market service controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your order records and policy text Order and shipment-record ingestion and source bindingWeek 1
02Representative past enquiries Answer baseline, source binding and the policy corpusWeek 2
03Your disclosure text and remedy caps Disclosure, policy-corpus and remedy-cap rule mappingWeek 1
04Access to relevant APIs, feeds or exports Order and carrier source assessment, then integration setupWeek 2
05Answers you would not want relied on Promise cases and the evaluation suiteWeek 4
06What an answer may never promise Confidence scoring, hand-off routing, guardrails and remedy controlsWeek 3
07Named people to take handed-over enquiries Person-on-request workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Each phase sits on the weeks it actually occupies, and week 5 carries both evaluation and launch work.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Enquiry workflow discovery, policy mapping and the boundary W2Order and carrier integration and the answer baseline W3Answering workflow, confidence logic and hand-off controls W4Evaluation suite, promise cases and failure-mode testing W5Service-desk integration, a pilot market and corrections W6One peak week answered under supervision, then Agent Care handover
Reading the bandEach bar covers only the weeks its work is named in. The week 5 overlap is real, not padding.
At the end of W6Validation closes on live enquiries, and Agent Care picks up monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Transportation AI agent

Build a WISMO agent that discloses first and promises nothing.

Show us your order records, your policy text and the dates you state to customers. Tell us who signs off the disclosure wording and who owns the remedy caps — those two names set the outer edge of what the bot may ever say.

Nestack Agents · Delivery and WISMO answersAGT-TR-12 · Agent Care available after launch