Nestack Agent Care
Industries / Advertising & Marketing / Marketing-mix agent

Advertising AI agent · MMM & incrementality

Marketing-Mix Modelling & Incrementality AI Agent

Assemble the spend and outcome series, fit and refit the model, design the geo or holdout experiment that tests it, and turn the result into a budget recommendation the analyst accepts, interval attached.

4–6 weeksTypical delivery
Your stackDeployment
Pre-decisionAnalyst approval
Agent CareAfter launch

What this agent does

Models the spend, not the decision to move it

In
01

Spend and outcome series arrive from ad platforms, the CRM and finance systems.

02

Series are aligned to one calendar, one currency and one channel taxonomy.

Reason
03

The model is fit and refit against the aligned series, one channel at a time.

04

Each channel's prior is logged with its source — an experiment, an assumption or a default.

05

A geo or holdout test is designed with its minimum detectable effect and power agreed first.

Decide
06

The model's estimate is checked against the experiment, and any gap is surfaced, not resolved.

07

Every recommendation is routed to the named analyst before a number leaves the model.

Out
08

The specification, the priors, the experiment and every edit are retained against the run.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent fits the model and reads the experiment; the analyst decides which specification to accept, and no modelled figure leaves the page without its interval.

Example workflow

One channel estimate, series to approval

AgentHuman
1Series receivedSpend and outcome series, a new experiment result or a scheduled refit
2Inputs assembledHistoric series, the prior log, calibration history and channel taxonomy, each with its source
3Estimate draftedChannel contributions, credible intervals, experiment status and confidence
4Controls appliedSpecification checks, prior-source checks, calibration-gap checks and confidence threshold
No human action required

Stages 1 to 4 run unaided, and no figure is recommended at any of them — the agent is modelling, and the analyst's lane opens at the confidence gate.

5DecisionBranches at the confidence threshold
High confidence

Goes to the analyst to approve.

Low confidence

Adds a second modeller review first.

Analyst approval

The estimate is held with its interval, its calibration status and the confidence.

Approve · Revise · Escalate
Approved — released to planning
6Reporting systems updatedOnly where write access and approval policy allow it
7Outcome evaluatedInterval accuracy, calibration gap, edits and post-decision corrections
Edits

Every analyst edit is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Reallocating budget between channels.
Publishing a modelled figure to the market.
Setting a channel's ROI prior on its own.
Declaring a test result before it is powered.
Automation boundaryAgent acts unaided
Assemble the spend and outcome series and fit the model.
Design the geo or holdout experiment and its power.
Attribute each channel's interval to its prior and source for the named owner.
Flag the gap when the model and the experiment disagree, and hold it for the analyst.
Any output leaves the model only inside the boundaries agreed at implementation, never as fact.
Presenting a modelled return as revenue already earned.
Reconciling attribution, incrementality and MMM into one figure.
Approving the budget shift the recommendation implies.
Changes to specification, priors or calibration rules.

Example output

One estimate, annotated

Everything the agent produces is attached to the run it was drawn from.

MMM output · single channelIllustrative example
Channel
Finding
Estimate status
Model run
Confidence
Attribution
Brand search, always-on
Large modelled contribution, no calibrating holdout run
Interval untested
This quarter's model run
68%
Analyst name on file
As receivedTaken from this quarter's model run and the series behind it — nothing on this side is asserted as fact by the agent.
Inputs checked Spend and KPI series Prior source logged Holdout test status
Why it's flaggedThe contribution is real in the model and a person still decides.
ActionApproveReviseEscalate
What the score decidesBelow the configured threshold the flag picks up a second modeller review before it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every channelFrom the assembled series
03Modelling

Fit against the series

Draw on the assembled series, the logged priors and the calibration history.

01Approved path

A model is not a measurement

Brand search and retargeting are the classic case: an observational read can show a strong return while a controlled test shows the outcome would have happened anyway — what the experiment exists to catch.

02Human review

Label the question each number

Attribution, incrementality and MMM measure different things; every output states which one produced it, and the three are never blended into a single figure.

04Build an evidence trail

The estimate, the interval around it and the analyst who accepted it stay on the plan.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Media & spend platformsGoogle Ads · Meta Ads
TikTok · Amazon Ads · DV360
Business & finance dataERP · POS systems
Revenue and conversion feeds
Modelling toolingMeridian · Robyn
Custom Bayesian MMM stacks

Agent

Marketing-mix modelling & incrementality

Fits the model
Designs the test
Holds for the analyst

Experimentation & warehouseGeoLift · Meridian GeoX
BigQuery · Snowflake
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the plan

Every control wraps the one within it. What the set does not catch is named in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeRevert to descriptive reporting when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, spec, prior and calibration changes.Track
L4TraceabilityRecord the series, the specification, the priors and the analyst's decision.Record
L3Analyst approvalHold the recommendation for the named analyst; it governs release, not whether the model is right.Gate
L2Policy guardrailsTest every run against configured prior-source and specification rules; a failure returns it.Restrict
L1Confidence thresholdsRoute low-confidence estimates to a second modeller review.Require review
Model coreEstimate produced — channel contributions, credible intervals, calibration status and confidence
L1 – L2Test whether an estimate may stand
L3Puts the release in the analyst's hands
L4 – L5Keep the estimate and the interval behind it
L6Reverts to descriptive reporting when signals degrade

How Nestack evaluates it

Evaluate the modelling workflow — not only the final number.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the number the analyst sees
Depth of coverage ▼
E1Final-output evaluationDid the channel estimate ship with its credible interval?
E2Step-level evaluationDid the agent use the current series, priors and calibration record?
E3Tool evaluationDid it read the correct series version and channel ID?
E4Confidence calibrationDo low-confidence estimates actually diverge more from the experiment?
E5Slice evaluationHow does interval width change across specific channel types?
E6Business outcomeHow many estimates needed an analyst edit or a later correction?
Floor — the outcome the plan answers for

Failure modes

Where each failure originates in the agent

Seven failure modes, placed at the stage each one originates.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
MM-03

Stale series pull

The spend or outcome series read is a cached pull, not this period's data.

Stage gathersSpend and outcome series, priors and calibration history
02 · Reasoning2 modes
MM-04

Prior source blurred

A default or assumed prior is shown as if an experiment set it.

MM-06

Mismatched calibration basis

An experiment's window or population doesn't match the model's, but calibrates it anyway.

Stage proposesChannel estimates, intervals and confidence
03 · Tool / write2 modes
MM-02

Low-confidence release

A recommendation reaches the analyst ahead of its confidence gate.

MM-05

Modelled return shown as real

A modelled contribution is presented as revenue already delivered.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
MM-01

Interval dropped

A channel estimate reaches the analyst with no credible interval attached.

Stage returnsThe estimate the analyst approves
05 · Change / Version1 mode
MM-07

Silent specification drift

A prior, adstock or specification change widens what the agent will report as fitted.

Stage tracksModel, priors, specification and calibration version
Sev-1 · plan acted on outside the boundary Sev-2 · a wrong estimate reaches the analyst Sev-3 · series degrades, run routes to review

Affected slices

One channel can carry most of the risk

A single model fit assigns every channel one number; some channels move far more than others when the specification changes. Nestack reports estimate instability by channel, not only in aggregate.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Channels with little spend variation7.5%3.9× Review
Brand search and retargeting5.7%3.0× Review
Channels changed mid-period3.3%1.7× Watch
Large channels with varied spend1.9%0.7× Normal
Bar: estimate-instability lift vs. the large-varied-spend baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

The loop closes on a case, not a story

The loop shuts when the wrong estimate is a regression case, not when it has been explained. That suite is what the next model run is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Interval width or edit rate rises in a channel slice.

02Diagnose

The point estimate everyone quoted is checked against the series, the priors and the specification until one cause explains the swing.

03Improve

Whatever changes ships against a version, with the runs that prompted it attached.

04Verify

Nothing ships until the affected specification cases pass a second time.

05Learn

The suite grows by one case; so does the calibration record.

Learn → DetectThe return edge. The next run starts against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, modelling workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Analytics discovery and boundary scoping and boundary definition.
02Media, KPI and warehouse source assessment.
03Prior-source and calibration rule mapping and rule mapping.
04Series ingestion and taxonomy mapping.
05Model-fitting logic and interval binding.
06Confidence scoring and gap routing.
07Analyst approval workflow.
08Warehouse and experiment-tool integration.
09Specification and prior cases.
10Guardrails and acceptance controls.
11Run-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne channel, one market ProductionProduction data and warehouse access AdvancedMultiple markets / brands
Introduced at Pilot
Model fit to your data
Analyst approval
Calibration baseline
Introduced at Production
Reporting by channel
Escalation workflow in your tools
Approved output actions
Warehouse integration
Introduced at Advanced
Multi-market calibration rules
Multi-stage analyst approvals
High channel volume
Multi-market model controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your spend, KPI and warehouse access Series ingestion and taxonomy mappingWeek 1
02Representative historical series and any past tests Model-fitting baseline and interval attributionWeek 2
03Your prior-source and calibration rules Prior-source and calibration-boundary mappingWeek 1
04Access to relevant APIs, feeds or exports Media, KPI and warehouse assessment, then integration setupWeek 2
05Estimates you would not want budgeted on Experiment cases and failure-mode testingWeek 4
06What no estimate may be read as Confidence scoring, gap routing, guardrails and approval controlsWeek 3
07Named analysts to review recommendations Analyst approval workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Bands follow the real work rather than the plan, which is why evaluation and pilot share week 5.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Analytics discovery, calibration mapping and the boundary W2Source integration and the model-fitting baseline W3Modelling workflow, confidence logic and approval controls W4Evaluation suite, guardrails and failure-mode testing W5Warehouse integration, pilot runs and targeted corrections W6One planning cycle modelled under the analytics lead, then handover
Reading the bandA bar covers only the weeks its work is named in — the week 5 overlap is real, not padding.
At the end of W6Once the cycle validates, Agent Care owns the running agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Advertising AI agent

Build a marketing-mix agent around your measurement stack.

Show us your spend and outcome data, your prior sources and who signs off on a number. The judgement that turns an experiment into a calibrated prior stays with the analyst, not the agent.

Nestack Agents · Marketing-mix modelling & incrementalityAGT-AM-15 · Agent Care available after launch