SmartInfer Lab. Research Note

Marketing Science · Identifiability

The Confident Wrong Number: Identifiability in Marketing Mix Modeling

SmartInfer Lab

Synthetic ground truth Bayesian MMM Honest limits

Abstract

Ask a marketing mix model how much each channel drove your revenue, and it will answer — confidently, for every channel, including the ones your data cannot actually measure. When two channels move together, or one spends at a steady level all year, the math cannot separate their effects; the model returns a number anyway, and it will not tell you which numbers it earned and which it invented to balance the equation. This note is about that gap — the distance between a number a model prints and a number the data can support — and it is where marketing budgets quietly go wrong.

We show it concretely, on synthetic data where the true answer is known. On a brand whose media truly drives 15% of revenue, a competent textbook fit attributed 97% to media — and reported the result as credible. We then describe the two disciplines Galileo's estimator adds — an identifiability gate and a reference-centered correction for always-on spend — and show the same case returning a corrected, and honestly caveated, answer that declines to claim what it cannot know.

The point is not a higher accuracy score. On our benchmarks, competent methods are close on accuracy and none dominates. The point is what a model does when the data is weak — whether it confesses, or guesses. This is a synthetic demonstration of a procedure, not a promise of accuracy on any particular real dataset, and we are careful throughout to separate the two.

01A confident number is not a known number

Every MMM tool will give you a contribution for every channel. That is the product. But some of those contributions are not identifiable from your data — no method, however sophisticated, can recover them — and a tool that hands you a confident number for an unidentifiable channel is not being more capable than one that refuses. It is being less honest.

The danger is specific. Budgets move toward the channels that look good. If a channel looks good only because the model could not tell it apart from something else, the money follows an artifact of the fit rather than a real effect. The failure is quiet: nothing in a typical report tells you which of its numbers are earned and which are manufactured to make the equation balance.

02What identifiability means, in plain terms

Identifiability is whether the data can, even in principle, tell two explanations apart. Two situations break it in marketing data.

03A demonstration

We can show this exactly, because on synthetic data we know the truth. We built a brand where media truly drives 15.1% of revenue — the rest is baseline and seasonal demand — with several channels spending at steady, always-on levels. Then we fit it two ways.

A standard regularized regression — a competent, textbook MMM — attributed 97.0% of revenue to media. Its average channel-ROI error was more than five-fold. And it stamped the result credible: nothing in the output hinted that a number six times too large was about to steer a budget.

Always-on synthetic brand · media share of revenue (true: 15.1%)

Naive fit

97.0%

✗ verdict: CREDIBLE — the 6× over-attribution is not flagged

Identifiability-gated fit

31.3%

✓ verdict: NOT CREDIBLE — the data can't support a confident split

Same synthetic data, true media share 15.1%. The gated fit is not the right answer either — and we say so. What changes is the verdict: the naive fit prints 97% with a clean bill of health; the gated fit more than halves the over-attribution and, crucially, returns NOT CREDIBLE rather than a confident artifact.

The identifiability-gated fit, on the same data, returned 31.3%. That is not the right answer either — 15% is — and we say so plainly: on this deliberately hard case the correction more than halves the over-attribution but does not eliminate it. Identifiability is a property of your data, not something a method can conjure away. What the gated fit does do is return a verdict of NOT CREDIBLE. It tells you the data cannot support a confident channel split here, instead of printing 97% with a clean bill of health. The difference that matters is not 97 versus 31. It is CREDIBLE versus NOT CREDIBLE. Most tools would rather be confidently wrong than admit a gap; we built one that would rather tell you the truth.

04The two disciplines that produce that behavior

Both feed a credibility certificate that closes every report: which channels are identified, which are not, what the confidence ranges are, and — when the data cannot support a confident answer — a NOT CREDIBLE verdict that gates the recommendations accordingly. These are not aspirations; they run as permanent regression tests, and a change to the engine does not ship if recovery regresses.

Ground-truth recovery gates (permanent tests)
GateResult
Channel-effect magnitude recoverywithin gate
Confidence-interval coverage (always-on benchmark)100%
Interaction detection precision1.00 (no false positives)
Out-of-sample error (held-out weeks)≈ 11% MAPE
Channel rank recovery (≈ 52 weeks)power-limited — reported as magnitude, not rank

05What the assistant does with it

This discipline is only useful if it survives contact with a user asking questions. Galileo's assistant runs the model and answers in plain language, but it is bound to the same honesty as the certificate.

Every figure it states is reconciled against the engine that computed it; it does not invent a plausible number to be helpful. Two public demonstrations show both ends of this behavior — a confident brand (American Heritage Leather) where the channels are identified, and a weak one (Bloom & Co.) where the honest answer is a refusal.

06When the answer is “not yet”

When Galileo returns NOT CREDIBLE, the natural reaction is that it's less capable than the tool — or the agency — that hands you a confident number for the same channel. It's the opposite. Confidence is not knowledge. If your data genuinely can't separate a channel's effect, a number that claims to have separated it isn't a better analysis — it's the confident wrong number, and someone is about to act on it. “This can't be known from the data you have” isn't the analysis failing. It's the analysis working.

And it isn't the end of the conversation — it's the start of a plan. When a channel can't be identified, the report says so and points to the fix: usually a small, deliberate experiment — a brief pause on the ambiguous channel, a dark test — that creates exactly the variation the observational data never had, so the effect becomes measurable. You can simulate the trade-offs first, with honest confidence ranges, to plan around the uncertainty instead of pretending it away. Designing the optimal experiment — the smallest test that reduces the uncertainty most — is an active research direction, written up in Exploration Is Bounded by Estimation. The principle is simple: when the data can't answer, you don't guess. You go get the data that can.

07Why this differs from running open-source MMM

Open-source marketing-mix libraries are powerful and flexible, and a skilled team can do excellent work with them. This note is not a claim to beat them on accuracy — we have run many methods against our own ground-truth benchmarks, and across competent methods accuracy is close; no approach dominates. The difference is in priorities — and in what a model does when it reaches the edge of what your data can support.

A general-purpose library will fit whatever you give it and return a coefficient for every channel, including the ones it cannot identify. Interpreting that output — knowing which numbers to trust, correcting the always-on trap, deciding when to refuse — is the job of the data-science team the library assumes you have. Galileo makes opinionated choices in the other direction: it would rather return NOT CREDIBLE than a confident artifact, it corrects always-on absorption by default, and it wraps the whole thing so a team without a staff econometrician can use it and interrogate it. Different product, different priority — honesty and deliverability before a confident-looking number.

08Limitations

These results should be read as what they are.

09Notes

The recovery gates behind this note — planted-truth recovery, the credibility invariants, and the honest out-of-sample error — run as permanent tests and are described on the How we validate page. The exploration side of this work — deciding where additional spend actually improves estimation under saturating response — is written up in the working paper Exploration Is Bounded by Estimation: Fisher-Directed Budget Allocation under Hill Saturation, listed on the research page.

Seeing this behavior on your own data happens in a private pilot — and the honest read on whether your data can support a confident answer is free. Start with Galileo.