Data-Driven Attribution: How Google's Model Works and What It Cannot See

David Lopes

TL;DR

  • Data-driven attribution assigns credit for a conversion across touchpoints using your own historical conversion paths instead of a fixed rule, and Google has made it the default model for new Google Ads conversion actions. At the job it was built for, splitting credit among Google ad interactions, it genuinely beats last click.
  • The limit is scope, not quality. Google Ads only scores Google touchpoints, so branded search absorbs credit created by channels the model never sees, sessions without marketing consent are estimated rather than observed, and the credit weights are never exposed to the advertiser. Attribution also only divides credit for conversions that happened, so it never establishes that the spend was incremental.
  • That gap is exactly what Polar closes. The Polar Pixel is first-party and matched directly to Shopify orders, so revenue reconciles instead of floating alongside platform numbers. Every paid, owned and organic channel lands in one journey, the attribution model stays your choice (first click, last click, linear, U-shaped or Full Impact), and offline touchpoints, post-purchase survey answers and refunds enter the same picture Google structurally cannot reach.

If you run Google Ads, you are already using data-driven attribution whether you chose it or not. Google made it the default model for new conversion actions and retired the rule-based alternatives from the interface. The number in your Google Ads column is now the output of a machine learning model you did not configure and cannot inspect.

That is not automatically a problem. The model is genuinely better than last click at the job it was built for. But almost every guide on this topic explains data-driven attribution from inside Google's own surfaces, using Google's own framing, and stops there. The question operators actually ask is narrower and harder: how much of my business does this model even see, and what do I do about the part it does not?

This piece covers the mechanics, then spends most of its length on the boundary. Not because the model is bad, but because knowing where a measurement stops is the difference between using it and being led by it.

What data-driven attribution actually is

Data-driven attribution assigns fractional credit for a conversion across the touchpoints that preceded it, using your own historical conversion paths to decide how much each one is worth. It sits opposite the rule-based models, which decide the split in advance: last click gives everything to the final touch, first click to the opening one, linear splits evenly, and position-based hands 40% each to the ends.

The distinction that matters is not "smart versus dumb". It is where the rule comes from. A rule-based model applies the same split to every advertiser on earth. A data-driven model derives the split from your account, so two brands running identical campaigns can get different answers, and the same brand can get a different answer in March than it got in January.

Google runs a version of this in two places. In Google Ads it distributes credit across ad interactions, including clicks and video engagements, on Search, Shopping, YouTube, Display and Demand Gen. In Google Analytics 4 it runs cross-channel, so organic search, email and referral traffic can receive credit alongside paid. OWOX has a useful breakdown of how the implementations differ across Google's products. They are related models, they are not the same model, and they will not agree with each other.

Before going further, here is the honest map of what the model covers. This is the part no ranking page publishes, and it is the thing worth reading twice.

Touchpoint or signalGoogle Ads DDAGA4 DDAIndependent attribution layer
Google paid clicksFullFullFull
YouTube engaged viewsFullPartialDepends on linking and consentPartial
Meta, TikTok, Pinterest clicksNonePartialOnly if tagged and not strippedFull
Email and SMSNonePartialFull
Organic and directNoneFullFull
Sessions without marketing consentModelledEstimated, not observedModelledServer-side recoverable
Offline, connected TV, direct mailNoneNoneIngestible as touchpoints
Post-purchase survey answersNoneNoneUsable as a touchpoint
Refunds and cancellationsNoneConversion stays creditedNoneNetted against revenue
The credit weights themselvesNot exposedNot exposedModel is selectable

Read down the first two columns and the pattern is clear. Google's model is excellent inside Google and blind outside it. That is not a defect, it is the design. The trouble starts when a number built on that footprint gets used to judge channels that were never in the footprint to begin with.

How Google's data-driven attribution model works

The counterfactual method

The model does not simply count touchpoints. It estimates how the probability of a conversion changes when a given ad interaction is added to or removed from the path. Google describes the method plainly in its Analytics documentation: the system builds conversion probability models from both converting and non-converting paths, then contrasts what happened with what could have happened, and assigns credit according to the difference.

An example makes it concrete. If a path with four ad exposures carries a 3% conversion probability, and removing the fourth exposure drops that to 2%, the model reads the fourth exposure as responsible for a 50% lift in probability and weights it accordingly. Repeat across every interaction and the weights become the split.

Google also calibrates these models against holdback experiments, which is a real methodological strength and one worth acknowledging before criticising anything. A model tuned against randomised holdouts is a genuinely different object from a heuristic someone picked in a meeting.

What the model actually looks at

The published signal list is short and specific: time between the ad interaction and the conversion, device type, the number of ad interactions on the path, the order of exposure, the format type, the creative asset, and query signals. Conversions can also be reattributed for up to seven days after they happen, which is why last week's numbers move after you have already read them.

Notice what is not on that list. Nothing about margin, nothing about whether the buyer was new or returning, nothing about whether the order was later refunded. The model optimises for the conversion event it was given, and it optimises for it well. It has no opinion on whether that conversion was worth having.

Where it runs, and why the two versions disagree

Google Ads DDA works across Google's own inventory. GA4 DDA works cross-channel across whatever GA4 can observe. Different inputs, different scope, different conversion definitions, so different outputs. Two numbers from the same vendor describing the same week will not match, and neither will match Shopify. Nobody is broken. They are three different measurements presented as though they were one.

What data-driven attribution cannot see

Four boundaries, in rough order of how much money they move.

Every touchpoint outside Google

Google Ads DDA distributes credit among Google ad interactions. A buyer who saw three Meta ads, opened two emails, then clicked a branded search ad produces one Google touchpoint and five invisible ones. The model does not undervalue the Meta ads. It never receives them.

The practical effect is that every channel Google can see competes for credit only against other channels Google can see. Branded search looks extraordinary in that contest, because everything that created the brand demand happened offstage. This is the single most expensive misreading in paid media, and it survives the switch from last click to data-driven attribution completely intact. Fixing it needs cross-channel attribution, not a better model inside one channel.

Sessions you never got consent for

When a visitor declines the marketing cookie banner, the browser-side signal that identifies their journey does not fire. Google fills that hole with modelled conversions, which are an estimate produced from the users who did consent. The estimate is reasonable in aggregate and unusable at the level operators actually make decisions: this campaign, this creative, this week.

Consent rates vary enormously by geography and by how the banner is built, so the size of the hole is specific to your store and worth measuring rather than assuming. What is consistent is the direction. As consent-gated tracking degrades, a larger share of the number you are reading is inferred rather than observed, and nothing in the interface distinguishes the two.

The weights themselves

You can see the output. You cannot see the function. There is no screen that shows what the model decided a YouTube view is worth relative to a Shopping click for your account, no export of the coefficients, and no changelog when the weights shift. Growth Method notes a related limitation on the Analytics side: the attribution values Google calculates stay inside the interface and are not exported to BigQuery, so even teams with a warehouse cannot rebuild the model's reasoning from their own data.

This is where the structural point sits, and it deserves stating without drama. The company scoring the media is the company selling the media. That does not make the model dishonest, and there is no evidence it is. It does mean the scorekeeper has a commercial interest in the score, the methodology is not auditable by the advertiser, and any disagreement between Google's number and yours is resolved by Google. Treating that as a fact about the arrangement rather than an accusation is the correct posture.

Whether the ad caused anything

Attribution divides credit for conversions that happened. It does not establish which of those conversions would have happened anyway. A retargeting campaign that reaches people already on their way to checkout will accumulate credit under any model, data-driven ones included, because those paths do convert at a high rate. Correlation is exactly what the model is built to measure.

Google's own calibration against holdback experiments is a partial answer to this, and it is why DDA outperforms last click. But calibration happens inside Google's model, on Google's inventory, for Google's purposes. The advertiser-side version of the question, whether this budget produced incremental revenue, is answered by incrementality testing, which deliberately withholds spend and measures the difference. Attribution and incrementality answer different questions. Neither substitutes for the other.

Data-driven attribution versus the rule-based models

Rule-based models get dismissed too quickly. Their weakness is that the rule is arbitrary. Their strength is that the rule is legible, stable and identical for everyone, which makes them a reasonable instrument for comparison even when they are a poor instrument for truth.

ModelHow credit is assignedGood forWhere it breaks
Last click100% to the final touchpointA stable baseline everyone can reproduceOverpays the closer, erases the opener
First click100% to the opening touchpointJudging prospecting and top of funnelIgnores everything that closed the sale
LinearSplit evenly across all touchpointsLong consideration cyclesTreats a stray visit as equal to a decisive one
Position-based40% first, 40% last, 20% sharedBrands that value discovery and close equallyThe weights are still a guess
Shapley or full impactMarginal contribution of each touchpoint across observed pathsAssessing paid channels against each other on one footingNeeds volume, and is only as good as its touchpoint coverage
Google data-drivenCounterfactual probability lift, learned from your accountBidding and budget decisions inside GoogleNot auditable, and blind to everything outside Google

If you want the longer treatment of each one, we have written up the nine attribution models separately. Neil Patel frames the choice as data-driven versus last click, which is the framing most guides use. It is the wrong axis. The real choice is not which model, it is whose model, running on whose data.

Worth crediting the upside honestly, because it is real and it is measured. Google publishes advertiser results including an 18% reduction in cost of sales against last click at one European retailer, and an 8% increase in incremental conversions at 8% lower cost per lead at another. Those are Google's own reported figures for Google's own model, which is worth holding in mind, but they point in a consistent direction: inside Google's inventory, data-driven attribution beats last click. The argument in this article is about scope, not quality.

Why your Google number will never match Shopify

This is the ticket that lands in every ecommerce team's inbox, so it is worth walking the arithmetic rather than restating that the tools differ.

Five mechanics produce the gap, and they compound:

  • Different attribution windows. Google credits a conversion to the click that started it, inside its lookback window. Shopify records the order on the day it was placed. A 30-day window means today's Google number contains orders from the past month.
  • Different credit rules. Google splits one order fractionally across several interactions. Shopify counts one order, once.
  • Double counting across platforms. Meta claims the same order Google claims. Add every platform's self-reported conversions together and the total exceeds your actual revenue, sometimes substantially.
  • Modelled conversions. A share of Google's count was estimated rather than observed, and it is not labelled.
  • Refunds. Shopify nets them. Ad platform conversion counts generally do not.

None of that is a bug in either system. It becomes a bug when the two numbers are put in the same report as though they measure the same thing. The fix is a reconciliation layer that starts from the order table, since orders are the only ledger that is not owned by a party with a stake in the answer, and works outward to the touchpoints.

How to read data-driven attribution without being misled

Five checks that cost nothing and change how the number reads.

  1. Check what share of conversions is modelled. Below a modest share, treat channel-level splits as directional. As it grows, stop making creative-level calls on that data.
  2. Never sum platforms. If Google claims 400 orders and Meta claims 350 on a week you shipped 500, the overlap is the story. Sum only from an order-level source.
  3. Compare like for like across time. Weights drift and conversions reattribute for up to seven days, so a week-old comparison is not stable. Pull the same window twice before concluding a channel moved.
  4. Separate branded from non-branded search. Branded search absorbs credit created elsewhere. Blending them makes Google Ads look better than it is and makes prospecting look worse.
  5. Judge budget shifts with a holdout, not a report. If a reallocation is worth making, it is worth proving. Withhold spend on one channel, watch total revenue, and see whether the model was describing cause or coincidence.

Feeding the model better data

The alternative to distrusting the model is improving its inputs. This is the underrated move, because it makes Google's bidding better at the same time as it makes your reporting more honest, and the two usually get treated as opposing goals.

Three things are worth doing, in order of return:

Recover the conversions the browser tag missed. A server-side conversion feed sends orders back to Google from your own order record rather than from the visitor's browser, so purchases lost to consent refusal, ad blockers and tracking prevention re-enter the model. The model then learns from a fuller picture of what converted, which improves the bidding as much as the reporting. Polar's Conversion API Enhancer does this by uploading offline conversions from Shopify order data back into the ad platforms.

Put the touchpoints Google cannot observe somewhere they count. Offline impressions, connected TV, direct mail and podcast exposure can be brought into an attribution journey as real touchpoints rather than being written off as unmeasurable. So can post-purchase survey answers, which for brands whose demand comes from word of mouth or practitioner recommendation are often the only honest signal available. None of that will ever appear in a Google report.

Fix the definitions before you fix the model. Decide whether you are counting on a cash or accrual basis, whether only paid touchpoints are eligible for credit, and how far back a touchpoint can sit and still qualify. These settings move the numbers more than the choice of model does, and most teams have never set them deliberately.

When data-driven attribution is the right model

The decision you are makingUse DDA?Why
Setting bids inside Google AdsYesIt is the model the bidding runs on, and it beats last click at this job
Comparing Google campaigns against each otherYesSame footprint, same rules, so the comparison is fair
Splitting budget between Google and MetaNoEach platform scores itself and cannot see the other
Reporting revenue to a board or investorNoFractional credit does not reconcile to the order ledger
Judging whether spend was incrementalNoAttribution divides credit, it does not establish cause
Diagnosing a channel with low click volumeCarefulThin paths make the modelled share larger and the split noisier
Measuring offline or word of mouth demandNoThose touchpoints never enter the model

The pattern: data-driven attribution is a good instrument for decisions inside Google and a poor instrument for decisions about Google. Most of the damage this model does comes from being used for the second while being validated on the first.

Building an attribution layer that does not belong to a bidder

The structural fix is not a better model. It is putting the scorekeeping somewhere the outcome does not benefit anyone selling media, and starting from the order table rather than from any platform's conversion count.

That is what Polar does. The Polar Pixel is first-party and matched directly against Shopify orders, so conversion value reconciles to gross sales instead of floating alongside it. Every paid, owned and organic channel lands in one journey, which means a Meta impression, an email open and a branded search click are finally scored against each other rather than each being scored by its own vendor.

The model stays your choice rather than the platform's. First click, last click, linear, U-shaped and Full Impact, a Shapley model that estimates each touchpoint's marginal contribution, all switchable on the same underlying data. Seeing one week under two models tells you more about the fragility of a channel's case than any single number does. Lookback windows, paid-only eligibility and cash versus accrual are settings you control, not defaults someone else picked.

Then the parts Google structurally cannot reach: offline and connected TV touchpoints ingested into the same journey, post-purchase survey answers treated as real signal, refunds netted against revenue, and conversions fed back into the ad platforms so their models learn from what actually happened rather than from what the browser managed to report.

None of that replaces Google's data-driven attribution. It gives you something to check it against, which is the thing operators have been missing since the model became the default and the alternatives disappeared from the interface.

FAQ

Data-driven attribution is an attribution model that assigns fractional credit for a conversion across touchpoints using your own historical conversion paths, rather than a fixed rule. Google runs it in Google Ads across its ad inventory and in Google Analytics 4 across channels. It is now the default model for new Google Ads conversion actions.
Data-driven attribution works by comparing converting and non-converting paths to estimate how each ad interaction changes the probability of a conversion. If removing an interaction drops that probability, the model credits the interaction with the difference. It uses signals including time to conversion, device, exposure order and format type.
Data-driven attribution is accurate at the job it was built for, which is splitting credit among Google ad interactions, and Google calibrates it against holdback experiments. It is not accurate as a picture of your whole marketing mix, because touchpoints outside Google never enter the model and unconsented sessions are estimated rather than observed.
For bidding and budget decisions inside Google, data-driven attribution is better than last click, and Google publishes advertiser results showing lower cost per acquisition after the switch. For deciding how much to spend on Google versus another platform, neither model helps, because both are scored on Google's footprint alone.
Google Ads conversions do not match Shopify orders because the two count different things. Google dates a conversion to the click inside its lookback window and splits one order fractionally across interactions, while Shopify records whole orders on the day they happen and nets refunds. Modelled conversions and cross-platform double counting widen the gap further.
Google removed the minimum data requirements for data-driven attribution and made it available for every conversion action. Removing the threshold does not remove the underlying effect: with thin conversion paths a larger share of the result is modelled and the credit split is noisier, so treat low-volume campaigns as directional.

Table of contents

Make strategic decisions in minutes

See every metric that matters, in one place.

Book a demo

Ecommerce Benchmark

4,000+ brands, refreshed weekly.

See the benchmark

Frequently asked questions

Ready to stop guessing and start growing?

Make strategic decisions in minutes, not weeks.

Book a demo