← All projects
Project · Machine learning · Attribution modeling

Customer journey
attribution.

Six standard models agree on who gets the credit. Tested against a known answer, only one holds up.

Splitting marketing budget across channels means knowing which touchpoints actually drive conversions — and real conversion data carries no ground truth to check that against. So this project builds one: a simulator generates 8,000 journeys with fixed, known per-channel effects, every method is scored on whether it recovers that known ranking, and only then is the surviving method pointed at 267,084 real journeys from the GA4 public e-commerce dataset. Six standard methods — five touch-position heuristics and a Markov chain — all fail the test; a data-driven, Shapley-value model is the only one that passes.

267,084Real journeys analyzed
1.63%Empirical conversion rate
8,000Simulated journeys, known effects
0.79Best rank recovery (ρ) vs. truth
01

You can't check attribution against reality. So build a reality you know.

8,000 synthetic
journeys
7 channels

A conversion never tells you which touchpoints caused it, so no attribution method can be scored for correctness on real data. The workaround is a simulator: 8,000 journeys generated from a fixed rule where each channel's true effect on conversion is set by construction — email strongest, display almost inert, the rest in between. Every method is then ranked on one question: does its channel ordering match the known truth? Spearman ρ = 1.0 is perfect recovery, 0 is chance, negative is backwards.

← backwards better → ρ = 0 (chance) Data-driven (Shapley) ρ = 0.79 First-touch ρ = 0.07 Last-touch, linear, time-decay, position ρ = −0.04 Markov removal effect ρ = −0.18
Rank recovery vs. known channel effects · 8,000 simulated journeys Spearman ρ between each method's channel ranking and the true generative effect
What the test shows

None of the five touch-position heuristics beats chance (ρ between −0.04 and 0.07). Markov removal effect does worse still at −0.18 — actively backwards: it ranks the rare-but-powerful email channel last, because removal effect is confounded by how often a channel appears, not just how much it moves conversion. Only the data-driven Shapley model recovers the true ordering (ρ = 0.79) — strong, not perfect.

One thing the Markov chain gets exactly right: its predicted overall conversion rate matches the empirical rate to four decimal places, on both simulated and real data. The chain arithmetic is sound; it's the per-channel credit split that fails — and it fails consistently, holding the same backwards ranking across bootstrap resamples. Stable and correct are different things.

“A method that can't recover an answer you planted yourself has no business ranking channels when nobody knows the answer.”

That's the license for everything below: the real GA4 data is read mainly through the one model that passed — with the caveat that clearing a simulation is necessary, not sufficient, and the real dataset carries problems of its own (section 06).

02

Every heuristic rewards volume. One model doesn't.

6 heuristics
+ 1 data-driven
6 channels

First-touch, last-touch, linear, time-decay, position-based, and Markov removal-effect all land in roughly the same place: organic search leads, direct is a clear second, paid search is reliably last. That agreement looks reassuring — until you notice it tracks touchpoint frequency almost exactly, not incremental conversion lift. It's the same exposure bias the simulator exposed in section 01.

Heuristic average Data-driven (Shapley)
Organic search 30.7% 12.8% Direct 22.6% 10.3% Referral 19.7% 25.5% Display 12.5% 6.0% Other 10.4% 43.0% Paid search 3.9% 2.3%
Heuristic average vs. data-driven credit, by channel navy = avg. of 6 heuristics · red = data-driven (Shapley)
The actual finding

Swap in the data-driven model and the ranking doesn't just shift — parts of it invert. Organic search drops from ~31% average heuristic credit to 12.8% (0.42×), direct drops from ~23% to 10.3% (0.45×), while “other” jumps from ~10% to 43% (4.1×) and referral rises from ~20% to 25.5% (1.29×). Paid search is the one channel every model agrees on — weak everywhere, 2–5% of credit regardless of method.

Worth the caveat: this is the model that passed the simulator test, but it passed at ρ = 0.79, not 1.0, and its AUC on the real data is only 0.712 — decent, not excellent. Treat the reshuffling as directional evidence that heuristic credit is inflated by exposure, not as a precise reallocation of budget.

ChannelFirst-touchLast-touchLinearTime-decayPosition-basedMarkov removalData-driven (Shapley)
Paid search0.0500.0330.0390.0370.0410.0360.023
Organic search0.3530.2800.3070.3000.3130.2900.128
Referral0.1660.2160.1980.2030.1940.2070.255
Direct0.2300.2250.2260.2240.2270.2260.103
Display0.1560.1070.1250.1190.1290.1160.060
Other0.0450.1400.1040.1170.0970.1240.430
03

Attribution ranking ≈ touchpoint volume ranking.

349,594 total
touchpoints
6 channels

Line this up against the heuristic attribution chart above and the match is almost exact: organic search is both the highest-credited channel and the most frequent touchpoint. That's the signature of models rewarding exposure rather than causal contribution.

Organic search 120,356 · 34.4% Direct 81,236 · 23.2% Referral 60,755 · 17.4% Display 51,452 · 14.7% Other 20,324 · 5.8% Paid search 15,471 · 4.4%
Touchpoint frequency by channel share of all touchpoints
04

Multi-touch paths barely register next to single-touch volume.

Top 10 paths
by journey count

Organic search alone (87,082 journeys) and direct alone (54,267) each dwarf the best multi-hop combination — organic search → referral, at 3,262. The five single-channel paths in navy are each an order of magnitude bigger than any two-step sequence in red.

Single touch Multi-touch
Organic search 87,082 Direct 54,267 Display 39,693 Referral 33,561 Paid search 11,894 Organic search → Referral 3,262 Organic search → Other 2,680 Organic search → Direct 2,670 Direct → Referral 2,147 Direct → Organic search 1,803
Top 10 journey paths navy = single touch · red = multi-touch
05

83% of journeys are one touch. There's no journey to attribute.

Mean length
1.31 touches
log scale

The attribution debate above is entirely about splitting credit across touchpoints in a journey — but for 83% of the 267,084 journeys, there's only one touchpoint. Multi-touch attribution modeling only meaningfully applies to the ~17% of journeys with two or more touches, and the long tail past 5 touches is just over 1% of total volume.

1M 100K 10K 1K 100 10 221,730 1 28,948 2 7,981 3 3,597 4 1,911 5 1,076 6 656 7 454 8 265 9 207 10 161 11 98 12
Journey length distribution log-scale y-axis · counts labeled
06

Every channel's biggest outcome is drop-off, not conversion.

First-order
transition rates

Looking at what immediately follows a visit to each channel — convert, continue to another channel, or drop — drop-off dominates everywhere — for every channel except “other” it's the outcome roughly four times in five. But the conversion rates that do differ are telling. “Other” converts at 3.3% per touch and referral at 1.7%, against roughly 1% for organic search, display and paid search. Those are the same two channels the data-driven model rewards and the heuristics ignore.

Continue Drop Convert
Organic search Continue 17.0% · Drop 81.9% · Convert 1.1% Direct Continue 18.9% · Drop 79.8% · Convert 1.3% Referral Continue 21.8% · Drop 76.6% · Convert 1.7% Display Continue 17.7% · Drop 81.4% · Convert 0.9% Other Continue 39.3% · Drop 57.4% · Convert 3.3% Paid search Continue 18.7% · Drop 80.3% · Convert 0.9%
Per-channel outcome breakdown share of each channel's first-order transitions · convert segments small but drawn to scale
A caveat on the numbers

These rates come from the full first-order transition matrix — every channel→outcome edge, including the small conversion edges the flow diagram above drops for readability. One caveat remains: “other” is a catch-all bucket for traffic the channel-grouping logic couldn't label, so its 3.3% rate blends together sources that may behave nothing alike. In the obfuscated GA4 sample that bucket is large and disputed — read the number as a reason to break the bucket apart, not as a channel to fund.

“The channel doing the most work isn't the one showing up most often — it's the one nobody built a heuristic to notice.”
07

Four conclusions worth acting on.

Findings
not recommendations
as gospel
01

Trust the method before the ranking.

On simulated data with a known answer, five touch-position heuristics scored no better than chance and Markov removal effect scored worse (ρ = −0.18). Only the data-driven Shapley model recovered the true order (ρ = 0.79). Last-touch dashboards are measuring exposure, not lift.

Method
02

Organic search and direct look strong mainly because they're loud.

Both collapse under the one model that passed the test (0.42× and 0.45× their heuristic credit). Treat the drop as directional, not exact — that model cleared the simulator at ρ = 0.79 and scores only 0.712 AUC on the real data.

Caution
03

Referral is the real signal; “other” is probably an artifact.

Both punch above their heuristic weight in the causal model, and both convert faster per touch (1.7% and 3.3% vs. ~1% elsewhere). But “other” is an unlabelled catch-all — its jump to 43% credit is more likely a channel-grouping artifact than a channel. Break that bucket apart before acting on it.

Investigate
04

The bigger lever might not be attribution at all.

83% of journeys are single-touch and the blended conversion rate is 1.63%. Most of this model debate concerns a minority of traffic — why single-touch conversion is so low is arguably the larger question.

Conversion