Where Data Stops in Marketing: A CTO's Notes on Data-Driven Marketing

Posted on Oct 10, 2026

In June 2026 I wrote an internal post at Buser with a title I liked more than I should have: "Marketing is now MarTech". Marketing moved under Technology, and I, the CTO, became accountable for it. The reason was practical. A good marketing idea lost speed waiting in the tech and data queue, and putting engineering and data science next to marketing closes the loop between a decision and its measurement. But the sentence I cared about most in that post was another one: marketing strategy decides what the data should chase. Four months in, I still think that sentence is right, and I now think it is the hardest part of data-driven marketing, because data will chase anything you point it at, with total confidence, and report success the whole way. This post is my attempt to organize what I have been reading and living since then, written by someone whose instinct is to trust the funnel and who is learning, slowly, where that instinct breaks.

Where Data Stops in Marketing: A CTO's Notes on Data-Driven Marketing

Growth as an experiment loop, and the door Ellis left open

The phrase "growth hacker" comes from a 2010 post by Sean Ellis, Find a Growth Hacker for Your Startup (the URL still carries its original title, "Where are all the growth hackers?"). His argument was aimed at startups that had already found product-market fit and an efficient conversion process. At that stage, he wrote, the next step is finding "scalable, repeatable and sustainable ways to grow the business. If you can't do this, nothing else really matters." Instead of hiring a VP of Marketing with a generic list of requirements, he proposed hiring or appointing "a person whose true north is growth", someone disciplined enough to run a loop of prioritizing ideas, testing them, keeping what works and cutting what doesn't, and doing it faster than anyone else. He even noted that some of the best growth hackers he knew came from engineering. Seven years later he and Morgan Brown turned the idea into Hacking Growth (Crown Business, 2017), which formalized the loop as a cross-functional process. The mechanism is easy to see from an engineering seat: the loop is the scientific method applied to a funnel, and it works because it replaces opinion with feedback.

What people forget is a single sentence in the original post. Ellis asks whether positioning matters and answers himself: "Only if a case can be made that it is important for driving sustainable growth (FWIW, a case can generally be made)." The man who coined growth hacking left a door open for the part of marketing that is not an experiment, and most growth teams I have seen walked right past it. The limit of the loop is structural, not a matter of discipline. An experiment can only test what can be randomized, measured inside a short window and attributed to a single change. Positioning, memory and reputation fail all three conditions. They build slowly, they spread across every touchpoint at once, and their effect shows up months later as people who "just came directly". So a team that runs only the loop is not neutral. It quietly shifts budget toward whatever the loop can see. For a data team, the practical implication is to treat the experiment loop as one instrument among several, and to write down explicitly which decisions the loop is not allowed to make, because otherwise it will make them by default.

The mind before the funnel: what an A/B test can and cannot see

Daniel Kahneman's Thinking, Fast and Slow (2011) describes two modes of thinking. System 1 is fast, automatic and associative. It recognizes a brand, feels that something is safe and picks without deliberating. System 2 is slow and effortful. It compares prices, reads conditions and calculates. A funnel is a System 2 artifact. It describes a person who searches, compares, clicks and pays, one observable step after another, and that is exactly why engineers trust it. The trouble is that a large part of the decision has already happened in System 1 before the first event fires. When someone opens a search box and types a brand name, or skips the search engine and opens the app directly, the choice of which company to look at was made somewhere the funnel never saw: an ad remembered from weeks ago, a friend's comment, a bus seen on the road. The funnel measures the execution of a decision, and we keep reading it as if it measured the decision itself.

Prospect theory, published by Kahneman and Amos Tversky in Econometrica in 1979, explains why that hidden part matters so much. In a series of choice experiments with hypothetical gambles, they showed that people do not evaluate outcomes in absolute terms. They judge them as gains or losses relative to a reference point, and losses weigh more than gains of the same size. That gives framing real power: the same price presented as a discount or as the absence of a surcharge is not the same price to the brain. Here is where testing gets subtle. An A/B test is very good at measuring which frame wins this week. It cannot tell you that every winning discount message also moves the reference point. If customers learn that the brand is always on sale, the full price starts to read as a loss, and the next test will once again say that a discount wins. Each test was correct. The sum of the tests can still be a brand that only sells on promotion. Robert Cialdini's Influence (first published in 1984) shows the same pattern from the persuasion side. Social proof, authority and scarcity work because they are shortcuts System 1 uses to save effort, and "only 2 seats left" will lift conversion in almost any test. When the scarcity is not real, the tactic becomes a dark pattern, and what it spends is trust, something no dashboard tracks. I wrote about the product version of this in When product breaks, it breaks trust.

The counter-argument is fair: you can test long-term effects with holdout groups and longer windows, and good teams do. In practice, though, the window is set by how long the company is willing to wait, and that is usually weeks. At Buser we feel this in pricing communication. Price is the easiest thing to put in an ad, the easiest to compare in a funnel and the easiest to test, so the data keeps confirming that price messages convert. But a bus trip is time, safety and arriving well at the destination, and the reason someone chooses a brand a second time is rarely the thing that made them click the first time. For a data team, the practical rule I took from this is to never let a test result be the only record of a decision. Write down the reference point the change could move and the trust it could spend, then check those again a quarter later, when the test has long been closed and celebrated.

Data against data: light buyers, time horizons and the water analogy

This is the section where the fight is not between data and intuition but between two kinds of data. Byron Sharp and the Ehrenberg-Bass Institute, in How Brands Grow (Oxford University Press, 2010), built their argument on decades of panel data about what people actually buy. The recurring pattern is that brands grow mostly by increasing penetration, meaning more buyers, and especially more light buyers who purchase rarely. Loyalty rises a little as a consequence of size, not as the cause of it. The mechanism they propose has two parts. Mental availability is the chance that a brand comes to mind in a buying situation, and physical availability is how easy it is to find and buy. Both are built across the whole market, among people who are not buying right now. This contradicts the instinct of performance marketing, which keeps narrowing toward the people most likely to convert today. The data that performance teams look at is real, but it only covers people already in the market, and so it keeps recommending more spend on the people who were already going to buy.

I use Sharp's data, but not his catechism. His laws were mostly measured on supermarket shelves, in categories where buying is habitual, low involvement and nobody remembers why they picked the blue box over the red one, and somewhere along the way they started being preached as the physics of every market. A 2007 paper by Jenni Romaniuk, Sharp and Andrew Ehrenberg found low perceived differentiation between competing brands and proposed putting distinctiveness, not differentiation, at the center of brand strategy. That is a fair finding about shampoo. In the hands of a lazy team it becomes a permission slip to have nothing to say, as long as you say it in a recognizable color. Critics like Felipe Thomaz go further and argue, in an interview with Contagious, that the model leaves differentiation out of branding altogether. I am less radical than that, but I share the discomfort. Treating loyalty as a byproduct of size fits a cereal aisle better than a service people trust with their safety on a ten hour overnight trip, where the trip itself, and what people tell their friends about it, is a large part of the brand. And there is an irony I can't resist: a school born from fighting marketing dogma now has its own dogma, its own followers and its own sermons. What I keep from it is the part that survives contact with our reality, mental availability and the reminder that growth needs people who are not buying today.

At Buser someone on the team explained this with an analogy I keep repeating: it is like selling water only to people who just finished running. Almost all of them buy, so the return looks amazing. But runners are few. To sell more water, you have to sell to people who are less thirsty, and the return per ad falls. A falling return in paid media is often read as the channel getting worse, and sometimes that is true. Often, though, it is the channel reaching the edge of the people who were already thirsty. Les Binet and Peter Field add the time dimension in The Long and the Short of It (IPA, 2013). They analyzed 996 campaigns from the IPA Effectiveness Databank, covering about 700 brands in more than 80 categories between 1980 and 2010, comparing top performers with weaker ones over short, medium and long horizons. Their finding is that sales activation produces short spikes that fade, while brand building produces slower growth that compounds over years, and that the most effective campaigns combine both. The famous 60:40 split comes from this work. It is an average optimum for brand building versus activation, which they say varies by category and follows how category spend itself divides, so it is a starting point to reason from, not a law.

The limits deserve honesty. Ehrenberg-Bass patterns come mostly from consumer goods, and a bus trip is a considered purchase. The IPA databank is made of award entries, which leans toward success stories. Still, there is a trap here that every company with a strong brand falls into, and it is the lens I now apply at Buser. When people already know you, many of them come directly or search for your name, and every channel that touches them on the way looks wonderfully efficient. The dashboard credits the channel for demand the brand created years before. Read through Sharp and Binet & Field, that efficiency is partly harvest: you are collecting mental availability that was planted earlier, by advertising, by word of mouth, by every trip that went well. The danger is not the harvest itself. The danger is reading it as proof that the channel alone can carry growth, because a dashboard has no way to tell efficiency apart from eating the seed stock. For a data team, the implication is to measure the stock and not only the flow: direct and brand-search traffic, unaided awareness and penetration among light buyers, tracked over years. When performance looks great and nobody can say where the demand came from, that is the moment to get suspicious, not the moment to celebrate.

Measurement: attribution, incrementality and the metric that becomes the target

Two papers changed how I read every marketing report. In the first, Thomas Blake, Chris Nosko and Steven Tadelis (Econometrica, 2015) ran a series of large field experiments at eBay. For brand keywords, eBay stopped buying paid search ads on the brand name in some search engines and watched what happened. Almost all of that traffic came back through organic results, and the authors found no measurable short-term benefit from brand-keyword ads. For non-brand keywords, they switched paid search off in a share of geographic markets and compared them with the rest. New and infrequent users were positively influenced by the ads, but frequent users, who were not influenced at all, accounted for most of the spend. The result was that the average return was negative, and that the causal return was a fraction of what non-experimental estimates suggested. The mechanism is selection. Ads show up in front of people who already intend to buy, so clicks and purchases correlate whether or not the ad caused anything. The second paper, by Randall Lewis and Justin Rao (Quarterly Journal of Economics, 2015), looked at 25 large field experiments with major US retailers and brokerages, most reaching millions of customers. Individual sales are so volatile compared with the cost of advertising per person that the median confidence interval on ROI was over 100 percentage points wide. An informative experiment can need more than 10 million person-weeks. In plain words, even honest experiments struggle to tell a profitable campaign from a useless one, and observational methods are badly biased by the targeting itself.

Living inside this is humbling. At some point our Google return started swinging from one day to the next, and on some days a result even showed up as negative. We went hunting for a bug in the data pipeline and the trail led somewhere else, to the weights of the attribution model. The model was doing exactly what it was built to do, redistributing credit across touchpoints, and it also rewrites the past as data matures, so a month that looks bad today can look fine three weeks later without anything changing in the world. Another time we found a cashback browser extension that popped up when the user was already on our site, at checkout, and took the attribution of a sale we would have made anyway. Both cases are the eBay paper at a smaller scale. The report credited activity that happened to sit next to a purchase, not activity that caused it. None of that is a reason to stop measuring. It is a reason to stop reading the attribution report as if it were the world. The report is a model, and a model has opinions.

Donald Campbell described what happens next in 1979, in Assessing the impact of planned social change: the more a quantitative indicator is used for social decision making, the more it is subject to corruption pressures and the more it distorts the process it was meant to monitor. Charles Goodhart made a similar point about monetary targets a few years earlier, and it became Goodhart's law. Years ago I wrote a note to myself that still holds: a big return on ad spend is the marketing team's own vanity metric. ROAS rewards spending exactly where people were already going to buy, which is the eBay result turned into an incentive. The more a team is judged by ROAS, the more it drifts toward brand terms, retargeting and frequent buyers, and the better the number looks while the business learns less. I made the same argument about engineering in Vanity Metrics in Engineering. For a data team, the implication is to change the question from "what did this channel get credit for" to "what would have happened without this money". That means holdouts and geo tests wherever they are affordable, ROI and incremental margin instead of ROAS, and treating any attribution model as a hypothesis to calibrate against experiments, not as an accounting truth.

Alchemy, the SEO scar and how I run data-driven marketing now

Rory Sutherland, in Alchemy (2019), argues that the opposite of a good idea can also be a good idea, and that much of what works in marketing fails a logic test. His mechanism is that people are not optimizing the variable the analyst measures. They are managing anxiety, status and the feeling of control, so a change that adds no measurable value can matter a lot, and a change that wins on paper can backfire. I have my own scar here. In 2024, before I owned marketing, I pushed an SEO idea: pages for bus companies, so people searching for a company's name would land on Buser and see a search box. Our competitors did the same thing, the traffic logic was clean and the pages brought visits. We later had to take them down. The traffic data was right about the traffic. It simply had no column for brand risk, or for how the companies on those pages would feel about it. I wrote about the SEO side in Topical authority: destination before ticket. What I did not write then is that the most dangerous metric is the one that is correct and incomplete, because nobody argues with it.

So this is how I try to run it now. Data is an instrument, not a judge. In September we had to choose between a few spend scenarios for paid search in a month where we had never invested that much. Our analyst's reading was by marginal return: each extra unit of money still returned more than it cost, but less than the unit before, so returns were diminishing but present. We kept the pace and I asked the team to watch the ceiling. The month closed below the forecast and above the floor we had set, because the forecast had asked for a return the channel could not deliver at that volume. I am fine with that outcome, and the reason is the lesson. The useful decision was the floor, a limit we agreed on before looking at the results, not the forecast, which was a model's guess dressed up as a target. Second, automate what is mechanical. We are developing mads, an open source tool that is still early and in active development, to generate search campaigns from route data, with keywords, negatives, ad groups and responsive ads. Building the structure of a campaign is engineering work and engineering should do it, but deciding what the campaign should say is not. Third, protect what does not show up in attribution. Brand budget is always the first line a dashboard suggests cutting, because its effect never shows up inside an attribution window, so it deserves its own line, set by strategy and not renegotiated every time a weekly report looks great. Fourth, keep the intelligence in-house. The models, the segments and the decisions about who gets which message are ours, built by our own engineering and data science, and we plug them into our CRM. The CRM is the delivery pipe, not the brain. That matters for the same reason as everything above: every vendor ships a default metric and a default way to optimize it, and if the intelligence lives inside the tool, that default quietly becomes your strategy. Owning the intelligence is how we make sure strategy, not a product setting, decides what the data chases. Underneath all four sits the sentence from June. Marketing decides what the data should chase. If you let the data decide, it will chase the people who were already thirsty and report success all the way down.

I would like to hear from people who made this move from the other side, engineers or data people who inherited marketing. What did the data tell you that turned out to be true, and what did it tell you that a good marketer knew was wrong?

References

Leave a comment

Comments on this blog live on X. The button opens a post with this page already linked — just write what you think.

Comment on X →