AB Tasty and VWO have united, because half the loop was never enough.

Learn about the merger

Blog

The Crossover Test: A Better Way to Turn Customer Data Into Incremental Revenue

11 Min Read
The Crossover Test: A Better Way to Turn Customer Data Into Incremental Revenue

Any optimization program can point to a list of wins. Far fewer can sit across from a CFO and explain how much revenue those wins actually created. That usually gets framed as a reporting problem, but in my experience the issue starts much earlier, with how teams decide which customer data is worth acting on in the first place.

Customer data only creates value when it changes a decision that would have gone differently otherwise, and the outcome is worth more than the cost to implement and maintain it. The real question is which parts of your customer data can pass that test, and whether you can identify them before you invest.

“You keep using that [number]. I do not think it means what you think it means.”

The Crossover Test feature image

McKinsey found that faster-growing companies generate roughly 40% more of their revenue from personalization than slower-growing ones. It’s easy to interpret that as proof that personalization drives more revenue, but that isn’t what the finding actually shows. It tells us who is doing well, not what caused them to do well. Fast-growing companies also tend to have more money, better data, and more resources. Personalization may be part of the reason they grow faster, but it may also be a result of already having the resources to invest in it.

Confusing those two interpretations can get expensive. Researchers at Yahoo studied whether running display ads on its sites increased the likelihood that users would search for the advertised brand. When they compared exposed and unexposed users observationally, the apparent lift ranged from roughly 870% to 1,200%, depending on how the groups were matched. But when compared to the randomized holdout, the estimated lift was only 5.4%.

Ron Kohavi and Stefan Thomke later used the same example in Harvard Business Review to illustrate the gap between what your data shows and what your actions actually cause.

Customer data is very good at telling you who already buys and how those customers behave. However, on its own, it can’t tell you how those same customers would respond to a different experience. That’s why so many well-researched segments end up creating more work than incremental revenue.

Different customers do not automatically need different experiences

That distinction matters when deciding where personalization is actually worth the effort.

Say you run a test on a new product page. Version B beats version A overall. When you segment the results by audience, B wins by 3% for most visitors and by 9% for returning customers.

At first glance, returning customers look like an obvious personalization opportunity, right? Not exactly. If B wins for both groups, ship B to everyone and you’ve captured the gain without the added cost of creating a separate experience.

When you are deciding whether two audiences deserve different versions of an experience, the case for personalization starts when the best decision actually changes by audience. B might win for one group while A wins for another. That is known as a crossover interaction.

If the same version wins for everyone, a more complicated targeting strategy has to overcome an immediate disadvantage: more rules, more variations, and more work to produce an outcome you could have achieved by simply shipping the winner.

The Crossover Test chart

Researchers at Wharton put some numbers behind this. Anya Shchetkina and Ron Berman studied when differences between customers are actually worth acting on and tested five common personalization methods across two large field studies. The two studies were remarkably similar, with each including roughly 20 variations, similar best-performing versions, and the same general problem.

The results were wildly different.

In one study, personalization outperformed simply rolling out the single best version by 18%. In the other, the gain was only 4%. Nearly identical setups produced more than a 4x difference in the value of personalization.

The difference wasn’t the sophistication of the personalization method. Rather, it was whether the available customer data helped identify groups that responded differently to experiences being tested.

That leads to a less intuitive conclusion: more customer data doesn’t automatically increase the value of personalization. Data that helps you identify meaningful differences in responses does. Unfortunately, many programs invest heavily in the first without ensuring they have enough of the second.

The Crossover Test

Before building a separate experience for a segment, use this simple four-question diagnostic. If the answer to any of these questions is no, the better decision is usually to ship the best-performing experience broadly and invest the personalization effort elsewhere.

1. Does the winner change, or does one group just respond more strongly?

You’re looking for a flip, not a bigger lift. A group where B wins by 9% instead of 3% is a reason to ship B faster, not a reason to personalize on its own.

2. Did you define the segment before looking at the results?

If you slice a completed test enough ways, you’ll find something interesting eventually. A segment you discover after the fact is worth investigating, but it should become a new hypothesis to test, not a conclusion to act on.

3. Can you identify the customer early enough to act on the difference?

A crossover based on lifetime value doesn’t help if you can’t identify the customer until after checkout. The signal has to be available before the experience or decision you want to change.

4. Will the incremental value cover the additional cost of the personalization?

Every separate experience creates incremental cost. There may be an upfront build, plus ongoing content updates, QA, reporting, platform, and maintenance work. If the expected incremental gain doesn’t justify the added cost by a meaningful margin, the simpler experience is probably the better investment.

The Crossover Test flowchart

Do the math before you build

That last question needs a number behind it. Before you commit to the personalization, calculate how much incremental profit the audience has to generate just to cover the additional cost of creating and operating the experience.

Minimum Lift Needed = Annualized incremental cost of the personalization ÷ annual profit generated by the eligible group

“Annualized incremental cost” should include the costs that exist because the personalization exists. Depending on the program, that may include implementation, content production, QA, platform costs, reporting, and ongoing maintenance.

For example:

InputValue
Annual visits to the page in question2,000,000
Share of visits from this group12% (240,000 visits)
Revenue per visit from this group$4.00
Annual revenue from this group$960,000
Profit margin40% ($384,000)
Annualized incremental cost to maintain the rule$18,000
Minimum win needed4.7%

A confirmed 3% revenue lift looks like a win in the test report, but once you account for the cost of the experience, it creates only about $11,500 in incremental annual profit. Compared to the $18,000 in annualized cost, the personalization loses roughly $6,500 per year.

Now, to be clear, that’s not an argument against personalization. It’s an argument for knowing the break-even point before you build anything. The best opportunities are usually the ones where the audience is large enough, the behavior is different enough, or the upside is significant enough to justify the added cost.

Most personalization programs I’ve reviewed carry a ton of personalizations that were never likely to clear that bar.

I once worked with a brand in travel and hospitality that was testing a personalized lodging page for repeat visitors who had previously browsed premium properties. The experience surfaced their recently viewed property, prioritized a relevant package offer, and let them resume their prior search.

The test won by 6% and the team was ready to ship it, but we crunched the numbers first. The audience generated about $900,000 in annual revenue, so the lift was worth roughly $54,000. At a 25% margin, that became $13,500 in incremental profit. The incremental cost of operating the experience across seasonal offers, properties, QA, and reporting was about $15,000 per year. Against $13,500 in incremental profit, the winning test would actually lose about $1,500 annually.

The test worked, but the economics did not. We ended up shelving the personalization and putting that capacity toward a broader booking-path improvement with more upside.

You may already have several crossover opportunities

Before commissioning new research, go back through the tests you’ve already completed and ask a simple question: did the winning experience actually change by audience, or did the same version just perform better for everyone?

Only look at segments you had a reason to evaluate in the first place, and treat anything you find as a hypothesis to confirm with another test. The goal isn’t to mine old results for significance, it’s to identify a short list of plausible personalization opportunities using evidence you’ve already paid to generate.

For most teams, that list is shorter than expected. That might feel discouraging, but it’s actually a very important signal.

A short list doesn’t necessarily mean your customers all want the same experience. More often, it means the data you currently collect isn’t doing a good job of distinguishing customers who respond differently. That tells you where the next investment should go. Running more tests against the same audience definitions won’t solve the problem. You need better signals.

That might mean bringing behavioral, purchase, CRM, and other customer data together so you can identify new audience characteristics worth testing. Platforms like Wingify Data Platform can help make those signals usable by creating unified customer profiles and audiences that can be carried into experimentation. The data still won’t tell you which customers deserve a different experience, but it will give you better hypotheses about where those differences may exist.

Test those hypotheses and look for cases where the winning experience genuinely changes by audience. Once that difference is confirmed, a tool like Wingify Personalization Suite can turn it into a targeted experience that continues to be measured over time, rather than another one-off personalization rule that gets launched and forgotten.

Measure the program, not the wins

Even if every personalization clears the Crossover Test, there is one more number worth reconciling.

Adding the reported lift from every winning experiment will most likely overstate what the optimization program actually contributed. Even when deploying only those experiences that achieved statistically significant results, some of those winners will still reflect random variation. The closer a result is to the significance threshold, the greater that risk.

Minyong Lee and Milan Shen at Airbnb studied this problem and proposed methods for estimating the cumulative impact of experimentation programs. One of the most practical options was to keep a persistent holdout group.

The idea is simple. Keep a small portion of eligible traffic from receiving the changes the program ships, then compare the long-term performance with traffic receiving the accumulated winners.

That comparison may be less impressive than adding up every reported win, but it gives finance something much more useful: a credible estimate of what the program actually contributed.

Where to start

Customer data becomes useful for personalization when you can identify a group that responds differently enough to justify a different decision, recognize that group in time to act, and generate incremental profits that exceed the cost of building and maintaining the experience. 

That’s a higher bar than simply having a segment available to target.

A simple place to start is with your own test history. Pull the last twenty completed experiments and look for cases where the winning experience actually changed for a segment you defined in advance.

The number of validated differences you find is a more useful measure of your personalization opportunity than the number of segments sitting in your CDP. If the same experience keeps winning across audiences, the answer may be simpler: ship the winner and keep looking for customer differences that actually change the decision.

FAQs

Q1. Does this mean personalization is not worth doing?

Personalization is still worth doing, but the value is often concentrated in fewer cases than teams expect. In the Wharton research, both studies found value in personalization, but one produced roughly four times the gain of the other from a very similar starting point. The difference was whether the audiences actually responded differently enough to justify separate experiences.

Q2. So how do you know whether a crossover is real?

A crossover is more credible when the audience is defined before the test and has enough traffic to support a reliable result. Treat anything you discover after the fact as a hypothesis until you can confirm it with a follow-up test.

Q3. What if we don’t have enough traffic to test at the group level?

Limited traffic means the minimum win needed to justify testing at the group level is higher. In that case, the better decision is often to ship the strongest experience to everyone. Smaller programs usually get more value from improving that experience and learning more about their customers than from dividing limited traffic across personalization rules they can’t reliably validate.


Trevor Aneson

Trevor Aneson is VP of Digital Experience at 85SIXTY, where he leads customer experience, experimentation, personalization, and UX strategy across enterprise brands. For nearly two decades, he has built and scaled optimization programs that bring together customer research, analytics, experimentation, and technology to improve digital performance. His focus is on turning customer data into better decisions, validated experiences, and measurable incremental growth.

You might also love to read these on Website Personalization

UX and AI: 10 Emotional Audience Segments That Transform Personalization
10 Min Read

UX and AI: 10 Emotional Audience Segments That Transform Personalization

AdaptiveCX: Real-Time AI Personalization for Every Visitor
12 Min Read

AdaptiveCX: Real-Time AI Personalization for Every Visitor

Beyond Static Personalization: How Real-Time Intent Is Transforming Customer Experience
9 Min Read

Beyond Static Personalization: How Real-Time Intent Is Transforming Customer Experience

The better way to test,
learn and grow

Get a Demo Start Trial