Conversion experimentation

Shopify A/B testing that ships winners

Shopify A/B testing for $100K+/month stores ready to scale conversion experiments without shipping blind. We design hypotheses, run product page and checkout testing, then ship statistically significant winners — not dashboards.

Experiment readoutRunning
Control2.8%
Variant B3.4%Lead
Confidence96%

Past 95% — ship it. Below that it stays running, or it gets reverted.

Hypothesis-led experiment design

Every test starts from a real problem and a falsifiable hypothesis — not a hunch or a trend.

Statistical significance reporting

We run tests to significance and report the truth — including the ones that lose — so you scale what actually works.

Ship-winners deployment workflow

Winners get shipped into native Liquid, not left living in a testing tool that slows your store down.

Read this before you brief a test

Most Shopify stores shouldn't be A/B testing yet.

A/B testing is a sample-size problem before it is a creativity problem. The maths does not care how good the idea is. If too few people see the two versions, the difference you measure is noise — and shipping on noise is worse than not testing at all, because you have now hard-coded a random number into your storefront and you believe it works.

The numbers below are not our results. They are what the statistics require before a difference can be told apart from chance. Read them as an entry price.

~3,800

Visitors per variant

What it takes to detect a 10% relative lift at 95% confidence and 80% power. A 10% lift is already a good outcome — this is the price of being able to see one.

~1,000

Visitors per variant

What a 20% relative lift needs at the same confidence and power. A swing that large is rare, and designing for it means testing big structural changes rather than button colours.

100 + 1–2 wks

Conversions per variant, minimum run

Sample size alone is not enough. You need at least 100 conversions in each variant and a run covering one to two full weekly cycles, because Tuesday shoppers and Saturday shoppers are not the same people.

In practice that puts the floor at roughly 10,000 monthly visitors before testing is worth starting, and closer to 50,000 a month before results are reliably trustworthy. Find your store below.

Whether A/B testing is worth starting at your traffic level
Store sessionsWhat the maths allowsWhat we would tell youTest?
Under 10,000 / monthNothing you can trust. Even a 20% swing needs months to reach significance, by which time the season has changed underneath the test.Do not test. Fix what is plainly broken and buy traffic.Not yet
10,000 – 30,000 / monthOne test at a time, large structural changes only, with run times measured in weeks rather than days.Test rarely. Put most of the effort into acquisition and merchandising.Selectively
30,000 – 50,000 / monthGenuine tests on your highest-traffic template, read at the overall level. Segment-level reads are still thin.Start a light programme. One test per cycle, product page first.Selectively
50,000+ / monthA continuous programme, smaller detectable effects, and enough volume to read results by device and source.Test properly. This is where experimentation starts to compound.Yes

This is the part of the pitch that costs us work.

We would rather say it now than four months into a retainer. If your store sits under the line, the highest-return work is not experimentation — it is fixing the things that are plainly broken. A slow product page. A cart that takes three taps to open. A mobile filter nobody can use. Then getting more people through the door. None of that needs a test to justify it, and a test at that traffic level cannot tell you whether it worked anyway. Come back to testing when the traffic can pay for it.

Test surface

What we test.

Six areas cover almost everything worth testing on a Shopify store. We start where the traffic and the money already are, which is nearly always the product page, and work outwards from there.

Product pages

Image order and gallery behaviour, review and trust badges, variant selectors, the position and wording of add to cart, and how far down the description sits.

Why it moves. The product page carries the most qualified traffic in the store and it is where the buying decision is actually made, so a change here reaches more of the right people than a change anywhere else.

Collection and category pages

Filter placement and defaults, sort order, card density, how many products load before pagination, and whether swatches appear on the grid.

Why it moves. Collections decide whether a visitor ever reaches a product page. Sort defaults in particular are one of the few one-line changes that affect the entire catalogue at once.

Cart and checkout

Drawer against a full cart page, shipping and returns content, upsell placement, and the wording of the final call to action. Deep checkout testing needs Shopify Plus — on standard plans the checkout itself is fixed, so we test everything up to it.

Why it moves. This is the last stretch before the money, and drop-off here is usually about doubt rather than lost interest. Answering the doubt in place beats sending people away to find the shipping policy.

Pricing and offer presentation

How bundles are framed, quantity breaks, how close the free-shipping threshold feels, subscription against one-off framing, and where the saving is shown.

Why it moves. You are not testing the price. You are testing how the offer reads — the same discount presented two ways is two different offers to a shopper, and only one of them is legible on a phone.

Landing pages for paid traffic

Message match against the ad that sent them, offer and form placement, where social proof sits, and page length.

Why it moves. Paid traffic arrives with volume you control and intent you already paid for, so these tests reach significance faster than anything else on the site — and the answer feeds straight back into media spend.

Navigation and search

Menu structure and labels, how many top-level items, mega-menu against simple dropdown, how prominent search is, and what the empty-results state does.

Why it moves. Navigation is the one component every visitor touches. Label changes are cheap to build and they move every page downstream of them, which is the whole funnel.

Process

How a test actually runs.

One test cycle takes two to six weeks end to end, depending on your traffic. More traffic shortens the run. It does not shorten the thinking.

01

Hypothesis

Written before anything gets built, in one sentence: the change, the expected direction, the reason we expect it, and the single metric that decides the outcome. If we cannot name the metric in advance, the test does not get built. Choosing the metric after seeing the data is how teams talk themselves into winners that were never there.

  • One-sentence hypothesis, dated
  • Primary metric named up front
  • Guardrail metrics agreed
  • Required sample size calculated
02

Build

The variant is built in native theme code — Liquid, CSS and a small amount of JavaScript — and served without a client-side rewrite. No visual editor overlay, no flash of the original before the variant paints. Both variants get the same QA as a production release: real devices, 375px first, tracking verified before a single visitor is bucketed.

  • Variant built in Liquid, not an overlay
  • No flicker, no render-blocking script
  • QA on real devices, mobile first
  • Event tracking verified pre-launch
03

Run

Fixed duration, agreed before launch — a minimum of one full week and usually two, so every day of the weekly cycle is represented in both variants. Nobody opens the dashboard hoping for a number. We check that data is arriving and that neither variant is broken. That is the whole of the mid-test job.

  • Run length locked before launch
  • No peeking, no early calls
  • Daily health check only
  • Traffic split and sample monitored
04

Decide

Ship, revert, or iterate. Winners are merged into the live theme as permanent code. Losers are reverted the same day and written up. Inconclusive results get labelled inconclusive — that is a real outcome, not a reporting failure, and pretending otherwise is how a backlog fills with fiction.

  • Result written up, whatever it says
  • Winner merged into the live theme
  • Loser reverted and documented
  • Next hypothesis queued from the finding

Failure modes

How tests go wrong.

Experiment programmes rarely fail on ideas. They fail on process. These six are the ones we see most, and every one of them produces a confident, wrong answer — which is more damaging than no answer at all.

Stopping the test early

Calling a winner the moment the dashboard turns green is the most expensive mistake in testing. Checking an in-flight test repeatedly and stopping at the first significant reading produces false positives roughly 30–50% of the time. The run length is set before launch and it is not negotiable afterwards.

Testing several variables at once

Change the hero image, the button copy and the badge in one variant and a win tells you nothing about which one did it. You cannot reproduce the effect and you cannot apply it anywhere else. One variable per test, unless you have the traffic for a proper multivariate design — and almost nobody does.

Judging a mobile store on desktop data

Most Shopify traffic is mobile, and the two devices frequently disagree. A variant that wins overall can be losing badly on the device carrying the majority of your sessions. Every result gets read by device before anything ships.

Running through an anomalous period

A sale, a press hit, a paid campaign switching on, a public holiday — each one changes who is on the site. The result then describes those visitors rather than your normal customer. We check the calendar before scheduling, and we discard runs that get contaminated mid-flight rather than quietly reporting them.

Not documenting the losers

Undocumented losing tests get retested. Someone new joins, has the same idea, and spends another three weeks disproving it again. A test that lost is a permanent fact about your customers, and it is only worth something if it is written down where the next person will look.

Flicker from client-side testing tools

Tools that rewrite the page in the browser show the original for a moment before the variant paints. Visitors see the flash, it changes their behaviour, and it biases the result. The same script is render-blocking, so it damages Core Web Vitals on every page it loads on — including pages you are not testing. This is why we build variants into the theme instead.

The deliverable

What you get after every test.

Every test ends in a written record, whatever the result. Dashboard access is not an answer — someone still has to read it and say what it means.

The hypothesis as written

The original sentence, unedited, with the date it was agreed — so the result is judged against what we actually predicted, not against a story assembled afterwards.

The result

Per-variant conversion rate, revenue per visitor and average order value, reported against the primary metric named at the start.

The confidence level

The actual confidence figure and the sample size reached — not a green tick. If the test did not reach significance, the report says so in the first line.

Segment breakdown

Mobile against desktop, new against returning, and by traffic source. An overall winner that loses on mobile is not a winner on a Shopify store.

The decision

Ship, revert or iterate, with the reasoning written by the person who ran the test rather than generated by the tool that ran it.

The code

Winners merged into the live theme in Liquid. Losers reverted the same day. Either way the diff sits in the repository and you own it.

Losing tests are worth as much as winners.

Most tests do not produce a winner — that is the nature of the method, not a sign the programme is failing. A loser removes an idea from the backlog permanently and tells you something true about your customers. The programmes that compound are the ones that keep the record.

Book a free technical call
function appendSignupLocation() {
  const SIGNUP_KEY = "signup";
  const uriParams = new URLSearchParams(
    window.location.search
  );
  if (signupCode) {
    uriParams.set(SIGNUP_KEY, signupCode);
  }
}
LANCER
Anti-AgingSkin Care
Start Your Test Drive
SHOPLIFT

Channels

RedBlueGreen

Collections

  • Collection list
  • Featured collection

Content

  • Accordion content
  • Blog posts

A/B Testing and Data-Driven Decisions

Optimize your Shopify store with data-driven decisions through A/B testing. Move beyond guesswork and drive growth by measuring the real impact of every change.

  • Strategize, Design, Develop, Deploy
  • Exclusively Plus Certified
  • In-Depth Analytics
Start the A/B Test Drive

Frequently asked questions

  • Yes—run theme tests for headers, drawers, app placements, or even entire themes.

Ready to run your first winning test?

Results vary by traffic and audience; all uplift figures are from real case studies and/or Shoplift reporting. Benchmarks are directional, not guarantees.

Start the A/B Test Drive