Shopify A/B testing that ships winners
Shopify A/B testing for $100K+/month stores ready to scale conversion experiments without shipping blind. We design hypotheses, run product page and checkout testing, then ship statistically significant winners — not dashboards.
Past 95% — ship it. Below that it stays running, or it gets reverted.
Hypothesis-led experiment design
Every test starts from a real problem and a falsifiable hypothesis — not a hunch or a trend.
Statistical significance reporting
We run tests to significance and report the truth — including the ones that lose — so you scale what actually works.
Ship-winners deployment workflow
Winners get shipped into native Liquid, not left living in a testing tool that slows your store down.
Read this before you brief a test
Most Shopify stores shouldn't be A/B testing yet.
A/B testing is a sample-size problem before it is a creativity problem. The maths does not care how good the idea is. If too few people see the two versions, the difference you measure is noise — and shipping on noise is worse than not testing at all, because you have now hard-coded a random number into your storefront and you believe it works.
The numbers below are not our results. They are what the statistics require before a difference can be told apart from chance. Read them as an entry price.
~3,800
Visitors per variant
What it takes to detect a 10% relative lift at 95% confidence and 80% power. A 10% lift is already a good outcome — this is the price of being able to see one.
~1,000
Visitors per variant
What a 20% relative lift needs at the same confidence and power. A swing that large is rare, and designing for it means testing big structural changes rather than button colours.
100 + 1–2 wks
Conversions per variant, minimum run
Sample size alone is not enough. You need at least 100 conversions in each variant and a run covering one to two full weekly cycles, because Tuesday shoppers and Saturday shoppers are not the same people.
In practice that puts the floor at roughly 10,000 monthly visitors before testing is worth starting, and closer to 50,000 a month before results are reliably trustworthy. Find your store below.
| Store sessions | What the maths allows | What we would tell you | Test? |
|---|---|---|---|
| Under 10,000 / month | Nothing you can trust. Even a 20% swing needs months to reach significance, by which time the season has changed underneath the test. | Do not test. Fix what is plainly broken and buy traffic. | Not yet |
| 10,000 – 30,000 / month | One test at a time, large structural changes only, with run times measured in weeks rather than days. | Test rarely. Put most of the effort into acquisition and merchandising. | Selectively |
| 30,000 – 50,000 / month | Genuine tests on your highest-traffic template, read at the overall level. Segment-level reads are still thin. | Start a light programme. One test per cycle, product page first. | Selectively |
| 50,000+ / month | A continuous programme, smaller detectable effects, and enough volume to read results by device and source. | Test properly. This is where experimentation starts to compound. | Yes |
This is the part of the pitch that costs us work.
We would rather say it now than four months into a retainer. If your store sits under the line, the highest-return work is not experimentation — it is fixing the things that are plainly broken. A slow product page. A cart that takes three taps to open. A mobile filter nobody can use. Then getting more people through the door. None of that needs a test to justify it, and a test at that traffic level cannot tell you whether it worked anyway. Come back to testing when the traffic can pay for it.
Test surface
What we test.
Six areas cover almost everything worth testing on a Shopify store. We start where the traffic and the money already are, which is nearly always the product page, and work outwards from there.
Product pages
Image order and gallery behaviour, review and trust badges, variant selectors, the position and wording of add to cart, and how far down the description sits.
Why it moves. The product page carries the most qualified traffic in the store and it is where the buying decision is actually made, so a change here reaches more of the right people than a change anywhere else.
Collection and category pages
Filter placement and defaults, sort order, card density, how many products load before pagination, and whether swatches appear on the grid.
Why it moves. Collections decide whether a visitor ever reaches a product page. Sort defaults in particular are one of the few one-line changes that affect the entire catalogue at once.
Cart and checkout
Drawer against a full cart page, shipping and returns content, upsell placement, and the wording of the final call to action. Deep checkout testing needs Shopify Plus — on standard plans the checkout itself is fixed, so we test everything up to it.
Why it moves. This is the last stretch before the money, and drop-off here is usually about doubt rather than lost interest. Answering the doubt in place beats sending people away to find the shipping policy.
Pricing and offer presentation
How bundles are framed, quantity breaks, how close the free-shipping threshold feels, subscription against one-off framing, and where the saving is shown.
Why it moves. You are not testing the price. You are testing how the offer reads — the same discount presented two ways is two different offers to a shopper, and only one of them is legible on a phone.
Landing pages for paid traffic
Message match against the ad that sent them, offer and form placement, where social proof sits, and page length.
Why it moves. Paid traffic arrives with volume you control and intent you already paid for, so these tests reach significance faster than anything else on the site — and the answer feeds straight back into media spend.
Navigation and search
Menu structure and labels, how many top-level items, mega-menu against simple dropdown, how prominent search is, and what the empty-results state does.
Why it moves. Navigation is the one component every visitor touches. Label changes are cheap to build and they move every page downstream of them, which is the whole funnel.
Process
How a test actually runs.
One test cycle takes two to six weeks end to end, depending on your traffic. More traffic shortens the run. It does not shorten the thinking.
Hypothesis
Written before anything gets built, in one sentence: the change, the expected direction, the reason we expect it, and the single metric that decides the outcome. If we cannot name the metric in advance, the test does not get built. Choosing the metric after seeing the data is how teams talk themselves into winners that were never there.
- One-sentence hypothesis, dated
- Primary metric named up front
- Guardrail metrics agreed
- Required sample size calculated
Build
The variant is built in native theme code — Liquid, CSS and a small amount of JavaScript — and served without a client-side rewrite. No visual editor overlay, no flash of the original before the variant paints. Both variants get the same QA as a production release: real devices, 375px first, tracking verified before a single visitor is bucketed.
- Variant built in Liquid, not an overlay
- No flicker, no render-blocking script
- QA on real devices, mobile first
- Event tracking verified pre-launch
Run
Fixed duration, agreed before launch — a minimum of one full week and usually two, so every day of the weekly cycle is represented in both variants. Nobody opens the dashboard hoping for a number. We check that data is arriving and that neither variant is broken. That is the whole of the mid-test job.
- Run length locked before launch
- No peeking, no early calls
- Daily health check only
- Traffic split and sample monitored
Decide
Ship, revert, or iterate. Winners are merged into the live theme as permanent code. Losers are reverted the same day and written up. Inconclusive results get labelled inconclusive — that is a real outcome, not a reporting failure, and pretending otherwise is how a backlog fills with fiction.
- Result written up, whatever it says
- Winner merged into the live theme
- Loser reverted and documented
- Next hypothesis queued from the finding
Failure modes
How tests go wrong.
Experiment programmes rarely fail on ideas. They fail on process. These six are the ones we see most, and every one of them produces a confident, wrong answer — which is more damaging than no answer at all.
Stopping the test early
Calling a winner the moment the dashboard turns green is the most expensive mistake in testing. Checking an in-flight test repeatedly and stopping at the first significant reading produces false positives roughly 30–50% of the time. The run length is set before launch and it is not negotiable afterwards.
Testing several variables at once
Change the hero image, the button copy and the badge in one variant and a win tells you nothing about which one did it. You cannot reproduce the effect and you cannot apply it anywhere else. One variable per test, unless you have the traffic for a proper multivariate design — and almost nobody does.
Judging a mobile store on desktop data
Most Shopify traffic is mobile, and the two devices frequently disagree. A variant that wins overall can be losing badly on the device carrying the majority of your sessions. Every result gets read by device before anything ships.
Running through an anomalous period
A sale, a press hit, a paid campaign switching on, a public holiday — each one changes who is on the site. The result then describes those visitors rather than your normal customer. We check the calendar before scheduling, and we discard runs that get contaminated mid-flight rather than quietly reporting them.
Not documenting the losers
Undocumented losing tests get retested. Someone new joins, has the same idea, and spends another three weeks disproving it again. A test that lost is a permanent fact about your customers, and it is only worth something if it is written down where the next person will look.
Flicker from client-side testing tools
Tools that rewrite the page in the browser show the original for a moment before the variant paints. Visitors see the flash, it changes their behaviour, and it biases the result. The same script is render-blocking, so it damages Core Web Vitals on every page it loads on — including pages you are not testing. This is why we build variants into the theme instead.
The deliverable
What you get after every test.
Every test ends in a written record, whatever the result. Dashboard access is not an answer — someone still has to read it and say what it means.
The hypothesis as written
The original sentence, unedited, with the date it was agreed — so the result is judged against what we actually predicted, not against a story assembled afterwards.
The result
Per-variant conversion rate, revenue per visitor and average order value, reported against the primary metric named at the start.
The confidence level
The actual confidence figure and the sample size reached — not a green tick. If the test did not reach significance, the report says so in the first line.
Segment breakdown
Mobile against desktop, new against returning, and by traffic source. An overall winner that loses on mobile is not a winner on a Shopify store.
The decision
Ship, revert or iterate, with the reasoning written by the person who ran the test rather than generated by the tool that ran it.
The code
Winners merged into the live theme in Liquid. Losers reverted the same day. Either way the diff sits in the repository and you own it.
Losing tests are worth as much as winners.
Most tests do not produce a winner — that is the nature of the method, not a sign the programme is failing. A loser removes an idea from the backlog permanently and tells you something true about your customers. The programmes that compound are the ones that keep the record.
function appendSignupLocation() {
const SIGNUP_KEY = "signup";
const uriParams = new URLSearchParams(
window.location.search
);
if (signupCode) {
uriParams.set(SIGNUP_KEY, signupCode);
}
}Channels
Collections
- Collection list
- Featured collection
Content
- Accordion content
- Blog posts
A/B Testing and Data-Driven Decisions
Optimize your Shopify store with data-driven decisions through A/B testing. Move beyond guesswork and drive growth by measuring the real impact of every change.
- Strategize, Design, Develop, Deploy
- Exclusively Plus Certified
- In-Depth Analytics
Explore next
Other things we do
Most stores need two or three of these working together. Book a call and we'll tell you which ones actually move your numbers.
- Shopify Theme DevelopmentCustom themes built from Figma to production Liquid on Online Store 2.0 — fast and merchant-editable.
- Shopify App DevelopmentCustom and public apps, Checkout UI Extensions and Shopify Functions built with React and Polaris.
- Shopify MigrationMove from WooCommerce, Magento or BigCommerce with a full 301 redirect map and zero downtime.
- Shopify SEOTechnical audits, Core Web Vitals, structured data and collection content built around buyer intent.
- Shopify Performance OptimizationFaster load times and green Core Web Vitals — theme refactors, image pipelines and app cleanup.
- Shopify Email Marketing & KlaviyoFlows, segmentation and campaigns — welcome, abandoned cart, winback and SMS, built and monitored.
- Shopify CRO & FunnelsValue-ladder, tripwire, quiz and post-purchase upsell funnels built to lift conversion and AOV.
- Shopify Maintenance & SupportRetainers covering uptime monitoring, security patching, bug fixes and proactive improvements.
Frequently asked questions
- Yes—run theme tests for headers, drawers, app placements, or even entire themes.
Ready to run your first winning test?
Results vary by traffic and audience; all uplift figures are from real case studies and/or Shoplift reporting. Benchmarks are directional, not guarantees.
Start the A/B Test Drive