CRO as a Pipeline: From Analytics to Revenue
In March 2025 the owner of an outdoor-gear shop in Porto called me about his checkout. His team sold tents, backpacks, and sleeping bags through a Shopify store plus Next.js landing pages. Traffic held steady near 9,000 visits per month. Checkout completion sat at 1.9 percent. Over the previous four months he had run eleven tests on the buy button: green, orange, larger padding, bolder label, sticky on mobile. Total revenue impact: zero. He asked me to run test twelve. I asked him to pause new tests for a month and open the event stream with me. Eight weeks later checkout completion reached 3.1 percent on the same traffic, worth about €14,000 in added monthly revenue at his €85 average order. Four stages produced that result. I run them in order on every revenue job: analytics, hypothesis, ship, measure. Each stage catches a failure the others miss.
The button-color trap
Button tests fail for arithmetic reasons before they fail for creative ones. At 9,000 visits per month and a 1.9 percent base rate, the shop saw about 171 orders per month. A 10 percent relative lift means 19 extra orders. Detecting that lift with 95 percent confidence needs about 30,000 visitors per variant. At his traffic that means seven months per test. He ran each test for ten days, read noise, shipped the winner, and watched revenue stay flat. I see this shape everywhere: small traffic, short reads, confident announcements, flat revenue.
Buttons sit at the end of a long chain. His funnel showed 62 percent of product-page visitors never started checkout. Of those who started, 41 percent left at the shipping step. A greener button touches neither group. It speaks to the 3 percent of visitors who already decided to buy and hesitate over color. That group barely exists. Tests at the narrow end of the funnel move the smallest audience. The pipeline below moves the test upstream, where the audience lives.
Stage 1 analytics that earns its keep
I closed the testing tool and opened PostHog with him. We mapped five events with strict names: view_item, add_to_cart, start_checkout, enter_shipping, purchase. Every event carried session_id, device class, and traffic source. I added the missing ones through the Shopify pixel plus a small Next.js route that forwarded server-side purchases with the same session_id. One week of clean data replaced four months of guesses.
Three leaks surfaced. First, the size-guide page swallowed 54 percent of its visitors. Session replay showed shoppers opening the guide, scrolling a dense table, and leaving. Second, shipping cost appeared late. International buyers reached step three, saw a €18 surcharge, and 41 percent left within a minute. Third, the mobile form asked for nine fields including company name and fax. Mobile completion ran at half the desktop rate.
I price this stage at three working days: one day to name events and add tracking, one day to watch forty replays and read the funnel by device and source, one day to write three numbered findings with screenshots. Analytics earns its keep when it names the exact screen where money stops.
Stage 2 hypotheses with a kill rule
Each leak became one written hypothesis with a mechanism, a metric, and a kill rule. Hypothesis one: shoppers who see a size finder inside the product page add to cart more often. Metric: add_to_cart rate on size-guide traffic. Kill rule: stop the test when 1,000 size-guide sessions pass with no lift, or when support contacts about sizing rise. Hypothesis two: buyers who see shipping cost on the product page start checkout more often. Metric: start_checkout rate for international traffic. Kill rule: stop when revenue per visitor drops more than 5 percent after 500 international sessions. Hypothesis three: mobile buyers who meet four fields finish payment more often. Metric: purchase rate on mobile. Kill rule: stop when payment errors rise above the baseline of 1.2 percent.
I rank the three by reach: shipping cost touched 38 percent of traffic, mobile fields touched 61 percent of checkouts, size finder touched 9 percent. Shipping went first.
Every hypothesis carries one sentence that kills it. I write the kill rule before I open the editor. A test with no kill rule runs until someone forgets it, and forgotten tests rot into permanent half-features. The rule protects the calendar and the codebase.
Stage 3 shipping without breaking trust
I ship every variant behind a flag in LaunchDarkly keyed by a hash of the customer id. The flag serves control or variant, fifty-fifty. I roll each change to 5 percent of traffic for one full day and watch three guardrails: checkout error rate in Sentry, failed payments in Stripe webhooks, and new support threads tagged checkout. Only then do I open the test to full traffic.
The size-finder build shows why. Version one overlaid the Apple Pay button on Safari and blocked taps. Sentry caught a spike in click errors within an hour. The flag killed the variant for Safari users in twenty minutes, and only 3 percent of sessions ever saw the bug. Without the flag that bug would have run all weekend.
I keep variants small enough to delete. Each variant lives in one component file with the flag name at the top. When a test ends, the losing branch leaves the codebase the same day. The store owner trusted the process because broken builds never reached his full audience. Trust comes from the flag, the guardrails, and the same-day cleanup.
Stage 4 reading results honestly
I read results on a fixed calendar: 21 days including three weekends, no early calls. Weekday buyers differ from weekend buyers, and his weekend share ran 44 percent. Peeking on day five and shipping the leader would have crowned noise twice during this job.
Primary metric stays revenue per visitor. Secondary metrics track checkout completion and average order value. Guardrails track refund rate and support contacts per hundred orders. A variant that lifts conversion and lifts refunds at the same time loses. I write that rule down before launch so the win does not seduce me on reading day.
Shipping-cost transparency won. Start_checkout rate for international traffic rose from 22 percent to 34 percent. Checkout completion moved from 1.9 percent to 3.1 percent site-wide. Revenue per visitor rose 41 percent. Refund rate held at 2.1 percent. Support contacts about shipping fell from eleven per week to three. The mobile form with four fields won next, adding another 0.4 points. The size finder drew even after 1,000 sessions, and the kill rule removed it on schedule.
I log every decision in one table: date, hypothesis, reading, action. Four rows cover this job. Each stage in the pipeline catches its own failure: analytics catches guessing, the hypothesis catches endless tests, shipping catches broken builds, honest reading catches vanity wins. Open your event stream this week and list the three exits that cost you most.