A/B testing ads and landing pages
Every technique across this series is a hypothesis about what will improve performance — a new headline, a different landing page structure, a bidding strategy change. A/B testing is how that hypothesis actually gets confirmed with real data, rather than assumed.
Google Ads Experiments: a real, controlled split test
Original campaign: "Search — Coffee Subscriptions"
Experiment: 50% of traffic → original landing page
50% of traffic → new landing page (redesigned CTA placement)
Duration: 3 weeksGoogle Ads' built-in Experiments feature splits real, live traffic between an original campaign and a modified copy — critically, running both variants simultaneously rather than testing one, then the other, sequentially. This controls for time-based factors (a day-of-week effect, a seasonal shift) that would otherwise contaminate a before-and-after comparison run at different times.
What's actually worth testing, roughly in order of impact
1. Landing page structure and CTA — often the single biggest lever (part 8)
2. Headlines and ad copy angle — meaningfully impactful, faster to iterate on
3. Bidding strategy — impactful, but needs a longer test window to read reliably
4. Extension variations — real, but typically the smallest individual impactThis ordering matters for where limited testing effort should actually go first — landing page changes routinely move conversion rate by a larger margin than a headline wording tweak, which is a real, if less exciting, insight: the biggest wins are rarely in the ad copy alone.
Statistical significance: not calling a result too early
Day 3: Variant B converting 40% better than originalA result like this, three days into a three-week test, is very likely noise, not a real signal — small early sample sizes produce large, misleading swings that regress toward a more modest, real difference (or disappear entirely) as more data accumulates. Google Ads Experiments reports a statistical significance indicator directly in the results — waiting for that, and for a reasonable minimum sample size (enough conversions on both sides to be meaningful, not just enough clicks), is what separates a real result from an artifact of small numbers.
A real example: testing Bright Leaf Coffee's CTA button copy
Original: "Start Your Subscription"
Variant: "Get My First Bag"After a full three-week run with sufficient volume on both sides: Variant converted at 4.1% versus the original's 3.6% — a real, statistically significant difference, not noise, confirmed by the experiment reaching Google's significance threshold rather than being called early on a promising first-week trend. This becomes the new default CTA copy going forward, and the next test builds from this new baseline rather than starting over from the original.
One variable at a time, when possible
Testing a landing page with a new headline, new CTA copy, and a new layout simultaneously against the original makes it impossible to know which specific change actually drove any observed difference — a genuinely useful test isolates one meaningful variable per experiment where practical, even though a full page redesign sometimes reasonably bundles related changes together as one cohesive variant, accepting some loss of individual attribution in exchange for testing a real, complete alternative.
Ending a test the moment one variant pulls ahead, without waiting for statistical significance or a reasonable minimum sample size. Early leads in a real experiment reverse more often than intuition expects — calling a winner too early is one of the most common ways a genuinely neutral or even losing change gets permanently adopted based on what was actually random noise.
Next: the mistakes that actually waste the most Google Ads budget in practice — most of them honest, not manipulative, and mostly things this series has already touched on individually.