A/B Testing in Marketing

Explore top LinkedIn content from expert professionals.

  • View profile for Arindam Paul
    Arindam Paul Arindam Paul is an Influencer

    Building Atomberg, Author-Zero to Scale

    159,507 followers

    Attribution is overrated. Incrementality is what actually matters Every new-age brand wants to know what’s working. Meta ROAS is looking good. CAC is steady. Revenue is growing But here’s the truth: Your Meta ad might get the conversion. But did it cause the conversion? That’s the difference between attribution and incrementality. Most dashboards, attribution tools, and agency reports stop at attribution. But if you’re a brand selling across Amazon, Flipkart, GT, MT, Q-com, and D2C—pure attribution will always lie to you Because the sale might happen on Amazon. But it might have been nudged by a Meta video or a YouTube bumper ad 4 days ago. You don’t need a full-blown Marketing Mix Model to get started. There are simpler, street-smart ways to directionally understand what’s working—and what’s not. Here are 4 that have worked for us at Atomberg: 1. Geo Split Testing Pick two similar markets. Run campaigns in one. Don’t run in the other. Then track: • Branded search volume • Sell-through on marketplaces • Secondary sales from GT counters If the test market moves faster than the control, you’re seeing true lift. That’s incrementality. 2. First-Time Buyer Growth vs Returning Buyer Growth Track whether your growth is coming from first-time buyers or repeats. If your campaigns are just bringing back old customers—you’re not creating net new demand. But if there’s a spike in new buyers across Amazon, Flipkart, D2C—your campaigns are likely working at an incremental level 3. Paid Traffic vs Organic Trend Lines If paid traffic, clicks and spends are going up—but your organic sales or branded search isn’t moving—you’re likely just harvesting demand that already existed. But if organic lifts alongside paid—your ads are creating interest. Not just closing it. Directionally, this is one of the simplest sanity checks most teams ignore. 4. Channel Crossover + Offline Signal Mapping Your Meta ad may not show up in last-click attribution. But it might have nudged the consumer to visit your store or buy on Amazon. You can detect this through: • Post-purchase surveys (Where did you first hear about us?) • Branded search + store footfall spikes in campaign-active cities • And most powerfully—offline signals passed back to Meta At Atomberg, we pass back data from installations and warranty registrations—including pincode and purchase timelines Sometimes, we’re even able to identify this at a unique customer level through their cookies for warranty registration This has helped us understand true incrementality of perf marketing campaigns even for offline sales If you’re only measuring ROAS, you might scale what’s only taking credit for sale about to happen anyway If you chase incrementality, you’ll scale what’s working. For more details, read the full post- link in first comment.

  • View profile for Malay Krishna

    Director of PM @ Vyapar | PM Coach - Helping you break into AI Product Management | 1:1 mentoring + portfolio-building products

    65,360 followers

    Why Every Product Manager Needs A/B Testing 🚀 Imagine cooking up a recipe for the perfect product feature. Would you trust your instincts blindly, or would you test different ingredients to get the best taste? That’s where A/B testing comes in. It’s the secret sauce that helps Product Managers make data-driven decisions with confidence. Here’s everything you need to know to master A/B testing: ❓ What is A/B Testing❓ A/B testing is the process of comparing two or more versions of a product to determine which one performs better. The versions might differ in small ways - a new button design, a revamped landing page, or an updated pricing structure but the impact on user behaviour can be monumental. This method helps you validate assumptions, optimize user experiences, and ensure every product decision adds value. ⚙️ How to Conduct a Successful A/B Test? ⚙️ 🔹 Set Clear Goals Ask yourself what are you trying to improve? It could be anything from conversion rates to user satisfaction. Your goal is your North Star. 🔹 Choose the Right Metrics Metrics like click-through rates (CTR), time spent on a page, or purchase frequency will guide you in evaluating success. 🔹 Hypothesize Frame your test with a simple prediction. Example: “I believe changing the CTA button color from blue to green will increase clicks by 15%.” 🔹 Design Your Experiment Define your control group (current version) and treatment group (variant to test), ensuring a large enough sample size for reliable results. Run the test for a sufficient duration to capture meaningful patterns and user behaviour. 🔹 Analyze & Implement Use tools like Google Optimize or Optimizely to analyze results and determine statistical significance. Roll out the winning variant confidently, or refine your hypothesis for future iterations if results are inconclusive. ♻️ Four Types of A/B Tests Every PM Should Know ♻️ 1️⃣ Feature Testing: Validate hypotheses for new features pre-launch. 2️⃣ Live Testing: Fine-tune existing features already in the wild. 3️⃣ Trapdoor Testing: Redirect traffic between variants dynamically. 4️⃣ Multi-Armed Bandit (MAB): Let machine learning allocate traffic to better-performing variants in real-time. ❌ Common Pitfalls to Avoid ❌ 1️⃣ Testing trivial changes that won’t move the needle. 2️⃣ Ignoring sample size requirements—small audiences lead to inaccurate conclusions. 3️⃣ Treating A/B testing as a one-off exercise. Optimization is an ongoing journey. What’s been your most surprising A/B testing discovery? Let’s discuss in the comments!👇 Ready to embark on an exhilarating journey into the heart of product management? I’ve recently launched a cohort that is focused on teaching end-to-end product management as well as providing career placement opportunities! 🧠 Fill in the form in the comments to register your interest in the cohort and I’ll reach out to you with further details. ✍️ #ProductManagement #ABTesting #PMTools #ContinuousOptimization

  • View profile for Rishabh Jain
    Rishabh Jain Rishabh Jain is an Influencer

    Co-Founder / CEO at FERMÀT - the leading commerce experience platform

    16,238 followers

    I've been vocal about this and I'll double down: 2025 is the year we finally have a consensus that attribution, particularly MTA, is not a healthy way to grow your business. The shift I'm seeing: ➝ Moving away from complex MTA models ➝ Embracing simple, real-world experiments ➝ Prioritizing measurable business impact over attribution modeling Why this matters: A basic geo-holdout test in Ohio outshines a complex MTA model. It provides clear, actionable insights about real impact—not perfection, but reality. That’s what teams increasingly value. The future isn't about attribution models. It's about running straightforward experiments that tell you if your marketing actually works. Whether you use sophisticated tools or basic holdout tests, measuring real impact beats attributing theoretical credit. Looking forward to seeing this transformation reshape how we all measure marketing effectiveness.

  • View profile for Peter Buckley

    Connection Planning Director, Meta

    17,533 followers

    Your most effective channel is losing you sales. You can often make campaigns more effective by moving money to less effective channels. What? Marketing Science maestro Simon Toms explains how: In the example image, the blue line represents a channel that’s 2x more effective than the pink one at every spend level. $1M invested in Channel 1 returns $2M in incremental revenue (A). But split the $1M between Channel 1 and 2 (50:50) and you’d drive $2.5M total incremental revenue (B + C). That’s 25% more revenue from investing in a “less effective” channel. So what? Don't accept average metrics alone, always look to understand the marginal returns. Ideally you should know the curves for all your investments. MMM can obviously help with this, but incrementality testing typically provides more detailed curves based on actual sales rather than modelled ones. Incrementality testing is not A/B testing. It's test and control - the test group see the ad, the control group (who match the ad audience but are withheld from the ads) don't. The difference is the incremental impact. (In an A/B test you do not withhold a segment of your audience from seeing the ad, so it can't measure incremental impact.) Here's where curves from incrementality testing can help: 1. Optimal Full Funnel  Different optimisations have very different curves. The curve for reach spend is very different to conversion spend which can be very different to ASC activity etc. Plotting curves helps you understand where you should pull back investment and where you should double down, critical insights for maximizing incremental returns. 2. Channel synergy The curve for one channel changes depending on your investment in others. Charlie Oscar found that social reach improves paid search performance by 32%, YouTube improves email by up to 25%, most crazy of all, 70% of the value from social and video channels is their impact on other channels with only 30% direct. 3. Plan at the margins  Don't use average ROIs to determine where to shift your budget. It depends on the curve, not the average. Incremental returns show which channels to invest in, marginal returns show how much. Your most effective channel isn't often where you should put your next $. Bottom line:   To make your campaigns work harder, you need to understand how each investment works at the margins. That's the route to higher returns across the mix.

  • View profile for Nick Babich

    Product Design | User Experience Design

    89,418 followers

    💡A/B Testing: 8 Essential Tips A/B testing is a powerful method for comparing two versions of a design against each other to determine which one performs better. Here are the top 8 tips for conducting effective A/B tests: 1️⃣ Define clear goals: Know what you want to achieve with your test. Whether it's increasing conversions, click-through rates, or user engagement, having clear goals is crucial. 2️⃣ Test one variable at a time: To understand the effect of a change, test only one variable at a time (i.e., color of a primary call to action button). Multiple changes can confound results. 3️⃣ Randomize your sample: Ensure your sample is randomly selected to avoid biases and ensure the test results are reliable. 4️⃣ Ensure sufficient sample size: Make sure your test runs long enough to gather a statistically significant sample size to make confident decisions. Use sample size calculator: https://lnkd.in/dCXpgv2Z 5️⃣ Segment your audience: Consider segmenting your audience to understand how different groups respond to the change. 6️⃣ Monitor metrics beyond primary goal: Track secondary metrics to ensure that the changes do not negatively impact other important aspects of user experience (i.e., you have a higher conversion rate but a lower user retention rate). 7️⃣ Check for statistical significance—you need to ensure that the data you collect cannot be attributed to pure chance. Use the calculator to check significance: https://lnkd.in/d5jcWa7N 8️⃣ Consider long-term effects: Assess whether the changes have a lasting positive impact or if they might lead to long-term negative consequences (this can happen if you use dark patterns: https://lnkd.in/dtztGgFW) 📕 Introduction to A/B testing for product designers (YouTube): https://lnkd.in/dxuW8-hq #testing #design #research #productdesign #design #abtesting

  • View profile for Jeffrey Bustos

    SVP Retail Media Analytics - Measurement Data AI - 🇨🇴

    26,896 followers

    How do you measure incrementality? 🧩 Every brand asks: did my media actually drive sales beyond what would have happened anyway? Here are three of the most effective approaches in play today: 📊 Randomized Controlled Trials (RCTs / A/B Tests) The cleanest measure of causal lift by splitting audiences into exposed vs. control. Still the gold standard, but scaling across hundreds of supplier campaigns is tough. 🗺️ Matched Market / Geo Experiments Turn media on in one region or store set, and off in another. Highly effective for brick-and-mortar, but sensitive to local dynamics like competitor promos or seasonality. 🧪 Synthetic Controls When true RCTs are not possible, synthetic controls create a statistical twin of the exposed group by weighting historical sku, sales, loyalty, and category data. This lets retailers simulate what performance would have been without media. 📈 Effective for enterprise networks running many campaigns at once 👥 Useful for comparing audience tiers such as new-to-brand, loyal, or lapsed 📊 More reliable than simple pre/post because it adjusts for seasonality and baseline shifts ⚖️ The right approach depends on balancing rigor, cost, and scalability. Without incrementality, retail media risks losing credibility and budget share. https://lnkd.in/egquzpJ8

  • View profile for Peter Quadrel

    Founder of Odylic Media | Profitable New Customer Growth for Premium & Luxury DTC Brands

    39,409 followers

    Meta, Google, TikTok, and other ad channels are misleading you. Third-party attribution tools like Triple Whale and North Beam aren't better—they’re flawed too. Tracking has always relied on estimated models, not hard numbers. After iOS 14, tracking became harder, leading to a surge in third-party solutions. But these also provide conflicting data, making it tough to find the truth. So, what is the truth? The only reliable way to measure your marketing efforts is through incrementality tests. These tests answer the question, "What if this channel or ad never existed?" By showing ads to one group and withholding from another, you can measure the true impact on revenue and profit. For example, if you're running Facebook ads and selling on Shopify and Amazon, incrementality tests reveal how Facebook ads impact Amazon sales. Without the initial Facebook touchpoint, an Amazon purchase might not have happened, even though traditional attribution wouldn’t show this. This is why ROAS and third-party attribution aren’t accurate. They use models that can be thwarted by privacy settings and cross-channel purchases. By running incrementality tests, you discover the true impact of your marketing efforts. We ran a 14-day Meta holdout test and found that zip codes shown ads generated 50% more Amazon revenue than those not shown ads, despite sending traffic to Shopify. Now is the perfect time to run these tests. Q3 is calm, free from major holidays that skew results. This is your chance to optimize before Q4. If your brand generates seven figures annually, this should be a top priority to grow profits in Q4.

  • View profile for Bahareh Jozranjbar, PhD

    UX Researcher at PUX Lab | Human-AI Interaction Researcher at UALR

    10,747 followers

    As UX researchers, we often encounter a common challenge: deciding whether one design truly outperforms another. Maybe one version of an interface feels faster or looks cleaner. But how do we know if those differences are meaningful - or just the result of chance? To answer that, we turn to statistical comparisons. When comparing numeric metrics like task time or SUS scores, one of the first decisions is whether you’re working with the same users across both designs or two separate groups. If it's the same users, a paired t-test helps isolate the design effect by removing between-subject variability. For independent groups, a two-sample t-test is appropriate, though it requires more participants to detect small effects due to added variability. Binary outcomes like task success or conversion are another common case. If different users are tested on each version, a two-proportion z-test is suitable. But when the same users attempt tasks under both designs, McNemar’s test allows you to evaluate whether the observed success rates differ in a meaningful way. Task time data in UX is often skewed, which violates assumptions of normality. A good workaround is to log-transform the data before calculating confidence intervals, and then back-transform the results to interpret them on the original scale. It gives you a more reliable estimate of the typical time range without being overly influenced by outliers. Statistical significance is only part of the story. Once you establish that a difference is real, the next question is: how big is the difference? For continuous metrics, Cohen’s d is the most common effect size measure, helping you interpret results beyond p-values. For binary data, metrics like risk difference, risk ratio, and odds ratio offer insight into how much more likely users are to succeed or convert with one design over another. Before interpreting any test results, it’s also important to check a few assumptions: are your groups independent, are the data roughly normal (or corrected for skew), and are variances reasonably equal across groups? Fortunately, most statistical tests are fairly robust, especially when sample sizes are balanced. If you're working in R, I’ve included code in the carousel. This walkthrough follows the frequentist approach to comparing designs. I’ll also be sharing a follow-up soon on how to tackle the same questions using Bayesian methods.

  • View profile for Mohsen Rafiei, Ph.D.

    Cognitive Psychologist

    12,195 followers

    Recently, someone shared results from a UX test they were proud of. A new onboarding flow had reduced task time, based on a very small handful of users per variant. The result wasn’t statistically significant, but they were already drafting rollout plans and asked what I thought of their “victory.” I wasn’t sure whether to critique the method or send flowers for the funeral of statistical rigor. Here’s the issue. With such a small sample, the numbers are swimming in noise. A couple of fast users, one slow device, someone who clicked through by accident... any of these can distort the outcome. Sampling variability means each group tells a slightly different story. That’s normal. But basing decisions on a single, underpowered test skips an important step: asking whether the effect is strong enough to trust. This is where statistical significance comes in. It helps you judge whether a difference is likely to reflect something real or whether it could have happened by chance. But even before that, there’s a more basic question to ask: does the difference matter? This is the role of Minimum Detectable Effect, or MDE. MDE is the smallest change you would consider meaningful, something worth acting on. It draws the line between what is interesting and what is useful. If a design change reduces task time by half a second but has no impact on satisfaction or behavior, then it does not meet that bar. If it noticeably improves user experience or moves key metrics, it might. Defining your MDE before running the test ensures that your study is built to detect changes that actually matter. MDE also helps you plan your sample size. Small effects require more data. If you skip this step, you risk running a study that cannot answer the question you care about, no matter how clean the execution looks. If you are running UX tests, begin with clarity. Define what kind of difference would justify action. Set your MDE. Plan your sample size accordingly. When the test is done, report the effect size, the uncertainty, and whether the result is both statistically and practically meaningful. And if it is not, accept that. Call it a maybe, not a win. Then refine your approach and try again with sharper focus.

  • View profile for Thomas B. Leonard

    Fractional Marketing Leader | Operationalizing MMM + Incrementality Testing

    2,941 followers

    Attribution is precise. The problem is, when budgets get cut, leaning on attribution leads to precisely the wrong decisions. Here's a recent client example: $75M/yr company. Recently bought back from PE. Renewed focus on being a profitable, family-owned company for the long run. After budget cuts last Q4, they saw an impact on revenue and new customer acquisition despite pulling from areas deemed unprofitable. I was brought in to assess things. I saw a common pattern: Based on attribution, they cut "poor performers" and kept what looked good, using CPA and ROAS as KPIs. They reduced NonBrand Search spend while maintaining PMax and Brand Search. To their credit, attributed metrics had improved, but that's not what mattered. The business was soft. Quick and dirty analysis revealed a correlation between NonBrand investment and new customers. Hypothesis: Branded Search was not incremental. Which, if true, we could use cost savings to invest in NonBrand. At this scale, simple incrementality testing is highly effective, especially for something like Brand Search. We split the US into two groups (image 1) and turned off Brand Search in half the markets while controlling how Brand appeared in PMax nationally. Doesn't need to be overly scientific at this budget level and for Brand Search. After 4 weeks, the data was clear: no impact on orders (image 2), nearly all the traffic flowed to Organic search (image 3). While these results justified cutting Brand Search, we examined the team's original rationale: competitive defense. They were using "max conversions" bidding, willing to pay whatever it took to win impressions. Rather than eliminate Brand Search entirely, we switched to manual bidding with low bids and budget, willing to give up some attributed conversions since they were not incremental. Results: 95% reduction in cost, 10% fewer clicks, and 25% fewer attributed conversions. Maintaining a minimal Brand Search presence made sense for SERP messaging control and competitor monitoring. This freed significant budget (20% of their budget) to test NonBrand in a handful of markets representative of the full US, which we'll do next. If we can drive profitable orders, we can confidently secure budget to scale nationally. Budget cuts happen, especially in uncertain economic environments. Attribution can lead to precisely the wrong decisions and is where incrementality shines. Incrementality isn't just about finding where to cut, but when budgets are finite, it's about freeing up capital to deploy elsewhere to drive growth.

Explore categories