CRO and experimentation with Jurni: a practical guide

Last updated: September 29, 2026

Learn how to identify conversion opportunities, choose meaningful tests, build better landing experiences, and turn experiment results into confident business decisions.

Conversion rate optimization (CRO) is the ongoing process of helping more of the right visitors take a valuable next step. In ecommerce, that usually means purchasing, but it can also mean finding the right product, starting a subscription, or becoming a qualified lead.

Experimentation makes this process measurable. Instead of assuming a new page is better, you compare it with an existing experience under similar conditions and measure what changes.

Jurni brings page creation, Smart Links, and experimentation into the same workflow. Use it to turn a customer insight into a landing experience, compare that experience with your current destination, and build on what you learn.

The goal is sustainable business improvement. A page that generates more clicks but fewer purchases is not necessarily better. A discount that increases orders but reduces contribution profit may not be a win. Every test should connect a customer problem to a measurable business outcome.

In this guide

  1. Understand the foundations of experimentation

  2. Find your biggest conversion opportunities

  3. Choose the right success metric

  4. Prioritize and scope your tests

  5. Match page concepts to shopper intent

  6. Build a practical testing backlog

  7. Set up a reliable experiment in Jurni

  8. Plan traffic, sample size, and test duration

  9. Read results and decide what to do

  10. Troubleshoot misleading results

  11. Roll out winners and build a testing habit

  12. Use the templates, examples, and checklists

1. Understand the foundations of experimentation

What an A/B test tells you

An A/B test compares a control with a challenger using traffic assigned to each experience during the same period.

  • Control: The baseline you want to improve, such as your current product page.

  • Variant: A different experience designed to address a specific opportunity.

  • Hypothesis: Your explanation of why the change should improve performance.

  • Primary metric: The outcome used to judge the test.

  • Guardrails: Other outcomes that must stay within acceptable limits.

For example, you might compare your current product page with a Jurni page that answers the main objection raised in customer reviews. Both receive a share of the same eligible traffic. You then compare purchase conversion and revenue, rather than judging the pages by appearance.

The result applies to the experience, audience, and conditions you tested. A winner for cold mobile traffic may not be the best destination for returning desktop customers.

Why before-and-after comparisons are weaker

If you launch a new page this week and compare it with last week's results, the page is only one of many things that changed. Ad spend, creative, discounts, stock availability, customer mix, and day of the week can all affect performance.

A simultaneous randomized comparison helps separate the effect of the experience from these shared influences. Before-and-after reports remain useful for monitoring the business, but they do not establish that the page caused the change.

Page tests and ad tests answer different questions

A page experiment asks: Given comparable incoming traffic, which destination performs better?

An ad experiment asks: Which creative, audience, or delivery strategy generates better outcomes?

If you send one ad to Page A and a different ad to Page B, the results reflect both acquisition and page differences. For a cleaner destination comparison, use the same eligible traffic stream and split it after the click through the experiment's Smart Link.

This does not freeze ad-platform behavior or guarantee identical visitors. It creates a more defensible comparison than assigning each destination a different acquisition strategy.

Concept tests versus focused tests

A concept test compares substantially different experiences: a product page versus an educational advertorial, or a general landing page versus a creator-led story. It can reveal a useful strategic direction quickly, but it does not isolate which individual element caused the result.

A focused test changes one meaningful factor: moving shipping information beside the CTA, replacing generic reviews with objection-specific reviews, or simplifying the bundle selector. It provides a narrower learning.

Use concept tests when the current journey may be fundamentally wrong for the audience. Use focused tests to improve a promising experience or investigate why it works. You do not need to limit every experiment to a single word or button color; you do need to know what question the comparison can answer.

Personalization needs its own experiment

Routing one audience to one page and another audience to another page is personalization, but it is not proof that personalization improved results. The audiences may have different baseline purchase intent.

To test a persona-specific page, split eligible traffic within that persona or campaign context between the existing page and the tailored page. Compare those randomized groups. Do not infer lift by comparing unrelated audience segments.

2. Find your biggest conversion opportunities

Start with evidence about where customers hesitate. A strong testing program combines behavioral data with the language customers use to describe their needs.

Useful inputs include customer reviews, support questions, post-purchase surveys, ad comments, search terms, checkout feedback, and analytics. Session recordings or heatmaps can add context if you use a separate tool that provides them. Treat them as clues about behavior, not proof of a cause.

Diagnose the journey before choosing a change

Observed pattern

Possible explanation

What to investigate or test

Strong ad engagement, weak purchase performance

The page does not continue the ad's promise, or the ad attracts low-intent clicks

Message match, offer consistency, and acquisition quality

Visitors leave before exploring

Slow loading, unclear value, accidental clicks, or a poor mobile first screen

Real-device loading and first-screen clarity

Visitors read but rarely interact with the purchase area

The offer is unclear, the buying section is hard to find, or questions remain unanswered

Buying-section placement, product demonstration, and objection handling

Many product interactions, few checkout starts

Confusing options, unexpected pricing, broken cart behavior, or discount failures

Buy-box usability and a complete cart test

Many checkout starts, few purchases

Shipping cost, payment friction, inventory, or tracking gaps

Checkout completion and order attribution before changing page copy

CVR improves while revenue per session declines

Smaller baskets or a more aggressive discount

Offer economics and order composition

One device type underperforms

Layout, performance, browser behavior, or traffic mix

Device-specific QA and sufficiently sized segment analysis

One variant records almost no orders

A weak experience or a measurement failure

End-to-end purchases and attribution in both variants

Several causes can produce the same pattern. Investigate enough to form a testable explanation rather than treating the table as a diagnosis.

Fix obvious failures before testing preferences

Broken CTAs, unreadable text, missing prices, overlapping elements, failed discounts, and incorrect product selections should be fixed directly. You do not need an experiment to decide whether checkout should work.

Once the experience is functional, test uncertain choices: how much education to provide, which proof matters, how to present options, and which page structure best serves the audience.

Use your own baseline

Compare like-for-like traffic wherever possible. A branded search visitor and a first-time social visitor arrive with different knowledge and intent. Category benchmarks can provide context, but your own comparable traffic is the more useful baseline for prioritizing improvements.

Bounce rate and time on page need particular care. A short visit can mean confusion or a quick purchase. A long visit can mean engagement or difficulty. Definitions also differ across analytics tools. Use these signals to form hypotheses, then judge the experiment using the business outcome.

3. Choose the right success metric

Define success before launching. Jurni's experiment goals include conversion rate (CVR), average order value (AOV), and revenue per session (RPS). Choose the goal that matches the question, then review the others for trade-offs.

Metric

Working definition

Best suited to

Main limitation

CVR

Attributed orders ÷ sessions

Clarity, trust, friction, and purchase completion

Can rise while order value or margin falls

AOV

Attributed revenue ÷ attributed orders

Basket size, bundles, and product mix

Describes purchasers; fewer people may buy

RPS

Attributed revenue ÷ sessions

Overall revenue yield from incoming traffic

Does not account for product, shipping, or fulfillment costs

Use the revenue and attribution definitions in your report consistently. Do not assume a different tool handles tax, shipping, refunds, repeat orders, currency, or attribution windows in the same way.

A simple example of conflicting metrics

The following numbers are illustrative, not a statistical result:

Metric

Control

Variant

Sessions

10,000

10,000

Orders

300

340

Revenue

$24,000

$23,800

CVR

3.00%

3.40%

AOV

$80.00

$70.00

RPS

$2.40

$2.38

The variant generates more orders but slightly less revenue from the same traffic. It leads on CVR and trails on RPS. Whether it is valuable depends on the stated objective and unit economics; “more conversions” alone is not enough to decide.

Add business guardrails

For an offer or bundle test, review contribution profit as well as revenue. For a subscription test, review retention, cancellation, and refund behavior when those outcomes have had time to mature. For a lead-generation test, review qualified leads and subsequent purchases rather than celebrating every email submission.

These checks may require Shopify, your subscription provider, CRM, or other reporting. Do not assume every guardrail is a native Jurni experiment goal.

Write guardrails as decision rules before launching. For example: “We will not roll out an offer if its additional discount and fulfillment costs erase the revenue benefit.” For operational risks, agree on a concrete acceptable limit with the team responsible for that outcome.

Use intermediate events to explain results

Product-section views, CTA clicks, quiz starts, quiz completion, add-to-cart, and checkout starts can help explain where an experience changes behavior, if those events are configured and verified.

They are usually diagnostic metrics. A higher click rate does not establish purchase lift. Some journeys also go directly to checkout, so a missing add-to-cart step is not automatically a problem.

4. Prioritize and scope your tests

Write a hypothesis that can be evaluated

Use this structure:

For [audience], we believe [observed problem] prevents [desired action]. Changing [experience] should improve [primary metric] because [customer reason]. We will also monitor [guardrails].

Example:

For first-time visitors from demonstration ads, we believe the product page does not explain how the product fits into a daily routine. A landing page with a short usage demonstration and a three-step routine should improve purchase CVR by reducing uncertainty. We will also monitor RPS and page performance.

“Try a new hero” is an activity. “Help first-time customers understand the outcome before introducing technical details” is a hypothesis.

Rank ideas using four practical questions

  1. Reach: How much eligible traffic encounters this problem?

  2. Impact: If the hypothesis is right, could it materially affect purchases or revenue?

  3. Evidence: What customer feedback or behavior supports the idea?

  4. Effort and risk: How much work, integration complexity, or operational exposure does it introduce?

Use a simple high/medium/low ranking if numerical scoring creates false precision. A high-traffic page with a repeated customer objection is usually a stronger starting point than a small visual preference on a rarely visited page.

Prefer meaningful differences

Early tests should generally explore the offer, message, product selection, buying experience, or page concept. Cosmetic changes can matter, but small effects are harder to distinguish from noise and often need more traffic.

Start with one control and one strong challenger. Add variants only when you have enough traffic and a distinct question for each. Producing ten pages quickly does not mean you can evaluate ten pages reliably at once.

Keep unrelated conditions stable

If you are testing page structure, keep price, products, offer, shipping terms, and purchase mechanics comparable unless those are deliberately part of the concept.

If the challenger changes several things together, describe it as a complete experience test. A win then supports that experience as a package; it does not prove that the headline, layout, or discount independently caused the improvement.

5. Match page concepts to shopper intent

Choose the page around the question the shopper needs answered next. The following concepts are testing ideas, not guarantees of performance or promises that every interaction works without additional setup.

Page concept

Useful when

Core content

Strong comparison

Direct-response landing page

An ad promotes one clear product benefit

Ad-matched headline, benefit, proof, offer, buying section

General PDP versus campaign-specific page

Product landing page / PDP

Visitors understand the category and need product details

Demonstration, options, specifications, reviews, shipping

Existing PDP versus clearer product presentation

Advertorial

The shopper needs education before evaluating the product

Problem, explanation, evidence, solution, purchase path

PDP versus educational narrative

Listicle

Several distinct reasons support consideration

Scannable reasons, concrete proof, product connection

Long narrative versus structured reasons

Comparison page

Shoppers are choosing between approaches or products

Relevant criteria, factual trade-offs, fit guidance

General benefits versus explicit comparison

Creator or founder story

The ad's credibility comes from a person

Authentic story, demonstration, relevant proof

Generic page versus continuity with the creator's ad

Persona or use-case page

One traffic cohort has a specific need

Relevant scenario, suitable benefits, matching reviews

General page versus tailored page within that cohort

Collection page

Visitors want to browse a category

Clear grouping, product distinctions, selection guidance

Broad collection versus curated selection

Quiz / guided selection

Choosing the right product is difficult

Necessary questions, useful recommendation, easy purchase

Direct shopping versus guided selection

Bundle / build-your-own bundle

Products work together or choice affects basket size

Contents, value, selection rules, total price

Single-product path versus bundle-focused path

Gifting page

Buyers need help choosing for someone else

Recipient fit, price ranges, delivery information

Standard collection versus gift-led selection

Lead-generation page

The immediate objective is a valuable signup

Clear value exchange, concise form, next-step expectations

Generic signup versus specific educational or early-access offer

Waitlist / coming-soon page

A launch or restock cannot yet convert to a purchase

Product value, availability expectations, signup

Generic notification versus launch-specific proposition

Interactive experience

An interaction helps shoppers understand or choose

Useful selector, demonstration, or comparison interaction

Static explanation versus purposeful interaction

Example page structures

For cold traffic that needs education:

Ad-matched opening → recognizable problem → explanation of the product's approach → demonstration or evidence → relevant customer proof → offer and buying section → practical objections and FAQs.

For high-intent product traffic:

Product and outcome → clear price and options → purchase action → delivery and reassurance → demonstration → detailed proof and FAQs.

For creator-led traffic:

Recognizable creator and message → their relevant experience → product use → supporting evidence → offer and buying section → remaining objections.

These are starting structures. Keep a purchase path accessible while providing enough information for the audience. Long pages are not automatically persuasive, and short pages are not automatically efficient.

Match the promise from ad to checkout

Check four kinds of continuity:

  • Message: The page answers the question or desire introduced by the ad.

  • Product: The featured product, variant, or bundle is available and easy to select.

  • Offer: The advertised price, discount, and conditions match the buying experience.

  • Visual identity: The page feels like the same brand and, where relevant, the same creator story.

For example, an ad about convenience should land on an experience that demonstrates convenience early. It should not require the visitor to search through an unrelated brand story before seeing how the product fits their routine.

6. Build a practical testing backlog

The ideas below are starting hypotheses. Choose the ones supported by your customers and traffic.

Messaging and education

Test idea

When it is relevant

What to watch

Outcome-led hero versus product-feature hero

Customers understand ingredients or features poorly

Purchase CVR and whether the promise remains accurate

Use-case headline versus broad brand headline

The campaign targets a specific problem

Performance within that same traffic cohort

Demonstration near the top versus later on the page

Customers ask how the product works

Purchase behavior and mobile loading

Concise explanation versus detailed education

Customers need category knowledge before buying

RPS, purchase CVR, and where visitors drop out

Objection-specific FAQ versus generic FAQ

Support repeatedly answers the same purchase question

Purchases, not only FAQ engagement

Trust and proof

Test idea

When it is relevant

What to watch

Relevant customer reviews beside the buying section

Visitors hesitate at the point of purchase

CVR and whether reviews address the actual concern

Demonstration video versus static product images

The result or usage is easier to show than describe

Mobile speed, accessibility, and purchase impact

Clear returns and delivery information near the CTA

Customers ask about risk, timing, or cost

Purchase completion and consistency with actual policies

Product-specific proof versus generic brand ratings

Shoppers need evidence about one product

Accuracy and relevance of the proof

Founder explanation versus a benefit summary

Expertise or product origin may influence trust

Whether the story earns its space on the page

Use authentic reviews and substantiated product claims. Do not invent customer quotes, independent endorsements, scarcity, or countdown deadlines. Label brand-created editorial content clearly rather than implying independent coverage.

Buying experience and merchandising

Test idea

When it is relevant

What to watch

Clearer option selector versus current selector

Customers struggle with size, flavor, quantity, or format

Selection errors, CVR, and returns

Recommended starter bundle versus individual products

New customers do not know where to start

RPS, CVR, and contribution profit

Fewer featured products versus a broad range

Choice appears to delay decisions

RPS and whether important needs become unsupported

Per-unit value explanation versus total price alone

Pack sizes make comparisons difficult

Accurate total cost and order mix

Subscription explanation versus minimal plan details

Customers do not understand recurring value or terms

First purchase, cancellation, and retention

Earlier buying section versus education-first structure

High-intent visitors must scroll past information they know

Purchases and effects on less-informed visitors

Make subscription frequency, recurring charges, and one-time purchase options understandable. A short-term signup increase caused by confusion is not a sustainable improvement.

Offers and incentives

Test idea

When it is relevant

What to watch

Percentage saving versus fixed-amount saving

Customers may respond differently to the same economic value

Keep actual discount value equivalent if testing presentation

Bundle value versus a single-item discount

A routine or set solves a broader need

Margin, inventory, AOV, and CVR

Gift with purchase versus discount

The gift has relevant customer value

Gift cost, fulfillment, eligibility, and checkout behavior

Free-shipping offer versus product discount

Shipping cost is a known purchase barrier

Net contribution after shipping subsidy

Introductory subscription offer versus standard offer

Acquisition and retention economics are understood

Cohort value after the introductory period

An offer shown on a page must also work in the cart and checkout. Validate discount stacking, eligibility, product exclusions, subscription compatibility, and currency behavior. Store-wide price or shipping changes require a setup that isolates the intended treatment; changing the store setting for everyone is not a valid variant-level test.

Mobile experience and performance

Test simpler first-screen layouts, compressed imagery, lighter video treatments, clearer touch targets, shorter forms, and more usable product options when evidence points to friction.

Compare the complete experience, including redirects and third-party scripts. A visually stronger page can lose if it loads slowly or makes buying harder. Check on real phones and typical mobile connections, including the in-app browser used by your main acquisition channel.

Sticky purchase controls can be useful to investigate where the page and hosting setup support them. Verify they do not cover consent controls, product options, or other essential content.

7. Set up a reliable experiment in Jurni

Step 1: Define the decision

Record the hypothesis, eligible traffic, control, challenger, primary metric, guardrails, and decision criteria. Give the test a descriptive name, such as:

Meta prospecting | Starter bundle | PDP vs routine page | RPS

The name should make the business question understandable without opening every variant.

Step 2: Build and publish the experiences

Create the challenger in Jurni using your existing page, a template, or an AI-generated starting point. Check the content against your product facts, brand settings, policies, and actual offer.

Keep the control representative of the experience you would otherwise use. When comparing a Jurni page with a Shopify destination, check that both routes have the tracking needed for a fair comparison.

Decide which hosting setup matches the intended experience. Shopify-hosted pages can use the Shopify theme context; a subdomain experience may differ in navigation, scripts, and cart behavior. If those differ between variants, the test measures the combined experience. Document that scope.

Step 3: Configure the experiment

Add the intended destinations, select the control, choose the success metric, and set your confidence requirement. For a straightforward two-variant comparison, a 50/50 allocation is a practical starting point because it balances data collection.

A smaller initial challenger allocation can be appropriate for operational validation. Decide in advance when that check ends and the measurement phase begins. Record allocation changes; do not repeatedly move traffic toward whichever variant currently appears ahead.

Step 4: Send eligible traffic through the Smart Link

Use the experiment's Jurni Smart Link as the destination for the traffic you intend to test. A page being published does not mean visitors automatically enter the experiment.

Inspect the actual destination configured in the ad or campaign. Do not assume a duplicated ad inherited the intended URL. Follow a real click through to the final page and check relevant campaign parameters survive the route.

Once traffic already uses the appropriate Smart Link, destination changes can be managed through that link. Ad-platform delivery is still affected by other campaign changes and market conditions; do not treat a stable URL as a guarantee that platform learning can never change.

Step 5: Check routing rules

If your setup includes campaign-parameter or audience-based routing, verify whether a rule forces traffic to a particular destination before interpreting the split as randomized.

Keep targeted routing and randomized comparison conceptually separate. A deterministic rule can be useful for delivering a tailored experience, but a comparison of differently targeted groups does not isolate the page's effect.

Check returning-visitor behavior as well. Browser storage, consent, cross-device visits, and entry paths can affect assignment continuity. Use an approved preview or QA approach to inspect both variants instead of assuming repeated refreshes should alternate them.

Step 6: Validate the full purchase journey

Test both variants from entry link to completed purchase. Confirm the correct product, options, price, discount, shipping behavior, and order attribution. Include relevant accelerated-payment routes such as Shop Pay or Apple Pay when enabled.

A successful browser purchase-event test is not the same as verifying every purchase route. Likewise, seeing an order in Shopify does not establish that it is attributed to the correct experiment variant. Check the systems and pathways you intend to use for the decision.

Step 7: Launch and record the baseline

Save the start time, configuration, page versions, and important campaign conditions. Check incoming traffic and tracking early. Once those are healthy, allow the experiment to collect evidence without continuous edits.

Jurni's connection or publication labels are operational signals. They are not, by themselves, proof that all intended traffic is entering the experiment or that orders are being attributed correctly. Review actual sessions and purchases.

8. Plan traffic, sample size, and test duration

There is no universal “enough data” number

Required sample size depends on your baseline, the smallest improvement worth detecting, variability, traffic allocation, number of variants, and statistical method. A small improvement generally needs more data than a large one.

Define a minimum worthwhile effect: the smallest improvement that would justify implementing and maintaining the change. This is a business input, not a target you can make the results reach.

For example, moving CVR from 3.0% to 3.3% is a 10% relative improvement and a 0.3 percentage-point absolute improvement. That is a much smaller signal than moving from 3% to 6%, despite both involving percentage figures.

Use a planning estimate

Before launching, estimate eligible traffic and expected orders for each variant. If a test receives 1,000 eligible sessions a day, splits 50/50, and has a 3% baseline order rate, each group would average roughly 15 orders a day. That estimate describes data volume; it does not tell you when the test will become conclusive.

If you use a sample-size calculator, ensure its metric and statistical assumptions match the question. A fixed-horizon conversion calculator is a planning aid, not a replacement for Jurni's own significance calculation. Revenue metrics need consideration of revenue variability, not just conversion counts.

Cover the business cycle

Plan to include complete weekday and weekend patterns when those affect the business, plus enough time for typical purchases to mature. A full week is a useful scheduling unit, not a universal stopping rule. Many tests need longer, and a short promotion may only support conclusions about that promotion.

Decide your minimum observation period and review schedule before launch. A high-confidence result during an unusual traffic spike should still be checked for representativeness.

As an external reference, Intelligems recommends at least seven days and 300 orders per group in its own significance guidance. Those are Intelligems recommendations, not Jurni requirements or universal guarantees. The important principle is to plan for data volume and elapsed time together. [1]

What to do with limited traffic

  • Run one challenger at a time.

  • Focus on larger, evidence-backed changes.

  • Use a meaningful amount of comparable traffic instead of fragmenting it across many small audiences.

  • Avoid testing minor details when a larger experience problem remains.

  • Set a maximum review date so weak tests do not run indefinitely.

  • Accept “inconclusive” as an honest outcome.

Do not switch from purchases to clicks simply because clicks produce a more exciting result. A click-focused test is valid when the decision itself concerns that action; it cannot establish purchase lift on its own.

Use A/A testing when validating measurement

An A/A test routes traffic to equivalent experiences to investigate assignment and measurement. It can be useful before a major migration or after a tracking concern.

The two groups will not produce identical results. Look for persistent routing, exposure, or attribution problems rather than expecting perfect equality. An A/A result also does not certify every future test or every payment route.

If the experiences look the same but use different hosting, scripts, or checkout paths, the comparison also measures those differences. It is not a pure identical-experience A/A test.

9. Read results and decide what to do

Read the result in layers

  1. Validity: Was traffic routed and measured correctly?

  2. Outcome: What happened to the preselected primary metric?

  3. Uncertainty: How strong is the evidence, and how stable is the estimate?

  4. Economics: Is the effect large enough to matter, with acceptable trade-offs?

  5. Scope: Which audience and operating conditions does the conclusion cover?

A large percentage lift from a handful of orders should not bypass these checks.

Understand lift

Relative lift is:

(Variant metric ÷ Control metric − 1) × 100

If RPS moves from $2.00 to $2.20, relative lift is 10%. The absolute improvement is $0.20 per session. If the control metric is zero, a conventional relative-lift calculation is undefined and should not drive the decision.

Observed lift is an estimate. A result that strongly suggests a positive effect can still leave substantial uncertainty about the effect's size. Do not treat the displayed point estimate as a guaranteed future return.

Use Jurni's significance status with the full result

Jurni's significance workflow uses your selected control and success metric to help assess the evidence. Read the statuses as a progression in evidence rather than an automatic instruction to deploy:

Status

Practical interpretation

Collecting data

Continue gathering evidence; check setup and data quality

Early lead

One experience is ahead, but the result is still preliminary

Likely to win

The evidence is stronger; review remaining uncertainty and your planned observation period

Winner

Review the qualifying result alongside validity, guardrails, and business impact before rollout

Set the confidence requirement before launch. Higher requirements generally need stronger evidence and can take longer to reach. Do not lower the setting simply to obtain a winner label.

Different platforms use different statistical models and definitions of confidence. Read a probability according to the metric and comparison the report actually names. “Probability of beating the control” and “probability of being best among all variants” are different questions when there are multiple challengers. Do not assume every platform uses the same calculation. [1]

Avoid chasing significance

Monitoring is useful for catching errors. Repeatedly stopping at the first favorable reading is a different practice and can lead to unreliable decisions, especially when the statistical method does not support that stopping rule.

Follow the planned review criteria. Do not assume that a live-updating dashboard makes every stopping strategy valid. More variants, more metrics, and more segment searches also create more opportunities to find a chance pattern. Treat unexpected discoveries as hypotheses to validate.

Do not turn metric shopping into a win

If the planned goal was RPS and only AOV improves, report that distinction. The result may motivate a useful follow-up, but it does not satisfy the original revenue-yield objective.

Changing the selected control or KPI changes the comparison and statistical interpretation; it does not mean that the underlying traffic, orders, or revenue disappear. Preserve the original question and explain any revised analysis.

Use a clear decision framework

Outcome

Recommended action

Credible improvement on the primary metric, acceptable guardrails

Roll out to the tested audience and monitor

Positive direction, insufficient evidence

Continue if the planned traffic and time make a useful answer feasible

No clear result at the review limit

Mark inconclusive; keep the control or make an explicitly non-statistical operational choice

Strong negative result with valid measurement

Stop the challenger and record the learning

Primary metric improves but economics worsen

Investigate the trade-off before rollout

Tracking, routing, or checkout failure

Fix the issue and evaluate a clean measurement period or new experiment

A promising result appears only in an unplanned small segment

Treat it as exploratory and run a focused follow-up

An inconclusive result is not proof that the experiences are equivalent. It means the evidence did not resolve the question at the level you needed.

Estimate impact without overstating it

For planning, multiply the observed absolute RPS difference by future eligible sessions. A $0.20 RPS increase across 50,000 comparable sessions suggests $10,000 in additional revenue under similar conditions.

Label this as a scenario, not guaranteed incremental revenue. The estimate assumes the effect persists and the audience, inventory, and economics remain comparable. It is revenue, not profit, and it should not be extrapolated to all store traffic without evidence.

10. Troubleshoot misleading results

Uneven traffic allocation

A 50/50 configuration does not imply an exact 50/50 observed split at every moment. Random variation, repeat sessions, and the difference between assignment counts and session counts can matter.

A large, persistent, unexplained imbalance deserves investigation. Check destination availability, eligibility rules, forced routing, blocked scripts, redirects, and missing measurements. Formal sample-ratio checks should use the relevant assignment unit and expected allocation, rather than automatically treating session totals as independently randomized visitors. [2]

Attribution differs across tools

Jurni, Shopify, ad platforms, and web analytics may disagree because they measure different populations, attribution windows, revenue components, or session definitions. Reporting delays, consent, browser restrictions, and timezone boundaries can add differences.

Reconcile a sample of known purchases and compare equivalent definitions. Perfect agreement is not required for a useful test, but an unexplained tracking difference between variants can invalidate the result. Intelligems' analytics FAQ provides examples of why experiment reporting and store reporting can diverge; its particular rules should not be assumed to apply to Jurni. [3]

Tracking scripts fire more than once

Multiple integrations can record the same action. Audit native integrations, GTM, and custom scripts for duplicate purchase or add-to-cart events. Confirm whether repeated events represent distinct customer actions or duplication before using the count diagnostically.

Checkout paths differ

Check standard checkout, accelerated payments, subscriptions, discount combinations, and external destinations that matter to the test. A journey can complete successfully while losing browser-side measurement or variant attribution.

If the variants use different purchase paths, identify whether the difference is part of the experience being tested or an unintended implementation mismatch.

You edited the experience during the test

Changing a headline, offer, product, or checkout behavior midway through the experiment changes what the variant represents. A single aggregate result may now combine different treatments.

Record meaningful changes and use a new experiment or a clearly separated analysis period when needed. Simply renaming a variant does not isolate historical exposure to an earlier version. Allow for purchases from visitors who encountered that earlier version.

You changed weights, added variants, or disabled variants

These actions change who receives each experience and when. They do not automatically erase existing data, but they can complicate interpretation, especially when traffic quality changes over time.

Retain a clear change log. Compare overlapping exposure periods where appropriate and confirm how the reporting method handles the change. Do not assume a new challenger has the same evidence base as a control that has been running for weeks.

Several experiments affect the same customer

A page test can interact with a price test, cart upsell, shipping offer, or subscription test. Random overlap is not automatically invalid, but interactions can complicate what the result means.

For a straightforward program, avoid concurrent changes to the same decision point. If overlapping tests are necessary, coordinate eligibility and analysis with the owners of each test. Third-party tools must also be verified to work with the page's rendering and cart setup.

A few large orders dominate revenue

Review whether unusually large orders are genuine, expected customer behavior, or known internal/test purchases. Do not remove legitimate orders only because they reverse the result.

Define exclusions consistently before analyzing. If large orders materially change the conclusion, document the sensitivity and collect more evidence or seek a suitable statistical review.

A result changes after rollout

A winner can still face lower overall performance later because acquisition, seasonality, competition, or stock changed. Compare current conditions with the experiment's scope. A declining store-wide CVR does not, by itself, disprove a previous relative improvement. [3]

11. Roll out winners and build a testing habit

Deploy the tested experience deliberately

Confirm that the destination you roll out matches the version that won. Recheck prices, inventory, discounts, and tracking. Update the intended Smart Link allocation or destination, then verify live traffic follows the new route.

For higher-impact changes, consider a staged rollout or a planned holdout when traffic and tooling allow it. Retain a working fallback destination and monitor purchases, revenue, and operational issues after deployment.

Capture the learning, including its limits

A useful experiment record includes the customer problem, hypothesis, screenshots or page versions, audience, dates, allocation, primary metric, guardrails, result, implementation decision, and next question.

Prefer a bounded conclusion:

For first-time mobile visitors from demonstration ads, the routine-led page improved the selected purchase metric during the tested period. The next test will examine whether a shorter routine explanation preserves the benefit.

Avoid universal conclusions such as “long pages always win” or “customers prefer bundles.” The experiment rarely supports a statement that broad.

Reuse principles, then validate new contexts

If a page wins, identify what it teaches you about the customer. Perhaps new visitors need help choosing a starter product, or shipping transparency reduces hesitation.

Use that insight to inform another product or audience, but treat the extension as a new hypothesis. A winning component inside one complete experience may behave differently elsewhere.

Use a sustainable cadence

Review active experiments for health regularly and maintain a scheduled decision review. Separate this from the work of choosing the next test. Keep a short, ranked backlog rather than a long list of unowned ideas.

Assign one owner to coordinate the hypothesis and decision, with clear responsibility for page creation, campaign links, measurement, and QA. A test is complete when the decision is implemented and the learning is recorded, not merely when the dashboard displays a result.

12. Templates, examples, and checklists

Copy-ready experiment brief

Test name:

Customer problem and supporting evidence:

Eligible audience and entry source:

Hypothesis:

Control destination and version:

Challenger destination and version:

What changes / what stays comparable:

Primary metric and minimum worthwhile effect:

Guardrails and acceptable limits:

Traffic allocation and assignment/routing checks:

Confidence requirement, minimum observation period, and review limit:

Measurement source and verified purchase paths:

Owner, launch date, and decision date:

Rollout and fallback plan:

Example 1: Beauty — answer a product-fit objection

Evidence: Customers repeatedly ask whether the product suits their skin type.

Hypothesis: Bringing clear suitability guidance and relevant, authentic reviews beside the buying section will reduce uncertainty.

Comparison: Existing product page versus the same page with the fit guidance moved into the purchase decision area. Keep the product, price, and offer the same.

Primary metric: CVR. Guardrails: RPS and subsequent return reasons.

Learning: Whether that guidance helps this audience purchase. The test does not establish that every audience needs a longer page.

Example 2: Wellness — connect the ad to the daily routine

Evidence: The ad demonstrates a routine, while the existing destination opens with a dense ingredient explanation.

Hypothesis: A routine-led landing page will better answer the expectations created by the ad.

Comparison: Existing PDP versus a Jurni page with usage, schedule, product demonstration, and substantiated benefits.

Primary metric: RPS. Guardrails: CVR and offer economics; review retention if the test shifts subscription mix.

Learning: Whether the complete routine-led experience improves revenue yield for that traffic. Do not attribute the entire result to one section or make unsupported health claims.

Example 3: Apparel — reduce fit uncertainty

Evidence: Support and reviews show recurring questions about sizing.

Hypothesis: Accessible sizing information and relevant model measurements will help visitors choose confidently.

Comparison: Current buying section versus the same section with clearer fit guidance.

Primary metric: CVR. Guardrails: Size-related returns and exchanges after sufficient time has elapsed.

Learning: Whether the change improves initial purchasing without increasing downstream dissatisfaction.

Example 4: Food or beverage — test the starter purchase

Evidence: New customers find the range difficult to navigate.

Hypothesis: A curated starter bundle will simplify the first purchase and improve revenue per session.

Comparison: Existing collection or product destination versus a starter-bundle page with clear contents, quantities, and total cost.

Primary metric: RPS. Guardrails: CVR, contribution profit, fulfillment complexity, and repeat purchase when measurable.

Learning: Whether the bundle experience creates sufficient value to justify any additional discount or fulfillment cost.

Example 5: Product selection — quiz versus direct shopping

Evidence: Customers struggle to identify which of several products fits their needs.

Hypothesis: A short, useful selection flow will improve product confidence enough to offset the extra steps.

Comparison: Direct product browsing versus a tested quiz and recommendation path.

Primary metric: RPS or purchase CVR. Diagnostic metrics: Quiz starts, completion, recommendation clicks, and abandonment.

Guardrails: Page responsiveness, recommendation accuracy, and a usable fallback path.

Learning: Whether guided selection improves the full journey. High quiz completion alone is not a purchase win.

Pre-launch checklist

  • The customer problem and hypothesis are specific.

  • The control reflects the real baseline.

  • The primary metric and guardrails are selected in advance.

  • Each variant has a distinct purpose and sufficient expected traffic.

  • The intended Smart Link is used in the actual campaign destination.

  • Routing rules and returning-visitor behavior have been checked.

  • Both experiences work on mobile and the main in-app browser.

  • Product selection, stock, prices, and offers are correct.

  • Discounts and shipping conditions work at checkout.

  • Relevant payment and subscription paths have been tested.

  • Purchases can be attributed appropriately in each variant.

  • Tracking is consistent and duplicate events have been investigated.

  • Minimum observation criteria and the review limit are recorded.

  • The team knows which overlapping changes to avoid or document.

  • The owner and fallback destination are clear.

Results checklist

  • The experiment collected the intended traffic.

  • No unresolved routing or measurement issue explains the difference.

  • Both variants were compared over appropriate exposure periods.

  • The original primary metric is still the basis of the decision.

  • The observation period covers relevant business and conversion cycles.

  • Evidence strength and effect size are evaluated separately.

  • Guardrails and unit economics are acceptable.

  • Segment findings are labeled as planned or exploratory.

  • Live edits, promotions, stockouts, and allocation changes are documented.

  • The outcome is recorded as a win, loss, or inconclusive result.

  • The rollout scope and next hypothesis are explicit.

Common questions

Should I always test one thing at a time?

Use one coherent hypothesis at a time. Focused changes isolate a mechanism; complete experience tests compare broader approaches. Both are useful when the conclusion matches the design.

Can I stop when Jurni shows a winner?

Review the evidence alongside your planned observation period, data quality, business impact, and guardrails. A statistical status is an input to the decision.

What if my store does not have enough traffic?

Reduce the number of variants, prioritize larger changes, and focus on a meaningful traffic stream. If a useful answer remains infeasible, improve known usability problems and gather customer evidence without presenting the result as a proven lift.

Can a variant win on CVR and lose on RPS?

Yes. More customers may purchase smaller or more heavily discounted baskets. Use the preselected primary metric and review economics before rollout.

Should every ad get its own page?

Only when the differences in intent or message justify a different experience. Group ads around coherent customer needs, then test whether a tailored page improves outcomes within that traffic. Avoid fragmenting traffic into experiments too small to interpret.

Can I compare a Jurni page with my current Shopify page?

Use the appropriate destinations in a Smart Link experiment and verify comparable tracking across both paths. If hosting, navigation, cart, or checkout differ, interpret the result as a comparison of complete experiences.

What if the result is inconclusive?

Record it honestly. Revisit the evidence, the size of the change, and the available traffic. Keep the control or make an operational choice with the uncertainty stated; do not manufacture a winner by changing the metric, audience, or date range after the fact.