Easton Digital · Test Plan

The Mallet
Free Shipping Test

Does free shipping sell more mallets, even at a higher sticker price? Here is the hypothesis, the test, and how we run it.
Gongs Unlimited·August 4, 2026·Prepared for Haedar Kasem
Section 01

The Hypothesis

"Do people just see free shipping and are more inclined to buy? Some of these mallets have 18 to 20 dollars in ship cost. If we can bake that into the cost and people are more inclined to buy it, then I think it is worth a shot." Haedar Kasem · June 16, 2026
Mallets are where this should bite hardest. Shipping is a rounding error on a gong and a quarter of the price on a mallet.
$19 SHIPPING AS A SHARE OF ITEM PRICE Gong $988 avg 1.9% Mallet $74 avg 25.7%
The bar is the item price. The filled portion is what shipping adds on top. On a gong nobody notices. On a mallet it is a quarter of the price, and it appears only after the customer has already decided to buy.
Mallet revenue · 12 mo
$513,336
10.3% of store revenue
Units sold
6,894
Highest volume category
Median price
$68
Across 171 SKUs
Shipping as % of item
26%
Highest in the catalog
Stated as a testable claim
Because shipping is roughly a quarter of the price on a mallet, we believe presenting that cost inside the item price instead of adding it at checkout will lift revenue per visitor on mallet pages. We will know it is true if revenue per visitor rises while conversion rate and mallet attach rate hold.
Section 02

The Test

One variable. We raise the item price by exactly what shipping costs, then set shipping to zero. Both bars below are the same length, because the customer pays the same either way. Only the split moves.
Control · Arm A$87 total
$68 mallet
+ $19 shipping
Variant · Arm B$87 total
$87 malletshipping $0
Identical totals by design. If we changed the price and the shipping by different amounts, a loss could be the price and a win could be the framing, with no way to tell them apart. Holding the total fixed means anything the test finds is the framing effect and nothing else. It also makes the test margin neutral per order, so only the order count can move.
Test parameters
Split50 / 50, cookie based so returning visitors stay in the same arm
DurationTen weeks for a 20% lift, five for a 30% lift. See the power maths below, this is not a guess
Primary metricRevenue per visitor. Prices differ between arms, so conversion rate alone can mislead
SecondaryConversion rate and add to cart rate, always split by device
GuardrailMallet attach rate on gong orders. Most mallets sell as add ons, and those buyers see the price in a cross sell, not on a landing page
Call it when95% confidence on revenue per visitor, after the four week minimum, device split checked
Stop early only ifVariant conversion drops sharply and holds for a full week
Does it have the traffic to be significant?
We checked this properly rather than assuming, and the answer changed the plan. A handful of hero SKUs cannot carry this test. It has to run across the whole mallet category.
The useful thing about a conversion test is that its power is governed by order count, not traffic. At the purchase rates this store actually runs at, roughly 2.5% on mobile and 4.1% on desktop, detecting a 20% lift needs about 420 orders in each arm, and that figure barely moves whether the true rate is 2.5% or 9%. We know mallet order counts exactly from Shopify, so we do not need a traffic estimate to answer this.
4 hero SKUs990 units/yr · 35 orders per arm per month
53 wks
$40 to $150 band · 115 SKUs4,526 units/yr · 160 orders per arm per month
11 wks
All mallets · 189 SKUs7,766 units/yr · 275 orders per arm per month
7 wksrecommended scope
Weeks needed to detect a 20% lift at 95% confidence. A 30% lift resolves faster: 24 weeks, 5 weeks and 3 weeks respectively. Order counts are Shopify units with a 15% haircut, since a multi unit cart is one order not two.
Why the narrow SKU list fails
The four mallets we first proposed sell 990 units a year between them. Split across two arms that is 35 orders per arm per month, so reaching 420 takes a year. Any result read before then is noise. Scoping tightly felt careful, but on this volume it just guarantees an unreadable test.
Why category wide works
All 189 selling mallet SKUs move 7,766 units a year, which is 275 orders per arm per month and a seven week test. Mallets are your highest volume category, and that volume is the only reason this test is viable at all.
What we give up by widening
The clean $40 to $150 logic gets diluted. Category wide sweeps in eGong Wands at $19, where $19 of shipping exceeds the item price, and Flumie sets at $300 where shipping is 6% and the framing barely registers. Two ways to handle it, and we would do both. Keep the very cheapest items out, since doubling a sticker price is a different test, which costs about 21% of the units and still lands near eight weeks. Then segment the read by price band afterwards so we learn where the effect actually lives instead of only whether it exists on average.
Section 03

How It Gets Built

We looked hard at avoiding a price testing plan by duplicating each mallet at the higher price and split URL testing between the two. At four SKUs that works and costs $39 a month. At 189 SKUs it falls apart, and the traffic maths above says 189 is what the test needs.
Duplicating means, for every SKU: a second product built and priced, reviews synced across so the variant is not handicapped, the duplicate excluded from the Shopping feed, a canonical set, and two SKUs sharing one physical stock pool. Multiply that by 189 and it is a multi day build carrying real oversell risk for the whole test window. A price testing app applies one rule to a collection in minutes.
Mallet page visitor All site traffic 50 / 50 split URL test Sees current price $68 · shipping calculated Sees test price $87 · FREE shipping Checkout $87 either way
One product URL, two price treatments, assigned by cookie so a returning visitor always sees the same one. The app holds the assigned price through cart and checkout, which is the part worth QA testing with a real order before launch.
What it costs
Elevate StandardPrice and shipping testing on one plan
$99≈ $250 for a ten week run · recommended
Shoplift Advanced
$299price testing still beta · shipping not documented
Shogun Unlimited · Intelligems Plus
$499≈ $1,250 for the same run
Duplicate products, split URLShogun Pro · viable under ~10 SKUs only
$39cannot reach significance at that scope
The $39 route is genuinely cheaper and we would take it if the test could run on a few SKUs. It cannot. Elevate Standard at $99 is the recommendation, still a fifth of the Shogun price testing tier. Shopify's free Rollouts feature cannot do this at any scope, since it swaps themes only and explicitly excludes price and shipping.
Apply it as a collection rule
Point the test at the mallet collection rather than a SKU list, so new mallets inherit it and nothing has to be maintained by hand across 189 products for ten weeks.
Not through Google Ads
Routing the two URLs through Shopping or PMax breaks the test. Both SKUs sit in the same feed and the same auction, and Google serves whichever it predicts converts better. That prediction is the exact thing we are measuring, so the split stops being random. Meta is valid but far too thin on mallet traffic.
Section 04

How The Free Shipping Works

With a shipping profile, not a discount code. The obvious route is the wrong one, and the difference shows up in mixed carts.
✕ AUTOMATIC DISCOUNT Cart Test mallet $87 Gong $988 applies to the whole order Mallet shipping$0 Gong shipping$0 Charged$0 You just gave away gong shipping. Any lift is partly free shipping on a $988 item. ✓ SHIPPING PROFILE Cart Test mallet $87 Gong $988 splits by profile Mallet profile $0 General profile $45 Charged $45 Only the mallet ships free. The gong pays normally.
Shopify free shipping discounts apply to the entire order and cannot be scoped to products, so in a mixed cart they zero out the gong too. A shipping profile splits the cart into sub carts, prices each separately, then adds them back together. That is native behaviour and it is exactly what the test needs.
Which means the free shipping has to come from the app, not from Shopify settings
A shipping profile is attached to a product, so it is identical for everyone who views that product. It cannot show free shipping to one arm and calculated shipping to the other on the same SKU. That was fine when each arm was a separate duplicate product. Now that both arms are the same product, only the app can vary shipping per visitor, which is exactly why the plan needs Elevate Standard rather than the cheaper Basic tier. Price testing alone would raise the price in arm B without removing the shipping, which is not the test, it is just a price rise.
  1. Configure price and shipping in one test. Both levers move together in arm B, or the total is no longer held constant and the result is unreadable.
  2. Scope it to the mallet collection, not a SKU list, so nothing needs hand maintenance across 189 products.
  3. Keep Hawaii, Alaska, Puerto Rico and international on calculated rates, matching how your existing free shipping SKUs already work.
  4. Add the badge to arm B only. The badge is the part of the test the customer actually sees, and it has to appear for the variant and not the control.
  5. Verify the mixed cart during the trial. Put a mallet and a gong in one cart on the variant and confirm the gong still pays its normal shipping. This is the behaviour the diagram above shows, and it is the one thing most likely to be implemented wrong.
One quirk to expect at checkout
Shopify adds rates with the same name together. When names differ across profiles it takes the cheapest from each and shows a single merged line labelled Shipping. So a mallet only cart visibly shows FREE Shipping, but a mixed cart shows one combined line and the customer never sees the word free. Another reason this design reads the standalone shopper cleanly and the attach buyer poorly.
Section 05

Running It

Week 1
Set the price
Shipping cost pull, then variant pricing per band
Week 2
Build in trial
Collection rule, badge, mixed cart QA on a real order
Weeks 3 to 12
Run · no changes
Ten weeks to a 20% read, five to a 30% read. No peeking decisions, no pricing or feed edits
Week 13
Read
RPV, device split, price band split, attach guardrail
Longer than the four to six weeks we first sketched, because the honest power maths says so. Ten weeks at $99 a month is about $250 of tooling.
The traffic question is now answered, and it did not need GA4
We flagged a GA4 pull as the blocking item. It turns out not to be needed. A conversion test's power depends on order count, not sessions, and Shopify gives us order counts exactly. That matters here because Easton does not currently have GA4 access on this account, so the traffic estimate was never going to be verifiable quickly. The verdict stands on measured sales data instead: category wide is viable, a narrow SKU list is not.
Read it in slices, not just in total
SliceWhy
Price bandThe whole thesis is that shipping matters more when it is a bigger share of price. If true, the lift should be largest on cheap mallets and fade on expensive ones. That gradient is the real finding, more than the average
DeviceMobile is 70% of clicks here and converts at 2.53% against desktop's 4.08%. The two can move in opposite directions and mobile dominates the blend
Standalone vs attachMost mallets ride along with a gong. Those buyers meet the price in a cross sell, not on a page, so the framing barely reaches them while the higher sticker still does
What we need from you
Real shipping cost
Charged versus actual carrier cost on recent mallet orders, ideally split by weight band. Sets the variant price. Blocking.
Theme migration status
Complete, or scheduled outside a ten week window. A migration mid test corrupts both arms. Blocking.
Andrew sign off
Now covering the whole mallet category rather than four SKUs, so worth raising as such. Still margin neutral, and you already run 17 gong SKUs with free shipping.