Magister Digital AI

CRO Experimentation

Turns a backlog of CRO ideas into an ICE-ranked, sample-sized, stop-gated test plan so traffic is never burned on under-powered or unprioritized tests.

Client: CGH Injury Lawyers

What it produced for CGH Injury Lawyers

An ICE-scored, sample-sized test backlog for CGH's money pages (the /denver/<practice>-lawyer/ pages). The deliverable ranks the tests, writes the hypothesis for the top one, sizes every test against the skill's sample-size table, and tells CGH which tests are even runnable given their traffic. For a law firm, the conversion event is a qualified call or form submission, and that drives every sizing decision below.

Number discipline note

CGH's real page traffic and conversion rate are not in the source files. The baseline conversion rate of 4 percent for a Denver PI money page and the monthly traffic figures are labeled vertical-benchmark planning assumptions, not CGH-reported numbers. The ICE scores, MDE choices, and sample sizes are real applications of the skill against those assumptions. Where a test cannot reach significance on plausible CGH traffic, that is called out explicitly, which is the whole point of sizing before running.

Step 1: ICE-scored test backlog

ICE = (Impact + Confidence + Ease) / 3, each axis 1 to 10. Ranked highest first.

RankTest idea (money page element)ImpactConfidenceEaseICE
1Add a sticky above-the-fold click-to-call bar on mobile8998.7
2Replace generic hero with "Free Denver Car Accident Case Review" + local trust line8888.0
3Add a short 3-field form ("name, phone, what happened") above the fold vs current long form7877.3
4Add a trust band (Denver case results context, attorney photos, review count) below hero6776.7
5Add an exit-intent "Talk to a lawyer now" prompt5665.7
6Reorder page: move FAQ and "what is my case worth" higher5565.3
7Swap CTA color and copy from "Contact Us" to "Get My Free Case Review"5797.0

Tests 1, 2, and 7 are the run-first set: high confidence, easy to ship, and they attack the same friction (a buried, generic call to action) that the Quality Score audit flagged on the same pages. Run them in priority order, not all at once, so the conversion signal stays clean.

Step 2: Hypothesis for the top test (full template)

OBSERVATION: Mobile session recordings and the Quality Score audit both show the
phone CTA on /denver/car-accident-lawyer/ is below the fold. PI visitors arrive in
crisis on mobile and want to call now, but must scroll to find the number.

HYPOTHESIS: We believe that adding a sticky above-the-fold click-to-call bar for
mobile visitors will increase qualified call conversions because injury victims
convert by phone in the moment of high intent, and removing the scroll-to-call
friction captures that intent before they bounce to a competitor.
  - CHANGE   = sticky click-to-call bar pinned to the top on mobile
  - AUDIENCE = mobile visitors to the Denver car accident money page
  - OUTCOME  = higher rate of click-to-call conversion events
  - RATIONALE = high-intent, in-the-moment behavior; friction removal; the same
                buried-CTA problem the Quality Score landing-page audit found

METRICS:
  Primary   = qualified phone call conversion rate (click-to-call to tracked call)
  Secondary = form submissions, bounce rate, scroll depth (watch for cannibalizing the form)

TEST TYPE: A/B

MINIMUM DETECTABLE EFFECT (MDE): 20% relative lift

SAMPLE SIZE REQUIRED: see Step 3

Step 3: Sample sizing (from the skill's table, 95% confidence, 80% power)

Visitors needed per variation, read off the skill's reference table at a 4 percent baseline (interpolated between the 2 percent and 5 percent rows, labeled vertical-benchmark baseline):

TestBaseline CRMDEVisitors per variationTotal (2 variations)
1. Sticky click-to-call (mobile only)~4%20%~9,500~19,000
2. Hero rewrite~4%20%~9,500~19,000
7. CTA color + copy~4%30%~4,250~8,500

Higher MDE means a much smaller sample. The CTA copy test (test 7) is sized at a 30 percent MDE because a button change can plausibly move conversions that much, and it makes the test reachable on realistic traffic.

Step 4: Traffic reality check (the gate that saves wasted tests)

Assume a vertical-benchmark 3,000 visitors per month to the Denver car accident money page, roughly half mobile.

TestEligible traffic / monthTotal sample neededWeeks to reachVerdict
1. Sticky click-to-call (mobile only ~1,500/mo)~1,500~19,000far over 8 weeksNOT runnable as a single-page A/B. Pool across all city money pages or accept a larger MDE
2. Hero rewrite (all traffic ~3,000/mo)~3,000~19,000~6+ weeks, borderlineRunnable only if pooled across multiple money pages
7. CTA copy (all traffic ~3,000/mo)~3,000~8,500~3 weeksRunnable

This is the most valuable output: at PI traffic levels, single-page A/B tests with a tight MDE often cannot reach significance inside the 4 to 8 week useful window. The plan therefore says: pool equivalent money pages (all /denver|aurora|lakewood/car-accident-lawyer/ pages share the variant) to combine traffic, OR test changes big enough to justify a 30 percent MDE, OR roll out high-confidence, high-ease changes (tests 1 and 7) as best-practice improvements rather than forcing an under-powered test. Do not run a test that cannot reach its sample size, that is how false winners ship.

Step 5: Stop / no-stop rules

Stop a test ONLY when all four are true:

  1. Statistical significance reached (95%+, p < 0.05).
  2. The sample size from Step 3 is achieved.
  3. A full business cycle has completed (PI inquiry volume varies by week, so at least one full week, preferably 2 to 4).
  4. No external event polluted the window (no major local news event, ad outage, or holiday).

Do NOT stop because results "look good," because a stakeholder is impatient, or because only part of the cycle has elapsed. Duration: minimum 1 week, preferred 2 to 4 weeks, maximum useful 4 to 8 weeks. Run an AA test first to confirm the test setup is unbiased, and apply a Bonferroni correction if multiple variants run at once.

Pre-launch checklist (applied to test 7, the runnable one)

  • Hypothesis written in full template format
  • ICE score calculated (7.0) and documented
  • Baseline conversion rate measured (vertical-benchmark 4%, replace with real CGH analytics before launch)
  • MDE decided (30% relative)
  • Required sample size calculated (~8,500 total)
  • Traffic check passed (reachable in ~3 weeks)
  • Duration window set with a scheduled end date (3 to 4 weeks)
  • Primary and secondary metrics defined
  • AA test run to verify setup (schedule before the live test)

Compliance note

CRO copy is still YMYL copy. Any variant that touches case-value, settlement, fee, or outcome language routes through the claims ledger before it goes into the test. No variant may promise a result or imply that prior outcomes predict future ones.