
Trial starts vs D30 revenue: which paywall metric should pick the winner?
Compare trial starts vs D30 revenue as experiment primary metrics. Guardrails, sample size, when each metric misleads.
Your paywall experiment finished. Variant B lifted trial starts 12%. Leadership cheers. Finance reviews D30 revenue four weeks later and Variant B is flat or negative. Variant A looked worse on the dashboard metric but would have made more money.
This happens because trial starts and D30 revenue answer different questions. Trial starts measure immediate intent at the paywall. D30 revenue measures whether that intent was real after refunds, cancellations, and plan mix settle.
Choosing the wrong primary metric for paywall A/B tests is how teams ship "winners" that hurt LTV. This guide compares both metrics, when each misleads, guardrails to use, and sample size realities for mobile subscription apps.
Definitions that must stay consistent
| Metric | Numerator | Denominator | Typical source |
|---|---|---|---|
| Trial start rate | Users who start a trial | Installs or paywall views | RevenueCat, Rheo funnel |
| D30 revenue per install | Cumulative gross revenue by day 30 | Installs in cohort | Warehouse, RevenueCat charts |
| Trial-to-paid | Paid conversions | Trial starters | RevenueCat |
| Paywall arrival rate | Users who reach paywall step | Flow starts | Rheo |
Lock denominator before the experiment starts. Per install is standard for growth teams. Per paywall view is valid when testing paywall template only with fixed placement.
Trial starts as primary metric
When trial starts are the right choice
| Situation | Why |
|---|---|
| Onboarding or placement hypothesis | Moving paywall affects arrivals and starts together |
| Short runtime (7 to 14 days) | Revenue has not matured |
| Low traffic | Faster signal than revenue |
| New app with thin LTV history | No reliable D30 baseline |
| Testing pre-paywall copy in Rheo | Billing event unchanged; intent is the lever |
Trial starts align with Day-0 optimization. RevenueCat benchmarks show most trials begin in the first session. If your experiment changes what users see before the paywall integration presents, trial starts per install are often the cleanest primary metric.
When trial starts mislead
| Trap | Example |
|---|---|
| Aggressive annual push | More trials, more Day-0 buyer's remorse refunds |
| Misleading trial copy | Starts spike, trial-to-paid collapses |
| Softer gate | Users start trial without intent to pay |
| Shorter path to paywall | Unqualified users hit the ask sooner |
| Discount-first win-back style handoff | Starts rise, retention falls |
A variant can win trial starts while losing money. That is not a statistics error. It is a metric mismatch.
D30 revenue as primary metric
When D30 revenue is the right choice
| Situation | Why |
|---|---|
| Paywall template and package mix test | Price, annual default, tier count affect ARPU |
| Mature app with stable conversion curves | Revenue signal is estimable |
| High traffic | Enough cohort mass by day 30 |
| Price or introductory offer change | Short-term starts are explicitly misleading |
| Leadership optimizes ROAS | Finance thinks in revenue per install |
D30 revenue per install captures trial-to-paid, plan mix, and partial refund behavior for monthly cohorts. It is the metric closest to unit economics for subscription apps with monthly billing.
When D30 revenue misleads or delays
| Trap | Example |
|---|---|
| Long annual plans | D30 undercounts LTV; use D60 or modeled LTV |
| Small sample | Revenue variance is huge week to week |
| Seasonality | Holiday cohorts skew compare |
| Delayed experiments | You wait a month to learn what onboarding could fix in a week |
| External budget shifts | UA mix changes during experiment |
D30 is slow. Product teams iterating onboarding weekly cannot wait 30 days for every decision. Use D30 for paywall pricing and packaging decisions; use faster metrics for journey decisions.
Side-by-side comparison
| Dimension | Trial starts | D30 revenue per install |
|---|---|---|
| Speed to readout | Days | 30+ days after cohort start |
| Traffic needed | Lower | Higher |
| Sensitive to placement | High | High (via qualified arrivals) |
| Sensitive to package design | Medium | High |
| Captures refund behavior | No | Partially at D30 |
| Good for Rheo flow tests | Yes | Secondary |
| Good for RC paywall tests | Secondary | Yes |
| Gaming risk | Medium | Lower |
Recommended metric hierarchy
Use a primary and guardrails. Never promote a variant that wins primary but breaches guardrails without explicit sign-off.
Journey experiments (Rheo)
| Role | Metric |
|---|---|
| Primary | Trial starts per install |
| Guardrails | Paywall arrival rate, onboarding completion, paywall dismiss rate |
| Secondary (if runtime allows) | Trial-to-paid at D14, D30 revenue per install |
Paywall template experiments (RevenueCat or paywall vendor)
| Role | Metric |
|---|---|
| Primary | D30 revenue per install (or ARPU at D30 among starters) |
| Guardrails | Trial start rate, refund rate, annual mix |
| Secondary | Trial-to-paid, chargeback rate |
Flow-level test that moves both placement and offering
Avoid if possible. If unavoidable:
| Role | Metric |
|---|---|
| Primary | D30 revenue per install |
| Guardrails | Trial starts, completion, support tickets |
| Runtime | 4+ weeks |
Worked example: two variants, two stories
Variant A (control): Longer onboarding, paywall after demo. Trial starts per install: 8.2%. D30 revenue per install: $1.14.
Variant B: Short onboarding, early paywall. Trial starts per install: 9.4% (+15% relative). D30 revenue per install: $0.98 (-14% relative).
Variant B wins trial starts. Variant A wins revenue. Diagnosis from Rheo step data: Variant B brings lower-intent users to the paywall faster. Trial-to-paid is 22% versus 31%. RevenueCat shows similar offering; the journey qualified users differently.
Decision: Promote Variant A for paid traffic. Optionally test Variant B only on high-intent organic segment via decision node. Primary metric for that segmented test: D30 revenue, not starts.
Sample size and runtime
Trial starts converge faster than revenue.
| Metric | Rough weekly installs per variant (50/50, 10% baseline trial rate, 20% MDE) |
|---|---|
| Trial start rate | ~4,000 |
| D30 revenue (high variance) | Often 2 to 3x trial-only requirement |
Revenue variance includes plan mix ($4.99 monthly vs $59.99 annual), currency, and refunds. Pre-register minimum runtime:
| Primary metric | Minimum runtime |
|---|---|
| Trial starts | 14 days |
| D30 revenue | 30 days after last enrolled user + 30 days |
| Trial-to-paid | 21 to 45 days depending on trial length |
Do not peek daily without rules. Peeking inflates false positive rate.
Guardrail table for paywall A/B test metrics
| Guardrail | Breach signal | Action |
|---|---|---|
| Paywall dismiss rate | +15% relative vs control | Pause variant |
| Refund rate in first 7 days | +20% relative | Pause variant |
| Onboarding completion | -10% relative | Investigate placement |
| Support "billing confusion" tickets | Spike | Review copy compliance |
| ATT opt-in (if in path) | -5 absolute points | Check permission order |
When to use hybrid decision rules
Mature teams sometimes promote winners with a compound rule:
- Trial starts not worse than -5% relative versus control, and
- D30 revenue per install +8% relative or more, and
- No guardrail breach
This works when traffic supports reading both signals in one runtime. Low-traffic apps should serialize: optimize journey on trial starts, then optimize paywall template on D30 revenue with a fixed journey winner pinned.
Tooling alignment
| Question | Rheo | RevenueCat |
|---|---|---|
| Did users reach paywall? | Step funnel | Optional |
| Trial started? | Via integration events | Yes |
| Which experiment arm? | Channel experiment | RC Experiments |
| D30 revenue by arm? | Join via warehouse or export | Charts by offering |
| Paywall UI variant | No (external) | Yes |
Rheo does not manage paywall UI. Attribute trial starts to flow variant in Rheo and paywall template in RevenueCat explicitly in your experiment doc to avoid ambiguous readouts.
Documentation template
Copy into every experiment brief:
Hypothesis:
Primary metric:
Guardrails:
Denominator (per install / per paywall view):
Runtime minimum:
Peeking policy:
Rheo flow version IDs:
RevenueCat offering / placement IDs:
Decision rule if primary and revenue disagree:
Thirty minutes of documentation prevents a month of arguing.
Summary
Paywall A/B test metrics are not interchangeable. Trial starts pick winners fast for journey and placement tests in Rheo. D30 revenue per install picks winners for packaging and paywall template tests in RevenueCat or your paywall integration. Use guardrails so a trial start win cannot silently destroy LTV.
Pre-register primary metric, runtime, and decision rules before launch. When metrics disagree, trust revenue for monetization tests and trust starts for onboarding tests, then validate the other within the next experiment cycle.