Control Chart & KPI Diagnostic Guide
Use control charts to decide when a live portfolio KPI deserves action, then diagnose the first credible cause before the team changes spend, product, or channel strategy.
Executive Summary
A control chart is an action gate for ongoing KPI monitoring. Set the baseline from the earliest comparable stable period, use regime-aware limits for skewed game metrics, investigate only when a signal rule fires, and document the cause before changing the operating plan. Use the statistical-significance guide for experiments; use this page for live process monitoring.
Reviewed May 20, 2026 against NIST and ASQ control-chart references, Transcend's ROAS pipeline, the control-chart-generator skill, and Fairy Dragon/Slam Clash diagnostic files. The page gives operating rules; final decisions require current company data, version history, source manifests, and GP judgment.
MECE Boundary
Keep this page narrow: it answers whether a recurring metric changed enough to investigate. It should not duplicate experiment inference, paid-media waterfalls, or product-marketing diagnosis.
| Question | Use This Surface | Why |
|---|---|---|
| Did a monitored KPI move out of its expected range? | Control Chart | Ongoing process monitoring against a stable or regime-specific baseline. |
| Which ROAS component broke first? | ROAS Diagnostic | Waterfall diagnosis across attribution, CPM, CTR, ARPU/AOV, and conversion. |
| Can we trust an A/B test or lift result? | Statistical Significance | Sample size, confidence, power, and experiment interpretation. |
| What is the product or lifecycle root cause? | Product Marketing Analysis | Retention, conversion, economy, paywall, lifecycle, and message-market diagnosis. |
Build the Chart
Start with the decision the chart will trigger. A good control chart says "investigate now," "keep monitoring," or "re-baseline because the process changed." It does not prove causality by itself.
| Setup Choice | Transcend Rule | Failure Mode Prevented |
|---|---|---|
| Metric | Use one decision metric per chart: D1 retention, D7 ROAS, D30 ROAS, CPI, ARPU, pay conversion, or qualified pipeline. | Blended dashboards where the action trigger is unclear. |
| Grain | Use the grain the team can act on: daily cohorts for F2P UA, weekly reads for low-volume metrics. | False alarms from too-small daily samples. |
| Baseline | Begin with the earliest comparable stable period, then show peak and current regimes separately when relevant. | Charts that start from a peak and create a fake degradation story. |
| Maturity | Exclude immature cohorts using milestone plus revenue-lag cutoff from the data manifest. | Reacting to D7/D30 data that has not finished arriving. |
| Limits | For ROAS/game metrics, prefer per-regime median plus 3.5 x MAD; formal SPC individual charts may use moving-range 3-sigma limits. | Skewed whale-driven data widening or shifting a single overall limit. |
Read the Signal
Control charts manage investigation risk. A point outside the limit is a trigger to search for a cause; points inside the limit can still be suspicious when the pattern stops looking random.
| Signal Pattern | Action | Before Acting |
|---|---|---|
| One point outside a control limit | Investigate immediately. | Check data completeness, source changes, and cohort maturity. |
| Two of three points near the same edge | Open a diagnostic thread. | Look for a release, channel, geo, bid, or tracking change near the first point. |
| Four of five points on the same weak side | Treat as an emerging drift. | Confirm volume is high enough and the metric is not seasonally biased. |
| Long run on one side of baseline | Reassess the operating regime. | Decide whether to re-baseline, not just escalate a one-day alert. |
| All points inside limits but visible trend | Monitor and annotate. | Do not call it solved; non-random patterns can still imply a process change. |
KPI Stack
Pair fast warning metrics with slower confirmation metrics. The fast chart buys response time; the mature chart confirms economic severity.
D1 Retention / D1 ROAS
Use daily for product releases, onboarding changes, FTUE issues, campaign setup errors, and sudden traffic-quality shifts.
D7 / D30 ROAS
Use when cohorts are mature enough to confirm whether the early signal actually changes payback and budget posture.
Spend, CPI, Geo, Platform
Use before blaming product. Channel mix, platform split, or spend collapse can move blended ROAS without a product break.
Diagnose and Act
When a signal fires, move from chart to cause in a fixed order. The goal is not to collect every possible explanation; the goal is to find the first explanation strong enough to guide the next operating move.
| Step | Decision | Evidence Required |
|---|---|---|
| 1. Confirm | Is the point real? | Export completeness, no zero-cost artifacts, no immature cohorts, no broken revenue mapping. |
| 2. Timestamp | When did the break begin? | First affected cohort date, not the day the lagging dashboard finally showed it. |
| 3. Segment | Where did it happen? | Channel, geo, platform, campaign, cohort, payer segment, and spend-volume splits. |
| 4. Correlate | What changed at the same time? | Version history, release notes, bid strategy, creative rotation, store listing, MMP settings, and platform policy changes. |
| 5. Attribute | Which cause best explains the movement? | Metric decomposition plus negative controls or unaffected segments when available. |
| 6. Decide | What action is reversible, targeted, and measurable? | Owner, action, expected metric response, check date, and rollback condition. |
Portfolio Lesson
The Fairy Dragon/Slam Clash diagnostic history is the pattern to remember: a chart that starts from a temporary peak can imply product degradation, while a baseline-plus-peak-plus-current view can show a more nuanced reality. The later work also found platform, channel, geo, spend-volume, and attribution confounds. Control charts should force the next diagnostic question, not become the final diagnosis.
Current Source References
| Reference | Use in This Guide |
|---|---|
| NIST: What are Control Charts? | General definition of center line, upper/lower control limits, and investigation logic. |
| NIST: Individuals Control Charts | Moving-range method for individual observations and formal 3-sigma limits. |
| ASQ: Control Chart | Common and special cause framing, run-rule examples, and documentation discipline. |
skills/skills/ROAS_ANALYSIS_PIPELINE.md |
Transcend rules for data maturity, dual baselines, per-regime statistics, and attribution caveats. |
skills/skills/control-chart-generator/SKILL.md |
Implementation standard for any-metric, regime-aware control charts. |
portfolio/FairyDragon/DIAGNOSTIC_FINDINGS.md |
Sanitized portfolio lesson on avoiding peak-baseline and channel/platform confounds. |
Decision Rule
Act when the chart produces a signal and the diagnostic identifies a plausible operational cause. Monitor when the chart is noisy but unsignaled. Re-baseline when the business has entered a genuinely new regime and the old comparison would now mislead the team.