Six Sigma Analyze Phase: Finding Root Causes with Data
The Six Sigma Analyze phase is where a DMAIC project stops describing the problem and starts proving why it happens. You take the data gathered in Measure, test the suspected causes, and narrow the list to the few inputs that truly drive defects, delays, variation, or waste. For professionals developing structured process improvement expertise, a Certified Six Sigma Expert pathway can complement practical work with the analytical principles used throughout DMAIC.
This is the phase where projects either become useful or become theater. A fishbone diagram full of guesses is not root cause analysis. It is only a list of suspects. Analyze is about evidence.

Where the Analyze Phase Fits in DMAIC
DMAIC stands for Define, Measure, Analyze, Improve, and Control. Analyze is the third step. Define clarifies the business problem and customer requirements. Measure confirms the current baseline. Analyze explains the gap between current and required performance.
For professionals who apply analytical thinking across teams and operational functions, Management Certifications can complement Six Sigma learning by strengthening broader management and decision-making capabilities.
The expected output is simple: a validated set of root causes that can be addressed in Improve. Not a vague theme. Not a department to blame. A specific, actionable cause tied to a measurable effect on CTQs, or critical to quality requirements.
Take an example. "Operator error" is usually too lazy to be useful. "Fill head 4 drifts below target after sanitation when residue is present on the contact surface" is actionable. You can test it, fix it, and monitor it.
Start with Process Thinking, Not Statistics
Good Analyze work usually begins on the floor, in the workflow, or inside the system where the defect is created. Before you run regression, walk the process. Review the value stream map. Look for rework loops, queues, handoffs, batching, inspection points, and places where people use workarounds.
Common tools include:
Cause and effect diagrams: useful for organizing suspected causes across methods, machines, materials, people, measurement, and environment.
5 Whys: helpful when used with evidence, risky when it turns into storytelling.
Pareto charts: good for finding the vital few defect categories that account for most of the loss.
Process maps: essential for spotting hidden steps that never appear in written procedures.
To be blunt, if your team skips process observation and jumps straight into software, you will often model the wrong problem very precisely.
Turn Root Causes into Testable Hypotheses
In the Analyze phase, suspected causes should be written like hypotheses. Each one should connect an input X to an output Y.
Examples:
If supplier lot moisture is higher, defect rate increases.
If claims are touched by more than two processors, cycle time rises.
If machine temperature exceeds the upper control band, dimensional variation increases.
Then you decide how to test each hypothesis. You may need a simple stratified Pareto. You may need ANOVA to compare shifts, machines, or suppliers. You may need regression to quantify the relationship between several inputs and one output. Pick the tool that fits the question. Do not reach for a complex model just because it looks impressive.
Statistical Tools That Matter Most
Several analytical methods show up again and again in Analyze:
Histograms and box plots: show distribution shape, spread, and outliers.
Scatter plots: reveal whether two variables move together.
Control charts: separate common cause variation from special cause signals.
Hypothesis tests: confirm whether an observed difference is likely to be real.
ANOVA: tests whether group means differ across machines, shifts, lines, or materials.
Regression analysis: estimates how strongly process inputs affect the output.
Multi vari analysis: separates variation across time, location, and product characteristics.
A practical rule: use the simplest tool that can answer the question with confidence. A p-value below 0.05 may support statistical significance, but it does not prove the cause matters operationally. Check effect size too. A statistically significant 0.02 percent improvement may not pay for the change.
Do Not Trust Bad Measurement Data
Measurement system analysis is often the unglamorous step that saves the project. If the gauge, form field, inspection method, or data extract is unreliable, your root cause analysis is built on sand.
In manufacturing, Gauge R and R is used to test whether observed variation comes from the part or the measurement system. A study with less than 10 percent measurement variation is generally treated as acceptable in industrial settings, while 10 to 30 percent is a judgment call. In service processes, the same principle applies. If two claims reviewers classify the same error differently, you do not yet have clean defect data.
Here is a mistake I see in the field. Teams compare day shift and night shift reject rates without removing startup, changeover, or rework samples. The night shift gets blamed, but the real driver is product mix. Stratify first. Argue later.
Metrics Used in the Analyze Phase
Analyze connects root causes to business and customer metrics. Useful measures include:
DPMO: defects per million opportunities. The formula is total defects divided by total opportunities, where total opportunities equals units produced times CTQ opportunities per unit, multiplied by 1,000,000.
Cpk: a process capability index that shows how well the process fits within specification limits.
First pass yield: the percentage of units or transactions completed without rework.
Cycle time: especially important in service, claims, onboarding, and software support processes.
Case evidence shows why this discipline matters. In a high speed food manufacturing line, analysis using multi vari methods, Gauge R and R, and multiple regression identified sanitation residue buildup as the main driver of underweight rejects. The reported model had an R squared near 0.89, and the validated corrective actions cut the reject rate sharply while lifting Cpk from below 0.2 to close to 2.0.
The point is not the exact figures. It is that those gains did not come from a brainstorming session alone. They came from narrowing the problem to verified inputs, then confirming the fix moved the metric.
For professionals working with technology-driven processes, data analysis can also extend beyond traditional manufacturing measurements into systems, automation, and digital workflows. A Deep Tech Certification pathway can provide complementary technology-focused knowledge for teams analyzing increasingly digital processes.
Common Analyze Phase Mistakes
Stopping at symptoms: "late orders" is an outcome, not a cause.
Choosing causes by seniority: the loudest manager is not a statistical method.
Ignoring measurement error: bad data makes neat charts dangerous.
Testing too many variables at once: keep the test plan disciplined.
Moving to Improve too early: unverified fixes often create new variation.
What Certification Candidates Should Master
If you are preparing for Six Sigma Green Belt or Black Belt work, expect Analyze to test both tool knowledge and judgment. You need to know when to use Pareto analysis, fishbone diagrams, 5 Whys, hypothesis testing, ANOVA, regression, and measurement system analysis.
The question type that trips candidates up is choosing the right test for the data. Two continuous variables point toward correlation or regression. Comparing three or more group means points toward ANOVA. Attribute data needs a chi-square or proportion test. Memorizing tool names is not enough. You have to match the tool to the data type and the process question.
Universal Business Council runs training paths in Six Sigma, Lean Six Sigma, quality management, data analysis, and process improvement that build this judgment through applied work rather than theory alone.
Your Next Step
Take one current process problem and list five suspected root causes. For each one, write the X and Y, identify the data needed, and choose the test you would run. If you cannot test a cause, rewrite it until you can. That is the real work of the Six Sigma Analyze phase.
As analytical work becomes increasingly connected with software, automation, dashboards, and technology systems, professionals can also benefit from broader technical awareness. A Tech Certification pathway can complement Six Sigma expertise with additional technology-focused learning.
FAQs
1. What is the Analyze phase in Six Sigma DMAIC?
The Analyze phase is the third stage of the Six Sigma DMAIC methodology:
Define → Measure → Analyze → Improve → Control
Its purpose is to determine why a process problem occurs by identifying and validating the factors that drive defects, delays, variation, cost, or other poor outcomes.
The team moves from measuring the problem to explaining it:
Problem → Potential Causes → Data Analysis → Validated Root Causes
Analyze matters because plausible explanations are plentiful. Evidence is less cooperative.
2. What is the main goal of the Analyze phase?
The main goal is to identify the critical few root causes responsible for the performance gap.
During Measure, the team establishes what is happening and how often. During Analyze, it asks why.
For example, if the Measure phase confirms a 6.5% defect rate, Analyze investigates whether defects are associated with machine settings, materials, shifts, suppliers, environmental conditions, process methods, or other factors.
The goal is not to create the longest possible list of causes. It is to determine which causes materially influence the outcome.
3. What is a root cause in Six Sigma?
A root cause is an underlying factor that contributes meaningfully to the process problem and can be addressed to improve performance.
Suppose customers receive late orders. “Orders wait too long” merely describes the symptom.
Further investigation might reveal that orders are released in large batches because approvals occur only twice per day. The approval and batching policy may therefore be a deeper process cause.
A useful root cause should have a defensible relationship to the outcome and ideally be actionable.
4. What is the difference between a symptom and a root cause?
A symptom is an observable result of a problem. A root cause explains why that result occurs.
For example:
Symptom: High defect rate.
Possible Cause: Incorrect machine settings.
Deeper Cause: Setup parameters are manually entered from uncontrolled documents.
Fixing symptoms often provides temporary relief. Addressing verified underlying causes has a better chance of producing sustainable improvement.
“Operators make errors” is frequently closer to an accusation than a root cause analysis.
5. What does Y = f(X) mean in the Analyze phase?
Six Sigma often represents process relationships as:
Y = f(X)
Here, Y is the process output being improved, while Xs are potential input variables that may influence it.
For example:
Y = Delivery Lead Time
Potential Xs could include:
Order volume + approval time + inventory availability + picking time + staffing + carrier performance
The Analyze phase attempts to determine which X variables significantly influence Y.
The objective is to move from dozens of suspected Xs to the critical few Xs that deserve improvement action.
6. What inputs are needed before starting the Analyze phase?
A team should enter Analyze with a clearly defined problem, reliable measurement system, representative baseline data, current-state process understanding, operational definitions, and relevant CTQs.
The Measure phase should have answered questions such as:
What is being measured?
How is it measured?
Can the measurement system be trusted?
What is current performance?
Analyzing unreliable data merely produces unreliable conclusions with nicer statistics attached.
7. What tools are commonly used in the Six Sigma Analyze phase?
Analyze-phase tools range from qualitative root cause techniques to statistical methods.
Common tools include Pareto charts, Fishbone diagrams, 5 Whys, process maps, scatter diagrams, box plots, histograms, hypothesis tests, confidence intervals, correlation, regression, ANOVA, and Design of Experiments where appropriate.
The tool should be selected according to the question and data.
A Six Sigma practitioner does not earn extra analytical virtue by using regression when a simple stratified plot already exposes the problem.
8. How is Pareto analysis used to find root causes?
A Pareto chart ranks categories according to frequency, cost, severity, or another relevant measure.
Suppose defects are:
Defect Type | Number |
|---|---|
Seal failure | 1,200 |
Label error | 600 |
Carton damage | 300 |
Missing insert | 180 |
Other | 120 |
Seal failures clearly deserve deeper investigation.
However, Pareto analysis identifies the largest categories of problems, not necessarily their underlying root causes.
The team still needs to determine why seal failures occur.
9. How is a Fishbone Diagram used during Analyze?
A Fishbone Diagram, also called an Ishikawa or Cause-and-Effect Diagram, organizes possible causes of a problem.
Common categories include:
People, Machine, Method, Material, Measurement, and Environment.
For example, possible causes of seal failure might include temperature settings, machine wear, material thickness, setup procedures, measurement error, or humidity.
Fishbone analysis is useful for generating hypotheses and organizing process knowledge.
Everything written on the diagram remains a potential cause until evidence validates it. Fish bones, regrettably, do not possess statistical authority.
10. How are the 5 Whys used in root cause analysis?
The 5 Whys technique repeatedly asks why a problem occurred to move beyond superficial explanations.
For example:
Why were orders shipped late?
Because picking started late.
Why did picking start late?
Because orders were released late.
Why were orders released late?
Because approval was delayed.
Why was approval delayed?
Because approvals were processed in scheduled batches.
The investigation suggests that the batching policy may be a deeper driver of delivery delay.
The number five is a guideline, not a sacred statistical constant. Stop when the causal chain is sufficiently understood and supported.
11. How does process mapping help identify root causes?
A detailed process map shows where work, information, materials, decisions, and handoffs occur.
Teams can annotate the map with defects, waiting time, rework, cycle time, queues, and variation.
For example:
Receive Application → Validate → Wait 8 hrs → Approve → Rework → Process
The map may reveal that most total lead time occurs around approval and rework rather than actual processing.
This helps focus data collection and root cause investigation on the parts of the process where performance deteriorates.
12. How are scatter plots and correlation used in Analyze?
A scatter plot visualizes the relationship between two quantitative variables.
For example, a team might plot:
Machine Temperature vs Defect Rate
If defect rates tend to increase as temperature changes, the pattern may justify further investigation.
Correlation can quantify the strength and direction of a relationship.
However:
Correlation ≠ Causation
Two variables moving together does not prove that one causes the other. Confounding factors, common drivers, or coincidence may explain the relationship.
13. How is hypothesis testing used to validate root causes?
Hypothesis testing helps determine whether observed differences or relationships are consistent with more than random sampling variation under a specified null model.
For example, the team might test:
H₀: Mean cycle time is the same for Shift A and Shift B.
H₁: Mean cycle time differs between the shifts.
If the evidence is sufficiently strong against H₀, shift-related conditions deserve further investigation.
Statistical significance should be considered alongside effect size, practical significance, data quality, and process knowledge. A microscopic difference can become statistically significant with enough observations and still be commercially irrelevant.
14. How is ANOVA used in the Analyze phase?
ANOVA, or Analysis of Variance, is commonly used to compare means across three or more groups.
Suppose a process uses four material suppliers. The team wants to know whether average product strength differs among suppliers.
ANOVA evaluates whether the between-group differences are large relative to within-group variation.
If significant differences are detected, post-hoc comparisons may help identify which groups differ.
ANOVA can therefore help determine whether categorical process factors such as supplier, machine, shift, or method are associated with differences in a continuous output.
15. How is regression analysis used to identify process drivers?
Regression analysis models the relationship between an outcome and one or more predictor variables.
Suppose:
Y = Processing Time
Potential predictors include:
X₁ = Order Size
X₂ = Staffing Level
X₃ = Rework Count
X₄ = System Downtime
Regression can estimate how these factors relate to processing time while accounting for other variables included in the model.
Teams should still examine model assumptions, residuals, interactions, multicollinearity, and practical significance before declaring an X a verified causal driver.
16. How should data be stratified during root cause analysis?
Stratification means dividing data into meaningful groups to reveal patterns hidden in aggregate results.
Useful categories may include:
Machine, shift, supplier, product, operator, location, customer segment, day, material lot, or transaction type.
Suppose overall defect rate is 5%. Stratification reveals:
Machine A = 2.1%
Machine B = 2.4%
Machine C = 11.0%
The overall average concealed an obvious concentration of defects.
Aggregated data is wonderfully capable of making very different processes look politely similar.
17. How do you distinguish a potential cause from a validated root cause?
A potential cause is a reasonable hypothesis generated from process knowledge, brainstorming, Fishbone analysis, 5 Whys, or exploratory data.
A validated root cause has supporting evidence showing a meaningful relationship with the process outcome.
The progression should be:
Potential Cause
↓
Define Expected Relationship
↓
Collect Appropriate Data
↓
Analyze Evidence
↓
Confirm Practical Importance
↓
Validate Root Cause
For example, “night shift causes more defects” is a hypothesis. Data may instead reveal that the real driver is a particular machine used disproportionately during the night shift.
18. What are common mistakes during the Six Sigma Analyze phase?
A common mistake is treating brainstorming output as proof. Other failures include using poor-quality data, confusing correlation with causation, ignoring process stability, choosing statistical tests without checking assumptions, overlooking confounding variables, and focusing only on p-values.
Teams also sometimes analyze only average performance while ignoring variation, segmentation, time patterns, or process changes.
Perhaps the most expensive mistake is confirming the cause management already believes and quietly ignoring contradictory evidence.
Analyze exists precisely to make that habit harder.
19. How do you know when the Analyze phase is complete?
Analyze is generally ready to close when the team has narrowed the initial list of possible causes to a defensible set of validated critical drivers that explain enough of the problem to support improvement decisions.
The team should be able to answer:
Which factors materially influence Y?
What evidence supports those relationships?
How large is their impact?
Are the causes actionable?
Do the findings make operational sense?
The objective is not perfect knowledge of every source of variation. It is sufficient evidence to design targeted improvements with acceptable risk.
20. How should Six Sigma teams find root causes with data?
A disciplined Analyze phase can follow this sequence:
Confirm the Problem and Baseline
↓
Review Measurement-System Reliability
↓
Map the Current Process
↓
Stratify Performance Data
↓
Use Pareto Analysis to Focus Investigation
↓
Generate Potential Causes
↓
Apply Fishbone and 5 Whys
↓
Translate Suspected Causes into Testable Hypotheses
↓
Collect Additional Data Where Needed
↓
Use Graphical Analysis
↓
Apply Appropriate Statistical Tests
↓
Evaluate Correlation, ANOVA or Regression Where Relevant
↓
Check Assumptions and Confounding Factors
↓
Evaluate Effect Size and Practical Significance
↓
Validate the Critical Xs
↓
Prioritize Root Causes for Improve
The essential transition is:
“We think X causes Y”
to
“The evidence shows X materially influences Y, and changing X is a credible way to improve Y.”
Suppose the team begins with:
Y = High defect rate
Brainstorming produces 25 possible X variables. Pareto analysis narrows the problem to one major defect category. Stratification shows that defects are concentrated on one machine and certain material lots. Further testing demonstrates that material thickness and a machine-temperature setting are strongly associated with failures, while operator and shift differences disappear after those factors are accounted for.
The Analyze phase has now done its job.
It has replaced a convenient story such as “operators need more training” with evidence about the process variables actually driving failure.
That is the point of root cause analysis in Six Sigma: not to find someone or something that can plausibly be blamed, but to identify which controllable factors genuinely explain poor performance well enough to justify changing them.
Otherwise, root cause analysis becomes root cause theater, complete with diagrams, meetings, colored sticky notes, and remarkably little causality.
Related Articles
View AllSix Sigma
Six Sigma Cost of Poor Quality: Finding Hidden Losses
Six Sigma Cost of Poor Quality explains how defects, rework, downtime, and lost customers hide real financial losses inside everyday operations.
Six Sigma
Six Sigma Process Mining: Finding Hidden Inefficiencies
Learn how Six Sigma process mining uses event logs and DMAIC to reveal hidden inefficiencies, reduce rework, and improve process control.
Six Sigma
Six Sigma and Data Analytics: Turning Process Data into Insights
Learn how Six Sigma and Data Analytics turn process data into practical insights that improve quality, speed, cost, and control across industries.
Trending Articles
The Role of Blockchain in Ethical AI Development
How blockchain technology is being used to promote transparency and accountability in artificial intelligence systems.
AWS Career Roadmap
A step-by-step guide to building a successful career in Amazon Web Services cloud computing.
Top 5 DeFi Platforms
Explore the leading decentralized finance platforms and what makes each one unique in the evolving DeFi landscape.