ISO 24481-1:2026
(Main)Statistical methods for implementation of Six Sigma — Exploratory data analysis — Part 1: General methodology
General Information
- Abstract
This document describes the concept and general principles of EDA for categorical and numerical data. It also provides some guidelines for conducting EDA and its place within Six Sigma projects. This document focuses on the graphical tools of EDA. It is applicable to organizations using manufacturing processes as well as service and transactional processes.
- Status
- Published
- Publication Date
- 13-Aug-2026
- Current Stage
- 6060 - International Standard published
- Start Date
- 14-Aug-2026
- Due Date
- 01-Jul-2026
- Completion Date
- 14-Aug-2026
Overview
ISO 24481-1:2026, titled Statistical methods for implementation of Six Sigma - Exploratory data analysis - Part 1: General methodology, establishes key concepts and methods for exploratory data analysis (EDA) within Six Sigma projects. This international standard is published by ISO Technical Committee 69/Subcommittee 7 and is applicable to organizations utilizing manufacturing, service, or transactional processes.
By focusing on both categorical and numerical data, ISO 24481-1 delivers structured guidelines and practical recommendations for effective EDA-an essential phase in the Six Sigma approach to data-driven decision making and process improvement.
Key Topics
- Exploratory Data Analysis (EDA) in Six Sigma: Emphasizes the importance of hypothesis generation and discovery in problem-solving and improvement cycles, contrasting with confirmatory data analysis (CDA) which focuses on hypothesis testing.
- Types of Data and EDA Challenges: Addresses different data scenarios-such as large or high-dimensional datasets-and highlights challenges around data visualization and dimensionality which are common in modern business environments.
- Role of Data Visualization: Outlines how graphical tools such as histograms, box plots, Pareto charts, scatter plots, and heat maps maximize data insights and communicate findings effectively.
- Key EDA Principles:
- Understand and visualize data before modeling.
- Apply critical thinking to formulate and refine hypotheses.
- Use a mix of graphical and quantitative techniques.
- Flexibly and iteratively explore data, especially to manage outliers, missing values, and high-dimensional datasets.
- Conducting EDA:
- Define the analysis purpose.
- Understand data sources and structure.
- Clean and curate data, documenting any changes.
- Explore data using visualization and descriptive statistics.
- Document and communicate findings.
- Iterate based on feedback or newly uncovered insights.
- Integration with Six Sigma DMAIC Phases: Provides guidance on employing EDA tools at each stage (Define, Measure, Analyze, Improve, Control) to support effective problem resolution and process optimization.
Applications
ISO 24481-1 aids organizations in leveraging EDA for:
- Manufacturing: Identifying process variations, optimizing quality, and uncovering root causes of defects.
- Service and Transactional Processes: Analyzing customer data, service response times, or transactional patterns to enhance efficiency and customer satisfaction.
- Business Analytics and Quality Improvement: Using EDA to extract actionable insights from large, complex datasets, and to facilitate continuous improvement efforts throughout the Six Sigma lifecycle.
Typical use cases include:
- Assessing if collected data aligns with Six Sigma project objectives.
- Rapidly visualizing trends, patterns, or anomalies that may influence project direction.
- Communicating complex data findings to stakeholders in a clear, accessible way.
- Guiding the selection of subsequent analytical or modeling techniques based on initial observations.
Related Standards
- ISO 3534-1: Statistics - Vocabulary and symbols: Provides foundational statistical terminology and definitions referenced throughout the EDA standard.
- ISO 16269-4: Statistical interpretation of data - Detection and treatment of outliers: Supplementary for handling outliers identified during EDA.
- ISO 24481-2/3 (in development): Will further detail graphical tools for low- and high-dimensional data, respectively, as part of this Six Sigma EDA series.
Summary
ISO 24481-1:2026 serves as a foundational standard for integrating exploratory data analysis into Six Sigma initiatives across industries. By promoting robust EDA practices-including visualization, critical evaluation, iterative exploration, and clear documentation-organizations can accelerate process optimization and data-driven decision making. This standard is highly relevant for quality professionals, Six Sigma teams, analysts, and anyone engaged in statistical analysis for process improvement. For optimal results, it is recommended to use ISO 24481-1 in conjunction with related statistical and quality management standards.
Get Certified
Connect with accredited certification bodies for this standard

BSI Group
BSI (British Standards Institution) is the business standards company that helps organizations make excellence a habit.

Bureau Veritas
Bureau Veritas is a world leader in laboratory testing, inspection and certification services.

DNV
DNV is an independent assurance and risk management provider.
Sponsored listings
Frequently Asked Questions
ISO 24481-1:2026 is a standard published by the International Organization for Standardization (ISO). Its full title is "Statistical methods for implementation of Six Sigma — Exploratory data analysis — Part 1: General methodology". This standard covers: This document describes the concept and general principles of EDA for categorical and numerical data. It also provides some guidelines for conducting EDA and its place within Six Sigma projects. This document focuses on the graphical tools of EDA. It is applicable to organizations using manufacturing processes as well as service and transactional processes.
This document describes the concept and general principles of EDA for categorical and numerical data. It also provides some guidelines for conducting EDA and its place within Six Sigma projects. This document focuses on the graphical tools of EDA. It is applicable to organizations using manufacturing processes as well as service and transactional processes.
ISO 24481-1:2026 is classified under the following ICS (International Classification for Standards) categories: 03.120.30 - Application of statistical methods. The ICS classification helps identify the subject area and facilitates finding related standards.
ISO 24481-1:2026 is available in PDF format for immediate download after purchase. The document can be added to your cart and obtained through the secure checkout process. Digital delivery ensures instant access to the complete standard document.
Standards Content (Sample)
International
Standard
ISO 24481-1
First edition
Statistical methods for
2026-08
implementation of Six Sigma —
Exploratory data analysis —
Part 1:
General methodology
Méthodes statistiques pour la mise en œuvre de Six Sigma —
Analyse exploratoire des données —
Partie 1: Méthodologie générale
Reference number
© ISO 2026
All rights reserved. Unless otherwise specified, or required in the context of its implementation, no part of this publication may
be reproduced or utilized otherwise in any form or by any means, electronic or mechanical, including photocopying, or posting on
the internet or an intranet, without prior written permission. Permission can be requested from either ISO at the address below
or ISO’s member body in the country of the requester.
ISO copyright office
CP 401 • Ch. de Blandonnet 8
CH-1214 Vernier, Geneva
Phone: +41 22 749 01 11
Email: copyright@iso.org
Website: www.iso.org
Published in Switzerland
ii
Contents Page
Foreword .iv
Introduction .v
1 Scope . 1
2 Normative references . 1
3 Terms and definitions . 1
4 Symbols and abbreviated terms. 3
5 Fundamental concepts . 3
5.1 General .3
5.2 Different types of data sets and possible challenges encountered within EDA .4
5.3 Relation and difference between data visualization (DV) and EDA .5
5.4 Principles and related steps for conducting EDA.5
6 Tools of EDA from different perspectives . 6
6.1 General .6
6.2 Certain tools for the different types of data .6
6.2.1 Discrete data .6
6.2.2 Continuous data .7
6.3 Sorted according to their objectives .8
6.4 EDA tools used in each phase of DMAIC .8
7 Emphases and cautions for certain tools and other aspects of EDA . 9
7.1 General .9
7.2 Importance of checking the metadata .9
7.3 Cautions about some graphical methods .10
7.3.1 Use graphical methods efficiently .10
7.3.2 Bar charts or dot charts instead of pie charts .11
7.3.3 Avoid misleading graphs and note the data Inking issue .11
7.3.4 Causation is more important than correlation .11
7.3.5 The relationship between EDA and CDA or other tools . 12
7.3.6 Choosing suitable data analysis software packages is important for EDA . 12
Annex A (informative) Comparison of data visualization and EDA from different aspects .13
Annex B (informative) Some concrete EDA principles in Six Sigma and quality improvement
projects . 14
Annex C (informative) Limitation of descriptive statistics and a typical example .16
Bibliography . 17
iii
Foreword
ISO (the International Organization for Standardization) is a worldwide federation of national standards
bodies (ISO member bodies). The work of preparing International Standards is normally carried out through
ISO technical committees. Each member body interested in a subject for which a technical committee
has been established has the right to be represented on that committee. International organizations,
governmental and non-governmental, in liaison with ISO, also take part in the work. ISO collaborates closely
with the International Electrotechnical Commission (IEC) on all matters of electrotechnical standardization.
The procedures used to develop this document and those intended for its further maintenance are described
in the ISO/IEC Directives, Part 1. In particular, the different approval criteria needed for the different types
of ISO document should be noted. This document was drafted in accordance with the editorial rules of the
ISO/IEC Directives, Part 2 (see www.iso.org/directives).
ISO draws attention to the possibility that the implementation of this document may involve the use of (a)
patent(s). ISO takes no position concerning the evidence, validity or applicability of any claimed patent
rights in respect thereof. As of the date of publication of this document, ISO had not received notice of (a)
patent(s) which may be required to implement this document. However, implementers are cautioned that
this may not represent the latest information, which may be obtained from the patent database available at
www.iso.org/patents. ISO shall not be held responsible for identifying any or all such patent rights.
Any trade name used in this document is information given for the convenience of users and does not
constitute an endorsement.
For an explanation of the voluntary nature of standards, the meaning of ISO specific terms and expressions
related to conformity assessment, as well as information about ISO's adherence to the World Trade
Organization (WTO) principles in the Technical Barriers to Trade (TBT), see www.iso.org/iso/foreword.html.
This document was prepared by Technical Committee ISO/TC 69, Applications of statistical methods,
Subcommittee SC 7, Applications of statistical methods and related techniques: transformation for excellence
in process, product and services.
A list of all parts in the ISO 24481 series can be found on the ISO website.
Any feedback or questions on this document should be directed to the user’s national standards body. A
complete listing of these bodies can be found at www.iso.org/members.html.
iv
Introduction
Like many other sciences and technologies, a core idea of Six Sigma is data-driven decision making, that
is, exploiting the data from measurements or simulations at various points in the life cycle of a product
or certain service. It is widely accepted that statistics plays a very important role in such fields. In any
situation where statistics is applied, the analyst will follow a process, more or less formal, to reach findings,
recommendations, and actions based on the data. There are two phases in this process:
a) exploratory data analysis (EDA);
b) confirmatory data analysis (CDA).
In 1970s, EDA was promoted by John Tukey to encourage statisticians to explore the data, and possibly
formulate hypotheses. Generally speaking, the emphasis in EDA is on letting data talk by themselves or
hypothesis generation. In EDA efforts, the analyst searches for clues in the data that help identify theories
about underlying behaviours. In contrast, the focus of CDA is hypothesis testing and inference. CDA consists
of confirming these theories and behaviours. CDA follows EDA, and together they make up statistical
modelling during data analysis.
Since the rapid development of big data and data science, there is a clear need to evolve the mechanics of
Six Sigma both to accommodate the greater availability of data and to address the fact that, historically,
approaches to analysing data were overly concerned with hypothesis testing, to the detriment of the
hypothesis generation and discovery needed for improvement. Six Sigma teams often find warehouses of
data that are relevant to their efforts. Discovery is largely supported by the generation of hypotheses—
conjectures about relationships and causality. Today’s Six Sigma teams, and data analysts in the business
world in general, are often trained with a heavy emphasis on hypothesis testing, with comparatively little
emphasis given to hypothesis generation and discovery. They are often hampered in their problem-solving
and improvement efforts by the inability to exploit exploratory methods, which could enable them to make
more rapid progress, often with less effort.
Though the concept of EDA has evolved and is still evolving, the main idea of it is unchanged. In most
data analysis process, EDA is an approach for data analysis that employs a variety of techniques (mostly
graphical) to
— maximize insight into a data set;
— identify data issues (e.g. outliers, missing values, anomalies);
— uncover underlying structure, patterns and relationship;
— extract important variables;
— determine optimal factor settings;
— enhance problem-solving, root cause analysis, and continuous improvement efforts in process
optimization.
In recent times, we have seen incredible advances in visualization methods which is very important and
useful to EDA, supported by phenomenal increases in computing power. Some software packages have
incorporated such methods. Other software tools are continuously developed, with richer graphical
functionalities. These techniques and approaches should be further promoted within Six Sigma projects to
fully leverage the benefits of EDA.
v
International Standard ISO 24481-1:2026(en)
Statistical methods for implementation of Six Sigma —
Exploratory data analysis —
Part 1:
General methodology
1 Scope
This document describes the concept and general principles of EDA for categorical and numerical data. It
also provides some guidelines for conducting EDA and its place within Six Sigma projects. This document
focuses on the graphical tools of EDA.
It is applicable to organizations using manufacturing processes as well as service and transactional
processes.
2 Normative references
The following documents are referred to in the text in such a way that some or all of their content constitutes
requirements of this document. For dated references, only the edition cited applies. For undated references,
the latest edition of the referenced document (including any amendments) applies.
ISO 3534-1, Statistics — Vocabulary and symbols — Part 1: General statistical terms and terms used in
probability
3 Terms and definitions
For the purposes of this document, the terms and definitions given in ISO 3534-1 and the following apply.
ISO and IEC maintain terminology databases for use in standardization at the following addresses:
— ISO Online browsing platform: available at https:// www .iso .org/ obp
— IEC Electropedia: available at https:// www .electropedia .org/
3.1
population
totality of items under consideration
[SOURCE: ISO 3534-1:2006, 1.1, modified — Notes 1, 2 and 3 deleted.]
3.2
sample
subset of a population (3.1) made up of one or more sampling units
[SOURCE: ISO 3534-1:2006, 1.3, modified — Notes 1 and 2 deleted.]
3.3
observed value
obtained value of a property associated with one member of a sample (3.2)
[SOURCE: ISO 3534-1:2006, 1.4, modified — Notes 1 and 2 deleted.]
3.4
descriptive statistics
summary statistics that capture information about the shape, centre or spread of a variable or a distribution
3.5
frequency distribution
empirical relationship between classes and their number of occurrences or observed values (3.3)
[SOURCE: ISO 3534-1:2006, 1.60]
3.6
histogram
graphical representation of a frequency distribution (3.5) consisting of contiguous rectangles, each with base
width equal to the class width and area proportional to the class frequency
Note 1 to entry: Care needs to be taken for situations in which the data arises in classes having unequal class widths.
[SOURCE: ISO 3534-1:2006, 1.61]
3.7
boxplot
horizontal or vertical graphical representation of the five-number summary
[SOURCE: ISO 16269-4:2010, 2.16, modified — Notes 1, 2 and 3 deleted.]
3.8
scatter plot
graphical representation in a cartesian coordinate system where each data point is plotted as an individual
marker and the position of each marker is determined by an ordered pair of numerical values (x , y )
i i
representing two measured or observed variables for a single observation
3.9
Q-Q plot
scatter plot (3.8) for theoretical quantiles and empirical quantiles
3.10
static visualization
fixed graphical representation of data that does not allow for user interaction or dynamic change, presenting
information exactly as it was designed at the moment of creation
3.11
dynamic visualization
visualization that changes over time or updates based on underlying data changes
Note 1 to entry: Sometimes dynamic visualization can also be interactive and vice versa.
3.12
interactive visualization
graphical representation of data that allows users to manipulate, explore or query the information through
direct actions such as filtering, zooming, selecting or highlighting, enabling adjustment of the displayed
view without altering the underlying data interactively
3.13
critical thinking
(objective, sceptical and unbiased) analysis and evaluation of available facts, evidence, observations, and
arguments to form a reasoned judgement
Note 1 to entry: Not thinking critically is relying on tradition or authority without questioning and jumping to
conclusion.
3.14
exploratory data analysis
EDA
iterative process for uncovering patterns, relationships, outliers, and key insights that will guide further
analysis or modelling, using techniques related to statistics in particular graphical tools
3.15
data visualization
DV
graphical representations to convey information and insights clearly and effectively
Note 1 to entry: While it plays a key role in EDA, it is broader in scope and is often employed in storytelling and
presenting results to stakeholders.
4 Symbols and abbreviated terms
CDA confirmatory data analysis
DMAIC define, measure, analyse, improve and control
EDA exploratory data analysis
X , X , ., X sample or observed values or data
1 2 n
X covariate or causes
Y variable/outcomes of interest
5 Fundamental concepts
5.1 General
It is widely accepted that decisions or improvements are made based on the analysis of data available. In
practice, data may be observational or gained by design of experiment (also called experimental data).
There are some differences between these two kinds of data. Exploratory data analysis is mostly used for
observational data. In the era of very old years, one either collected the data by himself or got it from a
friend or colleagues. In either case, the result was a relatively small dataset, and one had a pretty good idea
of what it contained. Nowadays, the typical datasets one encounters are both much larger and provided
by other companies or organizations. Even in cases where a researcher is analysing their own data, the
dataset is often collected with the aid of electronic data acquisition systems or Internet web scrapers. As
a consequence, a very useful first step when one obtains a dataset is to look at it carefully to see what it
contains. This is even true if we have collected the data ourselves, since this preliminary exploration allows
us to determine whether the data we have is what we were expecting to have.
Exploratory data analysis (EDA) refers to a set of techniques originally developed by John Tukey to display
data in such a way that interesting features/variables will become apparent. Unlike classical methods which
usually begin with an assumed model for the data, EDA techniques are used to encourage the data to suggest
models that might be appropriate.
Most of the time, people will analyse data under some assumptions about the population distribution,
linear relationship between response and independent variables and so on based on experience. However,
some experience maybe very useful and some may be misleading or not convincing. We are encouraged to
formulate the assumptions or hypothesis via exploring data itself. In other words, exploratory data analysis
is the philosophy that data should first be explored without assumptions about probabilistic models, error
distributions, number of groups, relationships between the variables for the purpose of discovering what
they can tell us about the phenomena we are investigating.
By introducing EDA into the overall data analysis process, it enables for the inclusion of additional steps,
thus changing the process from a purely linear to a more dynamic and iterative one. We give the flowchart of
data analysis with or without considering EDA in the Figure 1 as follows:
a) Flowchart of data analysis without consider- b) Flowchart of data analysis with considering
ing EDA EDA
Figure 1 — Flowchart of data analysis without/with considering EDA
So, for example the addition of an exploration activity between data collection and analysis, will provide
additional insights on the data that might lead to requiring for collecting further data and/or questioning
the validity of the problem. Similarly, the addition of step between interpretation and conclusions enables
for further justification of the model and/or avoid making hasting conclusions.
5.2 Different types of data sets and possible challenges encountered within EDA
In most situations, our data sets will be arranged in a matrix of dimension n × p. Here, n represents the
number of observations we have in our sample, and p is the number of variables or dimensions. To be more
specific, we divide the data set into different types and give possible challenges encountered within EDA for
certain types.
Table 1 — Challenges in EDA for different types of data set
p - small p - large
Mainly dimensionality
n - small
challenge
Visualization and dimension-
n - large
Visualization
ality challenges
challenge
For large data set with n is bigger enough, visualization challenges may be partially solved by clustering or
subsampling methods which is reflected in Figure 1. For high dimensional data, dimensionality challenges
may be solved by dimension reduction (for example principal components analysis and projection pursuit)
or pairwise methods.
5.3 Relation and difference between data visualization (DV) and EDA
Visualization is a foundational technique for EDA to work, however not all data visualization methods belong
to EDA. EDA focuses on discovering insights from the data through statistical methods including graphical
tools, whereas data visualization transforms those insights into visual representations that can be easily
interpreted and communicated to others. The techniques in both areas overlap, especially when EDA uses
visual tools (like histograms or scatter plots), but data visualization is broader and aimed at preparation
and communication, while EDA is about exploration and understanding. Some detailed differences between
them are given in Annex A.
5.4 Principles and related steps for conducting EDA
Since EDA is a critical step in revealing the structure, patterns, and relationships in a dataset, here we give
key principles and typical steps for conducting EDA from a very broad view of data analysis. Firstly, key
principles are given as follows:
a) Do not jump on data analysis and modelling before understanding and visualizing it.
This is often called data “crunching”. This is the most important principle of the EDA, since the core idea of
the EDA is to explore the data before dealing with it.
b) Ask questions, formulate, and requestion hypotheses.
Adopt critical thinking to the problem in hand in order to understand it from different aspects. This will be
achieved by asking questions or formulating and requestioning hypotheses.
c) Make use of multiple techniques to improve the understanding of the data:
1) Emphasize data visualization techniques over numerical ones.
2) Identify patterns and relationships among variables of interest.
3) Use dynamic visualization tools to reveal hidden, complex, causal relationships, or both.
d) Handle anomalous and unexpected data.
Most real original data sets are complex. From statistical point of view, there may be some anomalous or
unexpected data such as missing data, outliers, incorrect data, incomplete data and so on. It is suggested
to consult statistical experts for advices to properly handle these situations. For example, one can refer to
ISO 16269-4 for detection and treatment of outliers under parametric population cases.
e) Use EDA insight to decide data transformation.
To enhance interpretability or meet analytical assumptions without distorting truth, one prefers to
transform the data. EDA’s graphical methods can be used to suggest the proper transformation and to justify
the transformation.
f) Apply EDA flexibly and iteratively.
As presented by Figure 1, the data analysis process with EDA steps may involve some iterative steps or
approaches. For example, if one cannot verify the population assumption by the current data with enough
confidence, one will require more data, if possible, for further exploring the distribution.
Secondly, there are some steps or procedures associated with the principles given above accordingly and the
general data analysis with EDA.
Step 1: understand the problem
...



