AC9M7ST02 · Year 7 · Statistics

Describing data distributions

ACARA v9 CONTENT DESCRIPTION create different types of numerical data displays including stem-and-leaf plots using software where appropriate; describe and compare the distribution of data, commenting on the shape, centre and spread including outliers and determining the range, median, mean and mode

Once data has been collected, the next task is to make sense of it. A long list of numbers tells you very little until you can see its overall pattern, its centre and its spread. This year you learn to display numerical data and to describe and compare distributions using their shape, their range and a measure of centre, turning raw figures into a clear summary you can reason about.

Describing a distribution well means answering three questions: what shape is the data, where is its centre, and how spread out is it. Together these give a faithful picture of the whole data set, far more useful than any single number on its own.

The shape of the data

The shape of a distribution is the overall pattern you see when the data is plotted. A simple dot plot, with one dot for each value, reveals this immediately. The data might pile up in the middle and tail off at both ends, forming a symmetric mound, or it might lean to one side, or spread out evenly. Reading the shape is the first step, because it shows where the data clusters and where it is sparse.

The shape of a distribution
A dot plot reveals whether data clusters, spreads out or leans to one side.
A distribution is the pattern of how data is spread. A symmetric mound: the data clusters around a central value with fewer points at each end. Switch the shape to see how the same nine values can sit in very different patterns.

Different displays suit different purposes. Dot plots show every value and suit small data sets, while stem-and-leaf plots organise larger sets while keeping the individual figures visible. The right display makes the shape of the distribution easy to see, which is why choosing and creating a suitable display is part of the skill.

Centre and spread

Beyond shape, two summaries capture a distribution: a measure of centre and a measure of spread. The centre, such as the mean or the median, gives a single typical value that represents the data. For the values clustered around 6 in the example, a centre of about 6 tells you what is typical. The mean balances the data, while the median sits at its middle when the values are ordered.

Centre and spread
Summarise data by where it centres and how widely it spreads.
Two numbers summarise a distribution well: a measure of centre and a measure of spread. All three of these data sets share the same centre, a mean of 6, yet their spread differs: this one has range 3 to 9, a range of 6. Switch the spread and watch the centre hold while the range widens or tightens.

The spread tells you how varied the data is, and the simplest measure is the range, the largest value minus the smallest. A small range means the data is tightly bunched; a large range means it is widely scattered. This is what makes spread essential for comparison: two data sets can share the same centre yet behave very differently. If two classes have the same mean test score but one has a far larger range, that class has much more varied results. Describing data by its shape, centre and spread, and comparing distributions on all three, is how statistics turns a column of numbers into genuine understanding.

Reading a stem-and-leaf plot

A stem-and-leaf plot keeps every value visible while grouping the data. The tens digit is the stem, the units digit is the leaf, and the leaves are listed in order so the median can be read straight off the plot.

Stem-and-leaf plot
A stem-and-leaf plot groups data while keeping every value readable, so the median is easy to find.
A stem-and-leaf plot splits each value into a stem, the tens digit, and a leaf, the units digit, so the whole data set stays visible while it is grouped. With n = 10 values the median is 23.5, found as mean of the two middle values 23 and 24. Switch the data set to read a different median straight off the ordered leaves.

Three measures of centre

The mean, median and mode each describe a typical value in a different way. The mean balances all the values, the median sits at the middle of the ordered data, and the mode is the value that occurs most often, so they need not agree.

Mean, median and mode
Three measures of centre, each describing the typical value in a different way.
The three common measures of centre can each represent a data set. For this set the mean is 7.5, the median is 7 and the mode is 7. They can differ: the mean is pulled by every value, the median sits at the middle, and the mode is the most frequent. Switch the set or highlight a measure to compare them.

Outliers and the mean

An outlier is a value far from the rest of the data. Because the mean adds in every value, a single outlier can pull it noticeably, while the median, which depends only on the middle, barely moves.

How an outlier pulls the mean
One extreme value can drag the mean while the median holds steady.
An outlier is a value far from the rest of the data. Without the outlier the mean and median are both 7.0. Adding the value 20 lifts the mean to 7.00 while the median stays 7, because the mean uses every value but the median only depends on the middle. Toggle the outlier to see the mean move a lot while the median resists.

Teaching tip: collect a small real data set with the student, such as the number of letters in each family member name, and plot it as a dot plot together. Asking what is typical and how spread out the values are introduces centre and spread naturally, grounded in data they helped create.

Emphasise that centre and spread answer different questions. A common habit is to report only an average and stop. Prompt the student to also ask how varied the data is, so that describing a distribution always covers both the typical value and the variation around it.

Builds on: Range, median, mean and mode (AC9M7ST01). That unit collected and organised data; this unit describes the shape, centre and spread of a distribution.
Quick self-check
1. What does the shape of a data distribution show you?
2. A measure of centre, such as the mean or median, tells you
3. For the data 3, 5, 6, 6, 9, what is the range?
4. Two classes have the same mean score, but one has a much larger range. This tells you
5. Why describe data with both a centre and a spread?
Teaching pack: free to printReady-to-teach plans, student sheets, cut-outs and answers for this unit. Print or save as PDF.