Comparing Data Sets
A data set tells a story
A data set is a collection of answers to one question: how each student travels to school, how many goals each player scored, how tall everyone in the class is. On its own a long list of values says little, so we display it — as a column graph for categories, or a dot plot for numbers — and read the story it tells. Some variables are categories, like sport or travel method, while others are numbers, like height or score. This unit reads a single data set, then learns to compare two using three plain ideas: the mode, the range and the shape.
The mode is the most common value
The mode is the value that appears most often. In a column graph it is simply the tallest column; in a list of numbers it is the value that repeats the most. If five students walk, eight come by car and ten catch the bus, the mode is the bus, because it is the most common answer. The mode works for categories and for numbers alike, which makes it the first thing to read from almost any data set. It answers a natural question: what is the most usual result here?
The range measures spread
Where the mode points to the most common value, the range describes how spread out the numbers are. It is the largest value minus the smallest, a single number that captures the whole sweep of the data. Scores from four to twelve have a range of eight; scores from two to five have a range of three, and are clearly more tightly bunched. The range only makes sense for numerical data, where subtracting is meaningful, and it is the simplest way to say whether results are close together or far apart.
Comparing two data sets
The real power of these ideas appears when two groups are placed side by side. Two classes might share the same most common score yet spread very differently, one tightly clustered and one ranging widely. Comparing their modes says which result was most usual in each; comparing their ranges says which group was more variable. Reading the two displays together, rather than one at a time, is what lets you make a fair statement about how the groups differ, instead of guessing from a jumble of numbers.
The shape of a distribution
Beyond the mode and the range, a distribution has an overall shape. The values might pile up in the middle and tail off evenly on both sides, a symmetric shape; they might bunch at one end with a long tail stretching the other way, a skewed shape; or they might sit fairly level across the whole range. Shape is read from the outline of the graph, and it adds what mode and range alone cannot: a picture of how the data is distributed, not just where it centres or how far it reaches.
Reading a summary
Mode, range and shape are three different questions about the same data. The mode asks which value is most common; the range asks how far the data spreads; the shape asks how the values are arranged across that spread. Knowing which one answers a given question is the heart of interpreting data well: a question about the most usual result wants the mode, a question about consistency wants the range, and a question about symmetry or skew wants the shape. Keeping the three apart keeps your reading of a data set clear.
From one data set to many
With mode, range and shape in hand, a data set stops being a list and becomes something you can describe and compare. These three summaries let you say what was most common, how spread the results were, and what the overall pattern looked like, for one group or for two placed side by side. From here the same ideas grow into the mean and median, into larger surveys, and into the statistical investigations where data is gathered, displayed and questioned to settle real arguments.