Relative Frequency and Combined Events
Estimating probability from data
The previous unit listed every outcome of an experiment to find probabilities exactly. Often, though, we do not know the exact probabilities and must estimate them from data, whether collected ourselves or given to us. The tool for this is relative frequency, and once we can estimate probabilities this way, we can combine events using the ideas of and, inclusive or, and exclusive or. This unit covers both: estimating from data, and reasoning about combined events.
Relative frequency
Relative frequency is simply how often an event actually happened, divided by the total number of trials. If a spinner is spun two hundred times and lands on red eighty-six times, the relative frequency of red is eighty-six out of two hundred, which is nought point four three. This is an estimate of the true probability of red, based on evidence rather than theory. The more trials we run, the closer the relative frequency tends to settle towards the true probability, which is why a large set of data gives a more trustworthy estimate than a handful of trials.
A two-way table of data
To combine events, a two-way table is the clearest starting point. Suppose one hundred students are surveyed about whether they play a sport and whether they play a musical instrument. The table might show thirty who do both, twenty-five who play sport only, twenty who play music only, and twenty-five who do neither. Every student falls into exactly one of these four cells, and the cells add to one hundred, so relative frequencies read straight off the table as estimates of probability.
And: the overlap
The first combination is and, meaning both events happen at once. The probability that a student plays sport and plays music is the count in the both cell over the total, thirty out of one hundred, or nought point three. The word and points to the overlap between the two groups, the students counted in both at the same time.
Inclusive or: one, the other, or both
Then there is or, and here lies the unit's key subtlety, because the everyday word or hides two different meanings. Inclusive or means one event or the other or both. The probability that a student plays sport or music in the inclusive sense includes everyone except those who do neither: thirty plus twenty-five plus twenty, giving seventy-five out of one hundred, or nought point seven five. There is a neat rule behind this: add the probability of sport and the probability of music, then subtract the both group, because adding the two totals counts the overlap twice. Fifty-five plus fifty minus thirty is again seventy-five.
Exclusive or: one or the other, not both
Exclusive or means one event or the other, but not both. The probability that a student plays sport or music in the exclusive sense counts only those in exactly one group: twenty-five who play sport only plus twenty who play music only, which is forty-five out of one hundred, or nought point four five. The difference between the two kinds of or is exactly the both group: inclusive or at seventy-five minus exclusive or at forty-five leaves the thirty who do both. Keeping these two meanings apart is the central skill of this unit.
A dice example
A simple dice example shows the same structure cleanly. Rolling one die, let event A be an even number, the set two, four and six, and event B be a number greater than three, the set four, five and six. A and B is the overlap, four and six, a probability of two out of six. A or B in the inclusive sense is two, four, five and six, a probability of four out of six. A or B in the exclusive sense, in exactly one set, is two and five, a probability of two out of six. As before, the inclusive probability minus the both probability gives the exclusive one. Reading data this way, and being precise about which kind of or is meant, turns a table of counts into a clear set of estimated probabilities for combined events.