Data, and the Story Hidden Inside It
Quick answer Raw data is only a heap of numbers. Sorting it into a frequency table, squeezing it into one representative value and drawing it as dots and lines is how the story comes out.
Every day numbers arrive in bunches. The runs a batter scored in her last ten matches. The marks of forty students in a unit test. The price of one kilogram of onions on the first day of every month. The rainfall of a district in June, July and August. A bunch of numbers collected about one thing is called data, and each single number inside it is called an observation.
Collected raw, data is just a heap. Here are the marks out of 5 that twenty students scored in a class test, written down in the order the answer sheets came in:
3, 4, 5, 2, 3, 4, 4, 5, 1, 3, 4, 2, 5, 3, 4, 4, 2, 4, 5, 3
You can stare at that line as long as you like and it will tell you almost nothing. Now count how many times each mark appears and set it out in a frequency table. The frequency of a value is simply the number of times that value turns up.
Marks (out of 5) Number of students
1 1
2 3
3 5
4 7
5 4
Total 20The heap has become five tidy rows. Notice the check built into the last line: the frequencies must add up to the number of students, and 1 + 3 + 5 + 7 + 4 = 20, so no answer sheet has been lost or counted twice. Always do that addition.
Straight away the table shows you things the list hid. The most common mark is 4, because it has the biggest frequency, 7. A value that occurs most often is called the mode. The marks run from a lowest of 1 to a highest of 5, so the range is 5 − 1 = 4, which tells you how widely spread the class is.
But a table is still a list, and two questions come up again and again with data. This chapter answers both.
The first question is: what is a typical value? If a parent asks how the class did, you do not want to read out twenty numbers; you want one number that stands in for all of them. Such a number is called a representative value. The two you will use here are the mean, which is the ordinary average, and the median, which is the value standing in the middle.
The second question is: how is the quantity behaving? Is it climbing, falling, steady, or jumping about? Is everybody crowded together with one odd exception? A column of figures will not tell you, but a picture will. This chapter uses two pictures: the dot plot, which puts one dot for every single observation, and the line graph, which plots a quantity against time and joins the points. Dots and lines really are the tools that let a page of figures tell its tale.
One warning before you start. A representative value throws information away on purpose. The moment you say the class average is 3.5, you have deliberately forgotten who scored 1 and who scored 5. That is exactly what a summary is for, but it is also its danger, and a good part of this chapter is about knowing which summary deserves your trust.
- Data is a set of numbers collected about one thing, and each number inside it is called an observation.
- A frequency table records how many times each value occurs, and the frequencies must add up to the total number of observations.
- Range = highest observation − lowest observation, so marks running from 1 to 5 have a range of 4.
- The mode is the value with the highest frequency; in these 20 marks it is 4, which occurred 7 times.
- A representative value such as the mean or the median stands in for the whole set and deliberately forgets the details.
- Dot plots and line graphs reveal shape, crowding and direction that a plain list of numbers hides.
