Class 9Maths · StatisticsFull chapter

Statistics

The whole chapter in one place — read it, then test yourself. Clear notes, formula sheet, a practice quiz, and worked NCERT solutions & PYQs.

Collection and Presentation of Data

Quick answer Data is first collected as primary or secondary data, then organised into a frequency distribution table using tally marks, class intervals, class size and class marks so that large amounts of information become easy to read.

Statistics is the branch of mathematics that deals with the collection, presentation and interpretation of numerical facts, called data. Data can be of two kinds. Primary data is collected by the investigator personally for a specific purpose, such as recording the heights of students in your own class. Secondary data is data that already exists, collected by someone else for another purpose, such as figures published in a newspaper or a government report.

Data collected in its original form, before any arrangement, is called raw data (or ungrouped data). Raw data is hard to read at a glance, so it is organised using tally marks into a frequency distribution table, which shows how many times (the frequency) each value occurs.

Worked Example. The number of goals scored by a school football team in 12 matches were: 1, 2, 0, 1, 3, 2, 1, 0, 2, 1, 1, 2. Arrange this as an ungrouped frequency distribution table.

  • Goals = 0: frequency 2
  • Goals = 1: frequency 5
  • Goals = 2: frequency 4
  • Goals = 3: frequency 1

Check: 2 + 5 + 4 + 1 = 12, which matches the total number of matches, so the table is correct.

When the range of data (the difference between the maximum and minimum value) is large, individual values are grouped into class intervals of equal width, such as 0-10, 10-20, 20-30, and so on. Each class interval has a lower limit and an upper limit. In the commonly used exclusive form, the upper limit of a class is not included in that class — it is counted in the next class. The mid-value of a class is called its class mark, and it represents that class on a graph.

Worked Example. For the class interval 20-30, the class size is 30 − 20 = 10, and the class mark is (20 + 30) ÷ 2 = 25.

Class Mark Class Mark = (Upper Class Limit + Lower Class Limit) ÷ 2
Class Size Class Size = Upper Class Limit − Lower Class Limit
Range Range = Maximum value − Minimum value
Remember
  • Primary data is collected first-hand by the investigator; secondary data is collected by someone else.
  • Raw data must be organised into a frequency distribution table before it can be studied easily.
  • In grouped data, class size = upper limit − lower limit, and class mark = (upper limit + lower limit) ÷ 2.
  • In the exclusive form of class intervals, a value equal to the upper limit is placed in the next class.
  • The sum of all frequencies in a frequency table must equal the total number of observations.

Graphical Representation: Bar Graphs and Histograms

Quick answer Bar graphs display categorical or discrete data as bars of equal width with gaps between them, while histograms display continuous grouped data as adjoining bars with no gaps, where the height represents the frequency.

A bar graph is a graphical way of showing data using rectangular bars of equal width, with uniform gaps between consecutive bars. The height of each bar is proportional to the frequency or value it represents. Bar graphs are used for categorical data (such as blood groups) or discrete numerical data (such as the number of goals scored).

Worked Example. Using the goals data from before (0 goals: 2 matches, 1 goal: 5 matches, 2 goals: 4 matches, 3 goals: 1 match), a bar graph would have four bars of equal width along the horizontal axis labelled 0, 1, 2, 3, with heights 2, 5, 4 and 1 on the vertical (frequency) axis, each bar separated from the next by an equal gap.

A histogram is used only for continuous grouped data arranged in class intervals of equal width. Unlike a bar graph, the rectangles of a histogram touch each other, with no gap, because the class intervals are continuous. The height of each rectangle is simply the frequency of that class, and the width of each rectangle equals the class size, which is the same for every class.

Worked Example. For the grouped frequency table with equal class size 10 — 0-10: 3, 10-20: 5, 20-30: 6, 30-40: 7, 40-50: 4 — a histogram is drawn with five touching rectangles over these classes, with heights equal to 3, 5, 6, 7 and 4 respectively, since all class widths are equal.

Histogram rectangle height Height of rectangle = Frequency of the class
Histogram rectangle width Width of each rectangle = Class size (equal for every class)
Remember
  • Bar graphs are used for discrete/categorical data; bars have equal width with gaps between them.
  • Histograms are used only for continuous grouped data; rectangles touch each other with no gaps.
  • In a histogram, the height of each rectangle equals the frequency of that class, and all class widths are equal.
  • The horizontal axis of a histogram must show class intervals of a continuous variable.
  • Histograms with unequal (varying) class widths are not part of the current CBSE-rationalised Class 9 syllabus; Class 9 histograms use equal class widths only.

Frequency Polygons

Quick answer A frequency polygon is a line graph obtained by plotting the class marks against frequencies and joining the points with straight lines, closed at both ends using imaginary classes of zero frequency.

A frequency polygon is another way of representing grouped frequency data using a line graph. It can be drawn in two ways: (i) by first drawing a histogram, then joining the midpoints of the tops of consecutive rectangles with straight lines, or (ii) directly, by plotting the class mark of each class against its frequency and joining the plotted points with straight line segments.

To complete the polygon so that it forms a closed figure with the horizontal axis, one extra class of zero frequency is imagined immediately before the first class and one extra class of zero frequency immediately after the last class. The polygon is then joined to the class marks of these two imaginary classes on the horizontal axis.

Worked Example. For the grouped frequency table 0-10: 3, 10-20: 5, 20-30: 6, 30-40: 7, 40-50: 4, the class marks are 5, 15, 25, 35 and 45. To close the polygon, an imaginary class (−10)-0 (class mark −5, frequency 0) is added before, and an imaginary class 50-60 (class mark 55, frequency 0) is added after. The frequency polygon then joins the points (−5, 0), (5, 3), (15, 5), (25, 6), (35, 7), (45, 4) and (55, 0) in order with straight lines.

Frequency polygons are especially useful for comparing two or more sets of grouped data on the same axes, since several polygons can be drawn on one graph while histograms cannot easily overlap.

Point plotted Each plotted point = (Class Mark, Frequency)
Closing the polygon One class of frequency 0 is added immediately before the first class and one immediately after the last class
Remember
  • A frequency polygon plots class marks (x-axis) against frequencies (y-axis).
  • It can be drawn directly, or by joining the midpoints of the tops of histogram bars.
  • Two imaginary classes of zero frequency are added, one before the first class and one after the last, to close the figure.
  • Frequency polygons are useful for comparing two or more distributions on the same graph.

Mean of Ungrouped Data

Quick answer The mean (average) of raw data is the sum of all observations divided by the number of observations; for a frequency table of discrete values, the mean is the sum of (frequency × value) divided by the total frequency.

The mean (also called the average), denoted (read as "x-bar"), is the most common measure of central tendency. For raw (ungrouped) data with n observations x1, x2, ..., xn, the mean is found by adding all the observations and dividing by the number of observations.

Worked Example 1. A batsman scored 45, 60, 32, 78 and 50 runs in 5 innings. Find his mean score.

Sum of runs = 45 + 60 + 32 + 78 + 50 = 265.

Mean = 265 ÷ 5 = 53 runs.

When the same value occurs several times, the data can be given as an ungrouped frequency distribution — a table of distinct values xi along with how many times each occurs, fi. In this case, the mean is found using x̄ = Σ(fi × xi) ÷ Σfi, where Σfixi means "the sum of each value multiplied by its frequency", and Σfi is the total number of observations.

Worked Example 2. The number of children in 10 families of a locality is given below.

  • Children = 0: frequency 1
  • Children = 1: frequency 2
  • Children = 2: frequency 4
  • Children = 3: frequency 2
  • Children = 4: frequency 1

Check: Σfi = 1 + 2 + 4 + 2 + 1 = 10 families, which matches the given total.

Σ(fi × xi) = (0×1) + (1×2) + (2×4) + (3×2) + (4×1) = 0 + 2 + 8 + 6 + 4 = 20.

Mean = 20 ÷ 10 = 2 children per family.

Mean of raw data x̄ = (x1 + x2 + ... + xn) ÷ n = Σxi ÷ n
Mean of ungrouped frequency distribution x̄ = Σ(fi × xi) ÷ Σfi
Remember
  • Mean = sum of all observations ÷ number of observations.
  • For an ungrouped frequency table, mean = Σ(fi × xi) ÷ Σfi.
  • Always verify that Σfi equals the total number of observations before calculating the mean.
  • The mean can be a decimal even if all the original observations are whole numbers.
  • The mean is affected by every value in the data, including extreme (very high or very low) values.

Median and Mode of Ungrouped Data

Quick answer The median is the middle-most value of data arranged in order, while the mode is the value that occurs most often; both are measures of central tendency that, unlike the mean, are not distorted much by extreme values.

The median is the value of the middle-most observation when the raw data is arranged in ascending (or descending) order. To find it: arrange the n observations in order; if n is odd, the median is the value of the ((n + 1) ÷ 2)th observation; if n is even, the median is the average of the (n ÷ 2)th observation and the (n ÷ 2 + 1)th observation.

Worked Example 1 (odd n). The weights (in kg) of 7 students are 42, 38, 45, 33, 40, 36, 50. Arranged in ascending order: 33, 36, 38, 40, 42, 45, 50. Here n = 7 (odd), so the median is the ((7+1)÷2) = 4th observation, which is 40 kg.

Worked Example 2 (even n). The marks of 6 students are 55, 62, 48, 70, 58, 65. Arranged in ascending order: 48, 55, 58, 62, 65, 70. Here n = 6 (even), so the median is the average of the 3rd and 4th observations: (58 + 62) ÷ 2 = 60.

The mode of raw data is simply the observation that occurs the maximum number of times, that is, the value with the highest frequency. Data can have one mode (unimodal), more than one mode, or no repeated value at all.

Worked Example 3. The shoe sizes of 10 students are 6, 7, 8, 7, 6, 7, 9, 7, 8, 6. Counting how often each size occurs: size 6 occurs 3 times, size 7 occurs 4 times, size 8 occurs 2 times, and size 9 occurs 1 time. Since size 7 occurs the most often, the mode is 7.

The median and the mode are both measures of central tendency, like the mean, but unlike the mean, they are not affected much by one or two extremely high or low values in the data.

Median (n odd) Median = value of ((n+1) ÷ 2)th observation
Median (n even) Median = [(n ÷ 2)th observation + (n ÷ 2 + 1)th observation] ÷ 2
Mode Mode = observation with the highest frequency
Remember
  • First arrange the data in ascending order before finding the median.
  • For odd n, median = value of the ((n+1)/2)th term; for even n, median = average of the (n/2)th and (n/2 + 1)th terms.
  • Mode = the observation with the highest frequency; data can have one, more than one, or no mode.
  • Unlike the mean, the median and mode are not greatly affected by extreme (very high or low) values.
  • Mean, median and mode can all be different values for the same set of data.

The formula sheet

Every formula in this chapter, in one place — screenshot it before your exam.

Class Mark = (Upper Class Limit + Lower Class Limit) ÷ 2
Class Mark
Class Size = Upper Class Limit − Lower Class Limit
Class Size
Range = Maximum value − Minimum value
Range
Height of rectangle = Frequency of the class
Histogram rectangle height
Width of each rectangle = Class size (equal for every class)
Histogram rectangle width
Each plotted point = (Class Mark, Frequency)
Point plotted
One class of frequency 0 is added immediately before the first class and one immediately after the last class
Closing the polygon
x̄ = (x1 + x2 + ... + xn) ÷ n = Σxi ÷ n
Mean of raw data
x̄ = Σ(fi × xi) ÷ Σfi
Mean of ungrouped frequency distribution
Median = value of ((n+1) ÷ 2)th observation
Median (n odd)
Median = [(n ÷ 2)th observation + (n ÷ 2 + 1)th observation] ÷ 2
Median (n even)
Mode = observation with the highest frequency
Mode

Test yourself

Tap an answer to check it instantly — you'll see why it's right, and what to revise if it isn't.

0 correct · 0/12 answered
Q1 Collection of Data easy

Which of the following is an example of primary data?

Q2 Presentation of Data easy

For the class interval 10-20, what is the class mark?

Q3 Presentation of Data easy

What is the class size of the interval 25-40?

Q4 Graphical Representation medium

Which statement correctly describes a histogram?

Q5 Graphical Representation medium

In a frequency polygon, what are plotted on the graph and then joined by straight lines?

Q6 Mean easy

What is the mean of the first five natural numbers, 1, 2, 3, 4 and 5?

Q7 Median easy

Find the median of the data: 12, 15, 9, 21, 18.

Q8 Mode easy

Find the mode of the data: 2, 3, 3, 5, 3, 7, 2.

Q9 Mean hard

In an ungrouped frequency distribution, the observations 10, 20 and 30 occur with frequencies 2, 3 and 5 respectively. Find the mean.

Q10 Graphical Representation medium

Which of the following is a key feature of a bar graph?

Q11 Presentation of Data hard

Class intervals are formed in the exclusive form as 0-10, 10-20, 20-30, and so on. In which class interval does the value 20 lie?

Q12 Median medium

When the number of observations n is even, how is the median calculated?

NCERT solutions & previous-year questions

Step-by-step model answers — tap a question to reveal the full solution.

NCERT questions 6

1 The blood groups of 30 students of a class were recorded as: A, B, O, O, AB, O, A, O, B, A, O, B, A, O, O, A, AB, O, A, A, O, O, AB, B, A, O, B, A, B, O. Represent this data in the form of a frequency distribution table. Which blood group is the most common and which is the rarest among these students?Presentation of Data

Counting the occurrences of each blood group in the given list of 30 students:

  • Blood group A: 9 students
  • Blood group B: 6 students
  • Blood group O: 12 students
  • Blood group AB: 3 students

Check: 9 + 6 + 12 + 3 = 30, which matches the total number of students, so the frequency distribution table is correct.

Since blood group O has the highest frequency (12), it is the most common blood group. Since blood group AB has the lowest frequency (3), it is the rarest blood group.

2 The weekly pocket money (in Rs) of 20 students is: 10, 15, 22, 28, 5, 32, 18, 25, 40, 12, 20, 35, 8, 27, 33, 19, 24, 30, 45, 16. Construct a grouped frequency distribution table using class size 10, starting from 0-10 (in the exclusive form).Presentation of Data

Using class intervals of size 10 in the exclusive form and tallying each value:

  • 0-10: 5, 8 → frequency 2
  • 10-20: 10, 15, 18, 12, 19, 16 → frequency 6
  • 20-30: 22, 28, 25, 20, 27, 24 → frequency 6
  • 30-40: 32, 35, 33, 30 → frequency 4
  • 40-50: 40, 45 → frequency 2

Check: 2 + 6 + 6 + 4 + 2 = 20, which matches the total number of students, so the table is correct.

3 Find the mean of the first 10 odd natural numbers.Mean

The first 10 odd natural numbers are 1, 3, 5, 7, 9, 11, 13, 15, 17, 19.

Sum = 1 + 3 + 5 + 7 + 9 + 11 + 13 + 15 + 17 + 19 = 100.

Mean = Sum ÷ Number of observations = 100 ÷ 10 = 10.

4 Find the median of the following data: 25, 34, 31, 23, 22, 26, 35, 28, 20, 32.Median

Arranging the data in ascending order: 20, 22, 23, 25, 26, 28, 31, 32, 34, 35.

Here, the number of observations n = 10, which is even. So the median is the average of the (n/2)th and (n/2 + 1)th observations, that is, the 5th and 6th observations.

5th observation = 26, 6th observation = 28.

Median = (26 + 28) ÷ 2 = 27.

5 The marks (out of 10) obtained by 15 students in a class test are: 4, 6, 7, 5, 3, 9, 6, 5, 7, 6, 4, 8, 6, 6, 5. Find the mode of this data.Mode

Counting how many times each mark occurs:

  • Mark 3: 1 time
  • Mark 4: 2 times
  • Mark 5: 3 times
  • Mark 6: 5 times
  • Mark 7: 2 times
  • Mark 8: 1 time
  • Mark 9: 1 time

Check: 1 + 2 + 3 + 5 + 2 + 1 + 1 = 15, which matches the total number of students.

Since the mark 6 occurs the maximum number of times (5 times), the mode of the data is 6.

6 Using the grouped frequency distribution table obtained in the pocket-money question above (0-10: 2, 10-20: 6, 20-30: 6, 30-40: 4, 40-50: 2), state the class marks that would be plotted to draw a frequency polygon, and the two imaginary classes needed to close the polygon at both ends.Graphical Representation

The class marks of the given classes are:

  • 0-10 → class mark 5, frequency 2
  • 10-20 → class mark 15, frequency 6
  • 20-30 → class mark 25, frequency 6
  • 30-40 → class mark 35, frequency 4
  • 40-50 → class mark 45, frequency 2

To close the frequency polygon on both ends, one imaginary class of zero frequency is taken immediately before the first class and one immediately after the last class:

  • Imaginary class (−10)-0 → class mark −5, frequency 0
  • Imaginary class 50-60 → class mark 55, frequency 0

The frequency polygon is drawn by plotting the points (−5, 0), (5, 2), (15, 6), (25, 6), (35, 4), (45, 2) and (55, 0), and joining them in order with straight line segments.

Previous-year board questions 4

Q1 The class mark of a class interval is 10 and its class size is 6. Find the class limits of this interval. CBSE 2020 1 mark

Let the lower limit be l and the upper limit be u.

Class mark = (l + u) ÷ 2 = 10, so l + u = 20.

Class size = u − l = 6.

Adding the two equations: 2u = 26, so u = 13. Then l = 20 − 13 = 7.

The class limits are 7 and 13, that is, the class interval is 7-13.

Q2 Find the mean of the first 6 multiples of 4. CBSE 2019 2 marks

The first 6 multiples of 4 are 4, 8, 12, 16, 20, 24.

Sum = 4 + 8 + 12 + 16 + 20 + 24 = 84.

Mean = Sum ÷ Number of observations = 84 ÷ 6 = 14.

Q3 Find the median of the following data: 12, 17, 3, 14, 5, 8, 7, 15. CBSE 2022 3 marks

Arranging the data in ascending order: 3, 5, 7, 8, 12, 14, 15, 17.

Number of observations n = 8, which is even, so the median is the average of the (n/2)th and (n/2 + 1)th observations, that is, the 4th and 5th observations.

4th observation = 8, 5th observation = 12.

Median = (8 + 12) ÷ 2 = 10.

Q4 The marks (out of 50) obtained by 25 students of a class in a unit test are: 12, 25, 33, 8, 45, 19, 27, 38, 5, 41, 22, 30, 15, 48, 36, 10, 29, 43, 17, 34, 26, 39, 6, 21, 32. (i) Construct a grouped frequency distribution table using class size 10, starting from 0-10. (ii) State which class interval has the maximum frequency. CBSE 2023 5 marks

(i) Using class intervals of size 10 in the exclusive form and tallying each value:

  • 0-10: 8, 5, 6 → frequency 3
  • 10-20: 12, 19, 15, 10, 17 → frequency 5
  • 20-30: 25, 27, 22, 29, 26, 21 → frequency 6
  • 30-40: 33, 38, 30, 36, 34, 39, 32 → frequency 7
  • 40-50: 45, 41, 48, 43 → frequency 4

Check: 3 + 5 + 6 + 7 + 4 = 25, which matches the total number of students, so the table is correct.

(ii) The class interval 30-40 has the maximum frequency (7), meaning the largest number of students scored marks in this range.

Part of Priodemy for School

Interactive Maths & Science — free with every school on Priodemy EduSuite. Explore more chapters and labs on the Priodemy for School hub.

Ask AI