A researcher studies how the number of hours of sleep affects students' exam scores. Which variable is the independent variable?
Strand 4 · Making Sense of and Using Data
Mathematics Year 3 Learner Material, Section 5: Data Handling And Probability
In Year One Section 8, we learnt about the types of data and the methods for collecting data. We went further to learn how to organise and present data to make it more meaningful. Finally, we learnt about the sample space for simple and compound events.
In Year Two Section 8, we studied the different types of data, the measures of central tendency and dispersion and more ways of data presentation, including drawing a cumulative frequency curve. In all of these activities, we were interested in studying the effect of one variable, say, age, mass or height, on given data.
In this section, we will be interested in the relationship between two variables for a given set of data. That is, we will study bivariate data, rather than univariate data. As in the first two years, we will work on a project to collect and analyse data collected from our communities.
KEY IDEAS
• Distinguishing between univariate and bivariate data.
• Investigating the relationship between two variables.
• Drawing a scatter diagram to determine the relationship between two variables.
• Using the scatter diagram to determine the relationship between the variables.
• Applying the skills we have learnt to real-life situations, for example, solving problems in our communities.
We have learnt that data handling means collecting, organising, and interpreting information (data) to make it easier to understand. Data can be qualitative (descriptive, like colours or names) or quantitative (numerical, like test scores or ages).
To make sense of raw data, we usually:
1. Organise it in a frequency table to show how often each score appears.
2. Represent it using graphs such as bar charts or histograms.
3. Summarise it with measures of central tendency (mean, median, mode) and spread (range, standard deviation).
• Mean is the average value.
• Median is the middle value when data is ordered.
• Mode is the most frequent value.
• Range is the difference between the highest and lowest scores.
• Standard deviation tells us how spread out the scores are around the mean (a small standard deviation = data is close together, large = more spread out).
By doing this, we can judge whether the average (mean) truly reflects the data.
We have also learnt that probability is the branch of mathematics that deals with the likelihood of an event happening. It is measured as a number between 0 and 1.
• 0 means the event is impossible.
• 1 means the event is certain.
• Values in between represent different levels of chance.
To calculate probability:
P(Event) = For example, if we have a bag with 4 red, 3 blue and 5 green bottle tops (total = 12), the probability of drawing a red one is:
When two events are considered, we need to decide if they are:
• Independent events (the outcome of the first does not affect the second), for example, tossing a coin twice.
• Dependent events (the outcome of the first affects the second), for example, drawing one bottle top without replacing it, then drawing another.
In this activity, we will calculate probabilities for one draw and for two draws, then decide whether the events are dependent or independent.
Activity 5.1 Data Handling and Analysis
Instructions The table below shows the scores of 15 students in a mathematics test out of 20.
Scores 12, 15, 14, 13, 16, 12, 18, 15, 17, 13, 14, 20, 12, 16, 19 Tasks
1. Identify the type of data represented.
2. Construct a frequency table for the data.
3. Draw a bar chart to represent the data.
4. Calculate the:
• Mean
• Median
• Mode
5. Calculate the range and standard deviation of the scores.
6. Briefly explain whether the mean is a reliable measure of central tendency in this case.
Activity 5.2 Probability Concepts
Instructions A bag contains 4 red, 3 blue and 5 green bottle tops. One bottle top is drawn at random.
Tasks
1. What is the probability that the bottle top drawn is:
a. Red?
b. Blue?
c. Not green?
2. If one bottle top is drawn and not replaced, and a second one is drawn:
a. What is the probability that both bottle tops are red?
b. What is the probability that the first is red and the second is blue?
3. State whether the events are dependent or independent and explain your reasoning.
In our everyday life, we observe that certain quantities are related. For example:
• The amount of rainfall and crop yield
• A student’s study time and exam scores
• The speed of a car and the distance it covers Here, we will learn how to identify, analyse and represent relationships between two variables using tables and graphs.
Identifying Variables in a Given Data Set
A variable is any characteristic, number or quantity that can be measured or counted and can vary from one individual or observation to another. A variable is any quantity that can change or take different values. When doing a study or investigation, we often look at how one variable affects another.
Types of Variables
1. Independent Variable (Explanatory Variable): This is the variable that is deliberately changed or controlled in an experiment. The independent variable is the cause. Think of it as: “What I change on purpose.”
Example: In an experiment to study plant growth, the amount of water given to the plant is the independent variable.
2. Dependent Variable (Response Variable): This is the outcome that is measured based on changes in the independent variable.
Example: The height of the plant (dependent variable) depends on the amount of water (independent variable).
3. Control Variables: These are kept constant to ensure a fair test.
Example: Soil type, sunlight and pot size in the plant growth experiment.
Use these scenarios to further your understanding of the types of variables Scenario 1: Investigating the Effect of Study Time on Mathematics Exam Scores Mrs. Antwi, a teacher in a senior high school in Accra, wants to know if the amount of time learners spend studying mathematics affects their exam scores.
• Independent Variable: Time spent studying mathematics (e.g., 2 hours, 4 hours, 6 hours)
• Dependent Variable: Mathematics exam score (e.g., 65%, 75%, 90%) The teacher changes the amount of study time (independent variable), and then measures the exam score (dependent variable).
Scenario 2: Studying the Impact of Feeding on Learners’ Concentration in Class At a boarding school in Kumasi, a research group investigates whether learners who eat breakfast are more attentive in morning lessons.
• Independent Variable: Whether or not the learner ate breakfast
• Dependent Variable: Level of concentration in class (measured by participation, responses, or test) In conclusion
a. The independent variable is like the input or cause.
b. The dependent variable is the output or result.
c. We study how changing one variable (independent) affects another (dependent).
d. Always ask yourself:
• “What is being changed?” → Independent
• “What is being measured?” → Dependent
Activity 5.3 Identify the Variables
Read each of the scenarios below and identify the following.
• The independent variable
• The dependent variable Scenarios
1. A health researcher is investigating whether the number of hours teenagers sleep affects their academic performance.
2. A PE teacher observes whether the number of laps run during practice affects a learner’s heart rate.
3. School cafeteria staff are studying whether the number of learners eating from the canteen affects the amount of food wastage.
4. A scientist studies how temperature affects the rate of a chemical reaction.
5. A teacher investigates whether more practice questions lead to higher test scores.
Follow-Up Questions
• What is being changed in each scenario?
• What is being measured?
Univariate and Bivariate Data
Univariate Data
Univariate data involves the study of only one variable at a time. The prefix “uni-” means one.
We collect univariate data when we are interested in just one characteristic, such as:
• Number of siblings
• Ages of learners in your class
• Exam scores for a subject
• Weight of babies at birth
• Salary of teachers.
Common ways to analyse univariate data:
• Frequency tables
• Bar graphs or histograms
• Pie charts
• Measures of central tendency (mean, median, mode)
• Range The table below shows an example of univariate data. It shows the test scores of 10 students.
Table 5.1: A teacher in Cape Coast records the mathematics scores of 10 students in a class test.
Student A B C D E F G H I J Score (%) 55 70 65 80 75 60 85 90 55 65 This is univariate data because we are looking at only one variable: the mathematics scores.
Bivariate Data
Bivariate data involves two related variables. The prefix “bi-” means two.
We collect bivariate data when we want to study the relationship between two variables, for example, how study time affects exam scores, or how age relates to height.
Common ways to analyse bivariate data
• Scatter plots
• Correlation (to measure strength of relationship)
• Line of best fit
• Two-variable frequency tables A teacher at a senior high school in Ho wants to find out whether the number of hours students spend studying affects their math exam scores.
Table 5.2: Study Time vs. Math Scores Student Study Time (hours) Math Score (%) Kekeli 1 45 Eyram 2 50 Etse 3 55 Mawuli 4 65 Etornam 5 70 This is bivariate data because two variables are involved:
• Study time (independent variable)
• Math score (dependent variable) Difference between Univariate and Bivariate data
Table 5.3: Distinguishing features of univariate and bivariate data Univariate data Bivariate data Involves a single variable Involves two variables Does not deal with cause and relationship Deals with cause and relationship The purpose is to describe The purpose is to explain Does not have any dependent variable Contains only one dependent variable
Activity 5.4 Class Survey and Univariate Analysis
Instructions
1. Conduct a classroom survey by asking each learner their favourite fruit (e.g., mango, banana, orange, pineapple, apple).
2. Record the responses in a tally and frequency table.
3. Represent the data using a bar graph or a pie chart.
4. Calculate:
• The mode (most preferred fruit)
• The total number of learners surveyed
• Is this univariate or bivariate data? Why?
Example 5.1
The ages of 6 SHS students are: 16, 17, 18, 17, 16, 19.
Find the:
a. Mode
b. Median
c. Mean
Solution
a. Mode = 16 and 17 (both appear twice)
b. Arrange: 16, 16, 17, 17, 18, 19 Median = = 17
c. Mean (x ̅ ) = = 17.17 years Plotting and Interpreting Scatter Graphs A scatter diagram (graph or plot) is a type of graph used to display the relationship between two variables. It consists of a collection of points, where each point represents a pair of values, one from each variable.
• The horizontal axis (x-axis) represents one variable.
• The vertical axis (y-axis) represents the other variable.
Each plotted point shows the value of both variables for a single observation. By looking at how the points are arranged, we can see if there is a pattern or relationship (called a correlation) between the two variables.
The graph below shows a scatter diagram that represents a distribution centre. It shows that workers who pick more lines (items or orders) tend to work longer hours.
Figure 5.1: Relationship Between Lines Picked and Hours of Overtime Why do we use scatter diagrams?
• To find out if there is a relationship between two variables.
• To see how strong the relationship is.
• To identify patterns, trends, or unusual data points (outliers).
Let us imagine you are studying whether the amount of time learners spend revising for mathematics affects their test scores.
• The x-axis shows the number of hours spent revising. This is the independent variable.
• The y-axis shows the marks scored in a test. This is the dependent variable.
Figure 5:2: Scatter graph of hours study and test score From the graph you can see that learners who revise more tend to score higher Direction of the Relationship The direction of the relationship could be:
1. Positive correlation: As one variable increases, the other also increases
2. (e.g., Hours studied vs. Test score) A positive correlation A positive perfect correlation
Figure 5.3: A positive correlation Figure 5.4: A positive perfect correlation
3. Negative correlation: As one variable increases, the other decreases
4. (e.g., Alcohol consumption vs. Health level) A negative correlation A negative perfect correlation
Figure 5.5: A negative correlation Figure 5.6: A negative perfect correlation
5. No correlation: No clear pattern or relationship
6. (e.g., Shoe size vs. Intelligence)
Figure 5.7: A no correlation scatter plot How to construct a scatter plot manually Steps:
1. Identify the variables. Since we are dealing with bivariate data, it will involve two variables, an independent variable and a dependent variable. The independent variable goes on the x-axis, with the dependent variable on the y-axis.
2. Label the axes and scale them.
3. Convert each data point into (x, y) coordinates and plot them on the graph. If two or more points fall on the same point, place them side-by-side.
How to construct a scatter diagram using Excel
Step 1: Open a new document in the Excel spreadsheet. A section of the interface should look like the picture below
Figure 5.8: Excel spreadsheet page
Step 2: Input your data in the first and second columns of the spreadsheet. As an
example, let us input the data in Table 5.4.
Table 5.4: Temperature versus number of people wearing jacket.
Temperature (x) 35 6 17 15 21 20 10 7 Number wearing jackets (y) 11 45 22 30 20 25 32 42 Your spreadsheet should be similar to the one below
Figure 5.9: Excel spread sheet page showing data of temperature and number of people wearing jackets.
Step 3: Select the data, go to Insert, and choose Scatter from the drop-down menu.
Figure 5.10: Excel spreadsheet page showing bivariate data and a scatter plot of temperature and number of people wearing jackets.
Step 4: Insert the title and the description of the x and y axes. To do this, click on the plus sign beside the scatter graph. This will open a dropdown menu titled Chart Elements. Choose “Axis Titles” from the dropdown menu.
Figure 5.11: Excel spreadsheet page showing bivariate data and a scatter plot
Step 5: Your final plot should be similar to the chart below.
162
Figure 5.11: Excel spreadsheet page showing bivariate data and a scatter plot
Step 5: Your final plot should be similar to the chart below.
Figure 5.12: Scatter plot of temperature and number of people wearing jackets.
Example 5.2
Table 5.5: A farmer records the amount of fertilizer used and the crop yield:
Fertilizer (kg) 0 2 4 6 8 10 Crop Yield (kg) 100 150 200 250 300 450
a. Draw a scatter diagram to illustrate this data
b. If the farmer uses 5 kg of fertilizer, what is the expected crop yield?
Solution
Steps involved in drawing a scatter diagram
1. Draw two perpendicular lines to form the x – y plane on a sheet of graph paper.
2. Choose an appropriate scale for the axes (x – axis representing the independent variable – the fertiliser and y– axis representing the dependent variable, the crop yield)
3. The ordered pairs (0, 100), (2, 150), (4, 200), (6, 250), (8, 300) and (10, 450) are plotted in the
4. x – y plane.
0 5 10 15 20 25 30 35 40 45 50 0 5 10 15 20 25 30 35 40 Number of people wearing Jackets Temperature scatter plot of Temperature and number of jacket
Figure 5.12: Scatter plot of temperature and number of people wearing jackets.
Example 5.2
Table 5.5: A farmer records the amount of fertilizer used and the crop yield:
Fertilizer (kg) 0 2 4 6 8 10 Crop Yield (kg) 100 150 200 250 300 450
a. Draw a scatter diagram to illustrate this data
b. If the farmer uses 5 kg of fertilizer, what is the expected crop yield?
Solution
Steps involved in drawing a scatter diagram
1. Draw two perpendicular lines to form the x – y plane on a sheet of graph paper.
2. Choose an appropriate scale for the axes (x – axis representing the independent variable – the fertiliser and y– axis representing the dependent variable, the crop yield)
3. The ordered pairs (0, 100), (2, 150), (4, 200), (6, 250), (8, 300) and (10, 450) are plotted in the
4. x – y plane.
Figure 5.13: A scatter graph showing quantity of fertilizer vs crop yield
Activity 5.5 Social Media Use vs. Sleep Duration Instructions In pairs, or individually, create a scatter graph for the data in the table below which shows the number of hours of social media use against the duration of sleep.
Learner Social Media Use (hours/day) Sleep Duration (hours/night)
Ama 1 8 Kojo 2 7
Afia 3 6 Kwame 4 5
Esi 5 4 Use graph paper to plot the scatter graph:
• x – axis: Social media use (independent variable)
• y – axis: Sleep duration (dependent variable) Analyse the graph and discuss:
• What type of correlation does it show? (Positive, negative, or no correlation)
• Is there a visible trend or pattern?
• What happens to sleep time as social media time increases?
Discuss with your classmates:
• Is this bivariate data? Why?
• Can we assume a cause-and-effect relationship here?
Follow-Up Questions:
• What does your graph suggest about the effect of social media on sleep?
• What other variables might influence a student’s sleep besides social media?
Extension Task:
• Interview five classmates and collect real data (anonymously) on how many hours they spend on social media and how long they sleep.
• Plot the results and compare them with the data above.
In many scientific and social studies, we often want to find out whether a certain treatment or intervention has an effect. One common way of doing this is by using experimental studies, where we compare two groups:
• A treatment group, which receives the intervention or condition being studied.
• A control group, which does not receive the treatment.
Data from such experiments can be analysed and represented using scatter plots to observe the relationship between the variables.
What is an Experimental Study?
An experimental study is a method of collecting data by intentionally changing one variable (called the independent variable) to observe how it influences another variable (called the dependent variable).
Table 5.6: Examples of experimental study.
Experiment Topic Independent
Variable Dependent
Variable Variable that a researcher may choose to control Dropping a ball to find out the number of times it bounces.
Drop Height Number of bounces. Ball type, surface, drop height.
Finding out how different amounts of sunlight affect plant growth Hours of Sunlight Plant Height Water, soil, and plant type How study time affects test scores Study Time Test Score Test type, room conditions How temperature of water affect dissolving time?
Water Temperature
Dissolving Time Amount of sugar, volume of water How the weight applied to a rubber band affects the stretch.
Weight Applied Length Stretched Rubber band type, air temperature Structure of an Experimental Study The structure of an experimental study includes.
• Independent Variable: The variable you change.
• Dependent Variable: The variable you measure.
• Controlled Condition(s): Factor(s) kept constant for fairness.
• Repeated Trials: Performing an activity multiple times for accuracy.
Let us consider a simple experiment conducted in a Senior High School in Ghana. The purpose is to determine whether attending an extra one-hour mathematics revision session each week improves students’ performance in mathematics.
Experiment Title: Effect of Extra Study Time on Learners’ Mathematics Performance
Step 1: Collecting the Data
Twenty learners were selected randomly and divided into two equal groups:
• Treatment Group (10 students): Received an additional 1-hour mathematics revision lesson each week.
• Control Group (10 students): Did not receive any additional lessons.
After six weeks, both groups took the same mathematics test marked out of 100. The average number of hours each student studied per week (including class lessons, the additional revision lesson and private study) was also recorded.
Table 5.7: Effect of Extra Study Time on Learners’ Mathematics Performance Group Learner Hours Studied per Week (X) Test Score (Y) Treatment T1 7 82 Treatment T2 6 76 Treatment T3 8 85 Treatment T4 9 90 Treatment T5 5 72 Treatment T6 7 80 Treatment T7 6 78 Treatment T8 8 84 Treatment T9 7 81 Treatment T10 5 70 Control C1 5 66 Control C2 6 68 Control C3 4 60 Control C4 5 65 Control C5 6 69 Control C6 7 72 Control C7 6 67 Control C8 5 64 Control C9 4 58 Control C10 6 68
Step 2: Plotting the Scatter Graph
Plot the points (X, Y) for each student on the same scatter plot. Use different symbols or colours to represent the treatment group and the control group.
Use the horizontal axis (x – axis) for “Hours Studied per Week” and the vertical axis (y – axis) for “Test Scores”.
Figure 5.14: A scatter graph showing the number of hours leaners studied against their test scores. The red represents the control group, the blue the treatment group The blue points are for the treatment group and red for control group.
Step 3: Interpreting the Scatter Plot
a. Trend Observation: By examining the scatter plot, you may observe a general upward trend — as the number of study hours increases, test scores tend to increase.
b. Treatment vs Control Comparison: Students in the treatment group generally scored higher than those in the control group who studied for the same number of hours.
c. Relationship Between Variables: There is a positive correlation between hours studied and test scores in both groups. However, the treatment group shows slightly higher test scores for the same number of study hours, suggesting that the extra revision sessions had a positive effect on performance.
Describing The Relationship Between Two
Variables The scatter plot helps us to visualise the relationship between two variables. An alternative way to describe the relationship is to use a linear function (a straight line).
Note that this is used to approximate the relationship between the variables being compared, as the data points do not form a straight line in most cases. Hence, the approximated line is called a line of best fit. The line of best fit should either pass through most of the points or be closest to most of the points. Because of this, the line of best fit is usually used to predict data values within the given data. This is known as interpolation. When used to predict data values outside a given data, it is called extrapolation. Pairs of data values (points) relatively far from the line of best fit are called outliers. We will obtain the line of best fit through observation.
Data Collection in Experimental Studies
Planning your experiment Before collecting data, you need to:
1. Define your research question: What are you trying to find out?
2. Identify your variables:
• Independent variable (what you manipulate)
• Dependent variable (what you measure)
3. Select your sample: Choose participants for both treatment and control groups
4. Plan your experimental procedure: Ensure it is consistent for all participants Let us use the example below to enhance our understanding of how to conduct a simple experimental study.
Example 5.3
Research Question:
Investigating the Relationship Between Drop Height and Bounce Height of a Ball.
Suggested Solution:
Step 1: Identify the Research Question
Research question: How does the height from which a ball is dropped affect how high it bounces?
Step 2: Identify the Variables:
Independent variable = Drop height (we will determine the heights from which we drop the ball) Dependent variable = Bounce height (we have no control over the highest point the bounce will reach) So, in this research, we will control the height from which we will drop the ball, the type of floor we want to drop the ball on, the type of ball we want to drop and maybe the temperature.
When a ball is dropped on a hard surface, all things being equal, it will bounce more than once. We will be only interested in the first bounce, as this gives the highest bounce.
Step 3: Plan the Experiment
Materials:
A rubber ball, a metre rule, a flat hard surface and a notebook for recording data We will undertake 7 trials. We will drop the ball in intervals of 30cm and repeat each trial three times, and take the average bounce height. This is summarised in the procedure below.
Procedure:
• Drop the ball from different heights (30cm, 60cm, 90cm, 120cm, 150cm, 180cm, 210cm).
• Measure the bounce height each time using a ruler. Recording the results in a
table.
• Repeat each measurement 3 times and record the average.
Step 4: Collect the Data
Table 5.8: Drop height versus average bounce height data.
Drop Height(cm)
Bounce Height (cm) Average bounce height.
First trial Second trial Third trial 30 19 17 18 18 60 39 42 39 40 90 58 59 60 59 120 72 72 72 72 150 83 86 83 84 180 103 103 109 105 210 146 144 142 144
Step 5: Present the data using a scatter plot.
a. Label the x-axis as Drop Height (independent variable).
b. Label the y-axis as Bounce Height (dependent variable).
c. Plot the data points from your table.
d. Look for a trend or pattern.
Figure 5.15: Scatter plot of drop height and average bounce height
Step 6: Interpret the Scatter Plot
From the graph, we can observe:
• As the drop height increases, the bounce height also increases.
• This shows a positive correlation.
• A line of best fit can help identify the relationship and predict unknown values.
Step 7: Using the graph to solve problems, if required:
Problem: If a ball is dropped from 70 cm, what is the expected bounce height?
Solution: Use the graph or line of best fit to estimate To do this, read up from 70cm to the line of best fit and then read across to the bounce height for a predicted bounce height.
Conducting an Observational Study
Steps:
1. Choose two variables you want to study.
2. Decide when and where you will observe (e.g., during break time, in the school library).
3. Use a table to collect data at regular intervals.
4. Draw a scatter plot to visualise the relationship.
5. Interpret the pattern you observe.
Example 5.4
A reading test was given to 10 learners in Basic 6. The learners then took part in an extensive reading programme. After participating in the programme they were retested.
The data collected was organised and plotted as a scatterplot as follows.
Figure 5.16: Scatter plot of pre and post intervention data.
Find the relationship between Pre-intervention Reading Test Scores and Post- intervention Reading Test Scores. For this, do a comparison, draw a conclusion and justify your conclusion.
Solution
Construct a distribution table for the graph
Table 5.9: Pre and post intervention data.
Pre-intervention scores 20 20 30 30 40 46 48 50 56 70 Post-intervention scores 30 40 30 40 50 60 70 60 70 86 The scatter graph shows a positive correlation between pre-intervention scores and post-intervention scores. The intervention increased the score Observation:
• Generally, as pre-intervention scores increase, post-intervention scores also increase.
• This suggests a positive correlation.
• The values are not perfectly aligned, but the upward trend is clear.
Interpretation:
• Low pre-scores (e.g., 20, 30) had variable post-scores (30–40), showing some improvement.
• Higher pre-scores (e.g., 50, 56, 70) had higher post-scores (60, 70, 86), showing a consistent increase.
• Although there is a bit of spread in the lower scores, the overall pattern supports a moderate to strong positive correlation.
Conclusion:
There is a positive correlation between pre-intervention and post-intervention scores.
Learners with higher scores before the intervention tended to score even higher after the intervention. This suggests the intervention may have improved performance across all levels, but the effect appears more consistent among higher-performing learners.
Activity 5.6 Conduct Your Own Experiment
Materials
• Seeds (beans or maize)
• Containers
• Soil
• Measuring tape/ruler
• Graph paper or computer with spreadsheet software Procedure
1. Divide your seeds into two equal groups
2. Plant all seeds in similar containers with the same amount of soil
3. Apply your treatment to one group (e.g., different watering schedule, fertilizer, light exposure)
4. Measure and record the growth every 3-5 days in a data table
5. Draw a scatter plot of your results
6. Analyse the relationship between your variables Reflect on these questions, with a classmate, or in a small group
1. What is the purpose of having a control group in an experiment?
2. How would you describe the relationship shown in a scatter plot where points move from bottom-left to top-right?
3. In an experiment measuring the effect of study time on test scores, what would be the independent and dependent variables?
4. If a scatter plot shows points scattered randomly with no pattern, what can you conclude about the relationship between the variables?
5. How does the slope of a line of best fit help us understand the relationship between variables?
Comparing Different Data Sets
In statistics, we often collect data from different groups or situations and need to decide which one gives a better picture of reality. To do this, we use measures of central tendency (mean, median, mode) and measures of dispersion (range, interquartile range, mean deviation, variance, standard deviation).
• Central Tendency helps us know the “average” behaviour of the data.
• Dispersion tells us how spread out the data is. A data set with very high spread may be less reliable in representing a situation.
Example 5.5: Comparing Exam Scores
Consider the following exam scores from two different classes in a senior High school:
Class A: 65, 70, 72, 75, 78, 80, 82, 85 Class B: 50, 60, 75, 78, 80, 85, 90, 95 Compare these data sets to draw conclusions about the two classes.
Solution
Measures of Central Tendency
Class A:
• Mean =
• Median =
• Mode = None (no repeated values) Class B:
• Mean =
• Median =
• Mode = None (no repeated values) Measures of Dispersion Class A:
• Range = 85 – 65 = 20
• Standard Deviation ≈ 6.6 Class B:
• Range = 95 – 50 = 45
• Standard Deviation ≈ 15.1 Interpretation Although Class B has a slightly higher mean (76.6 vs 75.9), the scores in Class A are more consistent (smaller standard deviation). Class B shows greater variability with some very high scores but also some very low scores.
Example 5.6
Two bicycle sales representatives, Naa and Yaw, are being considered for a promotion.
Their weekly performance over 5 weeks is shown below.
Table 5.10: Naa’s sales performance Hours worked 30 35 40 45 50 Sales made 12 13 16 19 20
Table 5.11: Yaw’s sales performance Hours worked 30 35 40 45 50 Sales made 15 16 17 17 18 Based on this data who do you consider worthy of the promotion?
Solution
Analysis and Interpretation:
Step 1: Plot the data for both Naa and Yaw using scatter graphs.
Scatter plot showing Naa’s sales perfor- mance Scatter plot showing Yaw’s sales perfor- mance
Figure 5.17: Scatter plot of Naa’s sales perfor- mance
Figure 5.18: Scatter plot of Yaw’s sales performance
Step 2: Identify the type of relationship.
Naa: Shows a strong positive correlation, as hours increase, so do sales of bicycles.
Yaw: Shows a less strong correlation, sales do not increase that much despite working more hours.
Step 3: Compare the trends and consistency.
Naa has a steady and consistent increase in sales with more work.
Yaw starts high, but the increase in sales is minimal, showing reduced efficiency with more hours.
Step 4: Conclusion
Naa’s data better represents a productive relationship between work and results. The linear growth in sales with effort justifies selecting Naa for a promotion.
Example 5.7
Two mathematics classes wrote the same test. Their results were as follows:
• Arts 1: Mean score = 65%, Standard Deviation = 5
• Arts 2: Mean score = 65%, Standard Deviation = 15 Interpret this data.
Solution
Although both classes have the same mean score, Arts 1’s performance is more consistent (less spread out), which means the scores of most students are closer to the average. Arts 2 shows that they have a lot more spread in their data, so some of these students will have scored more highly, but some will also have more low scores.
Justifying Which Data Set Better Represents a Given Situation When deciding which data set better represents a situation, consider:
1. The purpose of the analysis
2. The context of the data
3. The relevant statistical measures
4. Potential outliers and their impact
Example 5.8
The tables below present data on tree growth observed after applying two distinct fertilizer brands: Surge and Boom
Table 5.12: Tree growth using Surge fertilizer Week 1 2 3 4 5 Height (mm) 75 100 100 145 155
Table 5.13: Tree growth using Boom fertilizer Week 1 2 3 4 5 Height (mm) 50 80 120 160 200 Which of the fertilizers will you recommend? Justify your choice of fertilizer.
Solution
Step 1: Plot the data for both Surge fertilizer and Boom fertilizer using scatter graphs.
A scatter graph showing the growth of trees after applying surge fertilizer A scatter graph showing the growth of trees after applying boom fertilizer
Figure 5.19: Scatter plot of tree growth using surge fertilizer
Figure 5.20: Scatter plot of tree growth using boom fertilizer
Step 2: Identify the type of relationship Surge: Shows a weak positive correlation – the growth of the trees is not that consistent.
Boom: Shows a strong positive correlation – it shows a consistent growth as the weeks go by.
Step 3: Compare the trends and consistency Surge: shows steady growth as it is consistent and effective over time.
Boom: gives quick but limited growth since growth stalls early.
Step 4: Conclusion
Surge Fertilizer is more effective for long-term plant development.
Example 5.9
Two farming techniques were tested in the Western Region of Ghana, with yields (in kg per hectare) recorded as follows:
Technique A: 1200, 1250, 1300, 1350, 1400, 1450, 1500 Technique B: 900, 1100, 1400, 1500, 1600, 1700, 2000 Which technique would you recommend to cocoa farmers?
Solution
Analysis:
Technique A:
• Mean = 1350 kg/ha
• Range = 300 kg/ha
• Standard Deviation ≈ 104 kg/ha Technique B:
• Mean = 1457 kg/ha
• Range = 1100 kg/ha
• Standard Deviation ≈ 370 kg/ha Conclusion If a farmer wants consistent, reliable yields with minimal risk, Technique A would be more appropriate as it shows less variability (smaller standard deviation).
If a farmer can tolerate risk and wants the possibility of higher yields, Technique B might be preferred as it has a higher mean yield, though with much greater variability.
For most small-scale farmers who cannot afford significant crop failures, Technique A would likely be the better representation of a sustainable farming approach.
Activity 5.7
The following data shows the daily sales (in Ghana Cedis) for two fruit vendors at Makola Market over a week:
Vendor A: 120, 150, 145, 160, 155, 140, 130 Vendor B: 100, 180, 90, 200, 110, 190, 130
a. Calculate the mean, median, range and standard deviation for each vendor’s sales.
b. Which vendor has more consistent sales? Justify your answer.
c. If you were advising someone who wants to start a fruit selling business, which sales pattern would you recommend they aim for? Why?
d. What factors might explain the differences in the sales patterns?
Interpreting Data Presented on Media Platforms
In today’s world, information is everywhere. Media platforms such as local and international TV stations, newspapers, online journals and social media publish statistical data about health, education, the economy, sports and the environment.
Examples of data in the media include:
• Election results published by the Electoral Commission of Ghana.
• Ghana Statistical Service reports on unemployment or inflation.
• West African Examinations Council (WAEC) reports on student performance.
• World Health Organization (WHO) reports on global health issues like malaria or COVID-19.
Key skills in interpreting published data
1. Check the source: Is it credible (e.g., Ghana Statistical Service, World Health Organization, Ministry of Education)?
2. Understand the representation: Is the data shown in tables, graphs, or percentages?
3. Look for patterns or trends: Is there an increase, decrease, or fluctuation?
4. Make inferences: What does the data suggest about the situation?
5. Draw conclusions and make recommendations: What action should be taken based on the data?
Example 5.10: Data from Local Media (Ghana) Interpret the newspaper report below, and consider what recommendations you should make.
• In 2023, 80% of students in Accra had access to internet-based learning resources.
• In 2023, only 35% of students in rural Northern Ghana had access.
Solution
Interpretation: This shows a digital divide between urban and rural students. The difference in access to digital learning may affect performance in online-based learning
activities.
Recommendation: Government and stakeholders should provide affordable internet services and ICT facilities to rural areas to bridge the gap.
Example 5.11: Data from International Media An international journal reports:
• Ghana’s inflation rate in 2022 was 31%, while the average inflation rate across West Africa was 15%.
What inferences can you draw from this and what recommendations would you make?
Solution
Inference: Ghana’s inflation was more than double the regional average, suggesting higher cost of living pressures on citizens compared with neighbouring countries.
Recommendation: Policymakers should focus on stabilizing the local currency and promoting local production to reduce dependency on imports.
Example 5.12: Screen Time and Sleep
An online health article reported data collected from teens aged 15–18 on average daily screen time and hours of sleep.
Table 5.14: Screen time and sleep time Screen time(hrs) 3 6 8 5 1 2 1 6 4 7 Sleep time (hrs) 7.5 6.5 5.2 6.7 8.2 8 8.5 6.1 7 5.7
a. Present the data using a scatter plot.
b. What kind of correlation exists between screen time and sleep duration?
c. What might be the implication of this trend?
d. Is this data likely observational or experimental? Explain your answer.
Solution
a.
Figure 5.21: Scatter plot of screen time versus sleep time.
b. There is a negative correlation between screen time and sleep duration. As screen time increases, sleep decreases.
c. Excess screen time may lead to reduced sleep time.
d. The data is observational as no variable was controlled.
Example 5.13
The number of goals scored versus the number of goals conceded by a sample of seven football clubs in the Ghana Premier League for the 2024/2025 season is shown in the bar chart.
Figure 5.22: Bar chat showing football clubs and number of goals scored in the 2024/2025 season.
a. Copy and complete the table
Table 5.15: A partially completed table showing goals scored versus goals conceded Football Club Bibiani Lions Kotoko Nations Dreams Samartex Bechem Goals Scored 38 33 Goals Conceded 25
b. Present the data using a scatter plot.
c. What kind of correlation exists between goals scored and goals conceded?
d. What might be the implication of this trend?
Solution
a.
Table 5.16: A completed table showing goals scored versus goals conceded.
Football Club Bibiani Lions Kotoko Nations Dreams Samartex Bechem Goals Scored 38 38 37 40 31 33 32 Goals Conceded 21 24 27 21 29 25 28
b. Present the data using a scatter plot.
Figure 5.23: Scatter of goals scored versus goals conceded.
c. What kind of correlation exists between goals scored and goals conceded?
There is a negative correlation between goals scored and goals conceded.
d. Football clubs that score more goals concede fewer goals.
Activity 5.8
In pairs, or small groups, work through these questions and compare your answers with others.
1. The table below shows the monthly rainfall (in mm) recorded in two towns, Kumasi and Tamale, over the same period.
Town Mean Rainfall (mm) Standard Deviation (mm)
Kumasi 120 10 Tamale 118 25
• Which town has more consistent rainfall?
• Which data set best represents a stable rainfall pattern?
• Give reasons for your choice.
2. A news report shows that:
• 65% of graduates from Technical Universities in Ghana are employed within two years of graduation.
• 82% of graduates from Teacher Training Colleges are employed within two years of graduation.
• Which group of graduates is more likely to get jobs faster?
• What policy recommendations would you suggest to reduce graduate unemployment in Ghana?
3. A newspaper publishes the following data about two Ghana Premier League teams after 10 matches:
Team Average Goals per Match Standard Deviation Hearts of Oak 2.3 0.5 Asante Kotoko 2.5 1.8
• Which team is more consistent in scoring goals?
• Which team is less predictable?
• What inference can you make about their performance?
4. According to the Ghana Statistical Service (GSS), the unemployment rate increased from 8.5% in 2021 to 13.4% in 2023.
• What does this trend suggest about job opportunities in Ghana?
• Suggest two measures the government can take to address this problem.
5. The World Bank reported that in 2022, 33% of Sub-Saharan Africans had access to electricity, while in Ghana, 74% of the population had access.
• What conclusion can you draw about Ghana compared to the Sub- Saharan average?
• What recommendation would you make to ensure 100% electricity access in Ghana?
We already know that probability is the measure of the likelihood that an event will occur. It is widely used in daily life, decision-making, business, and the media. Media platforms (such as newspapers, television, radio, and the internet) frequently report probability to help individuals, governments and companies make informed choices.
Probability in the Media
Weather Forecasts (Television, Radio, Internet)
Example: The Ghana Meteorological Agency announces on TV:
• “There is a 70% probability of rainfall in Accra tomorrow.”
Explanation:
• This means there is a high chance (7 out of 10 times) that rain will fall.
• People hearing this forecast may decide to carry umbrellas, change their travel plans, or postpone outdoor events.
Influence on Decisions:
• Farmers may delay planting or harvesting.
• Schools may cancel sports activities.
• Traders may prepare to protect goods in the market.
Sports Predictions (Newspapers, Internet, TV Sports Channels)
Example: A sports analyst reports in the Daily Graphic that:
• “Based on past performance, Hearts of Oak has a 65% probability of winning their next match against Kotoko.”
Explanation:
• This prediction is based on previous data such as goals scored, player form and head-to-head results.
Influence on Decisions:
• Supporters may place bets on betting platforms.
• Coaches may adjust strategies knowing their probability of success.
• Fans may decide whether to attend the match or watch from home.
Health Reports (Newspapers, WHO Reports, Online Platforms)
Example: A health report states:
• “There is a 30% probability of contracting malaria in areas without mosquito nets compared to a 10% probability in areas where nets are used.”
Explanation:
• This shows that the likelihood of malaria is three times higher without mosquito nets.
Influence on Decisions:
• Families are encouraged to use insecticide-treated nets.
• Government and NGOs plan health campaigns to reduce the risk.
Election Polls (TV, Newspapers, Internet)
Example: Before Ghana’s general elections, a poll published in the Daily Guide states:
• “Candidate A has a 55% chance of winning, while Candidate B has a 45% chance.”
Explanation:
• Probability is used here to estimate the likely winner, based on sampled voters’ opinions.
Influence on Decisions:
• Voters may be motivated to turn out in large numbers to support their candidate.
• Political parties may change campaign strategies in areas where they are less popular.
Business and Insurance (Internet, Newspapers)
Example: An insurance company advertises:
• “There is only a 5% probability that your car will be involved in an accident in a given year. Protect yourself with insurance.”
Explanation:
• The small but possible risk is highlighted to encourage car owners to buy insurance.
Influence on Decisions:
• Drivers choose to purchase insurance policies for security.
• Businesses prepare contingency plans for unexpected losses.
Why probability is important in media reports
• It helps individuals make informed choices (e.g., whether to carry an umbrella or buy insurance).
• It allows policymakers and leaders to plan ahead (e.g., preparing for floods or disease outbreaks).
• It provides businesses with risk assessments (e.g., banks assessing loan repayment probability).
• It gives the public a scientific basis for decision-making rather than relying on guesswork.
Carry out the following activities individually but compare your answers with a classmate.
Activity 5.7: Weather Report
Watch a local TV news weather forecast.
Write down one probability statement given (e.g., “60% chance of thunderstorms in Koforidua”).
• How might this forecast influence a farmer?
• How might it influence a taxi driver?
Activity 5.8 Sports Probability
A betting company reports
• Chelsea has a 40% chance of winning,
• Manchester United has a 35% chance,
• A draw has a 25% chance.
Consider which outcome is most likely?
How might this probability influence fans or bettors?
Activity 5.9 Health Probability
A health journal reports
• The probability of contracting COVID-19 without wearing a mask is 25%,
• With a mask, the probability reduces to 10%.
What decision should an individual make based on this data?
Explain why probability helps in health campaigns.
Dependent and Independent Events
a. Independent Events: Two events are independent if the occurrence of one does not affect the probability of the other.
Examples:
• Tossing a coin twice. (The first toss does not affect the second.)
• Rolling a die and then tossing a coin.
• Rolling a single die twice or two dice once.
• Drawing two or more items with replacement.
If A and B are independent, then:
P(A and B) = P(A) × P(B)
b. Dependent Events: Two events are dependent if the outcome of one event affects the probability of the other.
Examples:
• Picking a card from a deck without replacement.
• Selecting two students from a class for a task, one after the other, without returning the first.
• Selecting individuals to form a group or a committee.
• Selecting individuals for a position where one person cannot occupy multiple positions If A and B are dependent, then:
P(A and B) = P(A) × P(B∣A) where P(B|A) means the probability of B given that A has already happened.
Addition and Multiplication Laws
a. Addition Law: The addition law is used when we want the probability of either event A or event B happening.
P(A or B) = P(A) + P(B) − P(A and B) If events A and B are mutually exclusive (they cannot happen together), then:
P(A or B) = P(A) + P(B)
b. Multiplication Law: The multiplication law is used when we want the probability of both events happening.
For independent events:
P(A and B) = P(A) × P(B) For dependent events:
P(A and B) = P(A) × P(B∣A) Examples 5.14 A coin is tossed and a die is rolled. Find the probability of obtaining a Head and a 4.
Solution
• P(Head) =
• P(rolling a 4) = Since the events are independent:
P(Head and 4) = × =
Example 5.15
Two cards are drawn from a standard deck of 52 cards without replacement. Find the probability that both are Aces.
Solution:
• P(1st Ace) = =
• After one Ace is removed, P(2nd Ace) = = P(Two Aces) = × =
Example 5.16
In a class of 40 students, 25 like Mathematics, 18 like Science, and 10 like both subjects.
If a student is chosen at random, find the probability that the student likes Mathematics or Science.
Solution:
P(Math or Science) = P(Math) + P(Science) − P(Math and Science) = + − =
Example 5.17
A building contractor submitted bids for two independent contracts, X and Y. The probability that he will win contract X is 0.5, and the probability that he will not win contract Y is 0.3. Calculate the probability that the contractor will:
a. Win both contracts
b. Win exactly one of the contracts
c. Win neither of the contracts
Solution
P(Win X) = 0.5, P(Not Win X) = 1 − 0.5 = 0.5 P(Not Win Y) = 0.3, P(Win Y) = 1 − 0.3 = 0.7 Contracts X and Y are independent
a. Probability of winning both contracts:
Since the events are independent:
P(Win X and Win Y) = P(Win X) × P(Win Y) = 0.5 × 0.7 = 0.35 Therefore, the probability of winning both contracts is
b. Probability of winning exactly one contract:
There are two ways to win exactly one contract:
• P(Win X and not win Y) = 0.5 × 0.3 = 0.15
• P(Not win X and Win Y) = (1 − 0.5) × 0.7 = 0.5 × 0.7 = 0.35 P(Win exactly one contract) = 0.15 + 0.35 = 0.50
c. Probability of winning neither contract:
P(winning neither contract) = P(Not win X) × P(Not win Y) = 0.5 × 0.3 = 0.15
Example 5.18
A basket at the Makola Market contains 5 oranges and 3 mangoes. If two fruits are selected randomly without replacement, what is the probability of selecting an orange followed by a mango?
Solution
P (Orange then Mango) = P(Orange) × P(Mango | Orange) P(Orange) = After selecting an orange, there are 4 oranges and 3 mangoes left out of 7 total fruits:
P(Mango/Orange) = Therefore: P(Orange then Mango) = × = ≈ 0.268 There is approximately a 26.8% chance of selecting an orange followed by a mango.
Example 5.19
A social media quiz app offers a 30% chance to win a prize per spin. If you spin twice, what is the probability that you win both times?
Solution
The question is an example of an independent event. The 30% chance of winning does not change regardless of the number of spins.
P(Win twice)
Activity 5.10 Real-Life Probability Problem Solving
Work with a partner; consider this task.
A fruit seller at Kejetia Market has 6 apples and 4 pears.
If two fruits are selected one after the other without replacement:
a. What is the probability of selecting an apple followed by a pear?
b. What is the probability of selecting two pears?
Solve using the multiplication law for dependent events.
After solving, discuss how probability could help the seller in planning (e.g., predicting sales or fruit combinations).
1. The following are the daily sales (in cedis) made by a kenkey seller over 5 days:
80, 95, 100, 85, 90
a. Find the mean and median
b. Draw a bar chart to represent the data
2. The number of children in each of 8 households in Sunyani are: 2, 3, 4, 2, 5, 3, 2, 4
a. What is the mode?
b. What is the range?
3. A teacher recorded the number of hours learners studied and their scores:
Learner Study Time (hrs) Score (%)
A 2 55 B 4 65 C 6 75 D 8 85
a. Plot a scatter diagram
b. What is the relationship between study time and grades?
4. The Regional Director of education is studying the relationship between the number of teachers in a school and the number of learners:
School Teachers Learners
A 10 200 B 15 350 C 8 150 D 20 400
a. Which school has the highest learner-teacher ratio?
b. Is there a correlation between number of teachers and number of learners?
5. Identify the independent and dependent variables:
i. A researcher examines how sleep affects exam performance.
ii. A shopkeeper records daily temperature and sales of cold drinks.
6. Plot a scatter graph for the data below and describe the relationship:
Hours of Exercise 1 2 3 4 5 6 Weight Loss (kg) 0.5 1.0 1.5 2.5 3.0 5.0
7. The Ghana Meteorological Service reports a 65% probability of rain in Accra tomorrow.
a. What common misunderstanding might people have about this forecast?
b. What does this 65% probability actually mean?
c. Based on this information, what actions might someone choose to take?
8. A recent study found that a new malaria vaccine is 87% effective.
a. How would you explain this 87% effectiveness to a friend in simple terms?
b. Why is this information important for individuals and the public?
c. What is a possible misunderstanding someone might have about this result?
9. A sports website predicts that “Medeama FC has a 70% chance of winning their next match.”
a. What does this 70% chance mean in context?
b. How might fans or gamblers respond to this information?
c. What is the possible misunderstanding someone might have about the prediction?
10. A bag contains 6 red bottle tops and 4 blue bottle tops. Two are picked one after the other with replacement.
Find the probability that:
a. Both are red.
b. One is red and one is blue.
11. A survey shows that the probability a person reads the Daily Graphic is 0.4, and the probability a person reads the Ghanaian Times is 0.3. These events can be taken as independent.
What is the probability that a person reads at least one of the two newspapers?
12. The probability that a football team wins a home match is 0.7 and the probability that it wins an away match is 0.4. Assuming independence, find the probability that:
a. The team wins both an away and a home match.
b. The team wins at least one match from a home and an away match.
13. In a class at Achimota School, the probability that a student selected at random plays football is 0.35, and the probability that a student plays basketball is 0.25.
If no student plays both sports, what is the probability that a randomly selected student plays either football or basketball?
14. In a survey of 200 Ghanaian high school students, 120 enjoy fufu, 100 enjoy jollof rice, and 60 enjoy both.
If a student is selected at random, what is the probability that they enjoy either fufu or jollof rice or both?
15. Two cards are drawn at random from a standard deck of 52 playing cards without replacement.
Find the probability of selecting:
a. A king, followed by a red card. Give your answer as a fraction in its simplest form.
b. A diamond, followed by a black card. Give your answer as a fraction, in its simplest form.
c. Two odd-numbered cards. Give your answer as a fraction, in its simplest form.
A researcher studies how the number of hours of sleep affects students' exam scores. Which variable is the independent variable?
In a study, as the number of hours a learner studies increases, the learner's test score generally increases. What kind of relationship is this?
For six students, the hours of exercise were and their weight losses in kg were respectively. If these points are plotted on a scatter graph, what relationship is shown?
Class A scores are . Class B scores are . Both classes have mean . Which statement is correct?
Two schools, School P and School Q, have the same mean score of in a mathematics test. School P has standard deviation , and School Q has standard deviation . Which conclusion is valid?
A research team at the University of Education, Winneba, conducted an experiment to find out whether a new mathematics revision method improves performance. Ten students were put into two groups of five. The treatment group used the new method, while the control group used the usual method. Each student's revision time (in hours) and test score (%) were recorded.
| Treatment group | T1 | T2 | T3 | T4 | T5 |
|---|---|---|---|---|---|
| Revision time (hours) | 1 | 2 | 3 | 4 | 5 |
| Test score (%) | 50 | 55 | 65 | 70 | 78 |
| Control group | C1 | C2 | C3 | C4 | C5 |
|---|---|---|---|---|---|
| Revision time (hours) | 1 | 2 | 3 | 4 | 5 |
| Test score (%) | 45 | 48 | 55 | 60 | 65 |
State the independent variable and the dependent variable in the relationship between revision time and test score.
Describe the relationship between revision time and test score for each group as it would appear on a scatter diagram. Give a reason for your answer.
Calculate the mean test score for the treatment group and the mean test score for the control group.
Use the means to decide whether the new revision method improved performance. Justify your answer.
A radio station in Accra reports: 'There is a 70% probability of rain in Accra tomorrow.' A sports analyst also says that Hearts of Oak has a 65% probability of winning their next match against Asante Kotoko.
Explain what the 70% probability of rain means, and state two ways this information may influence people's decisions.
Distinguish between independent and dependent events, giving one example of each.
A bag contains 4 red, 3 blue and 5 green bottle tops. Two bottle tops are picked one after the other without replacement. Calculate the probability that both bottle tops are red.
Calculate the probability that the first bottle top is red and the second is green.