Which of the following sets of data is bivariate?
Strand 4 · Handling Data
Additional Mathematics Year 2 Learner Material, Section 7: Correlation
The ability to organise, analyse and present data is an important skill as this is essential in real life. In this section we will develop these skills whilst learning about the concept of correlation and statistics, along with details on univariate and bivariate data, scatter plots, analysing scatter plots and spearman rank’s correlation.
KEY IDEAS
• A measure of the correlation is called the correlation coefficient.
• A measure of the nature and strength of the relationship between two or more variables is called correlation.
• A scatter plot is used to visually show the relationship between variables of bivariate data.
• A variable is a characteristic or measurement that can be determined for a population.
• An estimate of a linear function for bivariate data is called the line of best fit.
• Univariate data involves only one variable while bivariate data involves exactly two variables.
Univariate Data
In statistics, a variable is a characteristic or measurement that can be determined for a population. Variables may describe values like weight in kg, height in metres or favourite food.
The word; “Uni” means “one” and “variate” is another word for “variable”. So, “univariate data” means data with only one variable or a single characteristic.
Examples are the age of learners, the height of trees, the weight of babies at birth or the wages of construction workers.
You can describe patterns in univariate data using central tendencies (mean, mode and median) and dispersion (range, interquartile range, standard deviation and variance). Frequency distribution tables, bar charts, histograms, pie charts and frequency polygons can be used to present univariate data.
Bivariate Data
“Bi” means “two”. So, “bivariate data” means data that involves exactly two variables. Examples:
1. Amount of money spent on advertisement and total revenue.
2. Hours of work and wages
3. Hours of study and grades
4. Years of schooling and annual income
5. Total rainfall and amount of cocoa harvested.
6. Income and Expenditure
The table below shows an example of bivariate data. It gives the hours of study and marks of ten students in Mathematics.
Name Abi Ato Fafa Bushra Akos Yaa Oko Mbo Kasi John Hours 7.5 4.5 6 9 4 6 5 7 8 7 Marks 65 50 75 95 45 52 70 83 90 50 This data contains two variables, hours and marks.
To present bivariate data, we can use a scatter plot. It allows us to visualise the relationship between the two variables.
Difference between Univariate and Bivariate data Univariate data Bivariate data Involves a single variable Involves two variables Does not deal with cause and relationship Deals with cause and relationship The purpose is to describe The purpose is to explain Does not have any dependent variable Contains only one dependent variable Concept of Correlation Correlation is a measure of the nature and strength of the relationship between two or more variables. It describes how well two variables go together. The relationship between two variables may be positive, negative or non-existent (no correlation). The strength of the relationship ranges from − 1 to 1 inclusive.
A positive correlation between two variables means that an increase in the value of one variable is likely to increase the value of the other. Likewise, a decrease in one of the variables will cause a reduction in the other. It is a relationship that moves in tandem (in the same direction). It shows a direct relation between two variables. An example of a positive correlation is the relationship between Temperature and human water consumption. Generally, as temperature increases, people consume more water. Another example is deforestation and erosion. As we cut down more trees, the likelihood of erosion increases. Other examples are hours of study and test scores, demand and price, cleanliness and life expectancy, prices of fuel and prices of transportation, investment and interest, hours of work and pay. Figure 7.1 below shows a scatter plot with positive correlation.
Figure 7.1: Positive Correlation
A positive perfect correlation means that 100% of the time, the variables in question move together by the same percentage and direction and all points on the scatter plot lie on a perfect straight line, as shown in Figure 7.2.
Figure 7.2: Perfect Positive Correlation
Negative Correlation
A negative correlation describes an inverse relationship. As one variable increases, the other decreases. It describes the relationship between two variables that change in opposite directions. An example is the relationship between age and agility. As your age increases, your agility decreases. Another example is speed and travel time. The higher the speed, the shorter the travel time. Other examples include preparation and mistakes, supply and price, exercise and body weight, inflation and purchasing power. Figure 7.3 shows negative correlation on a scatter plot, whilst Figure 7.4 shows a perfect negative correlation with all points on a perfect straight line.
Figure 7.3: Negative correlation
Figure 7.4: Perfect Negative Correlation
Zero or no correlation This occurs when there is no relationship between two variables. It means changes in one variable do not predict changes in the other. An example is weight and intelligence. Your weight has nothing to do with your intelligence and an increase in your weight has nothing to do with your intelligence. Other examples are blood type and prosperity, shoe size and favourite colour, weight and income. Figure 7.5 shows a random array of dots on a scatter plot which shows no correlation.
Figure 7.5: No Correlation
Scatter Plot
The graph below is an example of a scatter plot
Figure 7.6: Scatter plot
Figure 7.6 shows bivariate data of monthly salary and savings of a sample of 17 teachers. Such a graph helps us to see the relationship between the two variables under consideration. In this example, the variables being compared are monthly salary and savings.
One important thing to note is that correlation does not imply that one variable causes the other. For example, there is a strong correlation between accidents and speeding. This does not imply that accidents are caused by speeding. The correlation simply tells us that when speeding increases, we are likely to have more accidents. Likewise, data may give a strong negative correlation between inflation and purchasing power. Does this mean we can conclude that inflation causes a decrease in purchasing power? No! Having a strong correlation does not give us the yardstick to make that conclusion.
Activity 7.1: How to construct a scatter plot manually (Work in pairs) Use the steps below
Step 1: Identify the variables.
Since we are dealing with bivariate data, it will involve two variables, an independent variable and a dependent variable.
The independent variable is the variable that is being manipulated in the data. It is also called the input variable. This variable is plotted on the x-axis, the horizontal axis.
The dependent variable is the variable being measured. It is the output or the outcome. It is plotted on the y-axis, the vertical axis.
In the example above, the independent variable is Monthly salary and the dependent variable is Monthly savings. This is because savings depend on the monthly salary or savings in this instant are obtained from the monthly salary.
Step 2: Label the axes and scale them.
In the example above, we used an interval of 2cm on the graph to represent GH¢ 1000 on the x axis (monthly salary) and 1cm on the graph to represent GH¢ 200.00 on the y axis.
Step 3: Convert each data point into (x, y) coordinates and plot them on the graph. If two or more points fall on the same point, place them side-by-side.
Example 7.1
Determine the independent and dependent variables in the following scenarios.
a. The effect of exercise on agility
b. The effect of motivation on output.
c. Grades and hours of study
d. Life expectancy and amount of sleep
e. Grades of students in English and Mathematics.
f. Relationship between the rainfall and yield of crops.
g. Relationship between the working time and salary.
Solution
a. The effect of exercise on agility:
b. independent variable = Exercise
c. dependent variable = agility
d. The effect of motivation on output:
e. independent variable = Motivation
f. dependent variable = Output
g. Grades and hours of study:
h. independent variable = hours of study
i. dependent variable = Grades j. Life expectancy and amount of sleep:
k. independent variable = Amount of sleep l. dependent variable = Life expectancy m. Grades of students in English and Mathematics:
n. Either of the subjects can be taken as the independent variable and the other will be the dependent variable.
o. Relationship between the rainfall and yield of crops:
independent variable = Rainfall dependent variable = Yield of crops p. Relationship between the working time and salary:
independent variable = Working time dependent variable = Salary
Example 7.2
State the type of correlation in the following real-life scenarios:
a. Motivation and output of workers
b. Age and Intelligence test scores among adults
c. Being in the choir and academic performance
d. Education and life expectancy
e. Distance and time
f. Speed and distance
g. Distance and time
Solution
a. Motivation and output of workers.
b. Generally, an increase in motivation results in an increase in the output of workers
c. This is a positive correlation
d. Age and Intelligence test scores among adults
e. Generally, as age increases from about 20 years there is a gradual and continuous decline in intelligence test scores
f. This represents a negative correlation
g. Being in the choir and academic performance
h. There is no correlation between being in the choir and academic performance
i. Education and life expectancy j. Generally the higher the level of education the higher the life expectancy Positive correlation k. Income and number of children l. Generally, people who earn more have fewer children m. Negative correlation n. Speed and distance o. The higher the speed, the greater the distance covered p. Positive correlation q. Distance and time r. Generally, you need more time to cover longer distance, depending on the same mode of transport being used s. Positive correlation
Example 7.3
The test scores of ten students in English and Mathematics are as shown:
English (x) 20 25 15 28 10 35 10 24 23 15 Mathematics (y) 25 30 20 35 15 43 20 35 30 15 Draw a scatter diagram for the data.
Solution
Let us use the steps to answer the question
1. From the data provided, the independent variable is being taken as English Marks and dependent variable as Mathematics marks.
2. We will use 2cm to represent 10marks on both axes.
3. Finally plot the ordered pairs (20, 25), (25, 30), (15, 20) (28, 35), (10, 15), (35, 43), (10, 20), (24, 35), (23, 30) and (15, 15) The graph is as shown:
Figure 7.7: A scatter graph showing the correlation between student’s English and Maths scores
Example 7.4
The data below shows the temperature of the day and the number of people wearing jackets.
Temperature (x) 35 6 17 15 21 20 10 7 Number wearing jackets (y) 11 45 22 30 20 25 32 42 Draw a scatter plot for the data.
Solution
Steps involved in drawing a scatter diagram:
1. Draw two perpendicular lines to from the x-y plane on a sheet of graph paper.
2. Choose an appropriate scale for the axes (x-axis – independent variable
- representing temperature and y-axis – dependent variable - representing number wearing jackets)
3. The ordered pairs (35, 11), (6, 45), (17, 22), (15, 30), (21, 20), (20, 25), (10,
32) and (7, 42) are plotted in the x-y plane.
Figure 7.8: A scatter graph showing the temperature and the number of people wearing jackets
Activity 7.2: How to construct a scatter diagram using Excel In small groups, work through the following steps to produce a scatter diagram in Excel of the data above.
Step 1: Open a new document in the Excel spreadsheet.
A section of the interface should look like the picture below
Figure 7.9: Excel spreadsheet
Step 2: Input your data in the first and second columns of the spreadsheet.
We will use the data below Temperature (x) 35 6 17 15 21 20 10 7 Number wearing jackets (y) 11 45 22 30 20 25 32 42 Your spreadsheet should be similar to the one below
Figure 7.10
Step 3: Select the data, go to Insert and choose Scatter from the drop-down menu.
Figure 7.11
Step 4: Insert the title and the description of the x and y axes.
To do this, click on the plus sign beside the scatter graph. This will open a dropdown menu titled Chart Elements. Choose “Axis Titles” from the dropdown menu.
Figure 7.12
Step 5: Your final plot should be similar to the chart below.
Figure 7.13: A scatter graph showing the temperature and the number of people wearing jackets
The scatter plot helps us to visualise the relationship between two variables. An alternative way to describe the relationship is to use a linear function (a straight line). Note that this is used to approximate the relationship between the variables being compared as the data points do not form a straight line in most cases. Hence the approximated line is called a line of best fit. The line of best fit should either pass through most of the points or closest to most of the points. Because of this, the line of best fit is usually used to predict data values within the given data.
This is known as interpolation. When used to predict data values outside a given data, it is called extrapolation. Pairs of data values (points) relatively far from the line of best fit are called outliers. We will obtain the line of best fit through observation.
Example 7.5
The scatter graph shows a sample of students’ mid-term and end-of-term physics scores.
Figure 7.14: A scatter diagram showing student’s mid- term scores and their end of term scores in physics
a. Make a copy of the graph and draw a line that best fits the data.
b. Use your graph to find the equation of the line of best fit.
c. Do you identify any outliers in the data?
d. Use your line of best fit to estimate the End of term score of a student who scored 30 in the midterm.
Solution
a.
Figure 7.15: A scatter diagram showing student’s mid-term scores and their end of term scores in physics with a line of best fit The line of best fit in this example passes through points (40,35) and (10, 20). It looks close to most of the points except for (38, 12).
b. The equation of the line of best fit is of the form y = mx + c
c. Where x = mid-term scores and y = end-of-term scores, m is the slope or gradient of the line and c is the y-coordinate of the y-intercept.
Figure 7.16: Calculating the equation of the line of best fit From the graph c = 15 since the graph intersects the y-axis at (0, 15) The gradient, m = change in y_________ change in x = 15/30 = 1 __ 2 Therefore, the equation of the line of best fit is: y = 1/2 x + 15
d. (38, 12) is an outlier since it is far away from the line of best fit and it shows a student who performed very well in the mid terms, but poorly in the end of term exams. All the other students show positive correlation.
e. To do this, trace a vertical line from the mid-term axis at 30 to meet the line and then trace a horizontal line to the end-of-term axis and record the value.
Figure 7.17: Calculating an estimate for end of term score for a student who scored 30 in their mid term From the graph, a student who scored 30 in the midterm is expected to have a score of 30 at the end of the term. Note that this answer can vary depending on where you have drawn your line of best fit by eye, but it should be around this score.
We have learnt that the relationship between variables of bivariate data can be positive linear, negative linear, non-linear and no association. We will now apply this knowledge to a few examples.
Example 7.6
Study the graphs below and for each of them determine the type of correlation between the variables they represent.
1. Data showing hours 11 sportswomen hours of exercise and their weight.
Figure 7.18
2. Data showing the velocity-time plot of a vehicle.
Figure 7.19
3. Marks of 10 students in Mathematics and Physics
Figure 7.20
4. Age of 11 companies and their annual profit.
Figure 7.21
Solution
1.
This shows a negative correlation between hours of exercise and weight.
This suggests that an increase in hours of exercise is associated with a lower weight.
2.
This shows a perfect negative correlation between time and velocity. Every 20m/s decrease in velocity is associated with 1 minute increase in time.
3.
The scatter plot shows a positive correlation between mathematics scores and Physics scores. Generally, a higher mathematics score is associated with a higher Physics score. Likewise, lower Mathematics scores are associated with lower Physics scores.
4.
This shows no correlation between the age of a company and its yearly profit.
Example 7.7
The table gives information about the distance from a school and the monthly rent of 11 houses.
Distance from the school (km) Monthly rent (GH¢ ) 12.5 250 20.0 100 18.0 200 10.0 400 7.5 500 5.5 470 5.0 550 14.5 240 10.0 280 8.0 420 12.5 360
a. Draw a scatter graph for the data.
b. Draw a line of best fit for the data.
c. Find the gradient of the equation of the line of best fit and interpret your answer.
d. Find the equation of the line of best fit and use it to find the monthly rent of a building which is 15km from the school.
e. Determine the type of correlation between distance and monthly rent.
Solution
a.
Figure 7.22: A Scatter graph showing the distance of a house from a school and the monthly rent
b. The equation is given as R = mD + C where R=monthly rent in Ghana Cedis, D = distance from the school (kilometres) and C is the value of R when D = 0km
Figure 7.23: Calculating the equation of the line of best fit From the graph C =700
c. Gradient = − change in rent_____________ change in distance = − 600/20 = − 30 GH¢ per km
d. Interpretation of the gradient:
e. For each km further from the school, the cost of rent is reduced by GH¢ 30.00
f. Therefore, the equation of the line of best fit is: R = − 30D + 700
g. When the distance is 15km, Rent = − 30 × 15 + 700 = − 450 + 700 = GH¢ 250.00 Or you could read this from the graph directly. Remember that this answer can vary slightly depending on where you have drawn your line of best fit.
h. The graph shows a negative correlation between distance from the school and monthly rent. The further away from the school a house is located, the cheaper the cost of rent.
Example 7.8
The scatter graph below shows data on daily sales and profits of a sample of traders.
Figure 7.24: A scatter graph showing the daily sales against the daily profits of some traders
a. Use the graph to complete the data below Daily sales (GH¢ ) 400 600 600 880 960 880 1200 1640 Daily profit (GH¢ ) 120 180 260 230 390 380
b. Find the equation of the line of best fit and interpret the gradient of the equation.
c. Determine the type of correlation between daily sales and daily profit.
Solution
a.
Daily sales (GH¢ ) 400 600 600 880 960 880 1200 1400 1640 Daily profit (GH¢ ) 100 120 180 190 260 230 300 390 380
b. The equation passes through points (400, 100) and (1200, 300). We will use these points to calculate the gradient of the line of best fit.
c. The best-fit equation is: P − = mS + C where P=daily profit and S=daily sales Gradient = 300 − 100/1200 − 400 = 200/800 = 1/4 This means that for every GH¢ 4.00 of sales, the trader makes GH¢ 1.00 profit.
Or a quarter of sales is profit.
Or profit is equal to 25% of sales.
The equation of the line of best fit is P = 1/4 S + C C is the y-intercept and we can see that the line will meet at the origin.
Therefore, the equation of the line of best fit is P = 1 __ 4 S
d. The correlation between sales and profit is positive. Generally, higher sales are associated with higher profit.
Spearman’s rank correlation coefficient To understand this concept, you must know what a monotonic function is.
A monotonic function is a function which is either:
1. always increasing as the independent variable increases OR
2. always decreasing as the independent variable increases In a nutshell, a strictly increasing function or a strictly decreasing function is monotonic.
Let us use the following graphs to illustrate the monotonic function.
Figure 7.25 Figure 7.26
Figure 7.27 Figure 7.28
f(x) is monotonic increasing. As x increases, f(x) increases.
g(x) is monotonic decreasing. As x increases, g(x) decreases.
p(x) and k(x) are not monotonic functions. As the x variable increases, the y variable sometimes increases and sometimes decreases.
A Spearman’s rank correlation coefficient is a statistical measure of a monotonic relationship between ranked variables.
Steps to Calculate Spearman’s Rank Correlation Coefficient:
1. Rank the independent variables ( Rₓ)
2. Rank the dependent variables ( R_(y))
3. For each pair ranks (Rₓ, R_(y)), calculate the deviation, d = Rₓ − R_(y)
4. Calculate the squared deviation, d²5. Calculate Spearman’s Rank Correlation Coefficient using the formula:
rₛ = 1 − 6∑d²________ n( n²− 1) Where n = total number of paired ranks The coefficient of the spearman’s rank correlation coefficient (rₛ) is between the interval [ − 1, 1], that is to say, − 1 ≤ rₛ ≤ 1 . The coefficient ( rₛ) is interpreted as follows:
1. The absolute value of rₛ indicates the strength of the relationship between the variables being compared. The larger the magnitude of the value, the stronger the relationship.
2. The sign of rₛ indicates the direction of the relationship. A positive rₛ means that an increase in the value of one variable is likely to increase the value of the other. Also, a decrease in one of the variables will cause a reduction in the other.
3. A negative rₛ means that an increase in the value of one variable will likely result in a decrease in the other variable and a decrease will increase the other.
4. When rₛ = 0 , it means there is no relationship between the variables.
Table 7.1 shows the size of a correlation coefficient (r ) and their respective interpretations. Values close to -1 or 1 represent a very strong correlation and values close to 0 indicate negligible correlation.
Table 7.1: Size of a correlation coefficient (r ) and their interpretations Size of Correlation Interpretation 1 0.9 ≤ correlation coefficient < 1.0 Very high positive correlation 2 − 1.0 < correlation coefficient ≤ − 0.9 Very high negative correlation 3 0.7 ≤ correlation coefficient ≤ 0.9 High positive correlation 4 − 0.9 ≤ correlation coefficient ≤ − 0.7 High negative correlation 5 0.5 ≤ correlation coefficient ≤ 0.7 Moderate positive correlation 6 − 0.7 ≤ correlation coefficient ≤ − 0.5 Moderate negative correlation 7 0.3 ≤ correlation coefficient ≤ 0.5 Low positive correlation 8 − 0.5 ≤ correlation coefficient ≤ − 0.3 Low negative correlation 9 0.3 ≤ correlation coefficient < 0.0 Negligible positive correlation Size of Correlation Interpretation 10 0.0 < correlation coefficient ≤ − 0.3 Negligible negative correlation 11 Correlation Coefficient = 0 No correlation 12 Correlation Coefficient= 1 Perfect positive correlation 13 Correlation Coefficient= − 1 Perfect negative correlation The examples below will enhance our understanding.
Example 7.9
The data shows the test scores of 7 students in Maths and Costing Maths (x) 30 34 40 44 57 68 78 Costing (y) 42 54 57 68 83 82 90
a. Illustrate the relationship between the scores with a scatter plot,
b. Calculate the Spearman’s rank correlation
c. Describe the correlation between the scores
Solution
a. Scatter plot of Maths and Costing scores of a sample of 7 students.
Figure 7.29: Scatter plot
b. Table 7.2: Generated values for Spearman’s rank correlation Mathematics (x) Costing (y) Rₓ R_(y) Rₓ − R_(y) = d d²30 42 7 7 0 0 34 54 6 6 0 0 40 57 5 5 0 0 44 68 4 4 0 0 57 83 3 2 1 1 68 82 2 3 − 1 1 78 90 1 1 0 0 Total 2 Now using the formula:
rₛ = 1 − 6∑d²_ n(n²− 1) , n = 7, and ∑d²= 2 rₛ = 1 − 6(2)_ 7(7²− 1) = 1 − 1_ 28 = 0.964
c. This shows a strong positive correlation between mathematics and costing scores. This shows that a student scoring high in mathematics will likely score high in Costing. Likewise, a student scoring low in mathematics will score low in Costing.
Example 7.10
Mr. Senyo and Mrs Awudi ranked the paintings of 12 students. Table 7.3 gives information about their ranks.
Table 7.3: Ranks of students Student Mr. Senyo Mrs Awudi Addo 5 5 Nii 6 6 Aba 2 1 Wan 8 7 Afi 3 3 Fofo 1 4 Ali 4 2 Student Mr. Senyo Mrs Awudi Ata 9 9 Ago 10 8 Dua 7 12 Yaw 11 10 Ayi 12 11
a. Calculate the Spearman’s rank correlation
b. Describe the correlation between the scores
Solution
a. Table 7.4: Values for calculating Spearman’s rank correlation Student Mr. Senyo (Rₛ) Mrs Awudi( R_(A)) d = Rₛ − R_(A) d²Addo 5 5 0 0 Nii 6 6 0 0 Aba 2 1 1 1 Wan 8 7 1 1 Afi 3 3 0 0 Fofo 1 4 -3 9 Ali 4 2 2 4 Ata 9 9 0 0 Ago 10 8 2 4 Dua 7 12 -5 25 Yaw 11 10 1 1 Ayi 12 11 1 1 ∑d²= 46 Since there are 12 pairs of ranks, n = 12 rₓ = 1 − 6∑d²_______ n( n²− 1) rₓ = 1 − 6 × 46/12(12²− 1) = 1 − 6 × 46/12 × 143 = 1 − 0.160839 = 0.83916 = 0.84
b. A correlation coefficient of 0.84 shows a strong positive correlation between Mr. Senyo and Mrs Awudi’s scores of the students. This means that the paintings which Mr Senyo scored highly, Mrs Awudi was likely to score highly. Conversely, the paintings which Mr Senyo scored low were also likely to be scored low by Mrs Awudi.
Data with tied values This is when two or more observations have the same value. There are many ways of dealing with tied values. One method is to assign the mean rank of the tied values for each tied observation.
The next example will focus on tied values.
Example 7.11
The data below gives information about the price and mileage of 10 cars.
Mileage (in 1 000km) Price (in GH¢ 1 000) 80 20 70 40 40 50 30 60 20 70 90 20 50 46 20 80 40 60 20 100
a. Draw a scatter graph for the data
b. Calculate the Spearman’s rank correlation coefficient for the data.
c. Comment on the strength of the correlation
d. Interpret the correlation between price and mileage of the used cars.
Solution
Scatter plot of Mileage and Price of Used Cars
Figure 7.30
a. We will first rank the mileage. To do this, we will arrange the values in ascending order:
90 80 70 50 40 40 30 20 20 20 1ˢᵗ2ⁿᵈ3ʳᵈ4ᵗʰ7ᵗʰ40 is tied at the 5ᵗʰand the 6ᵗʰpositions. So, we will assign the average of the 5ᵗʰand the 6ᵗʰpositions = 5 + 6/2 th = 11/2 = 5.5th position to 40 Also, 20 is tied at the 8ᵗʰ, 9ᵗʰand 10ᵗʰpositions. We will assign 8 + 9 + 10/3 = 27/3 = 9ᵗʰposition to 20.
90 80 70 50 40 40 30 20 20 20 1ˢᵗ2ⁿᵈ3ʳᵈ4ᵗʰ5.5ᵗʰ5.5ᵗʰ7ᵗʰ9ᵗʰ9ᵗʰ9ᵗʰNext, let’s use the same method to rank the Price.
First, arrange the data values in descending order:
100 80 70 60 60 50 46 40 20 20 1ˢᵗ2ⁿᵈ3ʳᵈ6ᵗʰ7ᵗʰ8ᵗʰ60 is tied at the 4ᵗʰand the 5ᵗʰpositions. So, we will assign the average of the 4ᵗʰand the 5ᵗʰpositions = 4 + 5/2 th = 9/2 = 4.5th position to 60 Likewise, 20 is tied at the 9ᵗʰand 10ᵗʰpositions. We will assign 9 + 10/2 = 9.5th position to 20.
100 80 70 60 60 50 46 40 20 20 1ˢᵗ2ⁿᵈ3ʳᵈ4.5ᵗʰ4.5ᵗʰ6ᵗʰ7ᵗʰ8ᵗʰ9.5ᵗʰ9.5ᵗʰLet us update our table with the ranks. Take care when doing this that you are giving the correct ranks to the data. Then we will use it to calculate the deviations (d) and the squared deviations (d²) Mileage (in 1,000km) Price (in GH¢ 1,000) Rₘ Rₚ d = Rₘ − Rₚ d²80 20 2 9.5 2 − 9.5 = − 7.5 (− 7.5)²= 56.25 70 40 3 8 3 − 8 = − 5 (− 5)²= 25 40 50 5.5 6 5.5 − 6 = − 0.5 (− 0.5)²= 0.25 30 60 7 4.5 7 − 4.5 = 2.5 (2.5)²= 6.25 20 70 9 3 9 − 3 = 6 6²= 36 90 20 1 9.5 1 − 9.5 = − 8.5 (− 8.5)²= 72.25 50 46 4 7 4 − 7 = − 3 (− 3)²= 9 20 80 9 2 9 − 2 = 7 7²= 49 40 60 5.5 4.5 5.5 − 4.5 = 1 1²= 1 × 1 = 1 20 100 9 1 9 − 1 = 8 8²= 64 ∑d²= 319
b. Finally, use the formula:
c. rₓ = 1 − 6∑d²_______ n( n²− 1) rₓ = 1 − 6 × 319/10( 10²− 1) rₓ = 1 − 1914/10(99) rₓ = 1 − 1914/990 rₓ = 1 − 1.933 rₓ = − 0.93
d. A correlation coefficient of − 0.93 shows a strong negative correlation between the mileage or distance a car has covered and the price.
e. This means that as the mileage of the car increases, the price decreases.
Thus, a high mileage car will have a lower price, compared with a low mileage car which will have a higher price.
1. Determine the independent and dependent variables in the following scenarios.
a) The effect of absenteeism on academic performance
b) The effect of advertisement on sales.
c) Grades and hours of study
d) Expenditure and income
e) Midterm and end-of-term History scores.
f) Height and weight of students
g) Amount of rainfall and bags of cocoa harvested
2. State the type of correlation in the following real-life scenarios.
a) Relationship between husband’s age and wife’s age
b) Height and occupation
c) Age and life expectancy
d) Gender and intelligence
e) Alcohol consumption and driving ability
f) Water consumption and temperature
g) Hours of study and grades
h) Wages earned and number of days worked
3. The data below shows the mid-term and end-of-term accounting scores of a sample of students.:
Mid-term (x) 60 75 45 80 30 85 40 58 75 60 End-of- Term(y) 75 90 60 90 45 92 50 70 84 60
(a) Construct the data on a scatter plot manually.
(b) Use Excel to construct a scatter plot for the data.
4. Aba experimented to determine whether different drug dosages affect the duration of relief from malaria. She used a random sample of 12 patients and recorded the following observations.
Dosage 3 3 4 5 6 6 7 8 8 9 6 5 Duration of relief (hours) 7 5 10 7 12 14 20 16 22 20 13 8
(a) Manually draw a scatter plot for the data.
(b) Calculate the Spearman’s rank correlation coefficient for the data and interpret the value.
5. The data below shows the income and expenditure of 10 Ghanaian workers.
Income (in GH¢
100) 60 35 40 63 75 35 20 70 60 65 48 Expenditure (in GH¢ 100) 55 50 55 40 68 40 35 70 70 50 40 Draw a scatter plot for the data.
6. The scatter diagram below provides information about 15 cars. It shows their engine size (litres) and the distance (km) they can travel on one litre of fuel. [Take the blue line as a line of best fit]
(a) Use your graph to find the equation of the line of best fit.
(b) Use your line of best fit to find the distance a car with an engine size of 1.5 litres travels on a litre of fuel.
7. Study the graphs below carefully and for each of them determine the type of correlation between the variables they represent.
a Distance and time travelled by a black ant.
b Mathematics and history scores of students c Price and demand of laptops d Height above sea level and air temperature.
8. The table gives information about the distance from a city centre and the monthly rent of 11 houses.
Distance from the city centre (km) Monthly rent (GH¢ ) 14.5 400 22.0 300 20.0 400 Distance from the city centre (km) Monthly rent (GH¢ ) 12.0 600 19.5 700 7.5 670 7.0 750 16.5 440 12.0 480 10.0 520 14.5 560
a. Draw a scatter graph for the data.
b. Draw a line of best fit for the data.
c. Find the gradient of the equation of the line of best fit and interpret your answer.
d. Find the equation of the line of best fit and use it to find the monthly rent of a building which is 17km from the city centre.
e. Determine the type of correlation between distance from the city centre and monthly rent.
f. Calculate the Spearman’s rank correlation coefficient between distance from the city centre and monthly rent.
9. The scatter graph shows data on the amount of investment and returns earned by a sample of businesswomen. [Take the green line as a line of best fit]
a. Use the graph to complete the data below Investment (million GH¢ ) Returns (million GH¢ )
b. Find the equation of the line of best fit. Interpret the gradient of the equation.
c. Madam Malik made 1 million Ghana Cedis returns on his investment, estimate the amount she invested.
d. Determine the type of correlation between Investment and Returns.
e. Calculate the Spearman’s rank correlation coefficient between investment and returns.
f. Interpret your correlation coefficient.
10. Mrs Paintsil and Mrs Blay ranked the performance of 12 Adowa dancers.
The table below gives information about their ranks.
Student Mrs Paintsil Mrs Blay
Adwoa 2 4 Nhyiraba 4 1
Antoa 8 8 Asantewaa 3 2
Yaa 11 12 Dufie 1 3
Abena 7 7 Ataa 12 9
Akos 10 11 Ohenewaa 9 10
Naana 6 5 Morowa 5 6
(a) Calculate the Spearman’s rank correlation
(b) Describe the correlation between Mrs Paintsil and Mrs Blay’s ranks.
11. The data below gives information about the price and age of 10 used cars.
Age (years) Price (in GH¢ 1 000)
10 20 9 40 6 50 5 60 4 70 11 20 7 46 4 80 6 60 4 100
a. Calculate the Spearman’s rank correlation coefficient for the data.
b. Comment on the strength of the correlation.
c. Interpret the correlation between price and mileage of the used cars.
12. The scatter plot below shows the height and weight of 10 students.
13. Take the red line as the line of best fit.
a. Use your graph to complete the table below Height (cm) 170 150 130 153 145 155 Weight(kg) 75 67 72 65 60 67
b. Find the gradient of the line of best fit and interpret your answer.
c. Find the equation of the line of best fit and use it to estimate the height of a student who is 70kg.
d. Find the Spearman rank correlation between height and weight.
e. Describe the correlation between height and weight.
Which of the following sets of data is bivariate?
A scatter plot of the monthly salary and monthly savings of some teachers shows that as monthly salary increases, monthly savings generally increase. What type of correlation does this show?
In an investigation of the effect of advertisement on sales, which variable is the independent variable?
A researcher obtains Spearman’s rank correlation coefficient for two ranked variables. What is the best interpretation?
A researcher at the Kumasi Metropolitan Education Directorate collected data on the weekly hours of study and Additional Mathematics test scores of 10 Senior High School students. The results are shown below.
| Student | Kofi | Ama | Yaw | Efua | Kwame | Abena | Kojo | Adjoa | Fiifi | Esi |
|---|---|---|---|---|---|---|---|---|---|---|
| Hours of study () | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 |
| Test score () | 35 | 50 | 48 | 55 | 62 | 70 | 68 | 80 | 85 | 90 |
Distinguish between univariate and bivariate data. Give one example of each from the scenario.
Identify the independent and the dependent variables in the table. State the type of correlation you would expect between them.
Calculate Spearman’s rank correlation coefficient for the data. Show your working.
Interpret the value obtained in (c) in terms of the strength and direction of the relationship between hours of study and test scores.
Explain why the correlation found in (d) does not prove that more hours of study cause higher test scores. Give one other factor that may affect the test scores.
A mobile phone dealer in Accra recorded the monthly advertising expenditure and sales for eight months. Advertising expenditure is in hundreds of Ghana cedis (), and sales are in thousands of Ghana cedis ().
| Month | Jan | Feb | Mar | Apr | May | Jun | Jul | Aug |
|---|---|---|---|---|---|---|---|---|
| Advertising expenditure () | 8 | 10 | 12 | 15 | 20 | 25 | 30 | 35 |
| Sales () | 45 | 50 | 56 | 60 | 70 | 130 | 88 | 95 |
State whether the data in the table is univariate or bivariate. Give a reason for your answer.
Explain what a line of best fit is. State one use of the line of best fit in interpolation.
Ignoring the outlier, a line of best fit for the data passes through and . Find the equation of this line in the form .
Use the equation in (c) to estimate the sales when advertising expenditure is 25 (that is, GH¢2,500). State whether this is interpolation or extrapolation.
Identify the outlier in the table. Explain one way this outlier could affect conclusions drawn from the data.