Which of the following best distinguishes correlation from regression?
Strand 4 · Handling Data
Additional Mathematics Year 3 Learner Material, Section 7: Statistics
In statistics, analysing relationships between variables is crucial for making informed decisions and solving real-world problems. Regression and correlation are essential concepts that help us understand how variables are related and how one can be used to predict another. While correlation measures the strength and direction of a relationship, regression goes further by providing a mathematical model for predictions. By collecting data through suitable research methods and tools, we can apply statistical techniques to fit a linear function and identify the line of best fit. This line allows us to make accurate forecasts and solve contextual problems based on the data. Effective analysis, interpretation, and presentation of findings not only improve understanding but also support evidence- based decision-making in fields such as business, health, education, and science. By the end of this section, you should be able to:
• Correlation : A measure of the strength and direction of a relationship between two variables. • Regression : This provides a mathematical model to predict one variable based on another. • Line of Best Fit : A straight line that best represents the data in a scatter plot, found through regression analysis. • Data Collection : Use appropriate research tools such as questionnaires, interviews, observations, or online databases to gather accurate information. • Data Analysis : Apply suitable statistical techniques (e.g., scatter plots, regression equations) to examine the data. • Interpretation of Results : Draw conclusions from the analysis to answer questions or solve problems in the given context.
In statistics, we often want to know how two or more numeric variables are related. For example, a teacher may want to know if there is a relationship between the grades of students on a mid-term exams and on their end-of-term exams. If there is a relationship, what is the relationship, and how strong is it? In another example, a business owner may want to know how advertising expenditure affects sales, or a farmer may want to know how rainfall influences crop yield. The type of data described in the examples is bivariate data (two variables). In reality, statisticians use multivariate data, meaning many variables. In this chapter, we will study the simplest form of regression, linear regression, with one independent variable and one dependent variable . This involves data that fit a line in two dimensions. We will also revise correlation since it is related to regression. Before we learn the process of fitting a linear function to a set of data, we will first look at the meaning of variable as used in statistics and the difference between correlation and regression. Although correlation and regression are related, they serve different purposes.
In statistics, a variable is a characteristic or measurement that can be determined for a population. Variables may describe values like weight in kg, height in metres, favourite food, satisfaction level, age, and gender.
Given the table below, identify the variables involved. Table 7.1: Expenditure versus number of cars sold Expenditure on advertisement (GHc) Number of cars sold 25,000 25 30,000 40 35,000 43 40,000 48
Variable 1: Expenditure on advertisement Variable 2: Number of cars sold
This is the variable you manipulate to see how it affects another variable. It is also referred to as the Predictor or Explanatory variable. Example: Hours studied (predicts exam score), Price of an item (determines the quantity sold), Temperature (determines the amount of water consumed). Usually, the independent variable can be determined when used in a context. Ask yourself: “What is being changed, used, or controlled to see its effect?” This is the independent variable.
This is the variable that changes as a result of changes in the independent variable. It is also referred to as the Outcome or Response variable. Example: Exam scores (outcome of hours of study), Expenditure (determined by the income), Amount of water consumed (determined by the temperature). Usually, the dependent variable can be determined when used in a context. Ask yourself: “What changes because of what I did or something that was done?” This is the dependent variable.
State the independent and dependent variables in the following research scenarios
Corelation is a measure of the nature and strength of the relationship between two or more variables. It describes how well two variables go together. The relation between two variables may be positive, negative or non-existent (no correlation). • A positive correlation between two variables means that an increase in the value of one variable is likely to increase the value of the other. Likewise, a decrease in one of the variables will cause a reduction in the other. It is a relationship that moves in tandem (in the same direction). It shows a direct relation between two variables. An example of a positive correlation is the relationship between temperature and water consumption. • A negative correlation describes an inverse relationship. As one variable increases, the other decreases. It describes the relationship between two variables that change in opposite directions. An example is the relationship between age and agility. As your age increases, your agility decreases. Another example is speed and travel time. The higher the speed, the shorter the travel time. Other examples include preparation and mistake, supply and price, exercise and body weight, inflation and purchasing power.
• Zero or no correlation. This occurs when there is no relationship between the two variables. It also means changes in one variable do not predict changes in the other. An example is weight and intelligence. Your weight has nothing to do with your intelligence, and an increase in your weight has nothing to do with your intelligence. Other examples are blood type and prosperity, shoe size and favourite colour, weight and income, and race and intelligence.
Regression is a statistical method used to model the relationship between two variables and provide an equation that can be used to make predictions. In simple linear regression, we have: Where: • y = dependent variable (predicted value) • x = independent variable (predictor) • c = intercept (value of y when ) • m = slope (rate of change of y with respect to x ) Typically, you choose a value to substitute for the independent variable and then solve for the dependent variable. Table 7.2: Distinguishing Between Correlation and Regression Feature Correlation Regression Measures strength and direction of Predicts the value of one variable Purpose relationship from another. Single value (correlation An equation of a line Nature of Result coefficient, r) between -1 and 1 ( ) No distinction between dependent One variable is dependent, the Direction and independent variables other is independent Describes relationship and Interpretation Describes association enables prediction
Regression is useful when:
In year 2, we learned how to plot data points on a scatter plot which enabled us to visually observe the relationship between two variables. We have noticed that many of these scatter plots show a clear, linear trend, where the points seem to cluster around a straight line. This straight line is called the line of best fit , and it is a powerful tool in statistics that allows us to model the relationship between variables mathematically. Now we will look at the process of fitting a linear function to a set of data. We will learn how to find the equation of the line of best fit and, most importantly, how to use this line to make predictions and solve problems in real-world contexts.
Data rarely fits a straight line exactly. Usually, we must be satisfied with rough predictions. A straight line that represents the relationship between a dependent variable (Y) and one or more independent variables (X) in a linear regression model is called a line of best fit. It is a single straight line that best represents a trend in the data. The line of best fit is usually used to predict data values within the given data. This is known as interpolation . When used to predict data values outside a given data, it is called extrapolation . Pairs of data values (points) relatively far from the line of best fit are called outliers. Extrapolation is a lot less reliable that interpolation.
While we can visually draw a line of best fit by hand, a more accurate method is to use a statistical technique called linear regression . Here, we will focus on the least squares method , which finds the line that minimises the sum of the squared vertical distances (residuals) from each point to the line.
Least-Squares Criteria for Best Fit The process of fitting the best-fit line is called linear regression. We assume that the data are scattered about a straight line. To find that line, we minimise the sum of the squared errors (SSE), or make it as small as possible. Any other line you might choose would have a higher SSE than the best-fit line. This best-fit line is called the least-squares regression line. To understand the SSE, consider Figures 7.1 and 7.2, which show the same data with different lines fitted by hand.
[Figure] [Figure] Figure 7.1: Error squares of a fitted line Figure 7.2 : Error squares of a fitted line Which of the lines of best fit gives the least SSE? If the line of best fit in Figure 7.1 is used for predicting the values of y the results would be 14, 16, 18, 20, 22, and 24 as can be read from the graph. Comparing this to the actual y-values, this respectively gives errors of: , , , , If the line of best fit in Figure 7.2 is used for predicting the values of the results would be 11, 14, 17, 20, 23, and 26 as can be read from the graph. Comparing this to the actual data values, this respectively gives errors of: , , , , To calculate the SSE, we will first draw a vertical line from our supposed line of best fit to each of the plotted data points, find the square of each distance and sum the result.
For Figure 7.1 For Figure 7.2 SSE SSE This means that the line of best fit in Figure 7.1 gives the least SSE. The least squares method aims to obtain the line for which the total area of these squares is the smallest. In statistics, it is known (though we will not prove) that the best-fitting line for a set of data always passes through the point . Here, is the mean of all the x-values, and is the mean of all the y-values. Given a dataset with points: for , a line with equation can be found to model the data and predict other data points. Note that: The value of can be found by substituting into the general linear equation.
The table shows values of two related variables. Table 7.3: Value of two related variables, x and y 1 2 3 4 5 6 7 9 8 5 9 10 6 4 7 1 2 1 1 4 Use the least squares method to determine the equation of the line of best fit.
Let the equation of the line of best fit be
and Table 7.4: Least Squares Regression Computation 1 9 4.5 16 2 10 5.5 9 3 6 1.5 4 4 4 1 5 7 0 2.5 0 6 1 1 1 7 2 2 4 9 1 4 1 8 1 3 1 5 4 0 0 Since lies on the line of best fit, Therefore, the equation of the best-fit line is Figure 7.3 shows the graph of .
[Figure] Figure 7.3: Scatter plot for the data showing the line of best fit
Work through this activity in pairs. The diagram shows a scatter plot of the weight and height of a sample of 10 students. [Figure] Figure 7.4: Scatter plot of Height versus Weight
Step 1: Observation a. Examine the scatter plot provided and describe what it suggests about the relationship between height and weight. b. Indicate whether the relationship is positive, negative or shows no correlation. c. Copy and complete the table below Table 7.5: Partially completed table of Height and Weight. Height 130 135 138 145 150 152 155 160 165 170 Weight 60 63 60 62 67 69 68 72 74 75
Step 2: Fitting the model a. Using the least squares method, calculate the equation of the line of best fit for the given data. b. Express the equation in the form: where is the gradient and is the y-intercept. c. Interpret the gradient of the line of best fit. Step 3: Drawing the Line of Best Fit a. Make a copy of Figure 7.4 in your graph book. Using the equation you obtained, draw the line of best fit on the diagram. b. Ensure the line passes through the point Step 4: Prediction Using your linear model, a. Estimate the height of a student who weighs 70 kg. Show all working clearly. b. Estimate the weight of a student whose height is
Step 1: Observation a. The scatter plot shows a positive correlation between height and weight. This means that taller people in this data tend to weigh more. b. Find below the completed table. Table 7.6: Height and Weight
Height 130 135 138 145 150 152 155 160 165 170 Weight 60 63 60 62 67 69 68 72 74 75 Step 2: Fitting the model a. Let the equation of the line of best fit be and Table 7.7: Least Squares Regression Computation. 130 60 400 135 63 225 138 60 144 145 62 25 150 67 0 0 0 153 69 3 2 9 155 68 5 1 25 160 72 10 5 100 165 74 15 7 225 170 75 20 8 400 Since lies on the line of best fit, , we have:
Therefore, the equation of the best-fit line is c. The gradient for the line of best fit is approximately 0.41. This means that for every 1 cm increase in height, the weight increases by about 0.41 kg. Step 3: Drawing the Line of Best Fit [Figure] Figure 7.5: Scatter plot and regression line of Height versus Weight Step 4: Prediction Using your linear model, a. When y=70, we have This means that a person who weighs 70kg is expected to be approximately 157cm tall. b. Estimate the weight of a student whose height is 165cm From the model, when , kg This means that a person who is 165cm tall is expected to weigh approximately 73.32kg.
Working in pairs, or small groups carry out the following activity.
A group of students conducted an experiment to determine how the number of hours spent studying affects performance in a mathematics test. Eight students recorded the number of hours they studied during the week before the test and their corresponding scores out of 20 marks. The data below shows the number of hours studied and the test scores Table 7.8: Hours studied and Test Scores Hours Studied 2 5 8 3 7 6 1 4 0 Test Score 6 11 16 9 14 14 4 12 4
Tasks:
The scatterplot is shown in Figure 7.6 [Figure] Figure 7.6: Scatter plot for data in Table 7.8
The scatter plot shows a general upward trend: as the number of hours studied increases, test scores also increase. This suggests a positive linear relationship between hours of study and test performance.
Let the linear model of best fit be Table 7.9: Least Squares Regression Computation. 2 6 -2 -4 8 4 5 11 1 1 1 1 8 16 4 6 24 16 3 9 -1 -1 1 1 7 14 3 4 12 9 6 14 2 4 8 4 1 4 -3 -6 18 9 4 12 0 2 0 0 0 4 -4 -6 24 16 Since lies on the line of best fit, Therefore, from the least-squares method, the equation of the line of best fit is: or or . [Figure] Figure 7.7: Scatter plot and graph of
Prediction and Reliability Using the model, the estimated test score of a student who studied for 6 hours is: . However, from the data obtained from actual readings, the score of the student who studied for 6 hours is 14, which gives an error of So, the error committed in using the model is . which falls between the interval and 1.5. Therefore, the model is appropriate for the prediction.
Research is a process of collecting information to answer a specific question, solve a problem, or understand a situation. In statistics, research begins with identifying the type of data needed and the most effective tools to collect it. Undertaking research helps us discover concepts we may not have grasped in the classroom. It may also offer us the opportunity to expand our understanding of concepts. Here we will apply the knowledge we have gained previously to analyse the data we will collect. We will also have the opportunity to present our findings to our classmates. Steps in Data Collection:
that every step of the research contributes to answering that specific inquiry. 2. Guide Data Collection The research question determines the specific data to be collected and the tools required for the task. In the cocoa yield example, a well-defined question indicates that data should be gathered on both the amount of fertiliser used and the final cocoa yield. This precision helps prevent collecting irrelevant information. 3. Saves Time and Resources By knowing exactly what we are looking for means that we can avoid the time- consuming process of gathering unnecessary data. A clear question helps to create a lean and efficient research plan, ensuring all efforts are directed only at what’s essential. 4. Determines Analysis Methods The nature of the research question directly determines the statistical tools to be used. A question about the relationship between two variables, like fertiliser and yield, will probably need regression or correlation analysis. A question about a group’s average performance, on the other hand, might simply require calculating the mean. 5. Simplifies Interpretation When the research question is clear from the start, the findings can be presented as a direct and concise answer. This makes interpreting and discussing the results easier, allowing us to clearly explain what has been discovered and how it relates to the original problem.
Working in pairs, or individually, carry out the following activity. Objective: To design and carry out a small research project that collects bivariate data using appropriate data collection tools. Task Instructions
Define the Research Question: Example: “Is there a relationship between the number of hours students’ study in a week and their last mathematics test scores?”
Identify the Variables: For this particular research question, Independent variable = Hours of study per week Dependent variable = Mathematics test scores.
Decide the Type of Data: The data to be collected will be primary data because it will be collected directly from students.
Select an Appropriate Data Collection Tool: You can use: • Questionnaire: Design a simple questionnaire that will solicit the study hours and test scores in mathematics for a particular term or semester. • Interview: Have a face-to-face interaction with learners to collect the data. • Class record: You may ask learners about their study hours and obtain their scores in mathematics from class records.
Prepare the Data Collection Tool: • For a questionnaire, write clear, concise questions (e.g., “How many hours did you study mathematics last week?”). • For an interview, prepare a short list of guiding questions.
Collect the Data: • Survey at least 10–15 students. • Record responses neatly in a table with two columns: Hours Studied and Test Score.
Organise the Data: • Present the collected data in a table. • Make sure both variables are correctly labelled.
Class Discussion Share with the class the: • The research question. • The type of data collected (primary or secondary). • The tool used and the reason for choosing it. • The raw data table.
Data analysis is the process of examining, cleaning, transforming, and modelling data to discover useful information, draw conclusions, and support decision-making. Steps in Data Analysis:
Organise the data o Arrange data in tables or spreadsheets for easier handling.
Select appropriate analysis techniques o Descriptive statistics: Mean, median, mode, range, percentages, frequency tables. o Graphical methods: scatter plots. o Inferential statistics: Correlation, regression.
Use appropriate tools o Manual calculations with calculators. o Software tools such as Microsoft Excel, Google Sheets, SPSS, or Python.
Analyse the data o Apply formulas to find averages and other statistical measures. o Draw graphs to visualise patterns or relationships.
Interpret the findings o Explain what the results mean in relation to the original research question.
Draw conclusions and make recommendations o Summarise the key insights from the analysis. o Suggest possible actions based on findings. Here is an activity specifically using regression for the indicator
Working individually, or in pairs, undertake the following activity. Objective: Apply regression techniques to analyse data, find the equation of the line of best fit, and use it for predictions. Scenario A group of agricultural researchers collected data on the amount of fertiliser applied to a plot of land and the corresponding crop yield. Use regression analysis to determine how fertiliser quantity affects crop yield. Table 7.10: Fertiliser and Crop Yield Fertiliser (cups) (x) Crop Yield (bags) (y) 1 16 2 21 3 23 4 26 5 31 6 34 7 40 8 41
Instructions Step 1 – Plot the Data Draw a scatter plot with fertiliser quantity on the x-axis and crop yield on the y-axis.
Step 2 – Perform Regression Analysis a. Using the method of least squares, calculate: i. The gradient ii. The y-intercept b. Write the regression equation in the form: Step 3 – Interpretation a. Explain what the gradient means in the context of crop yield and fertiliser use. b. Explain what the y-intercept means in the context of crop yield c. State the relationship (positive/negative correlation) between fertiliser and yield. Step 4 – Prediction a. Use your regression equation to predict the yield when 7 cups of fertiliser are applied. State whether your prediction seems reasonable based on the plotted data. b. Use your regression equation to predict the yield when 18 cups of fertiliser are applied. State whether your prediction seems reasonable based on the plotted data.
Step 1 – Plot the Data [Figure]
Figure 7.8: Scatter plot of Fertiliser versus Yield
Step 2 – Perform Regression Analysis a. Let the equation of the line of best fit be and Table 7.11: Least Squares Regression Computation. 1 16 45.5 12.25 2 21 20 6.25 3 23 9 2.25 4 26 0.25 5 31 0.5 0.25 6 34 1.5 7.5 2.25 7 40 2.5 27.5 6.25 8 41 3.5 42 12.25 i. ii. Since lies on the line of best fit, b) Therefore, the regression line is or Figure 7.9 shows the graph of
[Figure] Figure 7.9: Scatter plot and graph of Step 3 – Interpretation a. The gradient of the regression line is This means that, on average, each additional cup of fertiliser is associated with a yield increase of about 3.67 bags. In other words, increasing fertiliser by 1 cup raises yield by approximately 3.67 bags. b. The y-intercept of the regression line is 12.50 This gives the model’s predicted yield when no fertiliser is applied (i.e., when ). This means that the model predicts about 12.5 bags when no cups of fertiliser are applied. (In practice, this is an extrapolated value and should be interpreted cautiously — it may reflect a baseline yield from other factors (soil, rainfall), but the experiment did not include , so the intercept is mainly a model parameter rather than a direct measured value.) c. Correlation: There is a clear positive correlation between cups of fertiliser applied and crop yield. As fertiliser increases, yield tends to increase. Step 4 – Prediction a. Using the regression equation ( ), the predicted crop yield for 7 cups of fertiliser is: The predicted yield is 38.19 bags. This prediction seems reasonable, as it is very close to the actual observed yield of 40 bags for 7 cups of fertiliser, as seen in the scatter plot. b. Using the regression line, when 18 cups of fertiliser are applied, the yield will be: sacks. The prediction is NOT reasonable because of the following:
i. The original data cover fertiliser amounts from 1 to 8 cups. Predicting at x=18 is far outside that range, so the linear trend may not continue — yields could level off, fall, or behave nonlinearly beyond the observed range. ii. Crop yield typically cannot increase linearly without bound. Too much fertiliser may damage plants or give diminishing returns.
Woking in pairs, or individually, carry out the following activity. Objective Apply regression techniques to analyse the price and mileage of used cars, find the equation of the line of best fit, and use it for predictions. Background In the automobile market, the value of a used car is influenced by several factors such as brand, age, fuel efficiency, and mileage. One common belief is that as the mileage of a car increases, its price decreases. Investigate whether this relationship holds using the secondary data below. Table 7.12: Mileage and Price Mileage (in 1,000km) Price (in 1,000.00) 80 30 70 40 40 50 30 60 20 70 90 20 50 46 20 80 40 60 10 84
Procedure Step 1 – Plot the Data Draw a scatter plot with: • Mileage on the x-axis • Price on the y-axis Step 2 – Perform Regression Analysis
Mean Mileage Mean Price and Table 7.13: Least Squares Regression Computation. Mileage Price (in (in 1,000km 1,000) 80 30 1225 70 40 625 40 50 25 30 60 225 20 70 625 90 20 2025 50 46 25 20 80 625 40 60 25 10 84 1225
(i) Gradient of the line is -0.746. This means that the price decreases by about Ghc 0.75 or 75 pesewas for each 1km driven. ii) Since lies on the line of best fit, The y-intercept is GHC87,564 b) Therefore, the regression line is Figure 7.11 shows the graph of
[Figure] Figure 7.11: Scatter plot and graph of
Step 3 – Interpret Your Findings • Gradient: –0.746 means that for every extra 1 km on the odometer, the price drops by about 75 pesewas. • Intercept: When mileage is 0, the estimated price is 87,564 (theoretical, since cars are never truly at zero mileage). • The relationship is negative — the higher mileage the lower the price. Step 4 – Correlation Coefficient Table 7.14: Correlation Coefficient Computation. Mileage (in Price (in GHC1,000) 1,000km) 80 30 2 9 70 40 3 8 40 50 5.5 6 30 60 7 4.5 20 70 8.5 3 90 20 1 10
50 46 4 7 20 80 8.5 2 40 60 5.5 4.5 10 84 10 1 A correlation coefficient of shows a strong negative correlation between mileage and price, confirming that as the mileage of the car increases, the price decreases. Step 5 – Prediction When the mileage is 60 000km, using the model, the price The prediction is reasonable since it is within the range of the observed values.
The first four questions are multiple choice. Choose the correct option.
Which of the following is an example of primary data? A. Data from last year’s school records B. Information from an online encyclopaedia C. Measurements you take during a science experiment D. Statistics from a published government report
Which data collection tool is most suitable for finding out the opinions of many students in a short time? A. Interview B. Observation C. Questionnaire D. Weighing scale
In data collection, ethical consideration means: A. Asking only your friends to respond B. Getting consent and protecting participants’ privacy C. Making sure data supports your opinion D. Only collecting data you like
A thermometer used to record daily temperatures is an example of: A. Interview B. Measuring instrument C. Observation checklist D. Questionnaire
State a difference between primary data and secondary data.
List three tools you can use to collect primary data.
Why is it important to define the purpose of your research before collecting data?
You have been assigned to find out how many hours students in your class spend on the internet each week. a. State whether you would use primary or secondary data. Give a reason. b. Name two appropriate tools you could use for data collection. c. Explain one advantage of using each tool.
A science club wants to know the effect of different amounts of sunlight on bean plant growth. a. Suggest a suitable data collection tool. b. Describe how you would use this tool in the experiment.
Explain the difference between correlation and regression.
Give two examples each of positive correlation and negative correlation.
A company finds that monthly sales can be predicted from advertising cost using the equation a. Interpret the values 1000 and 15. b. Predict sales when is spent on advertising.
Mention three real-life situations where regression analysis would be useful.
Define the term “line of best fit” and explain its importance.
Given the data below, find the equation of the line of best fit and use it to predict when . Table 7.15: Values of two related variables, and . 2 3 4 6 8 5 7 9 13 17
State three real-life situations where the line of best fit can be applied.
The data below gives information about the price and age of 10 used cars. Table 7.16: Age and Price of cars. Price (in Age (years) ) 10 20 9 40 6 50 5 60 4 70 11 20 7 46 4 80 6 60 4 100 Tasks: a. Construct a scatter plot using the data in Table 7.16.
b. Describe the trend shown by the scatter plot. What does it suggest about the relationship between age and the price of cars? c. Fit a linear model to the data using the method of least squares. Express the equation in the form: . Using the equation you obtained, draw the line of best fit on the scatter plot. d. Prediction and Reliability: If the margin of error for predicting used car prices is , use your model to estimate the price of a 7-year-old car. 18. The table shows values of two related variables. Table 7.17: Values of two related variables, and . 11 12 13 14 15 16 17 18 19 15 90 100 60 40 70 10 20 10 10 40
Tasks: a. Construct a scatter plot for the data. b. Describe the trend shown by the scatter plot. What does it suggest about the relationship between and values? c. Fit a linear model to the data using the method of least squares. Express the equation in the form: . Using the equation you obtained, draw the line of best fit on the scatter plot. Interpret the gradient of the line of best fit and the y-intercept in the context of your data. d. Calculate to measure the strength and direction of the relationship. Interpret whether the relationship is weak, moderate, or strong. e. If the margin of error is , use your model to estimate the value of y when . Will the model be reliable for this prediction? 19. The diagram shows a scatter plot of the drop height and bounce height of a tennis ball. [Figure] Figure 7.12: Scatter plot of Drop Height versus Bounce Height
Step 1: Observation a. Examine the scatter plot provided and describe what it suggests about the relationship between drop height and bounce height. b. Indicate whether the relationship is positive, negative, or shows no correlation. c. Copy and complete the table below Table 7.18: Drop height and Bounce height. Drop height 30 60 90 120 150 180 210 240 Bounce height 18 72 84 132 144 Step 2: Fitting the model a. Using the least squares method, calculate the equation of the line of best fit for the given data. b. Express the equation in the form: where is the gradient and is the intercept on the y axis. c. Interpret the gradient of the line of best fit. Step 3: Drawing the Line of Best Fit a. Make a copy of the scatter plot in your graph book. Using the equation you obtained, draw the line of best fit on the diagram. b. Ensure the line passes through the point Step 4: Prediction Using your linear model, a. Estimate the bounce height when the drop height is . b. Estimate the drop height if the bounce height is . 20. A group of students experimented to determine how the number of hours spent studying affects performance in a Physics test. Ten students recorded the number of hours they studied during the week before the test and their corresponding scores out of 30 marks. The data below shows the number of hours studied and the test scores Table 7.19: Hours Studied and Test Scores Hours Studied 2 5 8 3 7 6 1 4 9 5 Test 17 21 27 20 24 25 15 22 27 22 Tasks: a. Construct a scatter plot using the data above with the x-axis as hours studied and the y-axis as test scores. b. Describe the trend shown by the scatter plot. What does it suggest about the relationship between hours of study and test scores?
c. Fit a linear model to the data using the method of least squares. Express the equation in the form: . Using the equation you obtained, draw the line of best fit on the scatter plot. Interpret the gradient of the line of best fit and the y intercept in the context of your data. d. Calculate to measure the strength and direction of the relationship. Interpret whether the relationship is weak, moderate, or strong. e. If the margin of error for predicting test scores is marks, use your model to estimate the score of a student who studied for hours. Would the model be reliable for this prediction? 21. Screen Time and Sleep An online health article reported data collected from teens aged 15–18 on average daily screen time and hours of sleep. Table 7.20: Screen time and sleep time Screen 3 6 8 5 1 2 1 6 4 7 time(hrs.) Sleep time 7.5 6.5 5.2 6.7 8.2 8 8.5 6.1 7 5.7 (hrs.) Tasks: a. Construct a scatter plot using the data. b. Describe the trend shown by the scatter plot. What does it suggest about the relationship between screen time and sleep time? c. Fit a linear model to the data using the method of least squares. Express the equation in the form: . Using the equation you obtained, draw the line of best fit on the scatter plot. Interpret the gradient of the line of best fit and the y-intercept in the context of your data. d. Calculate to measure the strength and direction of the relationship. Interpret whether the relationship is weak, moderate, or strong. e. If the margin of error for predicting sleep time is hours, use your model to estimate the sleep time of a student whose screen time is 3 hours. Would the model be reliable for this prediction? Justify your answer. 22. Working in small groups carry out the following project. Project : Finding the relationship between two variables using a scatter plot Project Goal: To create and interpret scatter plots to analyse the strength and nature of the relationship between two quantitative variables. Question Select a real-world topic involving two variables, gather or find data for these related variables, create a scatter plot, and then analyse the correlation (or absence of one) between them. Present your findings and discuss the implications of your observations.
The materials you will need include: • Graph paper or graphing software (e.g., GeoGebra, Microsoft Excel, Google Sheets) • Rulers • Calculators (optional) • Access to the internet for data collection (optional) • Presentation tools (e.g., poster board, Google Slides, PowerPoint) Project Steps: Phase 1: Brainstorming and Data Selection a. Introduction to Scatter Plots: In your groups, review what scatter plots are, their purpose, and the concept of correlation (positive, negative, no correlation). b. Brainstorming Variables: Brainstorm pairs of quantitative variables that might have a relationship and are easily assessable. You may choose one of the examples listed below: • Hours studied and test scores in a particular subject. • Daily temperature and water sales. • Height of a student and his/her shoe size • Population of regional capitals in Ghana and the number of JHS. • Distance from school and travel time • Mid-term and end-of-term scores. • Weight of a student and his/her height. • Income and expenditure. c. Topic Selection: Choose ONE pair of variables you are interested in investigating. • Define the Research Question Example: “Is there a relationship between the number of hours students’ study in a week and their last mathematics test scores?” • Clearly define your independent (x-axis) and dependent (y-axis) variables. d. Select an Appropriate Data Collection Tool: You can use: • Questionnaire. Design a simple questionnaire that will solicit their study hours and test scores in mathematics for a particular term or semester. • Interview. Have face-to-face interactions with learners to collect the data. • Class record. You may ask permission from authorities to obtain records of students. e. Prepare the Data Collection Tool: • For a questionnaire, write clear, concise questions • For an interview, prepare a short list of guiding questions.
f. Data Collection Strategy: • Collect at least 20 data points for your chosen variables. • Option A: Self-Collected Data: Collect your own data • Option B: Found Data: You can research data online from reliable sources (e.g., government statistics websites, sports statistics, scientific studies, reputable news organisations). You must cite your sources. • Option C: Teacher-Provided Data: You may ask a teacher or a parent to provide you with data. Phase 2: Creating the Scatter Plot a. Organise your collected data in a table with two columns (one for each variable). b. Decide whether to create the scatter plot by hand on graph paper or using technology. • By Hand: Provide the scale used for each axis, label the axes, and give the graph a clear title. • Using Technology: Use an Excel Spreadsheet or any appropriate software but still ensure everything is correctly labelled. Phase 3: Analysis and Interpretation a. Analyse your generated scatter plot. b. Use your scatter plot to answer the following questions • What is your independent variable (x-axis)? What are its units? • What is your dependent variable (y-axis)? What are its units? • Describe the overall pattern you observe in the scatter plot. • Does there appear to be a correlation between your two variables? If so, what type (positive, negative, or no correlation)? • Calculate the spearman’s rank correlation coefficient between the variables and interpret your answer. • What conclusions can you draw about the relationship between your two variables based on your scatter plot? • Obtain the least square regression line for the data. Interpret the gradient and the y-intercept in the context of your data. • Calculate the sum of squared error. • Does correlation imply causation? Explain this concept in the context of your data. Phase 4: Presentation a. Prepare a presentation (e.g., poster, slide show, written report) that includes: • Project Title • Names of group members/learners • A brief introduction of the chosen variables and the research question. • A displayed data table. • Your scatter plot (large and clear).
• Answers to all analysis questions from Phase 3. • A concluding summary of your findings. • State the sources you used for the data and sources of your data, research materials (if any). • Provide a brief statement acknowledging any individual or institution other than your group members and your teacher who contributed to the successful completion of your project. (Optional) Example: We would like to express our sincere gratitude to [Name of Individual or institution], whose insights and support were invaluable to the successful completion of this project. b. Presentation: You will present your project to your class and answer any questions your peers ask. Assessment and Scoring: You will be assessed on: • Topic Selection and Relevance : It should be clear, realistic and relevant • Data Collection/Sourcing: Appropriateness of data, sufficient number of points, proper citation (if applicable). • Data Representation : Is the scatter plot accurate, with clear labels and accurate point plotting • Data Analysis and Interpretation: Thoroughness and accuracy of answers to analysis questions and understanding of correlation. • Conclusions and Implications: Ensure your findings are well explained with well-considered implications and/or solutions • Presentation: Clarity, organisation, engagement, ability to explain your findings.
Which of the following best distinguishes correlation from regression?
A teacher wants to use students' mid-term examination scores to predict their end-of-term scores. Which variable is the independent variable?
The line of best fit for the number of hours Ama studies () and her test score () is . If Ama studies for 5 hours, what score does the line predict?
In a scatter plot, data points that lie relatively far from the line of best fit are called
A business uses a line of best fit from advertising expenditure between GH₵100 and GH₵500 to predict sales when advertising expenditure is GH₵900. Which statement is most correct?
A mobile money vendor, Ama, in Kumasi recorded her weekly advertising expenditure and the number of transactions over five weeks. The data are shown in the table below.
| Advertising expenditure, (GH¢10) | 2 | 4 | 6 | 8 | 10 |
|---|---|---|---|---|---|
| Number of transactions, (hundreds) | 11 | 15 | 19 | 23 | 27 |
Ama wants to use the relationship between advertising expenditure and transactions to plan her business.
Describe the relationship between advertising expenditure and the number of transactions suggested by the table.
Use the method of least squares to calculate the gradient and the -intercept of the line of best fit for the data.
Write down the equation of the line of best fit. Use it to predict the number of transactions when Ama spends GH¢80 on advertising.
Explain why using this equation to predict the number of transactions for an advertising expenditure of GH¢500 may not be reliable.
State one advantage of using the least-squares regression line instead of drawing a line of best fit by hand.
A health centre in Tamale collected data from five patients on the number of hours they exercised per week and their systolic blood pressure. The results are shown below.
| Exercise, (hours per week) | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| Systolic blood pressure, (mmHg) | 150 | 145 | 140 | 135 | 130 |
The health centre wants to use the data to advise patients.
Distinguish between correlation and regression.
Describe the type and strength of the correlation between exercise hours and systolic blood pressure suggested by the table.
Use the method of least squares to calculate the gradient and the -intercept of the line of best fit.
Write down the equation of the line of best fit and interpret the gradient in the context of the data.
Use the equation to predict the systolic blood pressure of a patient who exercises 6 hours per week. Comment on the reliability of this prediction.