maths.free › Statistics & Probability › Confidence intervals
Confidence intervals
Estimate ± margin of error, and what “95% confident” actually claims.
A confidence interval is a point estimate plus and minus a margin of error: x̄ ± z·σ/√n for a mean when σ is known, with t in place of z when it is not. “95% confidence” describes the method — 95% of intervals built this way capture the true value — not the probability that this one does.
A Single Population Mean Using the Normal Distribution
A confidence interval for a population mean with a known standard deviation is based on the fact that the sample means follow an approximately normal distribution. Suppose that our sample has a mean of \(\overset{\bar}{x}\text{ = 10}\) and we have constructed the 90 percent confidence interval (5, 15), where the margin of error = 5.
Calculating the Confidence Interval
To construct a confidence interval for a single unknown population mean, μ, where the population standard deviation is known, we need \(\overset{\bar}{x}\) as an estimate for μ, and we need the margin of error. The margin of error for the population mean is called the error bound for a population mean (EBM). The sample mean, \(\overset{\bar}{x}\text{,}\) is the point estimate of the unknown population mean, μ.
The confidence interval (CI) estimate will have the form:
(point estimate – error bound, point estimate + error bound) or, in symbols, (\(\overset{\bar}{x}-EBM,\overset{\bar}{x}\text{+}EBM\)).
The margin of error (EBM) depends on the confidence level (CL). The confidence level is often considered the probability that the calculated confidence interval estimate will contain the true population parameter. However, it is more accurate to state that the confidence level is the percentage of confidence intervals that contain the true population parameter when repeated samples are taken. Most often, the person constructing the confidence interval will choose a confidence level of 90 percent or higher, because that person wants to be reasonably certain of his or her conclusions.
Another probability, which is called alpha \((\alpha )\) is related to the confidence level, CL. Alpha is the probability that the confidence interval does not contain the unknown population parameter. Mathematically, alpha can be computed as \(\alpha =1-CL\).
Example
- Suppose we have collected data from a sample. We know the sample mean, but we do not know the mean for the entire population.
- The sample mean is seven, and the error bound for the mean is 2.5.
The confidence interval is (7 – 2.5, 7 + 2.5), and calculating the values gives (4.5, 9.5).
If the confidence level is 95 percent, then we say, "We estimate with 95 percent confidence that the true value of the population mean is between 4.5 and 9.5."
A confidence interval for a population mean with a known standard deviation is based on the fact that the sample means follow an approximately normal distribution. Suppose that our sample has a mean of \(\overset{\bar}{x}\) = 10, and we have constructed the 90 percent confidence interval (5, 15) where EBM = 5.
Condensed — the full section is in OpenStax Statistics.
Working Backward to Find the Error Bound or Sample Mean
When we calculate a confidence interval, we find the sample mean, calculate the error bound, and use them to calculate the confidence interval. However, sometimes when we read statistical studies, the study may state the confidence interval only. If we know the confidence interval, we can work backward to find both the error bound and the sample mean.
- From the upper value for the interval, subtract the sample mean,
- Or, from the upper value for the interval, subtract the lower value. Then divide the difference by 2.
- Subtract the error bound from the upper value of the confidence interval,
- Or, average the upper and lower endpoints of the confidence interval.
Notice that there are two methods to perform each calculation. You can choose the method that is easier to use with the information you know.
Example
Suppose we know that a confidence interval is (67.18, 68.82) and we want to find the error bound. We may know that the sample mean is 68, or perhaps our source only gives the confidence interval and does not tell us the value of the sample mean.
- If we know that the sample mean is 68, EBM = 68.82 – 68 = 0.82.
- If we do not know the sample mean, EBM = \(\frac{(68.82-67.18)}{2}\) = 0.82. The margin of error is the quantity that we add and subtract from the sample mean to obtain the confidence interval. Therefore, the margin of error is half of the length of the interval.
- If we know the error bound, \(\overset{\bar}{x}\) = 68.82 – 0.82 = 68.
- If we do not know the error bound, \(\overset{\bar}{x}\) = \(\frac{(67.18+68.82)}{2}\) = 68.
Calculating the Sample Size
If researchers desire a specific margin of error, then they can use the error bound formula to calculate the required sample size. In this situation, we are given the desired margin of error, EBM, and we need to compute the sample size n.
The formula for sample size is n = \(\frac{{z}^{2}{\sigma }^{2}}{EB{M}^{2}}\), found by solving the error bound formula for n. Always round up the value of n to the closest integer.
In this formula, z is the critical value \({z}_{\frac{\alpha }{2}}\), corresponding to the desired confidence level. A researcher planning a study who wants a specified confidence level and error bound can use this formula to calculate the size of the sample needed for the study.
Example
The population standard deviation for the age of Foothill College students is 15 years. If we want to be 95 percent confident that the sample mean age is within two years of the true population mean age of Foothill College students, how many randomly selected Foothill College students must be surveyed?
- From the problem, we know that σ = 15 and EBM = 2.
- z = z0.025 = 1.96, because the confidence level is 95 percent.
- n = \(\frac{{z}^{2}{\sigma }^{2}}{EB{M}^{2}}\) = \(\frac{{(1.96)}^{2}{(15)}^{2}}{{2}^{2}}\) = 216.09 using the sample size equation.
- Use n = 217. Always round the answer up to the next higher integer to ensure that the sample size is large enough.
Therefore, 217 Foothill College students should be surveyed in order to be 95 percent confident that we are within two years of the true population mean age of Foothill College students.
A Single Population Mean Using the Normal Distribution
Use the following information to answer the next five exercises: The standard deviation of the weights of elephants is known to be approximately 15 lb. We wish to construct a 95 percent confidence interval for the mean weight of newborn elephant calves. Fifty newborn elephants are weighed. The sample mean is 244 lb. The sample standard deviation is 11 lb.
Use the following information to answer the next seven exercises: The U.S. Census Bureau conducts a study to determine the time needed to complete the short form. The bureau surveys 200 people. The sample mean is 8.2 minutes. There is a known standard deviation of 2.2 minutes. The population distribution is assumed to be normal.
Use the following information to answer the next 10 exercises: A sample of 20 heads of lettuce was selected. Assume that the population distribution of head weight is normal. The weight of each head of lettuce was then recorded. The mean weight was 2.2 lb, with a standard deviation of 0.1 lb. The population standard deviation is known to be 0.2 lb.
Use the following information to answer the next 14 exercises: The mean age for all Foothill College students for a recent fall term was 33.2. The population standard deviation has been pretty consistent at 15. Suppose that 25 winter students were randomly selected. The mean age for the sample was 30.4. We are interested in the true mean age for winter Foothill College students. Let X = the age of a winter Foothill College student.
Construct a 95 percent confidence interval for the true mean age of winter Foothill College students by working out and then answering the next eight exercises.
A Single Population Mean Using the Student's t-Distribution
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this unknown number did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close-enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the Guinness brewery in Dublin, Ireland, ran into this problem. His experiments with hops and barley produced very few samples. Just replacing σ with s did not produce accurate results when he tried to calculate a confidence interval. He realized that he could not use a normal distribution for the calculation; he found that the actual distribution depends on the sample size. This problem led him to discover what is called the Student's t-distribution. The name comes from the fact that Gosset wrote under the pen name Student.
Up until the mid-1970s, some statisticians used the normal distribution approximation for large sample sizes and used the Student's t-distribution only for sample sizes of at most 30. With graphing calculators and computers, the practice now is to use the Student's t-distribution whenever s is used as an estimate for σ.
If you draw a simple random sample of size n from a population that has an approximately normal distribution with mean μ and unknown population standard deviation σ and calculate the t-score t = \(\frac{\overset{\bar}{x}-\mu }{(\frac{s}{\sqrt{n}})}\), then the t-scores follow a Student's t-distribution with n – 1 degrees of freedom. The t-score has the same interpretation as the z-score: It measures how far \(\overset{\bar}{x}\) is from its mean μ. For each sample size n, there is a different Student's t-distribution.
The degrees of freedom (df), n - 1, are the sample size minus 1.
Calculators and computers can easily calculate any Student's t-probabilities. The TI-83, 83+, and 84+ have a tcdf function to find the probability for given values of t. The grammar for the tcdf command is tcdf(lower bound, upper bound, degrees of freedom). However, for confidence intervals, we need to use inverse probability to find the value of t when we know the probability.
- \(EBM=({t}_{\frac{\alpha }{2}})(\frac{s}{\sqrt{n}})\),
- \({t}_{\frac{\sigma }{2}}\) is the t-score with area to the right equal to \(\frac{\alpha }{2}\),
- use df = n – 1 degrees of freedom, and
- s = sample standard deviation.
Condensed — the full section is in OpenStax Statistics.
A Single Population Mean Using the Student's t-Distribution
Use the following information to answer the next five exercises: A hospital is trying to cut down on emergency room wait times. It is interested in the amount of time patients must wait before being called back to be examined. An investigation committee randomly surveyed 70 patients. The sample mean was 1.5 hr, with a sample standard deviation of 0.5 hr.
Use the following information to answer the next six exercises: One hundred eight Americans were surveyed to determine the number of hours they spend watching television each month. It was revealed that they watch an average of 151 hours each month, with a standard deviation of 32 hours. Assume that the underlying population distribution is normal.
Use the following information to answer the next 13 exercises: The data in are the result of a random survey of 39 national flags (with replacement between picks) from various countries. We are interested in finding a confidence interval for the true mean number of colors on a national flag. Let X = the number of colors on a national flag.
| X | Freq. |
| 1 | 1 |
| 2 | 7 |
| 3 | 18 |
| 4 | 7 |
| 5 | 6 |
Construct a 95 percent confidence interval for the true mean number of colors on national flags.
A Population Proportion
During an election year, we see articles in the newspaper that state confidence intervals in terms of proportions or percentages. For example, a poll for a particular candidate running for president might show that the candidate has 40 percent of the vote within 3 percentage points (if the sample is large enough). Often, election polls are calculated with 95 percent confidence, so the pollsters would be 95 percent confident that the true proportion of voters who favored the candidate would be between 0.37 and 0.43 (0.40 – 0.03, 0.40 + 0.03).
Investors in the stock market are interested in the true proportion of stocks that go up and down each week. Businesses that sell personal computers are interested in the proportion of households in the United States that own personal computers. Confidence intervals can be calculated for the true proportion of stocks that go up or down each week and for the true proportion of households in the United States that own personal computers.
The procedure to find the confidence interval, the sample size, the error bound for a population (EBP), and the confidence level for a proportion is similar to that for the population mean, but the formulas are different.
How do you know you are dealing with a proportion problem? First, the data that you are collecting is categorical, consisting of two categories: Success or Failure, Yes or No. Examples of situations where you are the following trying to estimate the true population proportion are the following: What proportion of the population smoke? What proportion of the population will vote for candidate A? What proportion of the population has a college-level education?
The distribution of the sample proportions (based on samples of size n) is denoted by P′ (read “P prime”).
The central limit theorem for proportions asserts that the sample proportion distribution P′ follows a normal distribution with mean value p, and standard deviation \(\sqrt{\frac{pq}{n}}\), where p is the population proportion and q = 1 - p.
The confidence interval has the form (p′ – EBP, p′ + EBP). EBP is error bound for the proportion.
\[p'\text{ = }\frac{x}{n}\]Condensed — the full section is in OpenStax Statistics.
A Population Proportion
There is a certain amount of error introduced into the process of calculating a confidence interval for a proportion. Because we do not know the true proportion for the population, we are forced to use point estimates to calculate the appropriate standard deviation of the sampling distribution. Studies have shown that the resulting estimation of the standard deviation can be flawed.
Fortunately, there is a simple adjustment that allows us to produce more accurate confidence intervals: We simply pretend that we have four additional observations. Two of these observations are successes, and two are failures. The new sample size, then, is n + 4, and the new count of successes is x + 2.
Computer studies have demonstrated the effectiveness of the plus-four confidence interval for p method. It should be used when the confidence level desired is at least 90 percent and the sample size is at least ten.
Calculating the Sample Size
If researchers desire a specific margin of error, then they can use the error bound formula to calculate the required sample size.
The margin of error formula for a population proportion is
- \(EBP={z}_{\frac{\alpha }{2}}\times \sqrt{\frac{p'q'}{n}}\), where p′ is the sample proportion, q′ = 1 – p′, and n is the sample size.
- Solving for n gives you an equation for the sample size.
- \(n=\frac{{({z}_{\frac{\alpha }{2}})}^{2}({p}^{'}{q}^{'})}{EB{P}^{2}}\). This formula tells us that we can compute the sample size n required for a confidence level of \(Cl=1-\alpha\) by taking the square of the critical value \({z}_{\frac{a}{2}}\), multiplying by the point estimate p′, and by q′ = 1 – p′ and finally dividing the result by the square of the margin of error. Always remember to round up the value of n.
Example
Try it.
Suppose a mobile phone company wants to determine the current percentage of customers ages 50+ who use text messaging on their cell phones. How many customers ages 50+ should the company survey in order to be 90 percent confident that the estimated (sample) proportion is within 3 percentage points of the true population proportion of customers ages 50+ who use text messaging on their cell phones? Assume that p′ = 0.5.
Solution
From the problem, we know that EBP = 0.03 (3 percent=0.03), and \({z}_{\frac{\alpha }{2}}\) z0.05 = 1.645 because the confidence level is 90 percent.
To calculate the sample size n, use the formula and make the substitutions.
\[n=\frac{{z}^{2}{p}^{'}{q}^{'}}{EB{P}^{2}}\ gives\ n=\frac{{1.645}^{2}(0.5)(0.5)}{{0.03}^{2}}=751.7\]Round the answer to the next higher value. The sample size should be 752 cell phone customers ages 50+ in order to be 90 percent confident that the estimated (sample) proportion is within 3 percentage points of the true population proportion of all customers ages 50+ who use text messaging on their cell phones.
A Population Proportion
Use the following information to answer the next two exercises: Marketing companies are interested in knowing the population percentage of women who make the majority of household purchasing decisions.
Use the following information to answer the next five exercises: Suppose a marketing company conducted a survey. It randomly surveyed 200 households and found that in 120 of them, the women made the majority of the purchasing decisions. We are interested in the population proportion of households where women make the majority of the purchasing decisions.
Use the following information to answer the next five exercises: Of 1,050 randomly selected adults, 360 identified themselves as manual laborers, 280 identified themselves as non-manual wage earners, 250 identified themselves as mid-level managers, and 160 identified themselves as executives. In the survey, 82 percent of manual laborers preferred trucks, 62 percent of non-manual wage earners preferred trucks, 54 percent of mid-level managers preferred trucks, and 26 percent of executives preferred trucks.
Use the following information to answer the next five exercises: A poll of 1,200 voters asked what the most significant issue was in the upcoming election. Sixty-five percent answered "the economy." We are interested in the population proportion of voters who believe the economy is the most important.
Use the following information to answer the next 16 exercises: The Ice Chalet offers dozens of different beginning ice-skating classes. All of the class names are put into a bucket. The 5 p.m., Monday night, ages 8 to 12, beginning ice-skating class is picked. In that class are 64 girls and 16 boys. Suppose that we are interested in the true proportion of girls, ages 8 to 12, in all beginning ice-skating classes at the Ice Chalet. Assume that the children in the selected class are a random sample of the population.
Gewerkte voorbeeld: 50 + 1.96*10/sqrt(100)
Stap met stap
- 51.96 = 51.96
Evaluate.
Openbaar die antwoord
Practice (40)
Try each one on paper first. Reveal the answer to check; verified ones can be opened in the solver for every step.
-
Suppose we have data from a sample. The sample mean is 15, and the error bound for the mean is 3.2.
What is the confidence interval estimate for the population mean?
Openbaar die antwoord
(11.8, 18.2)
-
Find a 90 percent confidence interval for the true (population) mean of statistics exam scores.
Openbaar die antwoord
- You can use technology to calculate the confidence interval directly.
- The first solution is shown step-by-step (Solution A).
- The second solution uses the TI-83, 83+, and 84+ calculators (Solution B).
To find the confidence interval, you need the sample mean, \(\overset{\bar}{x}\), and the EBM.
- \[\overset{\bar}{x}\text{ = 68}\]\[EBM\text{ = }({z}_{\frac{\alpha }{2}})(\frac{\sigma }{\sqrt{n}})\]\[\sigma \text{ = 3; }\ n\text{ = 36}\text{;}\]
- The confidence level is 90 percent (CL = 0.90).
The area to the right of z0.05 is 0.05 and the area to the left of z0.05 is 1 – 0.05 = 0.95.
\[{z}_{\frac{\alpha }{2}}\text{ = }{z}_{0.05}\text{ = 1}\text{.645}\]using invNorm(0.95, 0, 1) on the TI-83,83+, and 84+ calculators. This can also be found using appropriate commands on other calculators, using a computer, or using a probability table for the standard normal distribution.
EBM = (1.645)\((\frac{3}{\sqrt{36}})\) = 0.8225
\(\overset{\bar}{x}\) – EBM = 68 – 0.8225 = 67.1775
\(\overset{\bar}{x}\) + EBM = 68 + 0.8225 = 68.8225
The 90 percent confidence interval is (67.1775, 68.8225).
-
Find a 90 percent confidence interval estimate for the population mean delivery time.
Openbaar die antwoord
(34.1347, 37.8653)
-
Find a 98 percent confidence interval for the true (population) mean of the SARs for cell phones. Assume that the population standard deviation is σ = 0.337.
Openbaar die antwoord
To find the confidence interval, start by finding the point estimate: the sample mean,
\[\overset{\bar}{x}=1.024\text{.}\]This is calculated by adding the specific absorption rate for the 30 cell phones in the sample, and dividing the result by 30.
Next, find the EBM. Because you are creating a 98 percent confidence interval, CL = 0.98.
You need to find z0.01, having the property that the area under the normal density curve to the right of z0.01 is 0.01 and the area to the left is 0.99. Use your calculator, a computer, or a probability table for the standard normal distribution to find z0.01 = 2.326.
\[EBM=({z}_{0.01})\frac{\sigma }{\sqrt{n}}=(2.326)\frac{0.337}{\sqrt{30}}=0.1431\]To find the 98 percent confidence interval, find \(\overset{\bar}{x}\pm EBM\).
\[\overset{\bar}{x}\text{ - }EBM\text{ = 1}\text{.024 - 0}\text{.1431 = 0}\text{.8809}\]\[\overset{\bar}{x}\text{ + }EBM\text{ = 1}\text{.024 + 0}\text{.1431 = 1}\text{.1671}\]We estimate with 98 percent confidence that the true SAR mean for the population of cell phones in the United States is between 0.8809 and 1.1671 watts per kilogram.
-
shows a different random sampling of 20 cell phone models. Use these data to calculate a 93 percent confidence interval for the true mean SAR for cell phones certified for use in the United States. As previously, assume that the population standard deviation is σ = 0.337.
Phone Model SAR Phone Model SAR 450 1.48 1450 1.53 550 0.8 1550 0.68 650 1.15 1650 1.4 750 1.36 1750 1.24 850 0.77 1850 0.57 950 0.462 1950 0.2 1050 1.36 2050 0.51 1150 1.39 2150 0.3 1250 1.3 2250 0.73 1350 0.7 2350 0.869 Openbaar die antwoord
\(\overset{\bar}{x}=0.940\)
\(\frac{\alpha }{2}=\frac{1-CL}{2}=\frac{1-0.93}{2}=0.035\)
Z0.035 = 1.812
\(EBM=({z}_{0.035})(\frac{\sigma }{\sqrt{n}})=(1.812)(\frac{0.337}{\sqrt{20}})=0.1365\)
\(\overset{\bar}{x}\) – EBM = 0.940 – 0.1365 = 0.8035
\(\overset{\bar}{x}\) + EBM = 0.940 + 0.1365 = 1.0765
We estimate with 93 percent confidence that the true SAR mean for the population of cell phones in the United States is between 0.8035 and 1.0765 watts per kilogram.
-
Suppose we change the original problem in by using a 95 percent confidence level. Find a 95 percent confidence interval for the true (population) mean statistics exam score.
Openbaar die antwoord
To find the confidence interval, you need the sample mean, \(\overset{\bar}{x}\), and the EBM.
- \[\overset{\bar}{x}\text{ = 68}\]\[EBM\text{ = }({z}_{\frac{\alpha }{2}})(\frac{\sigma }{\sqrt{n}})\]\[\sigma \text{ = 3; }\ n\text{ = 36}\]
- The confidence level is 95 percent (CL = 0.95).
The area to the right of \(z\)0.025 is 0.025, and the area to the left of \(z\)0.025 is 1 – 0.025 = 0.975.
\[{z}_{\frac{\alpha }{2}}={z}_{0.025}=1.96\text{,}\]when using invnorm(0.975,0,1) on the TI-83, 83+, or 84+ calculators. (This can also be found using appropriate commands on other calculators, using a computer, or using a probability table for the standard normal distribution.)
\[EBM\text{ = (1}\text{.96)}(\frac{3}{\sqrt{36}})\text{ = 0}\text{.98}\]\[\overset{\bar}{x}\text{ - }EBM\text{ = 68 - 0}\text{.98 = 67}\text{.02}\]\[\overset{\bar}{x}\text{ + }EBM\text{ = 68 + 0}\text{.98 = 68}\text{.98}\]Notice that the EBM is larger for a 95 percent confidence level in the original problem.
We estimate with 95 percent confidence that the true population mean for all statistics exam scores is between 67.02 and 68.98.
95 percent of all confidence intervals constructed in this way contain the true value of the population mean statistics exam score.
The 90 percent confidence interval is (67.18, 68.82). The 95 percent confidence interval is (67.02, 68.98). The 95 percent confidence interval is wider. If you look at the graphs, because the area 0.95 is larger than the area 0.90, it makes sense that the 95 percent confidence interval is wider. For more certainty that the confidence interval actually does contain the true value of the population mean for all statistics exam scores, the confidence interval necessarily needs to be wider.
- Increasing the confidence level increases the error bound, making the confidence interval wider.
- Decreasing the confidence level decreases the error bound, making the confidence interval narrower.
-
Refer back to the pizza-delivery Try It exercise. The population standard deviation is six minutes and the sample mean deliver time is 36 minutes. Use a sample size of 20. Find a 95 percent confidence interval estimate for the true mean pizza-delivery time.
Openbaar die antwoord
(33.37, 38.63)
-
Leave everything the same except the sample size. Use the original 90 percent confidence level. What happens to the error bound and the confidence interval if we increase the sample size and use n = 100 instead of n = 36? What happens if we decrease the sample size to n = 25 instead of n = 36?
- \(\overset{\bar}{x}\) = 68
- EBM = \(({z}_{\frac{\alpha }{2}})(\frac{\sigma }{\sqrt{n}})\)
- σ = 3, the confidence level is 90 percent (CL = 0.90), \({z}_{\frac{\alpha }{2}}\) = z0.05 = 1.645.
Openbaar die antwoord
If we increase the sample size n to 100, we decrease the margin of error.
When n = 100, EBM = \(({z}_{\frac{\alpha }{2}})(\frac{\sigma }{\sqrt{n}})\) = (1.645)\((\frac{3}{\sqrt{100}})\) = 0.4935.
-
Refer back to the pizza-delivery Try It exercise. The mean delivery time is 36 minutes and the population standard deviation is six minutes. Assume the sample size is changed to 50 restaurants with the same sample mean. Find a 90 percent confidence interval estimate for the population mean delivery time.
Openbaar die antwoord
(34.6041, 37.3958)
-
Suppose we know that a confidence interval is (42.12, 47.88). Find the error bound and the sample mean.
Openbaar die antwoord
Sample mean is 45, error bound is 2.88
-
The population standard deviation for the height of high school basketball players is three inches. If we want to be 95 percent confident that the sample mean height is within one inch of the true population mean height, how many randomly selected students must be surveyed?
Openbaar die antwoord
35 students
-
Identify the following:
- \(\overset{\bar}{x}\) = _____
- σ = _____
- n = _____
Openbaar die antwoord
- 244
- 15
- 50
-
In words, define the random variables X and \(\overset{\bar}{X}\).
-
Which distribution should you use for this problem?
Openbaar die antwoord
\(N(244,\frac{15}{\sqrt{50}})\)
-
Construct a 95 percent confidence interval for the population mean weight of newborn elephants. State the confidence interval, sketch the graph, and calculate the error bound.
-
What will happen to the confidence interval obtained, if 500 newborn elephants are weighed instead of 50? Why?
Openbaar die antwoord
As the sample size increases, there will be less variability in the mean, so the interval size decreases.
-
Identify the following:
- \(\overset{\bar}{x}\) = _____
- σ = _____
- n = _____
-
In words, define the random variables X and \(\overset{\bar}{X}\).
Openbaar die antwoord
X is the time in minutes it takes to complete the U.S. Census short form. \(\overset{\bar}{X}\) is the mean time it took a sample of 200 people to complete the U.S. Census short form.
-
Which distribution should you use for this problem?
-
Construct a 90 percent confidence interval for the population mean time to complete the forms. State the confidence interval, sketch the graph, and calculate the error bound.
Openbaar die antwoord
CI: (7.9441, 8.4559)
EBM = 0.26
-
If the Census wants to increase its level of confidence and keep the error bound the same by taking another survey, what changes should it make?
-
If the Census did another survey, kept the error bound the same, and surveyed only 50 people instead of 200, what would happen to the level of confidence? Why?
Openbaar die antwoord
The level of confidence would decrease, because decreasing n makes the confidence interval wider, so at the same error bound, the confidence level decreases.
-
Suppose the Census needed to be 98 percent confident of the population mean length of time. Would the Census have to survey more people? Why or why not?
-
Identify the following:
- \(\overset{\bar}{x}\) = ______
- σ = ______
- n = ______
Openbaar die antwoord
- \(\overset{\bar}{x}\) = 2.2
- σ = 0.2
- n = 20
-
In words, define the random variable X.
-
In words, define the random variable \(\overset{\bar}{X}\).
Openbaar die antwoord
\(\overset{\bar}{X}\) is the mean weight of a sample of 20 heads of lettuce.
-
Which distribution should you use for this problem?
-
Construct a 90 percent confidence interval for the population mean weight of the heads of lettuce. State the confidence interval, sketch the graph, and calculate the error bound.
Openbaar die antwoord
EBM = 0.07
CI: (2.1264, 2.2736) -
Construct a 95 percent confidence interval for the population mean weight of the heads of lettuce. State the confidence interval, sketch the graph, and calculate the error bound.
-
In complete sentences, explain why the confidence interval in is larger than in .
Openbaar die antwoord
The interval is greater, because the level of confidence increased. If the only change made in the analysis is a change in confidence level, then all we are doing is changing how much area is being calculated for the normal distribution. Therefore, a larger confidence level results in larger areas and larger intervals.
-
In complete sentences, give an interpretation of what the interval in means.
-
What would happen if 40 heads of lettuce were sampled instead of 20 and the error bound remained the same?
Openbaar die antwoord
The confidence level would increase.
-
What would happen if 40 heads of lettuce were sampled instead of 20 and the confidence level remained the same?
-
\(\overset{\bar}{x}\) = _____
Openbaar die antwoord
30.4
-
n = _____
-
________ = 15
Openbaar die antwoord
σ
-
In words, define the random variable \(\overset{\bar}{X}\).
-
What is \(\overset{\bar}{x}\) estimating?
Openbaar die antwoord
μ
-
Is \({\sigma }_{x}\) known?
-
As a result of your answer to , state the exact distribution to use when calculating the confidence interval.
Openbaar die antwoord
normal
Symbols used here
The non-negative number whose square (n-th power) is x.
Typical distance from the mean; its square.
Average of the data; average of the whole population.
Both signs at once: x = 3 ± 2 means 5 and 1.
Equal to the precision shown, not exactly.
n × (n−1) × … × 1; the number of orderings of n things. 0! = 1.
Number of k-element subsets of n things: n!/(k!(n−k)!).
Add a_k for k = 1 up to n.
In either; in both; in A but not B.
Chance of A; chance of A given that B happened.
Probability-weighted average of X; its spread.
The bell curve with mean μ and variance σ²; (x − μ)/σ.
How to: Confidence intervals
- Evaluate.
Questions people ask
Mean or median — which should I use?
Median when the data have outliers or a long tail (incomes, house prices); mean when the data are roughly symmetric and you want every value to count. Report both if they disagree — the gap is itself information.
What does a p-value actually say?
The probability of seeing data at least this extreme if the null hypothesis were true. It is not the probability that the null hypothesis is true.
Why divide by n − 1 for the sample variance?
The sample mean sits closer to the sample than the true mean does, so squared deviations from it are slightly too small on average; dividing by n − 1 instead of n corrects the bias.
Probeer jou eie
Parts of this page are adapted from OpenStax Statistics (CC BY 4.0). Condensed and re-explained here; errors are ours.
Meer in Statistics & Probability
Sampling and dataDescribing data with graphsMean, median and modeProbabilityCounting: permutations and combinationsDiscrete random variablesContinuous random variablesThe normal distributionThe central limit theoremHypothesis testingComparing two samplesChi-square testsLinear regression and correlationANOVA and the F distribution