maths.freeStatistics & Probability › The central limit theorem

The central limit theorem

Why averages are normal: the sampling distribution of the mean.

Whatever the population looks like, the average of a large enough sample is approximately normal, centred on the population mean, with standard deviation σ/√n — the standard error. That single fact is what makes confidence intervals and hypothesis tests work.

The Central Limit Theorem for Sample Means (Averages)

Suppose X is a random variable with a distribution that may be known or unknown (it can be any distribution). Using a subscript that matches the random variable, suppose

  1. \({\mu }_{x}\) = the mean of X
  2. \({\sigma }_{x}\) = the standard deviation of X

If you draw random samples of size n, then as n increases, the random variable \(\overset{\bar}{X}\), which consists of sample means, tends to be normally distributed and

\[\overset{\bar}{X}∼N({\mu }_{x},\frac{{\sigma }_{x}}{\sqrt{n}})\]

The central limit theorem for sample means says that if you keep drawing larger and larger samples (such as rolling one, two, five, and finally, ten dice) and calculating their means, the sample means form their own normal distribution (the sampling distribution). The normal distribution has the same mean as the original distribution and a variance that equals the original variance divided by the sample size. The variable n is the number of values that are averaged together, not the number of times the experiment is done.

To put it more formally, if you draw random samples of size n, the distribution of the random variable \(\overset{\bar}{X}\), which consists of sample means, is called the sampling distribution of the mean. The sampling distribution of the mean approaches a normal distribution as n, the sample size, increases.

The random variable \(\overset{\bar}{X}\) has a different z-score associated with it from that of the random variable X. The mean \(\overset{\bar}{x}\) is the value of \(\overset{\bar}{X}\) in one sample.

\[z=\frac{\overset{\bar}{x}-{\mu }_{x}}{(\frac{{\sigma }_{x}}{\sqrt{n}})}\text{,}\]

μX is the average of both X and \(\overset{\bar}{X}\).

\(\sigma \overset{\bar}{x}\text{ = }\frac{\sigma x}{\sqrt{n}}\) = standard deviation of \(\overset{\bar}{X}\) and is called the standard error of the mean.

Condensed — the full section is in OpenStax Statistics.

The Central Limit Theorem for Sample Means (Averages)

Use the following information to answer the next six exercises: Yoonie is a personnel manager in a large corporation. Each month she must review 16 of the employees. From past experience, she has found that the reviews take her approximately four hours each to do with a population standard deviation of 1.2 hours. Let Χ be the random variable representing the time it takes her to complete one review. Assume Χ is normally distributed. Let \(\overset{\bar}{X}\) be the random variable representing the mean time to complete the 16 reviews. Assume that the 16 reviews represent a random set of reviews.

The Central Limit Theorem for Sums (Optional)

Suppose X is a random variable with a distribution that may be known or unknown (it can be any distribution) and suppose:

  1. μX = the mean of Χ
  2. σΧ = the standard deviation of X

If you draw random samples of size n, then as n increases, the random variable ΣX consisting of sums tends to be normally distributed and ΣΧ ~ N[(n)(μΧ), (\(\sqrt{n}\))(σΧ)].

The central limit theorem for sums says that if you keep drawing larger and larger samples and taking their sums, the sums form their own normal distribution (the sampling distribution), which approaches a normal distribution as the sample size increases. The normal distribution has a mean equal to the original mean multiplied by the sample size and a standard deviation equal to the original standard deviation multiplied by the square root of the sample size.

The random variable ΣX has the following z-score associated with it:

  1. Σx is one sum.
  2. \(z\text{ = }\frac{\Sigma x-(n)({\mu }_{X})}{(\sqrt{n})({\sigma }_{X})}\)
    1. (n)(μX) = mean of ΣX
    2. \((\sqrt{n})({\sigma }_{X})\) = standard deviation of \(\Sigma X\)
Example

An unknown distribution has a mean of 90 and a standard deviation of 15. A sample of size 80 is drawn randomly from the population.

Try it.

  1. Find the probability that the sum of the 80 values (or the total of the 80 values) is more than 7,500.
  2. Find the sum that is 1.5 standard deviations above the mean of the sums.
Solution

Let X = one value from the original unknown population. The probability question asks you to find a probability for the sum (or total of) 80 values.

ΣX = the sum or total of 80 values. Because μX = 90, σX = 15, and n = 80, \(\Sigma X\) ~ N[(80)(90),
(\(\sqrt{\text{80}}\))(15)]

  • mean of the sums = (n)(μX) = (80)(90) = 7200
  • standard deviation of the sums = \(\text{(}\sqrt{n}\text{)(}{\sigma }_{X}\text{) = (}\sqrt{\text{80}}\text{)}\)(15)
  • sum of 80 values = Σx = 7500

a. Find Px > 7500)

Px > 7500) = 0.0127

b. Find Σx where z = 1.5.

Σx = (n)(μX) + (z)\((\sqrt{n})\)(σΧ) = (80)(90) + (1.5)(\(\sqrt{80}\))(15) = 7401.2

Condensed — the full section is in OpenStax Statistics.

The Central Limit Theorem for Sums (Optional)

Use the following information to answer the next four exercises: An unknown distribution has a mean of 80 and a standard deviation of 12. A sample size of 95 is drawn randomly from the population.


Use the following information to answer the next five exercises: The distribution of results from a cholesterol test has a mean of 180 and a standard deviation of 20. A sample size of 40 is drawn randomly.


Use the following information to answer the next six exercises: A researcher measures the amount of sugar in several cans of the same type of soda. The mean is 39.01 with a standard deviation of 0.5. The researcher randomly selects a sample of 100.


Use the following information to answer the next four exercises: An unknown distribution has a mean 12 and a standard deviation of one. A sample size of 25 is taken. Let X = the object of interest.


Use the following information to answer the next three exercises:
A market researcher analyzes how many electronics devices customers buy in a single purchase. The distribution has a mean of three with a standard deviation of 0.7. She samples 400 customers.


Use the following information to answer the next three exercises:
An unkwon distribution has a mean of 100, a standard deviation of 100, and a sample size of 100. Let X = one object of interest.

Using the Central Limit Theorem

It is important for you to understand when to use the central limit theorem. If you are being asked to find the probability of the mean, use the clt for the means. If you are being asked to find the probability of a sum or total, use the clt for sums. This also applies to percentiles for means and sums.

Examples of the Central Limit Theorem

The law of large numbers says that if you take samples of larger and larger sizes from any population, then the mean \(\overset{\bar}{x}\) of the samples tends to get closer and closer to μ. From the central limit theorem, we know that as n gets larger and larger, the sample means follow a normal distribution. The larger n gets, the smaller the standard deviation gets. (Remember that the standard deviation for \(\overset{\bar}{X}\) is \(\frac{\sigma }{\sqrt{n}}\).) This means that the sample mean \(\overset{\bar}{x}\) must be close to the population mean μ. We can say that μ is the value that the sample means approach as n gets larger. The central limit theorem illustrates the law of large numbers.

Condensed — the full section is in OpenStax Statistics.

Using the Central Limit Theorem

Use the following information to answer the next 10 exercises: A manufacturer produces 25-pound lifting weights. The lowest actual weight is 24 pounds, and the highest is 26 pounds. Each weight is equally likely, so the distribution of weights is uniform. A sample of 100 weights is taken.


Use the following information to answer the next five exercises:
The length of time a particular smartphone's battery lasts follows an exponential distribution with a mean of ten months. A sample of 64 of these smartphones is taken.


Use the following information to answer the next six exercises:
A uniform distribution has a minimum of six and a maximum of ten. A sample of 50 is taken.

Gumagana halimbawa: 10/sqrt(25)

Evaluate 10/sqrt(25)

2

Hakbang-hakbang

  1. \frac{10}{\sqrt{25}} = 2

    Power: 1/(√(25)) = 1/5.

Ipahayag ang sagot
2

Practice (40)

Try each one on paper first. Reveal the answer to check; verified ones can be opened in the solver for every step.

  1. a. Find the probability that the sample mean is between 85 and 92.

    Ipahayag ang sagot

    a. Let X = one value from the original unknown population. The probability question asks you to find a probability for the sample mean.

    Let \(\overset{\bar}{X}\) = the mean of a sample of size 25. Because μx = 90, σx = 15, and n = 25,

    \[\overset{\bar}{X}∼N({\mu }_{x},\frac{{\sigma }_{x}}{\sqrt{n}})\]

    Find P(85 < \(\overset{\bar}{x}\) < 92). Draw a graph.

    P(85 < \(\overset{\bar}{x}\) < 92) = 0.6997

    The probability that the sample mean is between 85 and 92 is 0.6997.

  2. b. Find the value that is two standard deviations above the expected value, 90, of the sample mean.

    Ipahayag ang sagot

    b. To find the value that is two standard deviations above the expected value 90, use the following formula

    \[\text{value}={µ}_{x}+(\#\text{ofSTDEVs})(\frac{{\sigma }_{x}}{\sqrt{n}})\]\[\text{value = 90 + 2 }(\frac{15}{\sqrt{25}})=96\text{.}\]

    The value that is two standard deviations above the expected value is 96.

    The standard error of the mean is \(\frac{\sigma x}{\sqrt{n}}\) = \(\frac{15}{\sqrt{25}}\) = 3. Recall that the standard error of the mean is a description of how far (on average) that the sample mean will be from the population mean in repeated simple random samples of size n.

  3. An unknown distribution has a mean of 45 and a standard deviation of eight. Samples of size n = 30 are drawn randomly from the population. Find the probability that the sample mean is between 42 and 50.

    Ipahayag ang sagot

    P(42 < \(\overset{\bar}{x}\) < 50) = \((\text{42,50,45,}\frac{8}{\sqrt{30}})\) = 0.9797

  4. The length of time, in hours, it takes a group of people, 40 years old and older, to play one soccer match is normally distributed with a mean of 2 hours and a standard deviation of 0.5 hours. A sample of size n = 50 is drawn randomly from the population. Find the probability that the sample mean is between 1.8 hours and 2.3 hours.

    Ipahayag ang sagot

    Let X = the time, in hours, it takes to play one soccer match.

    The probability question asks you to find a probability for the sample mean time, in hours, it takes to play one soccer match.

    Let \(\overset{\bar}{X}\) = the mean time, in hours, it takes to play one soccer match.

    If μX = _________, σX = __________, and n = ___________, then X ~ N(______, ______) by the central limit theorem for means.

    μX = 2, σX = 0.5, n = 50, and X ~ N\((\text{2, }\frac{0.5}{\sqrt{50}})\)

    Find P(1.8 < \(\overset{\bar}{x}\) < 2.3). Draw a graph.

    \[P(1.8<\overset{\bar}{x}<2.3)=0.9977\]

    normalcdf\[(1.\text{8,2}\text{.3,2,}\frac{.5}{\sqrt{50}})\text{ = 0.9977}\]

    The probability that the mean time is between 1.8 hours and 2.3 hours is 0.9977.

  5. The length of time taken on the SAT exam for a group of students is normally distributed with a mean of 2.5 hours and a standard deviation of 0.25 hours. A sample size of n = 60 is drawn randomly from the population. Find the probability that the sample mean is between two hours and three hours.

    Ipahayag ang sagot

    P(2 < \(\overset{\bar}{x}\) < 3) = normalcdf\((\text{2, 3, 2}\text{.5, }\frac{0.25}{\sqrt{60}})\) = 1

  6. In a recent study reported Oct. 29, 2012, the mean age of tablet users is 34 years. Suppose the standard deviation is 15 years. Take a sample of size n = 100.

    1. What are the mean and standard deviation for the sample mean ages of tablet users?
    2. What does the distribution look like?
    3. Find the probability that the sample mean age is more than 30 years (the reported mean age of tablet users in this particular study).
    4. Find the 95th percentile for the sample mean age (to one decimal place).
    Ipahayag ang sagot
    1. Because the sample mean tends to target the population mean, we have μχ = μ = 34. The sample standard deviation is given by \({\text{\sigma }}_{\text{\chi }}\) = \(\frac{\sigma }{\sqrt{n}}\) = \(\frac{15}{\sqrt{100}}\) = \(\frac{15}{10}\) = 1.5.
    2. The central limit theorem states that for large sample sizes (n), the sampling distribution will be approximately normal.
    3. The probability that the sample mean age is more than 30 is given by P(Χ > 30) = normalcdf(30,E99,34,1.5) = 0.9962.
    4. Let k = the 95th percentile.
      k = invNorm\((0.\text{95,34,}\frac{15}{\sqrt{100}})\) = 36.5
  7. A gaming marketing gap for men between the ages of 30 to 40 has been identified. You are researching a startup game targeted at the 35-year-old demographic. Your idea is to develop a strategy game that can be played by men from their late 20s through their late 30s. Based on the article’s data, industry research shows that the average strategy player is 28 years old with a standard deviation of 4.8 years. You take a sample of 100 randomly selected gamers. If your target market is 29- to 35-year-olds, should you continue with your development strategy?

    Ipahayag ang sagot

    You need to determine the probability for men whose mean age is between 29 and 35 years of age wanting to play a strategy game.

    P(29 < \(\overset{\bar}{x}\) < 35) = normalcdf\((\text{29,35,28,}\frac{4.8}{\sqrt{100}})\) = 0.0186

    You can conclude there is approximately a 1.9% chance that your game will be played by men whose mean age is between 29 and 35.

  8. The mean number of minutes for app engagement by a tablet user is 8.2 minutes. Suppose the standard deviation is one minute. Take a sample of 60.

    1. What are the mean and standard deviation for the sample mean number of app engagement minutes by a tablet user?
    2. What is the standard error of the mean?
    3. Find the 90th percentile for the sample mean time for app engagement for a tablet user. Interpret this value in a complete sentence.
    4. Find the probability that the sample mean is between eight minutes and 8.5 minutes.
    Ipahayag ang sagot
    1. \({\mu }_{\overset{\bar}{x}}=\mu =8.2\ {\sigma }_{\overset{\bar}{x}}=\frac{\sigma }{\sqrt{n}}=\frac{1}{\sqrt{60}}=0.13\)
    2. This allows us to calculate the probability of sample means of a particular distance from the mean, in repeated samples of size 60.
    3. Let k = the 90th percentile.
      k = invNorm\((0.\text{90,8}\text{.2,}\frac{1}{\sqrt{60}})\) = 8.37. This values indicates that 90 percent of the average app engagement time for table users is less than 8.37 minutes.
    4. P(8 < \(\overset{\bar}{x}\) < 8.5) = normalcdf\((\text{8,8}\text{.5,8}\text{.2,}\frac{1}{\sqrt{60}})\) = 0.9293
  9. Cans of a cola beverage claim to contain 16 ounces. The amounts in a sample are measured and the statistics are n = 34, \(\overset{\bar}{x}\) = 16.01 ounces. If the cans are filled so that μ = 16.00 ounces (as labeled) and σ = 0.143 ounces, find the probability that a sample of 34 cans will have an average amount greater than 16.01 ounces. Do the results suggest that cans are filled with an amount greater than 16 ounces?

    Ipahayag ang sagot

    We have P(\(\overset{\bar}{x}\) > 16.01) = normalcdf\((\text{16}\text{.01,E99,16,}\frac{0.143}{\sqrt{34}})\) = 0.3417. Since there is a 34.17% probability that the average sample weight is greater than 16.01 ounces, we should be skeptical of the company’s claimed volume. If I am a consumer, I should be glad that I am probably receiving free cola. If I am the manufacturer, I need to determine if my bottling processes are outside of acceptable limits.

  10. What is the mean, standard deviation, and sample size?

    Ipahayag ang sagot

    mean = 4 hours, standard deviation = 1.2 hours, sample size = 16

  11. Complete the distributions.

    1. X ~ _____(_____, _____)
    2. \(\overset{\bar}{X}\) ~ _____(_____, _____)

  12. Find the probability that one review will take Yoonie from 3.5 to 4.25 hours. Sketch the graph, labeling and scaling the horizontal axis. Shade the region corresponding to the probability.

    1. P(________ < x < ________) = _______

    Ipahayag ang sagot

    a. Check student's solution.
    b. 3.5, 4.25, 0.2441

  13. Find the probability that the mean of a month’s reviews will take Yoonie from 3.5 to 4.25 hrs. Sketch the graph, labeling and scaling the horizontal axis. Shade the region corresponding to the probability.

    1. P(________________) = _______

  14. What causes the probabilities in and to be different?

    Ipahayag ang sagot

    The fact that the two distributions are different accounts for the different probabilities.

  15. Find the 95th percentile for the mean time to complete one month's reviews. Sketch the graph.

    1. The 95th percentile =____________
  16. Previously, De Anza's statistics students estimated that the amount of change daytime statistics students carry is exponentially distributed with a mean of $0.88. Suppose that we randomly pick 25 daytime statistics students.

    1. In words, Χ = ____________.
    2. Χ ~ _____(_____, _____)
    3. In words, \(\overset{\bar}{X}\) = ____________.
    4. \(\overset{\bar}{X}\) ~ ______ (______, ______)
    5. Find the probability that an individual had between $0.80 and $1.00. Graph the situation, and shade in the area to be determined.
    6. Find the probability that the average amount of change of the 25 students was between $0.80 and $1.00. Graph the situation, and shade in the area to be determined.
    7. Explain why there is a difference in part (e) and part (f).
    Ipahayag ang sagot
    1. Χ = amount of change students carry
    2. Χ ~ Exp(1/0.88) or approximately Χ ~ Exp(1.1364)
    3. \(\overset{\bar}{X}\) = average amount of change carried by a sample of 25 students.
    4. \(\overset{\bar}{X}\) ~ N(0.88, 0.176)
    5. 0.0819
    6. 0.4276
    7. The probability in part (e) represents the probability of an individual value. In part (f), the probability describes the mean of a sample of 25. Part (f) relies on the central limit theorem, so the distributions are different. Part (e) is exponential and part (f) is normal.
  17. Suppose that the distance of fly balls hit to the outfield (in baseball) is normally distributed with a mean of 250 feet and a standard deviation of 50 feet. We randomly sample 49 fly balls.

    1. If \(\overset{\bar}{X}\) = average distance in feet for 49 fly balls, then \(\overset{\bar}{X}\) ~ _______(_______, _______).
    2. What is the probability that the 49 balls traveled an average of less than 240 feet? Sketch the graph. Scale the horizontal axis for \(\overset{\bar}{X}\). Shade the region corresponding to the probability. Find the probability.
    3. Find the 80th percentile of the distribution of the average of 49 fly balls.
  18. According to the Internal Revenue Service, the average length of time for an individual to complete (keep records for, learn, prepare, copy, assemble, and send) IRS Form 1040 is 10.53 hours (without any attached schedules). The distribution is unknown. Let us assume that the standard deviation is two hours. Suppose we randomly sample 36 taxpayers.

    1. In words, Χ = _____________.
    2. In words, \(\overset{\bar}{X}\) = _____________.
    3. \(\overset{\bar}{X}\) ~ _____(_____, _____)
    4. Would you be surprised if the 36 taxpayers finished their Form 1040s in an average of more than 12 hours? Explain why or why not in complete sentences.
    5. Would you be surprised if one taxpayer finished his or her Form 1040 in more than 12 hours? In a complete sentence, explain why.
    Ipahayag ang sagot
    1. length of time for an individual to complete IRS form 1040, in hours
    2. mean length of time for a sample of 36 taxpayers to complete IRS form 1040, in hours
    3. N\((\text{10}\text{.53, }\frac{1}{3})\)
    4. Yes, I would be surprised, because the probability is almost 0.
    5. No, I would not be totally surprised because the probability is 0.2312.
  19. Suppose that a category of world-class runners are known to run a marathon (26 miles) in an average of 145 minutes with a standard deviation of 14 minutes. Consider 49 of the races. Let \(\overset{\bar}{X}\) be the average of the 49 races.

    1. \(\overset{\bar}{X}\) ~ _____(_____, _____)
    2. Find the probability that the runner will average between 142 and 146 minutes in these 49 marathons.
    3. Find the 80th percentile for the average of these 49 marathons.
    4. Find the median of the average running times.
  20. The length of songs in a collector’s online album collection is uniformly distributed from 2 to 3.5 minutes. Suppose we randomly pick five albums from the collection. There are a total of 43 songs on the five albums.

    1. In words, Χ = _________.
    2. Χ ~ _____________
    3. In words, \(\overset{\bar}{X}\) = _____________.
    4. \(\overset{\bar}{X}\) ~ _____(_____, _____)
    5. Find the first quartile for the average song length.
    6. The IQR for the average song length is _______–_______.
    Ipahayag ang sagot
    1. the length of a song, in minutes, in the collection
    2. U(2, 3.5)
    3. the average length, in minutes, of the songs from a sample of five albums from the collection
    4. N(2.75, 0.0220)
    5. 2.74 minutes
    6. 0.03 minutes
  21. In 1940, the average size of a U.S. farm was 174 acres. Let’s say that the standard deviation was 55 acres. Suppose we randomly survey 38 farmers from 1940.

    1. In words, Χ = _____________.
    2. In words, \(\overset{\bar}{X}\) = _____________.
    3. \(\overset{\bar}{X}\) ~ _____(_____, _____)
    4. The IQR for \(\overset{\bar}{X}\) is from _______ acres to _______ acres.
  22. Determine which of the following are true and which are false. Then, in complete sentences, justify your answers.

    1. When the sample size is large, the mean of \(\overset{\bar}{X}\) is approximately equal to the mean of Χ.
    2. When the sample size is large, \(\overset{\bar}{X}\) is approximately normally distributed.
    3. When the sample size is large, the standard deviation of \(\overset{\bar}{X}\) is approximately the same as the standard deviation of Χ.
    Ipahayag ang sagot
    1. True. The mean of a sampling distribution of the means is approximately the mean of the data distribution.
    2. True. According to the central limit theorem, the larger the sample, the closer the sampling distribution of the means becomes normal.
    3. The standard deviation of the sampling distribution of the means will decrease, making it approximately the same as the standard deviation of X as the sample size increases.
  23. The percentage of fat calories that a person in America consumes each day is normally distributed with a mean of about 36 and a standard deviation of about ten. Suppose that 16 individuals are randomly chosen. Let \(\overset{\bar}{X}\) = average percentage of fat calories.

    1. \(\overset{\bar}{X}\) ~ ______(______, ______)
    2. For the group of 16, find the probability that the average percentage of fat calories consumed is more than five. Graph the situation and shade in the area to be determined.
    3. Find the first quartile for the average percentage of fat calories.
  24. The distribution of income in some economically developing countries is considered wedge shaped (many very poor people, very few middle income people, and even fewer wealthy people). Suppose we pick a country with a wedge-shaped distribution. Let the average salary be $2,000 per year with a standard deviation of $8,000. We randomly survey 1,000 residents of that country.

    1. In words, Χ = _____________.
    2. In words, \(\overset{\bar}{X}\) = _____________.
    3. \(\overset{\bar}{X}\) ~ _____(_____, _____)
    4. How is it possible for the standard deviation to be greater than the average?
    5. Why is it more likely that the average salary of the 1,000 residents will be from $2,000 to $2,100 than from $2,100 to $2,200?
    Ipahayag ang sagot
    1. X = the yearly income of someone in a Third World country
    2. the average salary from samples of 1,000 residents of a Third World country
    3. \(\overset{\bar}{X}\) ∼ N\((\text{2,000, }\frac{\text{8,000}}{\sqrt{\text{1,000}}})\)
    4. Very wide differences in data values can have averages smaller than standard deviations.
    5. The distribution of the sample mean will have higher probabilities closer to the population mean.
      P(2,000 < \(\overset{\bar}{X}\) < 2,100) = 0.1537
      P(2,100 < \(\overset{\bar}{X}\) < 2,200) = 0.1317
  25. Which of the following is NOT true about the distribution for averages?

    1. The mean, median, and mode are equal.
    2. The area under the curve is 1.
    3. The curve never touches the x-axis.
    4. The curve is skewed to the right.
  26. The cost of unleaded gasoline in the Bay Area once followed an unknown distribution with a mean of $4.59 and a standard deviation of $0.10. Sixteen gas stations from the Bay Area are randomly chosen. We are interested in the average cost of gasoline for the 16 gas stations. The distribution to use for the average cost of gasoline for the 16 gas stations is:

    1. \(\overset{\bar}{X}\) ~ N(4.59, 0.10)
    2. \(\overset{\bar}{X}\) ~ N\((\text{4}\text{.59, }\frac{0.10}{\sqrt{16}})\)
    3. \(\overset{\bar}{X}\) ~ N\((\text{4}\text{.59, }\frac{16}{0.10})\)
    4. \(\overset{\bar}{X}\) ~ N\((\text{4}\text{.59, }\frac{\sqrt{16}}{0.10})\)

    Ipahayag ang sagot

    b

    1. Find the probability that the sum of the 80 values (or the total of the 80 values) is more than 7,500.
    2. Find the sum that is 1.5 standard deviations above the mean of the sums.
    Ipahayag ang sagot

    Let X = one value from the original unknown population. The probability question asks you to find a probability for the sum (or total of) 80 values.

    ΣX = the sum or total of 80 values. Because μX = 90, σX = 15, and n = 80, \(\Sigma X\) ~ N[(80)(90),
    (\(\sqrt{\text{80}}\))(15)]

    • mean of the sums = (n)(μX) = (80)(90) = 7200
    • standard deviation of the sums = \(\text{(}\sqrt{n}\text{)(}{\sigma }_{X}\text{) = (}\sqrt{\text{80}}\text{)}\)(15)
    • sum of 80 values = Σx = 7500

    a. Find Px > 7500)

    Px > 7500) = 0.0127

    b. Find Σx where z = 1.5.

    Σx = (n)(μX) + (z)\((\sqrt{n})\)(σΧ) = (80)(90) + (1.5)(\(\sqrt{80}\))(15) = 7401.2

  27. An unknown distribution has a mean of 45 and a standard deviation of 8. A sample size of 50 is drawn randomly from the population. Find the probability that the sum of the 50 values is more than 2,400.

    Ipahayag ang sagot

    0.0040

  28. In a recent study reported Oct. 29, 2012, the mean age of tablet users is 34 years. Suppose the standard deviation is 15 years. The sample size is 50.

    1. What are the mean and standard deviation for the sum of the ages of tablet users? What is the distribution?
    2. Find the probability that the sum of the ages is between 1,500 and 1,800 years.
    3. Find the 80th percentile for the sum of the 50 ages.
    Ipahayag ang sagot
    1. μΣx = x = 50(34) = 1,700 and σΣx = \(\sqrt{n}\)σx = \((\sqrt{\text{50}}\text{ )}\)(15) = 106.01
      The distribution is normal for sums by the central limit theorem.
    2. P(1500 < Σx < 1800) = normalcdf (1500, 1800, (50)(34), \((\sqrt{\text{50}}\text{ )}\)(15)) = 0.7974
    3. Let k = the 80th percentile.
      k = invNorm(0.80,(50)(34),\((\sqrt{\text{50}}\text{ )}\)(15)) = 1789.3
  29. In a recent study reported Oct.29, 2012, the mean age of tablet users is 35 years. Suppose the standard deviation is 10 years. The sample size is 39.

    1. What are the mean and standard deviation for the sum of the ages of tablet users? What is the distribution?
    2. Find the probability that the sum of the ages is between 1,400 and 1,500 years.
    3. Find the 90th percentile for the sum of the 39 ages.
    Ipahayag ang sagot
    1. μΣx = x = 1,365 and σΣx = \(\sqrt{n}{\sigma }_{x}\) = 62.4
      The distribution is normal for sums by the central limit theorem.
    2. P(1400 < Σx < 1500) = normalcdf (1400,1500,(39)(35),(\(\sqrt{39}\))(10)) = 0.2723
    3. Let k = the 90th percentile.
      k = invNorm(0.90,(39)(35),(\(\sqrt{39}\)) (10)) = 1445.0
  30. The mean number of minutes for app engagement by a tablet user is 8.2 minutes. Suppose the standard deviation is one minute. Take a sample size of 70.

    1. What are the mean and standard deviation for the sums?
    2. Find the 95th percentile for the sum of the sample. Interpret this value in a complete sentence.
    3. Find the probability that the sum of the sample is at least 10 hours.
    Ipahayag ang sagot
    1. μΣx = x = 70(8.2) = 574 minutes and σΣx = \((\sqrt{n})({\sigma }_{x})\) = \((\sqrt{\text{70}}\text{ )}\)(1) = 8.37 minutes
    2. Let k = the 95th percentile.
      k = invNorm (0.95,(70)(8.2),\((\sqrt{\text{70}}\text{)}\)(1)) = 587.76 minutes
      Ninety-five percent of the app engagement times are at most 587.76 minutes.
    3. 10 hours = 600 minutes
      Px ≥ 600) = normalcdf(600,E99,(70)(8.2),\((\sqrt{\text{70}}\text{)}\)(1)) = 0.0009
  31. The mean number of minutes for app engagement by a tablet user is 8.2 minutes. Suppose the standard deviation is one minute. Take a sample size of 70.

    1. What is the probability that the sum of the sample is between seven hours and 10 hours? What does this mean in context of the problem?
    2. Find the 84th and 16th percentiles for the sum of the sample. Interpret these values in context.
    Ipahayag ang sagot
    1. 7 hours = 420 minutes
      10 hours = 600 minutes
      normalcdf\(P(420\le \Sigma x\le 600)=normalcdf(420,600,(70)(8.2),\sqrt{70}(1))=0.9991\)
      This means that for this sample sums there is a 99.9% chance that the sums of usage minutes will be between 420 minutes and 600 minutes.
    2. \(invNorm(0.84,(70)(8.2),\sqrt{70}(1))=582.32\)
      \(invNorm(0.16,(70)(8.2),\sqrt{70}(1))=565.68\)
      Since 84% of the app engagement times are at most 582.32 minutes and 16% of the app engagement times are at most 565.68 minutes, we may state that 68% of the app engagement times are between 565.68 minutes and 582.32 minutes.
  32. Find the probability that the sum of the 95 values is greater than 7,650.

    Ipahayag ang sagot

    0.3345

  33. Find the probability that the sum of the 95 values is less than 7,400.

  34. Find the sum that is two standard deviations above the mean of the sums.

    Ipahayag ang sagot

    7833.92

  35. Find the sum that is 1.5 standard deviations below the mean of the sums.

  36. Find the probability that the sum of the 40 values is greater than 7,500.

    Ipahayag ang sagot

    0.0089

  37. Find the probability that the sum of the 40 values is less than 7,000.

  38. Find the sum that is one standard deviation above the mean of the sums.

    Ipahayag ang sagot

    7326.49

  39. Find the sum that is 1.5 standard deviations below the mean of the sums.

Symbols used here

\sqrt{x},\ \sqrt[n]{x}
square root, n-th root
The non-negative number whose square (n-th power) is x.
\sigma,\ s,\ \sigma^2
standard deviation, sample s.d., variance
Typical distance from the mean; its square.
\bar{x},\ \mu
sample mean, population mean
Average of the data; average of the whole population.
\pm
plus or minus
Both signs at once: x = 3 ± 2 means 5 and 1.
\approx
approximately equal
Equal to the precision shown, not exactly.
n!
factorial
n × (n−1) × … × 1; the number of orderings of n things. 0! = 1.
\binom{n}{k}
binomial coefficient, "n choose k"
Number of k-element subsets of n things: n!/(k!(n−k)!).
\sum_{k=1}^{n} a_k
summation
Add a_k for k = 1 up to n.
A \cup B,\ A \cap B,\ A \setminus B
union, intersection, difference
In either; in both; in A but not B.
P(A),\ P(A \mid B)
probability, conditional probability
Chance of A; chance of A given that B happened.
E[X],\ \operatorname{Var}(X)
expected value, variance
Probability-weighted average of X; its spread.
N(\mu, \sigma^2),\ z
normal distribution, z-score
The bell curve with mean μ and variance σ²; (x − μ)/σ.

How to: The central limit theorem

  1. Power: 1/(√(25)) = 1/5.

Questions people ask

Mean or median — which should I use?

Median when the data have outliers or a long tail (incomes, house prices); mean when the data are roughly symmetric and you want every value to count. Report both if they disagree — the gap is itself information.

What does a p-value actually say?

The probability of seeing data at least this extreme if the null hypothesis were true. It is not the probability that the null hypothesis is true.

Why divide by n − 1 for the sample variance?

The sample mean sits closer to the sample than the true mean does, so squared deviations from it are slightly too small on average; dividing by n − 1 instead of n corrects the bias.

Subukan ang iyong sarili

Parts of this page are adapted from OpenStax Statistics (CC BY 4.0). Condensed and re-explained here; errors are ours.

Higit pa sa Statistics & Probability