maths.freeStatistics & Probability › Sampling and data

Sampling and data

Populations and samples, kinds of data, how a sample can mislead.

Statistics starts before any number is computed: with what was measured, on whom, and how they were chosen. A sample only tells you about the population it was drawn from fairly. This page collects the vocabulary — population, sample, parameter, statistic, variable, and the levels of measurement — with the textbook sections below.

Definitions of Statistics, Probability, and Key Terms

The science of statistics deals with the collection, analysis, interpretation, and presentation of data. We see and use data in our everyday lives.

In this course, you will learn how to organize and summarize data. Organizing and summarizing data is called descriptive statistics. Two ways to summarize data are by graphing and by using numbers, for example, finding an average. After you have studied probability and probability distributions, you will use formal methods for drawing conclusions from good data. The formal methods are called inferential statistics. Statistical inference uses probability to determine how confident we can be that our conclusions are correct.

Effective interpretation of data, or inference, is based on good procedures for producing data and thoughtful examination of the data. You will encounter what will seem to be too many mathematical formulas for interpreting data. The goal of statistics is not to perform numerous calculations using the formulas, but to gain an understanding of your data. The calculations can be done using a calculator or a computer. The understanding must come from you. If you can thoroughly grasp the basics of statistics, you can be more confident in the decisions you make in life.

Statistical Models

Statistics, like all other branches of mathematics, uses mathematical models to describe phenomena that occur in the real world. Some mathematical models are deterministic. These models can be used when one value is precisely determined from another value. Examples of deterministic models are the quadratic equations that describe the acceleration of a car from rest or the differential equations that describe the transfer of heat from a stove to a pot. These models are quite accurate and can be used to answer questions and make predictions with a high degree of precision. Space agencies, for example, use deterministic models to predict the exact amount of thrust that a rocket needs to break away from Earth’s gravity and achieve orbit.

However, life is not always precise. While scientists can predict to the minute the time that the sun will rise, they cannot say precisely where a hurricane will make landfall. Statistical models can be used to predict life’s more uncertain situations. These special forms of mathematical models or functions are based on the idea that one value affects another value. Some statistical models are mathematical functions that are more precise—one set of values can predict or determine another set of values. Or some statistical models are mathematical functions in which a set of values do not precisely determine other values. Statistical models are very useful because they can describe the probability or likelihood of an event occurring and provide alternative outcomes if the event does not occur. For example, weather forecasts are examples of statistical models. Meteorologists cannot predict tomorrow’s weather with certainty. However, they often use statistical models to tell you how likely it is to rain at any given time, and you can prepare yourself based on this probability.

Probability

Probability is a mathematical tool used to study randomness. It deals with the chance of an event occurring. For example, if you toss a fair coin four times, the outcomes may not be two heads and two tails. However, if you toss the same coin 4,000 times, the outcomes will be close to half heads and half tails. The expected theoretical probability of heads in any one toss is \(\frac{1}{2}\) or .5. Even though the outcomes of a few repetitions are uncertain, there is a regular pattern of outcomes when there are many repetitions. After reading about the English statistician Karl Pearson who tossed a coin 24,000 times with a result of 12,012 heads, one of the authors tossed a coin 2,000 times. The results were 996 heads. The fraction \(\frac{996}{2,000}\) is equal to .498 which is very close to .5, the expected probability.

The theory of probability began with the study of games of chance such as poker. Predictions take the form of probabilities. To predict the likelihood of an earthquake, of rain, or whether you will get an A in this course, we use probabilities. Doctors use probability to determine the chance of a vaccination causing the disease the vaccination is supposed to prevent. A stockbroker uses probability to determine the rate of return on a client's investments.

Data, Sampling, and Variation in Data and Sampling

Data may come from a population or from a sample. Lowercase letters like \(x\) or \(y\) generally are used to represent data values. Most data can be put into the following categories:

  • Qualitative
  • Quantitative

Qualitative data are the result of categorizing or describing attributes of a population. Qualitative data are also often called categorical data. Hair color, blood type, ethnic group, the car a person drives, and the street a person lives on are examples of qualitative data. Qualitative data are generally described by words or letters. For instance, hair color might be black, dark brown, light brown, blonde, gray, or red. Blood type might be AB+, O–, or B+. Researchers often prefer to use quantitative data over qualitative data because it lends itself more easily to mathematical analysis. For example, it does not make sense to find an average hair color or blood type.

Quantitative data are always numbers. Quantitative data are the result of counting or measuring attributes of a population. Amount of money, pulse rate, weight, number of people living in your town, and number of students who take statistics are examples of quantitative data. Quantitative data may be either discrete or continuous.

All data that are the result of counting are called quantitative discrete data. These data take on only certain numerical values. If you count the number of phone calls you receive for each day of the week, you might get values such as zero, one, two, or three.

Data that are not only made up of counting numbers, but that may include fractions, decimals, or irrational numbers, are called quantitative continuous data. Continuous data are often the results of measurements like lengths, weights, or times. A list of the lengths in minutes for all the phone calls that you make in a week, with numbers like 2.4, 7.5, or 11.0, would be quantitative continuous data.

Data Sample of Quantitative Discrete Data

The data are the number of books students carry in their backpacks. You sample five students. Two students carry three books, one student carries four books, one student carries two books, and one student carries one book. The numbers of books, 3, 4, 2, and 1, are the quantitative discrete data.

Condensed — the full section is in OpenStax Statistics.

Qualitative Data Discussion

Below are tables comparing the number of part-time and full-time students at De Anza College and Foothill College enrolled for the spring 2010 quarter. The tables display counts, frequencies, and percentages or proportions, relative frequencies. For instance, to calculate the percentage of part time students at De Anza College, divide 9,200/22,496 to get .4089. Round to the nearest thousandth—third decimal place and then multiply by 100 to get the percentage, which is 40.9 percent.

So, the percent columns make comparing the same categories in the colleges easier. Displaying percentages along with the numbers is often helpful, but it is particularly important when comparing sets of data that do not have the same totals, such as the total enrollments for both colleges in this example. Notice how much larger the percentage for part-time students at Foothill College is compared to De Anza College.

De Anza CollegeFoothill College
NumberPercentNumberPercent
Full-time9,20040.90%Full-time4,05928.60%
Part-time13,29659.10%Part-time10,12471.40%
Total22,496100%Total14,183100%

Tables are a good way of organizing and displaying data. But graphs can be even more helpful in understanding the data.

Two graphs that are used to display qualitative data are pie charts and bar graphs.

In a pie chart, categories of data are shown by wedges in a circle that represent the percent of individuals/items in each category. We use pie charts when we want to show parts of a whole.

In a bar graph, the length of the bar for each category represents the number or percent of individuals in each category. Bars may be vertical or horizontal. We use bar graphs when we want to compare categories or show changes over time.

A Pareto chart consists of bars that are sorted into order by category size (largest to smallest).

Sometimes percentages add up to be more than 100 percent (or less than 100 percent). In the graph, the percentages add to more than 100 percent because students can be in more than one category. A bar graph is appropriate to compare the relative size of the categories. A pie chart cannot be used. It also could not be used if the percentages added to less than 100 percent.

Characteristic/CategoryPercent
Students studying technical subjects40.9%
Students studying non-technical subjects48.6%
Students who intend to transfer to a four-year educational institutional61.0%
TOTAL150.5%

Condensed — the full section is in OpenStax Statistics.

Marginal Distributions in Two-Way Tables

Below is a two-way table, also called a contingency table, showing the favorite sports for 50 adults: 20 women and 30 men.

FootballBasketballTennisTotal
Men208230
Women57820
Total25151050

This is a two-way table because it displays information about two categorical variables, in this case, gender and sports. Data of this type (two variable data) are referred to as bivariate data. Because the data represent a count, or tally, of choices, it is a two-way frequency table. The entries in the total row and the total column represent marginal frequencies or marginal distributions. Note—The term marginal distributions gets its name from the fact that the distributions are found in the margins of frequency distribution tables. Marginal distributions may be given as a fraction or decimal: For example, the total for men could be given as .6 or 3/5 since \[30/50\ =\ .6\ =\ 3/5.\] Marginal distributions require bivariate data and only focus on one of the variables represented in the table. In other words, the reason 20 is a marginal frequency in this two-way table is because it represents the margin or portion of the total population that is women (20/50). The reason 25 is a marginal frequency is because it represents the portion of those sampled who favor football (25/50). Note: The values that make up the body of the table (e.g., 20, 8, 2) are called joint frequencies.

Conditional Distributions in Two-Way Tables

The distinction between a marginal distribution and a conditional distribution is that the focus is on only a particular subset of the population (not the entire population). For example, in the table, if we focused only on the subpopulation of women who prefer football, then we could calculate the conditional distributions as shown in the two-way table below.

FootballBasketballTennisTotal
Men208230
Women57820
Total25151050

To find the first sub-population of women who prefer football, read the value at the intersection of the Women row and Football column which is 5. Then, divide this by the total population of football players which is 25. So, the subpopulation of football players who are women is 5/25 which is .2.

Similarly, to find the subpopulation of women who play football, use the value of 5 which is the number of women who play football. Then, divide this by the total population of women which is 20. So, the subpopulation of women who play football is 5/20 which is .25.

Presenting Data

After deciding which graph best represents your data, you may need to present your statistical data to a class or other group in an oral report or multimedia presentation. When giving an oral presentation, you must be prepared to explain exactly how you collected or calculated the data, as well as why you chose the categories, scales, and types of graphs that you are showing. Although you may have made numerous graphs of your data, be sure to use only those that actually demonstrate the stated intentions of your statistical study. While preparing your presentation, be sure that all colors, text, and scales are visible to the entire audience. Finally, make sure to allow time for your audience to ask questions and be prepared to answer them.

Example

Try it.

Suppose the guidance counselors at De Anza and Foothill need to make an oral presentation of the student data presented in Figures 1.5 and 1.6. Under what context should they choose to display the pie graph? When might they choose the bar graph? For each graph, explain which features they should point out and the potential display problems that might exist.

Solution

The guidance counselors should use the pie graph if the desired information is the percentage of each school’s enrollment. They should use the bar graph if knowing the exact numbers of students and the relative sizes of each category at each school are important points to be made. For the pie graph, they should point out which color represents part-time students and which represents full-time students. They should also be sure that the numbers and colors are visible when displayed. For the bar graph, they should point out the scale and the total numbers for each category, and they should be sure that the numbers, colors, and scale marks are all displayed clearly.

Sampling

Gathering information about an entire population often costs too much or is virtually impossible. Instead, we use a sample of the population. A sample should have the same characteristics as the population it is representing. Most statisticians use various methods of random sampling in an attempt to achieve this goal. This section will describe a few of the most common methods. There are several different methods of random sampling. In each form of random sampling, each member of a population initially has an equal chance of being selected for the sample. Each method has pros and cons. The easiest method to describe is called a simple random sample. In a simple random sample, each group has the same chance of being selected. In other words, each sample of the same size has an equal chance of being selected. For example, suppose Lisa wants to form a four-person study group (herself and three other people) from her pre-calculus class, which has 31 members not including Lisa. To choose a simple random sample of size three from the other members of her class, Lisa could put all 31 names in a hat, shake the hat, close her eyes, and pick out three names. A more technological way is for Lisa to first list the last names of the members of her class together with a two-digit number, as in .

IDNameIDNameIDName
00Anselmo11King22Roquero
01Bautista12Legeny23Roth
02Bayani13Lisa24Rowell
03Cheng14Lundquist25Salangsang
04Cuarismo15Macierz26Slade
05Cuningham16Motogawa27Stratcher
06Fontecha17Okimoto28Tallai
07Hong18Patel29Tran
08Hoobler19Price30Wai
09Jiao20Quizon31Wood
10Khan21Reyes

Lisa can use a table of random numbers (found in many statistics books and mathematical handbooks), a calculator, or a computer to generate random numbers. The most common random number generators are five digit numbers where each digit is a unique number from 0 to 9. For this example, suppose Lisa chooses to generate random numbers from a calculator. The numbers generated are as follows:

.94360, .99832, .14669, .51470, .40581, .73381, .04399.

Lisa reads two-digit groups until she has chosen three class members (That is, she reads .94360 as the groups 94, 43, 36, 60.) Each random number may only contribute one class member. If she needed to, Lisa could have generated more random numbers.

The table below shows how Lisa reads two-digit numbers form each random number. Each two-digit number in the table would represent each student in the roster above in .

Random numberNumbers read by Lisa
.9436094433660
.9983299988332
.1466914466669
.5147051144770
.4058140055881
.7338173333881
.0439904393999

Condensed — the full section is in OpenStax Statistics.

Critical Evaluation

We need to evaluate the statistical studies we read about critically and analyze them before accepting the results of the studies. Common problems to be aware of include the following:

  • Problems with samples: —A sample must be representative of the population. A sample that is not representative of the population is biased. Biased samples that are not representative of the population give results that are inaccurate and not reliable. Reliability in statistical measures must also be considered when analyzing data. Reliability refers to the consistency of a measure. A measure is reliable when the same results are produced given the same circumstances.
  • Self-selected samples—Responses only by people who choose to respond, such as internet surveys, are often unreliable.
  • Sample size issues—: Samples that are too small may be unreliable. Larger samples are better, if possible. In some situations, having small samples is unavoidable and can still be used to draw conclusions. Examples include crash testing cars or medical testing for rare conditions.
  • Undue influence—: collecting data or asking questions in a way that influences the response.
  • Non-response or refusal of subject to participate: —The collected responses may no longer be representative of the population.  Often, people with strong positive or negative opinions may answer surveys, which can affect the results.
  • Causality: —A relationship between two variables does not mean that one causes the other to occur. They may be related (correlated) because of their relationship through a different variable.
  • Self-funded or self-interest studies—: A study performed by a person or organization in order to support their claim. Is the study impartial? Read the study carefully to evaluate the work. Do not automatically assume that the study is good, but do not automatically assume the study is bad either. Evaluate it on its merits and the work done.
  • Misleading use of data—: These can be improperly displayed graphs, incomplete data, or lack of context.

If we were to examine two samples representing the same population, even if we used random sampling methods for the samples, they would not be exactly the same. Just as there is variation in data, there is variation in samples. As you become accustomed to sampling, the variability will begin to seem natural.

Condensed — the full section is in OpenStax Statistics.

Variation in Data

Variation is present in any set of data. For example, 16-ounce cans of beverage may contain more or less than 16 ounces of liquid. In one study, eight 16 ounce cans were measured and produced the following amount (in ounces) of beverage:

15.8, 16.1, 15.2, 14.8, 15.8, 15.9, 16.0, 15.5.

Measurements of the amount of beverage in a 16-ounce can may vary because different people make the measurements or because the exact amount, 16 ounces of liquid, was not put into the cans. Manufacturers regularly run tests to determine if the amount of beverage in a 16-ounce can falls within the desired range.

Be aware that as you take data, your data may vary somewhat from the data someone else is taking for the same purpose. This is completely natural. However, if two or more of you are taking the same data and get very different results, it is time for you and the others to reevaluate your data-taking methods and your accuracy.

Variation in Samples

It was mentioned previously that two or more samples from the same population, taken randomly, and having close to the same characteristics of the population will likely be different from each other. Suppose Doreen and Jung both decide to study the average amount of time students at their high school sleep each night. Doreen and Jung each take samples of 500 students. Doreen uses systematic sampling and Jung uses cluster sampling. Doreen's sample will be different from Jung's sample. Even if Doreen and Jung used the same sampling method, in all likelihood their samples would be different. Neither would be wrong, however.

Think about what contributes to making Doreen’s and Jung’s samples different.

If Doreen and Jung took larger samples, that is, the number of data values is increased, their sample results (the average amount of time a student sleeps) might be closer to the actual population average. But still, their samples would be, in all likelihood, different from each other. This is called sampling variability. In other words, it refers to how much a statistic varies from sample to sample within a population. The larger the sample size, the smaller the variability between samples will be. So, the large sample size makes for a better, more reliable statistic.

The size of a sample (often called the number of observations) is important. The examples you have seen in this book so far have been small. Samples of only a few hundred observations, or even smaller, are sufficient for many purposes. In polling, samples that are from 1,200–1,500 observations are considered large enough and good enough if the survey is random and is well done. You will learn why when you study confidence intervals.

Be aware that many large samples are biased. For example, internet surveys are invariably biased, because people choose to respond or not.

Frequency, Frequency Tables, and Levels of Measurement

Once you have a set of data, you will need to organize it so that you can analyze how frequently each datum occurs in the set. However, when calculating the frequency, you may need to round your answers so that they are as precise as possible.

Answers and Rounding Off

A simple way to round off answers is to carry your final answer one more decimal place than was present in the original data. Round off only the final answer. Do not round off any intermediate results, if possible. If it becomes necessary to round off intermediate results, carry them to at least twice as many decimal places as the final answer. Expect that some of your answers will vary from the text due to rounding errors.

It is not necessary to reduce most fractions in this course. Especially in Probability Topics, the chapter on probability, it is more helpful to leave an answer as an unreduced fraction.

Levels of Measurement

The way a set of data is measured is called its level of measurement. Correct statistical procedures depend on a researcher being familiar with levels of measurement. Not every statistical operation can be used with every set of data. Data can be classified into four levels of measurement. They are as follows (from lowest to highest level):

  • Nominal scale level
  • Ordinal scale level
  • Interval scale level
  • Ratio scale level

Data that is measured using a nominal scale is qualitative (categorical). Categories, colors, names, labels, and favorite foods along with yes or no responses are examples of nominal level data. Nominal scale data are not ordered. For example, trying to classify people according to their favorite food does not make any sense. Putting pizza first and sushi second is not meaningful.

Smartphone companies are another example of nominal scale data. The data are the names of the companies that make smartphones, but there is no agreed upon order of these brands, even though people may have personal preferences. Nominal scale data cannot be used in calculations.

Data that is measured using an ordinal scale is similar to nominal scale data but there is a big difference. The ordinal scale data can be ordered. An example of ordinal scale data is a list of the top five national parks in the United States. The top five national parks in the United States can be ranked from one to five but we cannot measure differences between the data.

Another example of using the ordinal scale is a cruise survey where the responses to questions about the cruise are excellent, good, satisfactory, and unsatisfactory. These responses are ordered from the most desired response to the least desired. But the differences between two pieces of data cannot be measured. Like the nominal scale data, ordinal scale data cannot be used in calculations.

Like data measured on an ordinal scale, data that are measured on an interval scale have a definite ordering. While ordinal scale data are categorical, or qualitative, interval scale data are numerical, or quantitative. It is possible to calculate differences between values measured on an interval scale. There is no minimum or "zero" value, however.

Temperature scales like Celsius (C) and Fahrenheit (F) are measured by using the interval scale. In both temperature measurements, 40° is equal to 100° minus 60°. Differences make sense. But 0 degrees does not represent a minimum value. In both scales, 0 is not the absolute lowest temperature. Temperatures like –10 °F and –15 °C exist and are colder than 0.

Condensed — the full section is in OpenStax Statistics.

Frequency

Twenty students were asked how many hours they worked per day. Their responses, in hours, are as follows: 5, 6, 3, 3, 2, 4, 7, 5, 2, 3, 5, 6, 5, 4, 4, 3, 5, 2, 5, 3.

lists the different data values in ascending order and their frequencies.

DATA VALUEFREQUENCY
23
35
43
56
62
71

A frequency is the number of times a value of the data occurs. According to , there are three students who work two hours, five students who work three hours, and so on. The sum of the values in the frequency column, 20, represents the total number of students included in the sample.

A relative frequency is the ratio (fraction or proportion) of the number of times a value of the data occurs in the set of all outcomes to the total number of outcomes. To find the relative frequencies, divide each frequency by the total number of students in the sample, in this case, 20. Relative frequencies can be written as fractions, percents, or decimals.

DATA VALUEFREQUENCYRELATIVE FREQUENCY
23\(\frac{3}{20}\) or .15
35\(\frac{5}{20}\) or .25
43\(\frac{3}{20}\) or .15
56\(\frac{6}{20}\) or .30
62\(\frac{2}{20}\) or .10
71\(\frac{1}{20}\) or .05

The sum of the values in the relative frequency column of is \(\frac{20}{20}\) , or 1.

Cumulative relative frequency is the accumulation of the previous relative frequencies. To find the cumulative relative frequencies, add all the previous relative frequencies to the relative frequency for the current row, as shown in .

In the first row, the cumulative frequency is simply .15 because it is the only one. In the second row, the relative frequency was .25, so adding that to .15, we get a relative frequency of .40. Continue adding the relative frequencies in each row to get the rest of the column.

DATA VALUEFREQUENCYRELATIVE
FREQUENCY
CUMULATIVE RELATIVE
FREQUENCY
23\(\frac{3}{20}\) or .15.15
35\(\frac{5}{20}\) or .25.15 + .25 = .40
43\(\frac{3}{20}\) or .15.40 + .15 = .55
56\(\frac{6}{20}\) or .30.55 + .30 = .85
62\(\frac{2}{20}\) or .10.85 + .10 = .95
71\(\frac{1}{20}\) or .05.95 + .05 = 1.00
  • 59.95–61.95 inches
  • 61.95–63.95 inches
  • 63.95–65.95 inches
  • 65.95–67.95 inches
  • 67.95–69.95 inches
  • 69.95–71.95 inches
  • 71.95–73.95 inches
  • 73.95–75.95 inches

Condensed — the full section is in OpenStax Statistics.

Experimental Design and Ethics

Does aspirin reduce the risk of heart attacks? Is one brand of fertilizer more effective at growing roses than another? Is fatigue as dangerous to a driver as speeding? Questions like these are answered using randomized experiments. In this module, you will learn important aspects of experimental design. Proper study design ensures the production of reliable, accurate data.

The purpose of an experiment is to investigate the relationship between two variables. In an experiment, there is the explanatory variable which affects the response variable. In a randomized experiment, the researcher manipulates the explanatory variable and then observes the response variable. Each value of the explanatory variable used in an experiment is called a treatment.

You want to investigate the effectiveness of vitamin E in preventing disease. You recruit a group of subjects and ask them if they regularly take vitamin E. You notice that the subjects who take vitamin E exhibit better health on average than those who do not. Does this prove that vitamin E is effective in disease prevention? It does not. There are many differences between the two groups compared in addition to vitamin E consumption. People who take vitamin E regularly often take other steps to improve their health: exercise, diet, other vitamin supplements. Any one of these factors could be influencing health. As described, this study does not prove that vitamin E is the key to disease prevention.

Additional variables that can cloud a study are called lurking variables. In order to prove that the explanatory variable is causing a change in the response variable, it is necessary to isolate the explanatory variable. The researcher must design her experiment in such a way that there is only one difference between groups being compared: the planned treatments. This is accomplished by the random assignment of experimental units to treatment groups. When subjects are assigned treatments randomly, all of the potential lurking variables are spread equally among the groups. At this point the only difference between groups is the one imposed by the researcher. Different outcomes measured in the response variable, therefore, must be a direct result of the different treatments. In this way, an experiment can prove a cause-and-effect connection between the explanatory and response variables.

Condensed — the full section is in OpenStax Statistics.

Ethics

The widespread misuse and misrepresentation of statistical information often gives the field a bad name. Some say that “numbers don’t lie,” but the people who use numbers to support their claims often do.

A recent investigation of famous social psychologist, Diederik Stapel, has led to the retraction of his articles from some of the world’s top journals including, Journal of Experimental Social Psychology, Social Psychology, Basic and Applied Social Psychology, British Journal of Social Psychology, and the magazine Science. Diederik Stapel is a former professor at Tilburg University in the Netherlands. Over the past two years, an extensive investigation involving three universities where Stapel has worked concluded that the psychologist is guilty of fraud on a colossal scale. Falsified data taints over 55 papers he authored and 10 Ph.D. dissertations that he supervised.

Stapel did not deny that his deceit was driven by ambition. But it was more complicated than that, he told me. He insisted that he loved social psychology but had been frustrated by the messiness of experimental data, which rarely led to clear conclusions. His lifelong obsession with elegance and order, he said, led him to concoct results that journals found attractive. “It was a quest for aesthetics, for beauty—instead of the truth,” he said. He described his behavior as an addiction that drove him to carry out acts of increasingly daring fraud. Bhattacharjee, Y. (2013, April 26). The mind of a con man. The New York Times. Retrieved from http://www.nytimes.com/2013/04/28/magazine/diederik-stapels-audacious-academic-fraud.html?_r=3&src=dayp&.

The committee investigating Stapel concluded that he is guilty of several practices including

  • creating datasets, which largely confirmed the prior expectations,
  • altering data in existing datasets,
  • changing measuring instruments without reporting the change, and
  • misrepresenting the number of experimental subjects.

Clearly, it is never acceptable to falsify data the way this researcher did. Sometimes, however, violations of ethics are not as easy to spot.

Many types of statistical fraud are difficult to spot. Some researchers simply stop collecting data once they have just enough to prove what they had hoped to prove. They don’t want to take the chance that a more extensive study would complicate their lives by producing data contradicting their hypothesis.

Condensed — the full section is in OpenStax Statistics.

Ишлатилган мисол: mean of 4, 8, 15, 16, 23, 42

Mean of 4, 8, 15, 16, 23, 42

4,\ 8,\ 15,\ 16,\ 23,\ 42

Қадамма-қадам

  1. 4, 8, 15, 16, 23, 42

    6 values.

  2. \bar{x} = \frac{4 + 8 + 15 + 16 + 23 + 42}{6} = \frac{108}{6} = 18

    Mean: add them up and divide by how many there are.

Жавобни кўрсатиш
18

Practice (40)

Try each one on paper first. Reveal the answer to check; verified ones can be opened in the solver for every step.

  1. Determine what the population, sample, parameter, statistic, variable, and data referred to in the following study.

    We want to know the mean amount of extracurricular activities in which high school students participate. We randomly surveyed 100 high school students. Three of those students were in 2, 5, and 7 extracurricular activities, respectively.

    Жавобни кўрсатиш

    The population is all high school students.

    The sample is the 100 high school students interviewed.

    The parameter is the mean amount of extracurricular activities in which all high school students participate.

    The statistic is the mean amount of extracurricular activities in which the sample of high school students participate.

    The variable could be the amount of extracurricular activities by one high school student. Let X = the amount of extracurricular activities by one high school student.

    The data are the number of extracurricular activities in which the high school students participate. Examples of the data are 2, 5, 7.

  2. Find an article online or in a newspaper or magazine that refers to a statistical study or poll. Identify what each of the key terms—population, sample, parameter, statistic, variable, and data—refers to in the study mentioned in the article. Does the article use the key terms correctly?

    Жавобни кўрсатиш

    The population is all families with children attending Knoll Academy.

    The sample is a random selection of 100 families with children attending Knoll Academy.

    The parameter is the average (mean) amount of money spent on school uniforms by families with children at Knoll Academy.

    The statistic is the average (mean) amount of money spent on school uniforms by families in the sample.

    The variable is the amount of money spent by one family. Let X = the amount of money spent on school uniforms by one family with children attending Knoll Academy.

    The data are the dollar amounts spent by the families. Examples of the data are $65, $75, and $95.

  3. Determine what the key terms refer to in the following study.

    A study was conducted at a local high school to analyze the average cumulative GPAs of students who graduated last year. Fill in the letter of the phrase that best describes each of the items below.

    1. Population ____ 2. Statistic ____ 3. Parameter ____ 4. Sample ____ 5. Variable ____ 6. Data ____

    • a) all students who attended the high school last year
    • b) the cumulative GPA of one student who graduated from the high school last year
    • c) 3.65, 2.80, 1.50, 3.90
    • d) a group of students who graduated from the high school last year, randomly selected
    • e) the average cumulative GPA of students who graduated from the high school last year
    • f) all students who graduated from the high school last year
    • g) the average cumulative GPA of students in the study who graduated from the high school last year
    Жавобни кўрсатиш

    • 1. f
    • 2. g
    • 3. e
    • 4. d
    • 5. b
    • 6. c

  4. Determine what the population, sample, parameter, statistic, variable, and data referred to in the following study.

    As part of a study designed to test the safety of automobiles, the National Transportation Safety Board collected and reviewed data about the effects of an automobile crash on test dummies (The Data and Story Library, n.d.). Here is the criterion they used.

    Speed at which Cars CrashedLocation of Driver (i.e., dummies)
    35 miles/hourFront seat

    Cars with dummies in the front seats were crashed into a wall at a speed of 35 miles per hour. We want to know the proportion of dummies in the driver’s seat that would have had head injuries, if they had been actual drivers. We start with a simple random sample of 75 cars.

    Жавобни кўрсатиш

    The population is all cars containing dummies in the front seat.

    The sample is the 75 cars, selected by a simple random sample.

    The parameter is the proportion of driver dummies—if they had been real people—who would have suffered head injuries in the population.

    The statistic is proportion of driver dummies—if they had been real people—who would have suffered head injuries in the sample.

    The variable X = whether driver dummies—if they had been real people—would have suffered head injuries.

    The data are either: yes, had head injury, or no, did not.

  5. Determine what the population, sample, parameter, statistic, variable, and data referred to in the following study.

    An insurance company would like to determine the proportion of all medical doctors who have been involved in one or more malpractice lawsuits. The company selects 500 doctors at random from a professional directory and determines the number in the sample who have been involved in a malpractice lawsuit.

    Жавобни кўрсатиш

    The population is all medical doctors listed in the professional directory.

    The parameter is the proportion of medical doctors who have been involved in one or more malpractice suits in the population.

    The sample is the 500 doctors selected at random from the professional directory.

    The statistic is the proportion of medical doctors who have been involved in one or more malpractice suits in the sample.

    The variable X records whether a doctor has or has not been involved in a malpractice suit.

    The data are either: yes, was involved in one or more malpractice lawsuits; or no, was not.

  6. Below is a two-way table showing the types of college sports played by men and women.

    SoccerBasketballLacrosseTotal
    Women88420
    Men412420
    Total1220840

    Given these data, calculate the marginal distributions of college sports for the people surveyed.

    Жавобни кўрсатиш

    soccer = 12/40 = ;

    basketball = 20/40 = ;

    lacrosse = 8/40 = 0.2

  7. Below is a two-way table showing the types of college sports played by men and women.

    SoccerBasketballLacrosseTotal
    Women88420
    Men412420
    Total1220840

    Given these data, calculate the conditional distributions for the subpopulation of women who play college sports.

    Жавобни кўрсатиш

    women who play soccer = 8/20 = ;

    women who play basketball = 8/20 = ;

    women who play lacrosse = 4/20 = ;

  8. population

    Жавобни кўрсатиш

    patients with the virus

  9. parameter

    Жавобни кўрсатиш

    The average length of time (in months) patients live after treatment.

  10. statistic

  11. variable

    Жавобни кўрсатиш

    X = the length of time (in months) patients live after treatment

  12. For each of the following situations, indicate whether it would be best modeled with a mathematical model or a statistical model. Explain your answers.

    1. driving time from New York to Florida
    2. departure time of a commuter train at rush hour
    3. distance from your house to school
    4. temperature of a refrigerator at any given time
    5. weight of a bag of rice at the store
    Жавобни кўрсатиш
    1. statistical model: The time any journey takes from New York to Florida is variable and depends on traffic and other driving conditions.
    2. statistical model: Although trains try to leave on time, the exact time of departure differs slightly from day to day.
    3. mathematical model: The distance from your house to school is the same every day and can be precisely determined.
    4. statistical model: The temperature of a refrigerator fluctuates as the compressor turns on and off.
    5. statistical model: The fill weight of a bag of rice is different for each bag. Manufacturers spend considerable effort to minimize the variance from bag to bag.
  13. A fitness center is interested in the mean amount of time a client exercises in the center each week.

  14. Ski resorts are interested in the mean age that children take their first ski and snowboard lessons. They need this information to plan their ski classes optimally.

    Жавобни кўрсатиш
    1. all children who take ski or snowboard lessons
    2. a group of these children
    3. the population mean age of children who take their first snowboard lesson
    4. the sample mean age of children who take their first snowboard lesson
    5. X = the age of one child who takes his or her first ski or snowboard lesson
    6. values for X, such as 3, 7, and so on
  15. A cardiologist is interested in the mean recovery period of her patients who have had heart attacks.

  16. Insurance companies are interested in the mean health costs each year of their clients, so that they can determine the costs of health insurance.

    Жавобни кўрсатиш
    1. the clients of the insurance companies
    2. a group of the clients
    3. the mean health costs of the clients
    4. the mean health costs of the sample
    5. X = the health costs of one client
    6. values for X, such as 34, 9, 82, and so on
  17. A politician is interested in the proportion of voters in his district who think he is doing a new good job.

  18. A marriage counselor is interested in the proportion of clients she counsels who stay married.

    Жавобни кўрсатиш
    1. all the clients of this counselor
    2. a group of clients of this marriage counselor
    3. the proportion of all her clients who stay married
    4. the proportion of the sample of the counselor’s clients who stay married
    5. X = the number of couples who stay married
    6. yes, no
  19. Political pollsters may be interested in the proportion of people who will vote for a particular cause.

  20. A marketing company is interested in the proportion of people who will buy a particular product.

    Жавобни кўрсатиш
    1. all people (maybe in a certain geographic area, such as the United States)
    2. a group of the people
    3. the proportion of all people who will buy the product
    4. the proportion of the sample who will buy the product
    5. X = the number of people who will buy it
    6. buy, not buy
  21. What is the population she is interested in?

    1. all Lake Tahoe Community College students
    2. all Lake Tahoe Community College English students
    3. all Lake Tahoe Community College students in her classes
    4. all Lake Tahoe Community College math students
  22. Consider the following

    \(X\) = number of days a Lake Tahoe Community College math student is absent.

    In this case, X is an example of which of the following?

    1. variable
    2. population
    3. statistic
    4. data
    Жавобни кўрсатиш

    a

  23. The instructor’s sample produces a mean number of days absent of 3.5 days. This value is an example of which of the following?

    1. parameter
    2. data
    3. statistic
    4. variable
  24. The data are the number of machines in a gym. You sample five gyms. One gym has 12 machines, one gym has 15 machines, one gym has 10 machines, one gym has 22 machines, and the other gym has 20 machines. What type of data is this?

    Жавобни кўрсатиш

    quantitative discrete data

  25. The data are the areas of lawns in square feet. You sample five houses. The areas of the lawns are 144 sq. ft., 160 sq. ft., 190 sq. ft., 180 sq. ft., and 210 sq. ft. What type of data is this?

    Жавобни кўрсатиш

    quantitative continuous data

  26. Name data sets that are quantitative discrete, quantitative continuous, and qualitative.

    Жавобни кўрсатиш

    A possible solution

    • One example of a quantitative discrete data set would be three cans of soup, two packages of nuts, four kinds of vegetables, and two desserts because you count them.
    • The weights of the soups (19 ounces, 14.1 ounces, 19 ounces) are quantitative continuous data because you measure weights as precisely as possible.
    • Types of soups, nuts, vegetables, and desserts are qualitative data because they are categorical.
  27. The data are the colors of houses. You sample five houses. The colors of the houses are white, yellow, white, red, and white. What type of data is this?

    Жавобни кўрсатиш

    qualitative data

  28. Work collaboratively to determine the correct data type: quantitative or qualitative. Indicate whether quantitative data are continuous or discrete. Hint: Data that are discrete often start with the words the number of.

    • the number of pairs of shoes you own
    • the type of car you drive
    • the distance from your home to the nearest grocery store
    • the number of classes you take per school year
    • the type of calculator you use
    • weights of sumo wrestlers
    • number of correct answers on a quiz
    • IQ scores (This may cause some discussion.)
    Жавобни кўрсатиш

    Items a, d, and g are quantitative discrete; items c, f, and h are quantitative continuous; items b and e are qualitative or categorical.

  29. Determine the correct data type, quantitative or qualitative, for the number of cars in a parking lot. Indicate whether quantitative data are continuous or discrete.

    Жавобни кўрсатиш

    quantitative discrete

  30. A statistics professor collects information about the classification of her students as freshmen, sophomores, juniors, or seniors. The data she collects are summarized in the pie chart . What type of data does this graph show?

    Жавобни кўрсатиш

    This pie chart shows the students in each year, which is qualitative or categorical data.

  31. A large school district keeps data of the scores students earn on an end of the year standardized exam. The data he collects are summarized in the histogram. The class boundaries are 50 to less than 60, 60 to less than 70, 70 to less than 80, 80 to less than 90, and 90 to less than 100.

    Жавобни кўрсатиш

    A histogram is used to display quantitative data: the numbers of credit hours completed. Because students can complete only a whole number of hours (no fractions of hours allowed), this data is quantitative discrete.

  32. Suppose the guidance counselors at De Anza and Foothill need to make an oral presentation of the student data presented in Figures 1.5 and 1.6. Under what context should they choose to display the pie graph? When might they choose the bar graph? For each graph, explain which features they should point out and the potential display problems that might exist.

    Жавобни кўрсатиш

    The guidance counselors should use the pie graph if the desired information is the percentage of each school’s enrollment. They should use the bar graph if knowing the exact numbers of students and the relative sizes of each category at each school are important points to be made. For the pie graph, they should point out which color represents part-time students and which represents full-time students. They should also be sure that the numbers and colors are visible when displayed. For the bar graph, they should point out the scale and the total numbers for each category, and they should be sure that the numbers, colors, and scale marks are all displayed clearly.

  33. Suppose you were asked to give an oral presentation of the data graphed in the pie chart in Figure 1.11(b). What features would you point out on the graph? What potential display problems with the graph should you check before giving your presentation?

  34. Determine the type of sampling used (simple random, stratified, systematic, cluster, or convenience).

    1. A soccer coach selects six players from a group of boys aged eight to ten, seven players from a group of boys aged 11 to 12, and three players from a group of boys aged 13 to 14 to form a recreational soccer team.
    2. A pollster interviews all human resource personnel in five different high tech companies.
    3. A high school educational researcher interviews 50 high school female teachers and 50 high school male teachers.
    4. A medical researcher interviews every third cancer patient from a list of cancer patients at a local hospital.
    5. A high school counselor uses a computer to generate 50 random numbers and then picks students whose names correspond to the numbers.
    6. A student interviews classmates in his algebra class to determine how many pairs of jeans a student owns, on average.
    Жавобни кўрсатиш

    a. stratified b. cluster c. stratified d. systematic e. simple random f. convenience

  35. Determine the type of sampling used (simple random, stratified, systematic, cluster, or convenience).

    A high school principal polls 50 freshmen, 50 sophomores, 50 juniors, and 50 seniors regarding policy changes for after school activities.

    Жавобни кўрсатиш

    stratified

  36. a. Do you think that either of these samples is representative of (or is characteristic of) the entire 10,000 part-time student population?

    Жавобни кўрсатиш

    a. No. The first sample probably consists of science-oriented students. Besides the chemistry course, some of them are also taking first-term calculus. Books for these classes tend to be expensive. Most of these students are, more than likely, paying more than the average part-time student for their books. The second sample is a group of seniors who are, more than likely, taking courses for health and interest. The amount of money they spend on books is probably much less than the average part-time student. Both samples are biased. Also, in both cases, not all students have a chance to be in either sample.

  37. b. Since these samples are not representative of the entire population, is it wise to use the results to describe the entire population?

    Жавобни кўрсатиш

    b. No. For these samples, each member of the population did not have an equally likely chance of being chosen.

  38. c. Is the sample biased?

    Жавобни кўрсатиш

    c. The sample is unbiased, but a larger sample would be recommended to increase the likelihood that the sample will be close to representative of the population. However, for a biased sampling technique, even a large sample runs the risk of not being representative of the population.

  39. A local radio station has a fan base of 20,000 listeners. The station wants to know if its audience would prefer more music or more talk shows. Asking all 20,000 listeners is an almost impossible task.

    The station uses convenience sampling and surveys the first 200 people they meet at one of the station’s music concert events. Twenty-four people said they’d prefer more talk shows, and 176 people said they’d prefer more music.

    Do you think that this sample is representative of (or is characteristic of) the entire 20,000 listener population?

    Жавобни кўрсатиш

    The sample probably consists more of people who prefer music because it is a concert event. Also, the sample represents only those who showed up to the event earlier than the majority. The sample probably doesn’t represent the entire fan base and is probably biased towards people who would prefer music.

  40. Number of times per week is what type of data?

    • a. qualitative (categorical)
    • b. quantitative discrete
    • c. quantitative continuous

Ўзингизни синаб кўринг

Parts of this page are adapted from OpenStax Statistics (CC BY 4.0). Condensed and re-explained here; errors are ours.

Кўпроқ Statistics & Probability