maths.free › Statistics & Probability › 8. Statistics › Gathering and Organizing Data
Gathering and Organizing Data
Distinguish among sampling techniques.
Learning Objectives
After completing this section, you should be able to:
- Distinguish among sampling techniques.
- Organize data using an appropriate method.
- Create frequency distributions.
Sampling and Gathering Data
The Digest's failure highlights the need for what is now considered the most important criterion for sampling: randomness. This randomness can be achieved in several ways. Here we cover some of the most common.
A simple random sample is chosen in a way that every unit in the population has an equal chance of being selected, and the chances of a unit being selected do not depend on the units already chosen. An example of this is choosing a group of people by drawing names out of a hat (assuming the names are well-mixed in the hat).
A systematic random sample is selected from an ordered list of the population (for example, names sorted alphabetically or students listed by student ID). First, we decide what proportion of the population will be in our sample. We want to express that proportion as a fraction with 1 in the numerator. Let’s call that number D. Next, we’ll choose a random number between one and D. The unit at that position will go into our sample. We’ll find the rest of our sample by choosing every Dth unit in the list, starting with our random number.
To walk through an example, let’s say we want to sample 2% of the population: \(2\%=\frac{2}{100}=\frac{1}{50}\). (Note: If the number in the denominator isn’t a whole number, we can just round it off. This part of the process doesn’t have to be precise.) We can then use a random number generator to find a random number between 1 and 50; let's use 31. In our example, our sample would then be the units in the list at positions 31, 81 (31 + 50), 131 (81 + 50), and so forth.
A stratified sample is one chosen so that particular groups in the population are certain to be represented. Let’s say you are studying the population of students in a large high school (where the grades run from 9th to 12th), and you want to choose a sample of 12 students. If you use a simple or systematic random sample, there’s a pretty good chance that you’ll miss one grade completely. In a stratified sample, you would first divide the population into groups (the strata), then take a random sample within each stratum (that’s the singular form of “strata”). In the high school example, we could divide the population into grades, then take a random sample of three students within each grade. That would get us to the 12 students we need while ensuring coverage of each grade.
Condensed — the full section is in OpenStax Contemporary Mathematics.
Organizing Data
Once data have been collected, we turn our attention to analysis. Before we analyze, though, it’s useful to reorganize the data into a format that makes the analysis easier. For example, if our data were collected using a paper survey, our raw data are all broken down by respondent (represented by an individual response sheet). To perform an analysis on all the responses to an individual question, we need to first group all the responses to each question together. The way we organize the data depends on the type of data we’ve collected.
There are two broad types of data: categorical and quantitative. Categorical data classifies the unit into a group (or category). Examples of categorical data include a response to a yes-or-no question, or the color of a person’s eyes. Quantitative data is a numerical measure of a property of a unit. Examples of quantitative data include the time it takes for a rat to run through a maze or a person’s daily calorie intake. We’ll look at each type of data in turn when considering how best to organize.
The best way to organize categorical data is using a categorical frequency distribution. A categorical frequency distribution is a table with two columns. The first contains all the categories present in the data, each listed once. The second contains the frequencies of each category, which are just a count of how often each category appears in the data.
Creating a Categorical Frequency Distribution
Try it.
A teacher records the responses of the class (28 students) on the first question of a multiple choice quiz, with five possible responses (A, B, C, D, and E):
| A | A | C | A | B | B | A | E | A | C | A | A | A | C |
| E | A | B | A | A | C | A | B | E | E | A | A | C | C |
Create a categorical frequency distribution that organizes the responses.
Solution
Step 1: For each possible response, count the number of times that response appears in the data. In the responses for this class, “A” appears 14 times, “B” 4 times, “C” 6 times, “D” 0 times, and “E” 4 times.
Step 2: Make a table with two columns. The first column should be labeled so that the reader knows what the responses mean, and the second should be labeled “Frequency.”
| Response to First Question | Frequency |
| A | 14 |
| B | 4 |
| C | 6 |
| D | 0 |
| E | 4 |
Step 3: Check your work. If you add up your frequencies, you should get the same number as the total number of responses. Twenty-eight students answered that first question, and \(14+4+6+0+4=28\).
Condensed — the full section is in OpenStax Contemporary Mathematics.
Key Concepts
- Categorical data places units into groups (categories), while quantitative data is a numerical measure of a property of a unit.
- The sampling method for a study depends on the way that randomization is used to select units for the sample.
- Frequency distributions help to summarize data by counting the number of units that fall into a particular category or range of quantitative values.
Practice (3)
Try each one on paper first. Reveal the answer to check; verified ones can be opened in the solver for every step.
-
A teacher records the responses of the class (28 students) on the first question of a multiple choice quiz, with five possible responses (A, B, C, D, and E):
A A C A B B A E A C A A A C E A B A A C A B E E A A C C Create a categorical frequency distribution that organizes the responses.
@ action
Step 1: For each possible response, count the number of times that response appears in the data. In the responses for this class, “A” appears 14 times, “B” 4 times, “C” 6 times, “D” 0 times, and “E” 4 times.
Step 2: Make a table with two columns. The first column should be labeled so that the reader knows what the responses mean, and the second should be labeled “Frequency.”
Response to First Question Frequency A 14 B 4 C 6 D 0 E 4 Step 3: Check your work. If you add up your frequencies, you should get the same number as the total number of responses. Twenty-eight students answered that first question, and \(14+4+6+0+4=28\).
-
Attendees of a conflict resolution workshop are asked how many siblings they have. The responses are as follows:
1 0 1 1 2 0 3 1 1 4 1 2 0 1 3 1 2 1 2 4 1 0 1 3 0 1 2 2 1 5 Create a frequency distribution to organize the responses.
@ action
Step 1: Count the number of times you see each unique response: “0” appears 5 times, “1” appears 13 times, “2” appears 6 times, “3” appears 3 times, “4” appears twice, and “5” appears once.
Step 2: Make a table with two columns. The first column should be labeled so that the reader knows what the responses mean, and the second should be labeled “Frequency.” Then fill in the results of our count.
Number of Siblings Frequency Number of Siblings Frequency 0 5 3 3 1 13 4 2 2 6 5 1 Step 3: Check your work. If you add up your counts, you should get the same number as the total number of responses. Looking back at the raw data, there were 30 responses, and \(5+13+6+3+2+1=30\).
-
The GPAs of students enrolled in an advanced sociology class are listed in the following table. At this institution, 4.00 is the maximum possible GPA.
3.93 3.43 2.87 2.51 2.70 1.91 2.32 2.85 3.06 3.03 3.49 1.84 3.72 2.56 1.99 3.40 3.74 3.23 1.98 3.05 1.43 2.90 1.20 3.72 3.56 3.07 2.58 4.00 2.79 3.81 2.60 3.69 2.88 3.34 1.51 3.63 3.45 1.89 2.30 2.98 3.04 2.70 Create a binned frequency distribution for the data.
@ action
Step 1: Identify the max and min values in your bins. Looking at the dataset, you can see that the lowest value is 1.20, and the highest is 4.00.
Step 2: Get a rough idea of bin widths. Aim for seven or eight bins, give or take a couple. For eight bins, the minimum width can be found by taking the difference between the largest and smallest data values and dividing by the number of bins:
\[\frac{\text{maximum}-\text{minimum}}{\text{\# of bins}}=\frac{4.00-1.20}{8}=0.35\text{.}\]If we use 0.35 for our widths, starting at our minimum value of 1.20, we’ll get bins with these boundaries: 1.20, 1.55, 1.90, 2.25, 2.60, 2.95, 3.30, 3.65, 4.00.
Step 3: Consider the context of the values. Because these are GPAs, there are natural breaks at 2.00 and 3.00 that are important. (People like whole numbers!) Since 0.35 is very close to \(\frac{1}{3}\), let’s use that for our bin width instead, and make sure that whole numbers fall on the boundaries. That means our first bin needs to start at 1.00 and go up to 1.33 to make sure our minimum value is included. The next bin will run from 1.34 to 1.66, and so forth.
Step 4: Create the distribution table. We start our distribution table by filling in the bins:
GPA Range Frequency GPA Range Frequency GPA Range Frequency 1.00–1.33 2.00–2.33 3.00–3.33 1.34–1.66 2.34–2.66 3.34–3.66 1.67–1.99 2.67–2.99 3.67–4.00 Notice that the last bin doesn’t follow the pattern; since our maximum data value is right on the upper boundary of that last bin, this is a case where we can bend that rule just a little to avoid creating a bin for 4.00–4.33 (which wouldn’t really make sense in the context of these GPAs anyway, since 4.00 is the maximum possible GPA).
Step 5: Complete the table with the frequencies. Finish the table by counting the number of data values that fall in each bin, and recording them in the frequency column:
GPA Range Frequency GPA Range Frequency GPA Range Frequency 1.00–1.33 1 2.00–2.33 2 3.00–3.33 6 1.34–1.66 2 2.34–2.66 4 3.34–3.66 7 1.67–1.99 5 2.67–2.99 8 3.67–4.00 7 Step 6: Check your work. Add up the frequencies to make sure all the data values are included. We started with forty-two data values, and \(1+2+5+2+4+8+6+7+7=42\).
Symbols used here
Both signs at once: x = 3 ± 2 means 5 and 1.
Equal to the precision shown, not exactly.
n × (n−1) × … × 1; the number of orderings of n things. 0! = 1.
Number of k-element subsets of n things: n!/(k!(n−k)!).
Add a_k for k = 1 up to n.
In either; in both; in A but not B.
Average of the data; average of the whole population.
Typical distance from the mean; its square.
Chance of A; chance of A given that B happened.
Probability-weighted average of X; its spread.
The bell curve with mean μ and variance σ²; (x − μ)/σ.
How to: Gathering and Organizing Data
- Distinguish among sampling techniques.
- Organize data using an appropriate method.
- Create frequency distributions.
- A postal inspector wants to check on the performance of a new mail carrier, so she chooses four streets at random among those that the carrier serves. Each household on the selected streets receives a survey.
- A hospital wants to survey past patients to see if they were satisfied with the care they received. The administrator sorts the patients into groups based on the department of the hospital where they were treated (ICU, pediatrics, or general), and selects patients at random from each of those groups.
- A quality control engineer at a factory that makes smartphones wants to figure out the proportion of devices that are faulty before they are shipped out. The phones are currently packed in boxes for shipping, each of which holds 20 devices. The engineer wants to sample 100 phones, so he selects five crates at random and tests every phone in those five crates.
- A newspaper reporter wants to write a story on public perceptions on a project that will widen a congested street. She stands on the side of the street in question and interviews the first five people she sees there.
- An executive at a streaming video service wants to know if her subscribers would support a second season of a new show. She gets a list of all the subscribers who have watched at least one episode of the show, and uses a random number generator to select a sample of 50 people from the list.
Questions people ask
Mean or median — which should I use?
Median when the data have outliers or a long tail (incomes, house prices); mean when the data are roughly symmetric and you want every value to count. Report both if they disagree — the gap is itself information.
What does a p-value actually say?
The probability of seeing data at least this extreme if the null hypothesis were true. It is not the probability that the null hypothesis is true.
Why divide by n − 1 for the sample variance?
The sample mean sits closer to the sample than the true mean does, so squared deviations from it are slightly too small on average; dividing by n − 1 instead of n corrects the bias.
QDialogButtonBox
Parts of this page are adapted from OpenStax Contemporary Mathematics (CC BY-NC-SA 4.0). Condensed and re-explained here; errors are ours.
@ action Statistics & Probability
Sampling and dataDescribing data with graphsMean, median and modeProbabilityCounting: permutations and combinationsDiscrete random variablesContinuous random variablesThe normal distributionThe central limit theoremConfidence intervalsHypothesis testingComparing two samplesChi-square testsLinear regression and correlation