maths.freeStatistics & Probability › Describing data with graphs

Describing data with graphs

Stemplots, histograms, box plots — and the quartiles behind them.

A picture of the data comes before any summary of it. Histograms show shape, box plots show the five-number summary (min, Q1, median, Q3, max), and both make outliers visible. The bar chart the solver draws for any list of numbers is the simplest of these.

Stem-and-Leaf Graphs (Stemplots), Line Graphs, and Bar Graphs

One simple graph, the stem-and-leaf graph or stemplot, comes from the field of exploratory data analysis. It is a good choice when the data sets are small. To create the plot, divide each observation of data into a stem and a leaf. The stem consists of the leading digit(s), while the leaf consists of a final significant digit. For example, 23 has stem two and leaf three. The number 432 has stem 43 and leaf two. Likewise, the number 5,432 has stem 543 and leaf two. The decimal 9.3 has stem nine and leaf three. Write the stems in a vertical line from smallest to largest. Draw a vertical line to the right of the stems. Then write the leaves in increasing order next to their corresponding stem. Make sure the leaves show a space between values, so that the exact data values may be easily determined. The frequency of data values for each stem provides information about the shape of the distribution.

Example

For Susan Dean's spring precalculus class, scores for the first exam were as follows (smallest to largest):
33, 42, 49, 49, 53, 55, 55, 61, 63, 67, 68, 68, 69, 69, 72, 73, 74, 78, 80, 83, 88, 88, 88, 90, 92, 94, 94, 94, 94, 96, 100

StemLeaf
33
42 9 9
53 5 5
61 3 7 8 8 9 9
72 3 4 8
80 3 8 8 8
90 2 4 4 4 4 6
100

The stemplot shows that most scores fell in the 60s, 70s, 80s, and 90s. Eight out of the 31 scores or approximately 26 percent \((\frac{8}{31})\) were in the 90s or 100, a fairly high number of As.

The stemplot is a quick way to graph data and gives an exact picture of the data. You want to look for an overall pattern and any outliers. An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes, for example, writing 50 instead of 500, while others may indicate that something unusual is happening. It takes some background information to explain outliers, so we will cover them in more detail later.

Example

In a survey, 40 mothers were asked how many times per week a teenager must be reminded to do his or her chores. The results are shown in and in .

Number of Times Teenager Is RemindedFrequency
02
15
28
314
47
54

Condensed — the full section is in OpenStax Statistics.

Stem-and-Leaf Graphs (Stemplots), Line Graphs, and Bar Graphs

For each of the following data sets, create a stemplot and identify any outliers.


For the next three exercises, use the data to construct a line graph.

Histograms, Frequency Polygons, and Time Series Graphs

For most of the work you do in this book, you will use a histogram to display the data. One advantage of a histogram is that it can readily display large data sets.

A histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is more or less a number line, labeled with what the data represents, for example, distance from your home to school. The vertical axis is labeled either frequency or relative frequency (or percent frequency or probability). The graph will have the same shape with either label. The histogram (like the stemplot) can give you the shape of the data, the center, and the spread of the data. The shape of the data refers to the shape of the distribution, whether normal, approximately normal, or skewed in some direction, whereas the center is thought of as the middle of a data set, and the spread indicates how far the values are dispersed about the center. In a skewed distribution, the mean is pulled toward the tail of the distribution.

The relative frequency is equal to the frequency for an observed value of the data divided by the total number of data values in the sample. Remember, frequency is defined as the number of times an answer occurs. If

  • f = frequency,
  • n = total number of data values (or the sum of the individual frequencies), and
  • RF = relative frequency,

then \[\text{RF}=\frac{f}{n}\text{.}\]

For example, if three students in Mr. Ahab's English class of 40 students received from ninety to 100 percent, then f = 3, n = 40, and RF = \(\frac{f}{n}\) = \(\frac{3}{40}\) = 0.075. Thus, 7.5 percent of the students received 90 to 100 percent. Ninety to 100 percent is a quantitative measures.

To construct a histogram, first decide how many bars or intervals, also called classes, represent the data. Many histograms consist of five to 15 bars or classes for clarity. The width of each bar is also referred to as the bin size, which may be calculated by dividing the range of the data values by the desired number of bins (or bars). There is not a set procedure for determining the number of bars or bar width/bin size; however, consistency is key when determining which data values to place inside each interval.

IntervalFrequencyRelative Frequency
59.95–61.9555/100 = 0.05
61.95–63.9533/100 = 0.03
63.95–65.951515/100 = 0.15
65.95–67.954040/100 = 0.40
67.95–69.951717/100 = 0.17
69.95–71.951212/100 = 0.12
71.95–73.9577/100 = 0.07
73.95–75.9511/100 = 0.01

Condensed — the full section is in OpenStax Statistics.

Frequency Polygons

Frequency polygons are analogous to line graphs, and just as line graphs make continuous data visually easy to interpret, so too do frequency polygons.

To construct a frequency polygon, first examine the data and decide on the number of intervals and resulting interval size, for both the x-axis and y-axis. The x-axis will show the lower and upper bound for each interval, containing the data values, whereas the y-axis will represent the frequencies of the values. Each data point represents the frequency for each interval. For example, if an interval has three data values in it, the frequency polygon will show a 3 at the upper endpoint of that interval. After choosing the appropriate intervals, begin plotting the data points. After all the points are plotted, draw line segments to connect them.

Example

A frequency polygon was constructed from the frequency table below.

Frequency Distribution for Calculus Final Test Scores
Lower BoundUpper BoundFrequencyCumulative Frequency
49.559.555
59.569.51015
69.579.53045
79.589.54085
89.599.515100

Notice that each point represents frequency for a particular interval. These points are located halfway between the lower bound and upper bound. In fact, the horizontal axis, or x-axis, shows only these midpoint values. For the interval 49.5–59.5 the value 54.5 is represented by a point, showing the correct frequency of 5. For the interval occurring before 49.5–59.5, (as well as 39.5–49.5), the value of the midpoint, or 44.5, is represented by a point, showing a frequency of 0, since we do not have any values in that range. The same idea applies to the last interval of 99.5–109.5, which has a midpoint of 104.5 and correctly shows a point representing a frequency of 0. Looking at the graph, we say that this distribution is skewed because one side of the graph does not mirror the other side.

Frequency polygons are useful for comparing distributions. This comparison is achieved by overlaying the frequency polygons drawn for different data sets.

Example

We will construct an overlay frequency polygon comparing the scores from with the students’ final numeric grades.

Frequency Distribution for Calculus Final Test Scores
Lower BoundUpper BoundFrequencyCumulative Frequency
49.559.555
59.569.51015
69.579.53045
79.589.54085
89.599.515100
Frequency Distribution for Calculus Final Grades
Lower BoundUpper BoundFrequencyCumulative Frequency
49.559.51010
59.569.51020
69.579.53050
79.589.54595
89.599.55100

Condensed — the full section is in OpenStax Statistics.

Constructing a Time Series Graph

To construct a time series graph, we must look at both pieces of our paired data set. We start with a standard Cartesian coordinate system. The horizontal axis is used to plot the date or time increments, and the vertical axis is used to plot the values of the variable that we are measuring. By using the axes in that way, we make each point on the graph correspond to a date and a measured quantity. The points on the graph are typically connected by straight lines in the order in which they occur.

Example

Try it.

The following data show the Annual Consumer Price Index each month for 10 years. Construct a time series graph for the Annual Consumer Price Index data only.

YearJanFebMarAprMayJunJul
2003181.7183.1184.2183.8183.5183.7183.9
2004185.2186.2187.4188.0189.1189.7189.4
2005190.7191.8193.3194.6194.4194.5195.4
2006198.3198.7199.8201.5202.5202.9203.5
2007202.416203.499205.352206.686207.949208.352208.299
2008211.080211.693213.528214.823216.632218.815219.964
2009211.143212.193212.709213.240213.856215.693215.351
2010216.687216.741217.631218.009218.178217.965218.011
2011220.223221.309223.467224.906225.964225.722225.922
2012226.665227.663229.392230.085229.815229.478229.104
YearAugSepOctNovDecAnnual
2003184.6185.2185.0184.5184.3184.0
2004189.5189.9190.9191.0190.3188.9
2005196.4198.8199.2197.6196.8195.3
2006203.9202.9201.8201.5201.8201.6
2007207.917208.490208.936210.177210.036207.342
2008219.086218.783216.573212.425210.228215.303
2009215.834215.969216.177216.330215.949214.537
2010218.312218.439218.711218.803219.179218.056
2011226.545226.889226.421226.230225.672224.939
2012230.379231.407231.317230.221229.601229.594
Solution

Time series graphs are important tools in various applications of statistics. When a researcher records values of the same variable over an extended period of time, it is sometimes difficult for him or her to discern any trend or pattern. However, once the same data points are displayed graphically, some features jump out. Time series graphs make trends easy to spot.

Box Plots

Box plots, also called box-and-whisker plots or box-whisker plots, give a good graphical image of the concentration of the data. They also show how far the extreme values are from most of the data. As mentioned previously, a box plot is constructed from five values: the minimum value, the first quartile, the median, the third quartile, and the maximum value. We use these values to compare how close other data values are to them.

To construct a box plot, use a horizontal or vertical number line and a rectangular box. The smallest and largest data values label the endpoints of the axis. The first quartile marks one end of the box, and the third quartile marks the other end of the box. Approximately the middle 50 percent of the data fall inside the box. The whiskers extend from the ends of the box to the smallest and largest data values. A box plot easily shows the range of a data set, which is the difference between the largest and smallest data values (or the difference between the maximum and minimum). Unless the median, first quartile, and third quartile are the same value, the median will lie inside the box or between the first and third quartiles. The box plot gives a good, quick picture of the data.

Consider, again, this data set:

1, 1, 2, 2, 4, 6, 6.8, 7.2, 8, 8.3, 9, 10, 10, 11.5

The first quartile is two, the median is seven, and the third quartile is nine. The smallest value is one, and the largest value is 11.5. The following image shows the constructed box plot.

The two whiskers extend from the first quartile to the smallest value and from the third quartile to the largest value. The median is shown with a dashed line.

Condensed — the full section is in OpenStax Statistics.

Box Plots

Sixty-five randomly selected car salespersons were asked the number of cars they generally sell in one week. Fourteen people answered that they generally sell three cars, 19 generally sell four cars, 12 generally sell five cars, nine generally sell six cars, and 11 generally sell seven cars.

Isibonelo esisebenza: median of 3, 1, 4, 1, 5, 9, 2, 6

Median of 3, 1, 4, 1, 5, 9, 2, 6

3,\ 1,\ 4,\ 1,\ 5,\ 9,\ 2,\ 6

Isigaba

  1. 3, 1, 4, 1, 5, 9, 2, 6

    8 values.

  2. 1, 1, 2, 3, 4, 5, 6, 9

    Sort the values.

  3. \text{median} = \frac{7}{2}

    The middle value (or the mean of the two middle values).

Bonisa impendulo
\frac{7}{2}

Practice (40)

Try each one on paper first. Reveal the answer to check; verified ones can be opened in the solver for every step.

  1. For the Park City basketball team, scores for the last 30 games were as follows (smallest to largest):
    32, 32, 33, 34, 38, 40, 42, 42, 43, 44, 46, 47, 47, 48, 48, 48, 49, 50, 50, 51, 52, 52, 52, 53, 54, 56, 57, 57, 60, 61
    Construct a stemplot for the data.

    Bonisa impendulo
    StemLeaf
    32 2 3 4 8
    40 2 2 3 4 6 7 7 8 8 8 9
    50 0 1 2 2 2 3 4 6 7 7
    60 1
  2. Do the data seem to have any concentration of values?

    Bonisa impendulo

    The value 12.3 may be an outlier. Values appear to concentrate at 3 and 4 kilometers.

    StemLeaf
    11 5
    23 5 7
    32 3 3 5 8
    40 2 5 5 7 8
    55 6
    65 7
    7
    8
    9
    10
    11
    123
  3. The data below show the distances (in miles) from the homes of high school students to the school. Create a stemplot using the following data and identify any outliers.

    0.5, 0.7, 1.1, 1.2, 1.2, 1.3, 1.3, 1.5, 1.5, 1.7, 1.7, 1.8, 1.9, 2.0, 2.2, 2.5, 2.6, 2.8, 2.8, 2.8, 3.5, 3.8, 4.4, 4.8, 4.9, 5.2, 5.5, 5.7, 5.8, 8.0

    Bonisa impendulo
    StemLeaf
    05 7
    11 2 2 3 3 5 5 7 7 8 9
    20 2 5 6 8 8 8
    35 8
    44 8 9
    52 5 7 8
    6
    7
    80

    The value 8.0 may be an outlier. Values appear to concentrate at one and two miles.

  4. The table shows the number of wins and losses a sports team has had in 42 seasons. Create a side-by-side stem-and-leaf plot of these wins and losses.

    LossesWinsYearLosses WinsYear
    34481968–196941411989–1990
    34481969–197039431990–1991
    46361970–197144381991–1992
    46361971–197239431992–1993
    36461972–197325571993–1994
    47351973–197440421994–1995
    51311974–197536461995–1996
    53291975–197626561996–1997
    51311976–197732501997–1998
    41411977–197819311998–1999
    36461978–197954281999–2000
    32501979–198057252000–2001
    51311980–198149332001–2002
    40421981–198247352002–2003
    39431982–198354282003–2004
    42401983–198469132004–2005
    48341984–198556262005–2006
    32501985–198652302006–2007
    25571986–198745372007–2008
    32501987–198835472008–2009
    30521988–198929532009–2010
    Bonisa impendulo
    Atlanta Hawks Wins and Losses
    Number of WinsNumber of Losses
    319
    9 8 8 6 5 25 5 9
    8 7 6 6 5 5 4 3 1 1 1 1 030 2 2 2 2 4 4 5 6 6 6 9 9 9
    8 8 7 6 6 6 3 3 3 2 2 1 1 040 0 1 1 2 4 5 6 6 7 7 8 9
    7 7 6 3 2 0 0 0 051 1 1 2 3 4 4 6 7
    69
  5. In a survey, 40 people were asked how many times per year they had their car in the shop for repairs. The results are shown in . Construct a line graph.

    Number of Times in ShopFrequency
    07
    110
    214
    39
  6. By the end of 2011, a social media site had more than 146 million users in the United States. shows three age-groups, the number of users in each age-group, and the proportion (percentage) of users in each age-group. Construct a bar graph using this data.

    Age-GroupsNumber of Site UsersProportion (%) of Site Users
    13–2565,082,28045%
    26–4453,300,20036%
    45–6427,885,10019%
  7. The population in Park City is made up of children, working-age adults, and retirees. shows the three age-groups, the number of people in the town from each age-group, and the proportion (%) of people in each age-group. Construct a bar graph showing the proportions.

    Age-GroupsNumber of PeopleProportion of Population
    Children67,05919%
    Working-age adults152,19843%
    Retirees131,66238%
  8. The columns in contain the race or ethnicity of students in U.S. public schools for the class of 2011, percentages for the Advanced Placement (AP) examinee population for that class, and percentages for the overall student population. Create a bar graph with the student race or ethnicity (qualitative data) on the x-axis and the AP examinee population percentages on the y-axis.

    Race/EthnicityAP Examinee PopulationOverall Student Population
    1 = Asian, Asian American, or Pacific Islander10.3%5.7%
    2 = Black or African American9.0%14.7%
    3 = Hispanic or Latino17.0%17.6%
    4 = American Indian or Alaska Native0.6%1.1%
    5 = White57.1%59.2%
    6 = Not reported/other6.0%1.7%
  9. Park City is broken down into six voting districts. The table shows the percentage of the total registered voter population that lives in each district as well as the percentage of the entire population that lives in each district. Construct a bar graph that shows the registered voter population by district.

    DistrictRegistered Voter PopulationOverall City Population
    115.5%19.4%
    212.2%15.6%
    39.8%9.0%
    417.4%18.5%
    522.8%20.7%
    622.3%16.8%
  10. is a two-way table showing the types of pets owned by men and women.

    DogsCatsFishTotal
    Men4228
    Women46212
    Total88420

    Given these data, calculate the marginal distributions of pets for the people surveyed.

    Bonisa impendulo

    \[\text{Dogs = 8/20 = }\text{0.4}\]

    \[\text{Cats = 8/20 = }\text{0.4}\]

    \[\text{Fish = 4/20 = }\text{0.2}\]

    Note—The sum of all the marginal distributions must equal one. In this case, \[0.4\ +\ 0.4\ +\ 0.2\ =\ 1;\] therefore, the solution checks.

  11. is a two-way table showing the types of pets owned by men and women.

    DogsCatsFishTotal
    Men4228
    Women46212
    Total88420

    Given these data, calculate the conditional distributions for the subpopulation of men who own each pet type.

    Bonisa impendulo

    \[\text{Men who own dogs = 4/8 = }\text{0.5}\]

    \[\text{Men who own cats = 2/8 = }\text{0.25}\]

    \[\text{Men who own fish = 2/8 = }\text{0.25}\]

    Note—The sum of all the conditional distributions must equal one. In this case, \[0.5\ +\ 0.25\ +\ 0.25\ =\ 1;\] therefore, the solution checks.

  12. The miles-per-gallon ratings for 30 cars are shown below (lowest to highest):
    19, 19, 19, 20, 21, 21, 25, 25, 25, 26, 26, 28, 29, 31, 31, 32, 32, 33, 34, 35, 36, 37, 37, 38, 38, 38, 38, 41, 43, 43.

    Bonisa impendulo
    StemLeaf
    19 9 9
    20 1 1 5 5 5 6 6 8 9
    31 1 2 2 3 4 5 6 7 7 8 8 8 8
    41 3 3
  13. The height in feet of 25 trees is shown below (lowest to highest):
    25, 27, 33, 34, 34, 34, 35, 37, 37, 38, 39, 39, 39, 40, 41, 45, 46, 47, 49, 50, 50, 53, 53, 54, 54.

  14. The data are the prices of different laptops at an electronics store. Round each value to the nearest 10.
    249, 249, 260, 265, 265, 280, 299, 299, 309, 319, 325, 326, 350, 350, 350, 365, 369, 389, 409, 459, 489, 559, 569, 570, 610

    Bonisa impendulo
    StemLeaf
    25 5 6 7 7 8
    30 0 1 2 3 3 5 5 5 7 7 9
    41 6 9
    56 7 7
    61
  15. The following data are daily high temperatures in a town for one month:
    61, 61, 62, 64, 66, 67, 67, 67, 68, 69, 70, 70, 70, 71, 71, 72, 74, 74, 74, 75, 75, 75, 76, 76, 77, 78, 78, 79, 79, 95.

  16. In a survey, 40 people were asked how many times they visited a store before making a major purchase. The results are shown in .

    Number of Times in StoreFrequency
    14
    210
    316
    46
    54
  17. In a survey, several people were asked how many years it has been since they purchased a mattress. The results are shown in .

    Years Since Last PurchaseFrequency
    02
    18
    213
    322
    416
    59
  18. Several children were asked how many TV shows they watch each day. The results of the survey are shown in .

    Number of TV ShowsFrequency
    012
    118
    236
    37
    42
  19. The students in Ms. Ramirez’s math class have birthdays in each of the four seasons. shows the four seasons, the number of students who have birthdays in each season, and the percentage of students in each group. Construct a bar graph showing the number of students.

    SeasonsNumber of StudentsProportion of Population
    Spring824%
    Summer926%
    Autumn1132%
    Winter618%
  20. Using the data from Mrs. Ramirez’s math class supplied in , construct a bar graph showing the percentages.

  21. David County has six high schools. Each school sent students to participate in a county-wide science competition. shows the percentage breakdown of competitors from each school and the percentage of the entire student population of the county that goes to each school. Construct a bar graph that shows the population percentage of competitors from each school.

    High SchoolScience Competition PopulationOverall Student Population
    Alabaster28.9%8.6%
    Concordia7.6%23.2%
    Genoa12.1%15.0%
    Mocksville18.5%14.3%
    Tynneson24.2%10.1%
    West End8.7%28.8%
  22. Use the data from the David County science competition supplied in . Construct a bar graph that shows the county-wide population percentage of students at each school.

  23. Student grades on a chemistry exam were 77, 78, 76, 81, 86, 51, 79, 82, 84, and 99.

    1. Construct a stem-and-leaf plot of the data.
    2. Are there any potential outliers? If so, which scores are they? Why do you consider them outliers?
  24. contains the 2010 rates for a specific disease in U.S. states and Washington, DC.

    StatePercent (%)StatePercent (%)StatePercent (%)
    Alabama32.2Kentucky31.3North Dakota27.2
    Alaska24.5Louisiana31.0Ohio29.2
    Arizona24.3Maine26.8Oklahoma30.4
    Arkansas30.1Maryland27.1Oregon26.8
    California24.0Massachusetts23.0Pennsylvania28.6
    Colorado21.0Michigan30.9Rhode Island25.5
    Connecticut22.5Minnesota24.8South Carolina31.5
    Delaware28.0Mississippi34.0South Dakota27.3
    Washington, DC22.2Missouri30.5Tennessee30.8
    Florida26.6Montana23.0Texas31.0
    Georgia29.6Nebraska26.9Utah22.5
    Hawaii22.7Nevada22.4Vermont23.2
    Idaho26.5New Hampshire25.0Virginia26.0
    Illinois28.2New Jersey23.8Washington25.5
    Indiana29.6New Mexico25.1West Virginia32.5
    Iowa28.4New York23.9Wisconsin26.3
    Kansas29.4North Carolina27.8Wyoming25.1
    1. Use a random number generator to randomly pick eight states. Construct a bar graph of the rates of a specific disease of those eight states.
    2. Construct a bar graph for all the states beginning with the letter A.
    3. Construct a bar graph for all the states beginning with the letter M.
    Bonisa impendulo
    1. Example solution for using the random number generator for the TI-84+ to generate a simple random sample of eight states. Instructions are as follows.
      • Number the entries in the table 1–51 (includes Washington, DC; numbered vertically)
      • Press MATH
      • Arrow over to PRB
      • Press 5:randInt(
      • Enter 51,1,8)

      Eight numbers are generated (use the right arrow key to scroll through the numbers). The numbers correspond to the numbered states (for this example: {47 21 9 23 51 13 25 4}. If any numbers are repeated, generate a different number by using 5:randInt(51,1)). Here, the states (and Washington DC) are {Arkansas, Washington DC, Idaho, Maryland, Michigan, Mississippi, Virginia, Wyoming}.

      Corresponding percents are {30.1, 22.2, 26.5, 27.1, 30.9, 34.0, 26.0, 25.1}.


  25. The following data are the shoe sizes of 50 male students. The sizes are continuous data since shoe size is measured. Construct a histogram and calculate the width of each bar or class interval. Use six bars on the histogram.
    9, 9, 9.5, 9.5, 10, 10, 10, 10, 10, 10, 10.5, 10.5, 10.5, 10.5, 10.5, 10.5, 10.5, 10.5,
    11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11, 11.5, 11.5, 11.5, 11.5, 11.5, 11.5, 11.5,
    12, 12, 12, 12, 12, 12, 12, 12.5, 12.5, 12.5, 12.5, 14

    Bonisa impendulo

    Smallest value: 9

    Largest value: 14

    Convenient starting value: 9 – 0.05 = 8.95

    Convenient ending value: 14 + 0.05 = 14.05

    \(\frac{14.05-8.95}{6}=0.85\)

    The calculations suggests using 0.85 as the width of each bar or class interval. You can also use an interval with a width equal to one.

  26. Calculate the width of each bar/bin size/interval size.

    Bonisa impendulo

    The smallest data value is 1, and the largest data value is 6. To make sure each is included in an interval, we can use 0.5 as the smallest value and 6.5 as the largest value by subtracting and adding 0.5 to these values. We have a small range here of 6 (6.5 –– 0.5), so we will want a fewer number of bins; let’'s say six this time. So, six divided by six bins gives a bin size (or interval size) of one.

  27. The following data are the number of sports played by 50 student athletes. The number of sports is discrete data since sports are counted.

    1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1,
    2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2, 2,
    3, 3, 3, 3, 3, 3, 3, 3
    Twenty student athletes play one sport. Twenty-two student athletes play two sports. Eight student athletes play three sports. Calculate a desired bin size for the data. Create a histogram and clearly label the endpoints of the intervals.

    Bonisa impendulo

    1.5
    1.5 to 2.5
    2.5 to 3.5

  28. Using this data set, construct a histogram.

    Number of Hours My Classmates Spent Playing Video Games on Weekends
    9.95102.2516.750
    19.522.57.51512.75
    5.5111020.7517.5
    2321.92423.7518
    201522.918.820.5
    Bonisa impendulo

    Some values in this data set fall on boundaries for the class intervals. A value is counted in a class interval if it falls on the left boundary but not if it falls on the right boundary. Different researchers may set up histograms for the same data in different ways. There is more than one correct way to set up a histogram.

  29. The following data represent the number of employees at various restaurants in New York City. Using this data, create a histogram.

    22, 35, 15, 26, 40, 28, 18, 20, 25, 34, 39, 42, 24, 22, 19, 27, 22, 34, 40, 20, 38, 28

  30. Construct a frequency polygon of U.S. presidents’ ages at inauguration shown in .

    Age at InaugurationFrequency
    41.5–46.54
    46.5–51.511
    51.5–56.514
    56.5–61.59
    61.5–66.54
    66.5–71.52
    Bonisa impendulo

    The first label on the x-axis is 39. This represents an interval extending from 36.5 to 41.5. Since there are no ages less than 41.5, this interval is used only to allow the graph to touch the x-axis. The point labeled 44 represents the next interval, or the first real interval from the table, and contains four scores. This reasoning is followed for each of the remaining intervals with the point 74 representing the interval from 71.5 to 76.5. Again, this interval contains no data and is used only so that the graph will touch the x-axis. Looking at the graph, we say that this distribution is skewed because one side of the graph does not mirror the other side.

  31. The following data show the Annual Consumer Price Index each month for 10 years. Construct a time series graph for the Annual Consumer Price Index data only.

    YearJanFebMarAprMayJunJul
    2003181.7183.1184.2183.8183.5183.7183.9
    2004185.2186.2187.4188.0189.1189.7189.4
    2005190.7191.8193.3194.6194.4194.5195.4
    2006198.3198.7199.8201.5202.5202.9203.5
    2007202.416203.499205.352206.686207.949208.352208.299
    2008211.080211.693213.528214.823216.632218.815219.964
    2009211.143212.193212.709213.240213.856215.693215.351
    2010216.687216.741217.631218.009218.178217.965218.011
    2011220.223221.309223.467224.906225.964225.722225.922
    2012226.665227.663229.392230.085229.815229.478229.104
    YearAugSepOctNovDecAnnual
    2003184.6185.2185.0184.5184.3184.0
    2004189.5189.9190.9191.0190.3188.9
    2005196.4198.8199.2197.6196.8195.3
    2006203.9202.9201.8201.5201.8201.6
    2007207.917208.490208.936210.177210.036207.342
    2008219.086218.783216.573212.425210.228215.303
    2009215.834215.969216.177216.330215.949214.537
    2010218.312218.439218.711218.803219.179218.056
    2011226.545226.889226.421226.230225.672224.939
    2012230.379231.407231.317230.221229.601229.594
  32. The following table is a portion of a data set from a banking website. Use the table to construct a time series graph for CO2 emissions for the United States.

    CO2 Emissions
    UkraineUnited KingdomUnited States
    2003352,259540,6405,681,664
    2004343,121540,4095,790,761
    2005339,029541,9905,826,394
    2006327,797542,0455,737,615
    2007328,357528,6315,828,697
    2008323,657522,2475,656,839
    2009272,176474,5795,299,563
  33. 65 randomly selected car salespersons were asked the number of cars they generally sell in one week. 14 people answered that they generally sell three cars, 19 generally sell four cars, 12 generally sell five cars, nine generally sell six cars, and 11 generally sell seven cars. Complete the table.

    Data Value (Number of Cars)FrequencyRelative FrequencyCumulative Relative Frequency
  34. What does the frequency column in sum to? Why?

    Bonisa impendulo

    65

  35. What does the relative frequency column in sum to? Why?

  36. What is the difference between relative frequency and frequency for each data value in ?

    Bonisa impendulo

    The relative frequency shows the proportion of data points that have each value. The frequency tells the number of data points that have each value.

  37. What is the difference between cumulative relative frequency and relative frequency for each data value?

  38. To construct the histogram for the data in , determine appropriate minimum and maximum x- and y-values and the scaling. Sketch the histogram. Label the horizontal and vertical axes with words. Include numerical scaling.

    Bonisa impendulo

    Answers will vary. One possible histogram is shown below.

  39. Construct a frequency polygon for the following.

    1. Pulse Rates for WomenFrequency
      60–6912
      70–7914
      80–8911
      90–991
      100–1091
      110–1190
      120–1291
    2. Actual Speed in a 30-MPH ZoneFrequency
      42–4525
      46–4914
      50–537
      54–573
      58–611
    3. Tar (mg) in Nonfiltered CigarettesFrequency
      10–131
      14–170
      18–2115
      22–257
      26–292
  40. Construct a frequency polygon from the frequency distribution for the 50 highest-ranked countries for depth of hunger.

    Depth of HungerFrequency
    230–25921
    260–28913
    290–3195
    320–3497
    350–3791
    380–4091
    410–4391
    Bonisa impendulo

    Find the midpoint for each class. These will be graphed on the x-axis. The frequency values will be graphed on the y-axis values.

Symbols used here

\pm
plus or minus
Both signs at once: x = 3 ± 2 means 5 and 1.
\approx
approximately equal
Equal to the precision shown, not exactly.
n!
factorial
n × (n−1) × … × 1; the number of orderings of n things. 0! = 1.
\binom{n}{k}
binomial coefficient, "n choose k"
Number of k-element subsets of n things: n!/(k!(n−k)!).
\sum_{k=1}^{n} a_k
summation
Add a_k for k = 1 up to n.
A \cup B,\ A \cap B,\ A \setminus B
union, intersection, difference
In either; in both; in A but not B.
\bar{x},\ \mu
sample mean, population mean
Average of the data; average of the whole population.
\sigma,\ s,\ \sigma^2
standard deviation, sample s.d., variance
Typical distance from the mean; its square.
P(A),\ P(A \mid B)
probability, conditional probability
Chance of A; chance of A given that B happened.
E[X],\ \operatorname{Var}(X)
expected value, variance
Probability-weighted average of X; its spread.
N(\mu, \sigma^2),\ z
normal distribution, z-score
The bell curve with mean μ and variance σ²; (x − μ)/σ.

How to: Describing data with graphs

  1. 8 values.
  2. Sort the values.
  3. The middle value (or the mean of the two middle values).

Questions people ask

Mean or median — which should I use?

Median when the data have outliers or a long tail (incomes, house prices); mean when the data are roughly symmetric and you want every value to count. Report both if they disagree — the gap is itself information.

What does a p-value actually say?

The probability of seeing data at least this extreme if the null hypothesis were true. It is not the probability that the null hypothesis is true.

Why divide by n − 1 for the sample variance?

The sample mean sits closer to the sample than the true mean does, so squared deviations from it are slightly too small on average; dividing by n − 1 instead of n corrects the bias.

Zama wena

Parts of this page are adapted from OpenStax Statistics (CC BY 4.0). Condensed and re-explained here; errors are ours.

Okuningi Statistics & Probability