Rycal.
← All AP courses
AP Statistics · Cram sheet

Unit 1 · Exploring One-Variable Data and Collecting Data

114 key terms

● Core concept  ·  ○ Supporting concept

1.1 Introducing Statistics: What Can We Learn from Data?

Statistical study ● (core concept) — A statistical study is a study in which data are collected from a sample to answer an investigative question about a larger population. Statistical studies are necessary when the population is too large or it is too difficult to collect data from every item or individual.

Investigative question ● (core concept) — An investigative question is the question a statistical study sets out to answer; it should have a defined purpose, be posed so that the required data can be collected and analyzed, and should not be changed based on the data analysis or results.

Datum ● (core concept) — A datum (singular form of data) is a piece of information about an item or individual.

Data set ● (core concept) — A data set is a collection of data.

Population ● (core concept) — A population consists of all items or individuals of interest; the population size is represented by the symbol N.

Sample ● (core concept) — A sample selected for study is a subset of the population from which data are obtained; the number of items in the sample, called the sample size, is represented by the symbol n.

In context ● (core concept) — Reporting 'in context' means identifying each component of a statistical study and the resulting calculations with the corresponding real-world aspect from which they were derived.

1.2 Variables

Observational unit ● (core concept) — An observational unit is an item or individual from which a datum is collected.

Variable ● (core concept) — A variable is a characteristic that may change from one observational unit to another.

Parameter ● (core concept) — A parameter is a numerical attribute or summary of the variable of interest for a population.

Statistic ● (core concept) — A statistic is a numerical attribute or summary of the variable of interest for a sample. The value of a statistic from a certain sample is often not equal to the unknown value of the population parameter but may provide the basis for making inferences about the population parameter.

Categorical variable ● (core concept) — A categorical variable, also called a qualitative variable, takes on values that are category names or group labels.

Quantitative variable ● (core concept) — A quantitative variable, also called a numerical variable, takes on numerical values for a measured or counted quantity and generally has units of measure.

Discrete quantitative variable ● (core concept) — A discrete quantitative variable can take on a countable number of values; the number of values may be finite or countably infinite, as with the whole numbers.

Continuous quantitative variable ● (core concept) — A continuous quantitative variable can take on an infinite number of possible values within a given interval; it can take on all possible values between any pair of values.

1.3 Tabular Representation and Summary Statistics for One Categorical Variable

Frequency table ● (core concept) — A frequency table shows the number of observational units in each category of a categorical variable.

Relative frequency table ● (core concept) — A relative frequency table shows the proportion of observational units in each category of a categorical variable.

Relative frequency ● (core concept) — A relative frequency is a proportion (equivalently expressible as a percentage or ratio) of observational units in a category. Counts and relative frequencies of categorical variables reveal information that can be used to justify claims about the variables in context.

1.4 Graphical Representations for One Categorical Variable

Bar chart ● (core concept) — A bar chart, also called a bar graph, displays frequencies (counts) or relative frequencies (proportions) for the categories of a single categorical variable. The height or length of each bar corresponds to the frequency or relative frequency of the observational units in each category.

Pie chart ● (core concept) — A pie chart displays frequencies (counts) or relative frequencies (proportions) for categorical data as slices of a circle. The area of each slice, as a fraction of the total area, corresponds to the relative frequency of observational units in that category, and the slices sum to 1 (100%) of the total area.

Comparing multiple categorical data sets ● (core concept) — Frequency and relative frequency tables, bar charts, and pie charts can be used to compare two or more data sets in terms of the same categorical variable.

1.5 Graphical Representations for One Quantitative Variable

Histogram ● (core concept) — A histogram places the observed values of the quantitative variable into ordered intervals, or bins, along the horizontal axis. Each bar represents an interval or bin, and the height of each bar shows the frequency or relative frequency of the observations within that interval; altering the bin widths can change the appearance of the histogram.

Stem-and-leaf plot ● (core concept) — A stem-and-leaf plot splits each value of the quantitative variable into two parts: a stem (the first digit or digits) and a leaf (usually the single digit after the stem digit or digits). Both stems and leaves are ordered from smallest to largest.

Dotplot ● (core concept) — A dotplot represents each value of the quantitative variable by a dot placed above the horizontal axis (or beside the vertical axis) corresponding to the value of that observation, with nearly identical values stacked on top of each other.

Distribution ○ — The distribution of a variable is the pattern of its values — which values occur and how often — as shown by a histogram, stem-and-leaf plot, dotplot, or other representation.

1.6 Descriptions for One Quantitative Variable Distributions

Shape (of a distribution) ● (core concept) — The shape of a distribution describes its overall pattern, including whether it is skewed or symmetric and how many peaks it has.

Center (of a distribution) ● (core concept) — The center of a distribution is a typical or middle value of the data, commonly measured by the mean or median.

Variability (spread) ● (core concept) — Variability, also called spread, describes how spread out the values of a distribution are.

Skewed right (positively skewed) ● (core concept) — A distribution is skewed to the right (positively skewed) if the right tail, toward larger values, is longer than the left tail.

Skewed left (negatively skewed) ● (core concept) — A distribution is skewed to the left (negatively skewed) if the left tail, toward smaller values, is longer than the right tail.

Approximately symmetric ● (core concept) — A distribution is approximately symmetric if the left half is approximately the mirror image of the right half.

Unimodal ● (core concept) — A distribution with one main peak is called unimodal.

Bimodal ● (core concept) — A distribution with two prominent peaks is called bimodal.

Approximately uniform ● (core concept) — A distribution in which each frequency or each relative frequency is approximately the same, with no prominent peaks, is approximately uniform.

Outlier ● (core concept) — An outlier is a data point that is unusually small or large relative to the rest of the data.

Gap ● (core concept) — A gap is a region in a distribution between two values in which there are no observed data.

Cluster (of values) ● (core concept) — Clusters are concentrations of values usually separated by gaps.

1.7 Summary Statistics for One Quantitative Variable

Mean ● (core concept) — The mean is the sum of all the values divided by the number of values. For a sample, the mean is denoted by x-bar: x̄ = (1/n)Σx_i, where x_i is the ith data point and n is the number of data values in the sample.

Median ● (core concept) — The median is the middle value when the data set is ordered from smallest to largest. One common method for an even number of values is to use the mean of the two middle values; for an odd number of values, a common method is to use the value in the middle.

Minimum value ● (core concept) — In an ordered data set, the minimum value is the smallest value.

Maximum value ● (core concept) — In an ordered data set, the maximum value is the largest value.

First quartile (Q1) ● (core concept) — The first quartile, denoted Q1, is the median of the lower half of the ordered data set (from the minimum value to the position of the median); approximately 25% of the values are less than or equal to Q1.

Third quartile (Q3) ● (core concept) — The third quartile, denoted Q3, is the median of the upper half of the ordered data set (from the position of the median to the maximum value); approximately 75% of the values are less than or equal to Q3.

Second quartile (Q2) ● (core concept) — The second quartile, Q2, is the same as the median of the data set. Q1 and Q3 form the boundaries for the middle 50% of values in an ordered data set.

Percentile ● (core concept) — The pth percentile is the value that has p% of the data less than or equal to it when the data set is ordered from smallest to largest; Q1 and Q3 are the 25th and 75th percentiles.

Range ● (core concept) — The range is the difference between the maximum data value and the minimum data value.

Interquartile range (IQR) ● (core concept) — The interquartile range (IQR) is the difference between the third and first quartiles: Q3 − Q1. It measures the variability of the middle 50% of the data.

Standard deviation ● (core concept) — The standard deviation is a typical deviation of the data values from their mean. The sample standard deviation, denoted s, is calculated as the square root of Σ(x_i − x̄)²/(n − 1).

Sample variance ● (core concept) — The sample variance is the square of the sample standard deviation, denoted s².

Resistant (robust) ● (core concept) — A summary statistic is resistant (robust) if outliers do not greatly (if at all) affect its value. The median and IQR are resistant measures of center and variability.

Nonresistant (non-robust) ● (core concept) — A summary statistic is nonresistant (non-robust) if outliers can greatly affect its value. The mean is a nonresistant measure of center, and the range and standard deviation are nonresistant measures of variability.

1.5 × IQR rule for outliers ● (core concept) — Under the 1.5 × IQR rule, an outlier is a value located more than 1.5 × IQR above the third quartile or more than 1.5 × IQR below the first quartile.

2-standard-deviation rule for outliers ● (core concept) — Under the 2-standard-deviation rule, an outlier is a value located more than 2 standard deviations above, or below, the mean.

Changing units of measurement ● (core concept) — Changing units of measurement affects the values of the calculated statistics.

1.8 Graphical Representations of Summary Statistics for One Quantitative Variable

Five-number summary ● (core concept) — The five-number summary is made up of the minimum data value, the first quartile (Q1), the median, the third quartile (Q3), and the maximum data value.

Boxplot ● (core concept) — A boxplot is a graphical representation of the five-number summary: the box represents the middle 50% of data, with a line at the median and the ends of the box at the quartiles, and whiskers extending to the minimum and maximum.

Whiskers ● (core concept) — Whiskers are the lines in a boxplot extending from the first quartile to the minimum and from the third quartile to the maximum. If there are outliers, the whiskers extend only to the most extreme values that are not outliers, and the outliers are usually denoted with an asterisk or other symbol.

Mean–median relationship to shape ● (core concept) — If a distribution is relatively symmetric, the mean and median are relatively close to each other; if it is skewed right, the mean is usually larger than the median; if it is skewed left, the mean is usually smaller than the median.

1.9 Comparisons of the Distributions for One Quantitative Variable

Back-to-back stem-and-leaf plot ● (core concept) — A back-to-back stem-and-leaf plot places the leaves of two distributions on either side of a shared stem, so center, variability, shape, outliers, clusters, or gaps can be compared directly.

Standardized score ● (core concept) — A standardized score measures the number of standard deviations a data value falls above or below the mean.

z-score ● (core concept) — A z-score is calculated as (x_i − μ)/σ, where x_i is the data value, μ is the population mean, and σ is the population standard deviation. A positive z-score means the value is that many standard deviations above the mean; a negative z-score, below. When μ and σ are unknown, the sample mean and standard deviation may be used. z-scores may be used to compare relative positions of values within a distribution or between distributions.

1.10 The Investigative Question Revisited and Data Collection

Components of an investigative question ● (core concept) — An investigative question has three components: the first guides the data collection process (phrased in terms of the variable(s) of interest), the second guides the data analysis choice, and the third indicates the type(s) of conclusion applicable from the study, including the population to which conclusions apply and, for an experiment with random assignment, a cause-and-effect conclusion.

Investigative question for a hypothesis test ● (core concept) — A hypothesis test is a data analysis in which the investigative question should make clear the parameter and the direction of the alternative hypothesis (not equal, greater than, less than, association, or not independent).

Investigative question for a confidence interval ● (core concept) — A confidence interval is a data analysis in which the investigative question should make clear the parameter and the goal of estimating that parameter within a range of potential values.

Cause-and-effect conclusion ● (core concept) — A cause-and-effect conclusion attributes differences in the response variable to the explanatory variable (treatments); it is applicable from a well-designed experiment that uses random assignment of treatments to experimental units.

Census ● (core concept) — A census consists of recording information from all items or individuals in a population.

Experiment ● (core concept) — An experiment is a study in which a researcher assigns conditions, or treatments, to experimental units to explore an investigative question of interest about the population.

Experimental unit ● (core concept) — The experimental unit is the observational unit to which the treatment is assigned.

Subject (participant) ● (core concept) — When experimental units consist of people, they are sometimes referred to as subjects or participants.

Explanatory variable (factor) ● (core concept) — An explanatory variable, or factor, is a variable whose different categories, or levels, are imposed on the experimental units.

Treatment (level) ● (core concept) — Treatments are the different categories, or levels, of the explanatory variable imposed on experimental units; when there is more than one explanatory variable, the combinations of the levels of the explanatory variables are called treatments.

Response variable (experiment) ● (core concept) — A response variable is an outcome measured on each experimental unit after the treatment has been administered.

Observational study ● (core concept) — An observational study is a study where treatments are not imposed; the researcher records the values of the variables of interest in order to explore an investigative question of interest.

Prospective study ● (core concept) — A prospective study selects the observational units of study at a point in time, and data are gathered both at that time and into the future.

Retrospective study ● (core concept) — A retrospective study selects the observational units of study at a point in time and gathers data from the past.

Survey ● (core concept) — A survey is an observational study in which the data are collected from humans using a standard set of questions.

Confounding variable (observational study) ● (core concept) — In an observational study, a confounding variable provides an alternative explanation for the observed relationship between the explanatory and response variables, reducing the possibility of concluding a causal relationship. To be a confounding variable, it must be associated with both the explanatory variable and the response variable.

Random sample ● (core concept) — A sample is considered random when all observational units in the sample are selected from the population using some type of random mechanism, such as a random number generator. A sample is not randomly selected when units are deliberately chosen or volunteer themselves.

Scope of generalization ● (core concept) — When observational or experimental units are randomly selected from a population, it is appropriate to generalize about the entire population from which the sample was selected. When they are not randomly selected, generalizations are appropriate only about a population of individuals similar to those used in the study.

1.11 Random Sampling

Sampling without replacement ● (core concept) — In sampling without replacement, an observational unit from a population can be selected only once; it is not returned to the population before subsequent selections, so it cannot be selected again.

Sampling with replacement ● (core concept) — In sampling with replacement, an observational unit from the population can be selected more than once; it is returned to the population before subsequent selections, so it could be selected again.

Simple random sample (SRS) ● (core concept) — In a simple random sample of size n, every sample of size n has the same chance of being selected. It can be obtained using a random number generator or by randomly selecting numbered slips of paper.

Stratified random sample ● (core concept) — A stratified random sample divides all individuals in a population into non-overlapping groups, called strata, based on shared attributes (a homogeneous grouping); a simple random sample is selected within each stratum, and the selected individuals are combined to form one sample.

Stratum (strata) ● (core concept) — A stratum is one of the non-overlapping, homogeneous groups formed by dividing a population by shared attributes in stratified random sampling.

Cluster random sample ● (core concept) — A cluster random sample divides the population into smaller groups, called clusters; a simple random sample of clusters is selected, and data are collected from all observational units in each of the selected clusters.

Cluster (sampling group) ● (core concept) — A cluster is a smaller group into which the population is divided for cluster sampling. Ideally, each cluster mirrors the heterogeneity of the population, with clusters similar to one another.

Systematic random sample ● (core concept) — A systematic random sample selects sample members from a population according to a random starting point and a fixed, periodic interval between successive sampling units.

Random number generator ○ — A random number generator is a tool (such as a computer or calculator function) used to select observational units by random chance, for example to obtain a simple random sample.

Justifying an appropriate sampling method ● (core concept) — Each random sampling method has different characteristics that make it more appropriate for sampling populations depending on the question being investigated.

1.12 Potential Problems with Sampling

Bias (in a sampling method) ● (core concept) — Bias in a sampling method is a systematic error in the sampling procedure that results in a statistic being consistently larger or consistently smaller than the parameter the statistic is used to estimate.

Voluntary response bias ● (core concept) — Voluntary response bias is a bias that may occur when a sample consists entirely of volunteers.

Undercoverage bias ● (core concept) — Undercoverage bias may occur when the sampling method fails to include part of the population, or a part of the population is less likely to be selected.

Nonresponse bias ● (core concept) — Nonresponse bias may occur because of a failure to obtain responses from some individuals chosen to be sampled; the respondents and nonrespondents could differ significantly in ways that are important for the study.

Response bias ● (core concept) — Response bias may occur when responses to a survey or measurements of observational units tend to differ from the 'true' value in one direction.

Question wording bias ● (core concept) — Question wording bias is response bias caused by confusing or leading survey questions.

Nonrandom sampling methods ● (core concept) — Nonrandom sampling methods — for example, samples chosen by convenience or voluntary response — introduce potential bias because they do not use random chance to select the individuals.

1.13 Experimental Design

Well-designed experiment ● (core concept) — A well-designed experiment includes: comparisons of at least two treatment groups (one of which could be a control group), random assignment of treatments to experimental units, replication, and direct control of potential extraneous sources of variation in the response.

Control group ● (core concept) — A control group is a collection of experimental units created for comparison. It may be given a treatment different from the treatment of interest to determine if the treatment of interest has an effect — for example, a treatment with an inactive substance (a placebo).

Placebo ● (core concept) — A placebo is an inactive substance given as a treatment, used in a control group for comparison with the treatment of interest.

Placebo effect ● (core concept) — The placebo effect is the difference between the average response to a placebo and the average response to no treatment.

Single-blind experiment ● (core concept) — In a single-blind, also called single-masked, experiment, participants do not know which treatment they are receiving but the research team members who interact with them do — or vice versa.

Double-blind experiment ● (core concept) — In a double-blind, also called double-masked, experiment, neither the participants nor the research team members who interact with them know which treatment each participant is receiving.

Extraneous variable ● (core concept) — An extraneous variable, also called an extraneous source of variation, is a variable known or believed to affect the response that is not an explanatory variable being studied.

Random assignment ● (core concept) — The purpose of random assignment is to create treatment groups that are as similar as possible with respect to extraneous sources of variation; if successful, the respective distributions of each extraneous variable will be approximately the same for all treatment groups.

Confounding variable (experiment) ● (core concept) — A confounding variable in an experiment is related to the explanatory variable in such a way that it is difficult to determine which variable — explanatory or confounding — is influencing the change in the response variable. In a well-designed experiment, the potential for confounding variables is reduced.

Replication ● (core concept) — Replication within an experiment means more than one experimental unit is assigned to each treatment.

Direct control ● (core concept) — Direct control in an experiment means keeping the settings of certain potential extraneous sources of variation in the response variable the same from experimental unit to experimental unit.

Completely randomized design ● (core concept) — In a completely randomized design, treatments are assigned to experimental units completely at random. Often the number of units assigned to each treatment is the same, but the sample sizes in each treatment do not have to be the same.

Blocking variable ● (core concept) — A blocking variable is a source of extraneous variation in the response variable, used to group experimental units before random assignment.

Randomized block design ● (core concept) — In a randomized block design, experimental units are first grouped according to similar values of a blocking variable; these groups are called blocks, and units within a block are homogeneous with respect to the blocking variable. After the blocks are formed, treatments are randomly assigned to experimental units within each block so that all treatments occur within every block.

Block ● (core concept) — A block is a group of experimental units formed by grouping according to similar values of a blocking variable; units within the same block are homogeneous with respect to the blocking variable.

Matched pairs design ● (core concept) — A matched pairs design is a randomized block design with only two treatments: experimental units are arranged in pairs by matching on one or more extraneous sources of variation, and each pair receives both treatments (one treatment randomly assigned to each member of the pair). Alternatively, each experimental unit may get both treatments while the order of the treatments is randomized.

Purpose of blocking ● (core concept) — The purpose of blocking is to separate the variation in the response caused by the blocking variable from the rest of the extraneous variation in the response. Blocking allows for more precise comparisons of the response across the treatments.

Justifying an appropriate experimental design ● (core concept) — One experimental design may be more appropriate than another experimental design based on the goals of the investigative study, the characteristics of the population, and the sample and variables involved.