Unit 2 · Probability, Random Variables, and Probability Distributions
● Core concept · ○ Supporting concept
2.1 Tabular and Graphical Representations for the Distributions of Two Categorical Variables
Two-way table (contingency table) ● (core concept) — A two-way table, also called a contingency table, summarizes and compares data for two categorical variables. The entries in the cells can be frequencies (counts) or relative frequencies (proportions).
Side-by-side bar chart ● (core concept) — Side-by-side bar charts display the frequency or relative frequency of each category of one categorical variable for each category of the other categorical variable, so the relationship between the two variables can be compared.
Segmented bar chart ● (core concept) — Segmented bar charts display the frequency or relative frequency of each category (level) of one categorical variable for each category of the other categorical variable, so the relationship between the two variables can be compared.
Mosaic plot ● (core concept) — A mosaic plot displays the frequency or relative frequency of each category (level) of one categorical variable for each category of the other categorical variable, so the relationship between the two variables can be compared.
Association (between categorical variables) ● (core concept) — Two categorical variables are associated if the relationship of one variable differs across the levels of the other variable. Graphical representations of two categorical variables can be used to compare the relationship of one variable across the levels of the other and determine whether the variables are associated.
2.2 Summary Statistics for Two Categorical Variables
Joint relative frequency ● (core concept) — A joint relative frequency in a two-way table is a cell frequency divided by the total for the entire table.
Marginal relative frequency ● (core concept) — A marginal relative frequency in a two-way table is a row total divided by the total for the entire table, or a column total divided by the total for the entire table.
Conditional relative frequency ● (core concept) — A conditional relative frequency is a relative frequency computed by restricting to a particular level (category) of interest: a cell frequency in a row divided by the total for that row, or a cell frequency in a column divided by the total for that column.
2.3 Estimating Probabilities Using Simulation
Random process ● (core concept) — A random process generates results that are determined by chance.
Trial ● (core concept) — One trial is the single performance of a random process; an outcome is the result of one trial of a random process.
Outcome ● (core concept) — An outcome is the result of one trial of a random process.
Event ● (core concept) — An event is a collection of outcomes.
Simulation ● (core concept) — Simulation is a way to model random events such that simulated outcomes closely match real-world outcomes. All possible outcomes are associated with a value to be determined by chance; record the counts of simulated outcomes and the count total.
Law of large numbers ● (core concept) — The law of large numbers states that for independent trials, as the number of trials increases, the long-run relative frequency of the outcome or event gets closer and closer to a single value.
Empirical probability ○ — An empirical probability is the relative frequency of an outcome or event determined from data, used to estimate the actual (true) probability of that outcome or event. The probability of an outcome or event is its long-run relative frequency — its relative frequency over a large number of trials.
2.4 Introduction to Probability
Sample space ● (core concept) — The sample space of a random process is the set of all possible nonoverlapping outcomes; the probability of the sample space is 1.
Theoretical probability ● (core concept) — If all outcomes in the sample space are equally likely, the theoretical probability that an event E will occur is the number of outcomes in E divided by the total number of outcomes in the sample space. The probability of an event is a number between 0 and 1, inclusive.
Complement of an event ● (core concept) — The complement of an event E (written E′, Ē, or Eᶜ) is the event that E does not occur ('not E'). Its probability is 1 − P(E).
2.5 Mutually Exclusive Events
Joint probability ● (core concept) — The joint probability of events A and B is the probability that both will occur — the probability of the intersection of A and B, written P(A ∩ B).
Mutually exclusive (disjoint) events ● (core concept) — Two events are mutually exclusive, or disjoint, if they cannot occur at the same time. If two events are mutually exclusive, then P(A ∩ B) = 0.
2.6 Conditional Probability
Conditional probability ● (core concept) — The conditional probability P(A | B) is the probability that event A will occur given that event B has occurred. It is calculated as P(A ∩ B) / P(B).
General multiplication rule ● (core concept) — The general multiplication rule states that the probability that events A and B both will occur equals the probability that event A will occur multiplied by the conditional probability that event B will occur given that event A has occurred: P(A ∩ B) = P(A) · P(B | A).
2.7 Independent Events and Unions of Events
Independent events ● (core concept) — Events A and B are independent if and only if knowing whether event A has occurred (or will occur) does not change the probability that event B will occur. When A and B are independent, P(A | B) = P(A), P(B | A) = P(B), and P(A ∩ B) = P(A) · P(B).
Union of events ● (core concept) — The probability that event A or event B (or both) will occur is the probability of the union of A and B, written P(A ∪ B).
General addition rule ● (core concept) — The general addition rule for the union of two events is P(A ∪ B) = P(A) + P(B) − P(A ∩ B).
2.8 Introduction to Random Variables and Probability Distributions
Random variable ● (core concept) — A random variable is a variable whose values have numerical outcomes that result from a random phenomenon.
Probability distribution ● (core concept) — A probability distribution for a discrete random variable shows the probability associated with every possible value of the random variable, and the sum of the probabilities over all possible values is 1. It can be represented as a graph, table, or function, and can be determined using the rules of probability or estimated with a simulation.
Cumulative probability distribution ● (core concept) — A cumulative probability distribution, represented as a table or function, shows the probability of being less than or equal to each value of the discrete random variable.
2.9 Parameters of Random Variables
Parameter (of a probability distribution) ● (core concept) — A numerical value measuring a characteristic of a probability distribution of a random variable, or of a population, is a parameter; the value of a parameter is a single, fixed value.
Expected value (mean of a random variable) ● (core concept) — The expected value (mean) of a probability distribution, denoted E(X) or µ_X, is a parameter. For a discrete random variable X it is calculated as µ_X = Σ x_i · P(x_i) and interpreted as the long-run average outcome of the random variable.
Standard deviation of a random variable ● (core concept) — The standard deviation of a probability distribution, denoted SD(X) or σ_X, is a parameter. For a discrete random variable X it is calculated as σ_X = √Σ (x_i − µ_X)² · P(x_i) and interpreted as the typical deviation of the values of the random variable from the mean over the long run.
Variance of a random variable ● (core concept) — The square of the standard deviation of a random variable is called the variance of the random variable, denoted V(X) or σ_X².
Discrete random variable ● (core concept) — A discrete random variable can take on only values that are countable or finite.
2.10 The Binomial Distribution
Binomial random variable ● (core concept) — A binomial random variable X is a discrete random variable that counts the number of successes in repeated independent trials, n, that have only two possible outcomes (success or failure), with the probability of success p and the probability of failure 1 − p.
Mean of a binomial distribution ● (core concept) — If a random variable is binomial, its mean µ_X is np.
Standard deviation of a binomial distribution ● (core concept) — If a random variable is binomial, its standard deviation σ_X is √(np(1 − p)).
Binomial probability function ● (core concept) — The binomial probability function gives the probability that a binomial random variable X has exactly x successes in n independent trials when the probability of success is p: P(X = x) = C(n, x) · p^x · (1 − p)^(n−x), for x = 0, 1, 2, …, n.
2.11 The Normal Distribution
Continuous random variable ● (core concept) — A continuous random variable is a variable that can take on any value within a specified domain; every interval within the domain has a probability associated with it.
Normal distribution ● (core concept) — A normal distribution is a continuous, unimodal, bell-shaped, and symmetric curve. It can model a distribution of data or a continuous random variable, and it is identified by two parameters: the mean µ and the standard deviation σ. A smaller standard deviation makes the normal curve taller and more concentrated around its mean; a larger one makes it shorter and less concentrated.
Standard normal distribution ● (core concept) — A standard normal distribution is a normal distribution with mean μ = 0 and standard deviation σ = 1.
Empirical rule (68–95–99.7 rule) ● (core concept) — The empirical rule estimates the area under a normal distribution curve: approximately 68% of observations are within 1 standard deviation of the mean, approximately 95% within 2 standard deviations, and approximately 99.7% within 3 standard deviations. It is also called the 68–95–99.7 rule.
Normal probability as area under the curve ● (core concept) — If the distribution of a random variable is approximately normal, the probability that the random variable takes on values within a particular interval is determined by the area under the normal curve within that interval; the total area under the normal curve is 1.
Determining interval boundaries from a given area (inverse normal) ● (core concept) — The boundaries of an interval associated with a given area in a normal distribution can be determined using technology or using z-scores and a standard normal table.
2.12 Sampling Distributions and the Central Limit Theorem
Sampling distribution ● (core concept) — A sampling distribution of a statistic is the distribution of values of the statistic for all possible samples of a given size from a given population.
Randomization distribution ● (core concept) — A randomization distribution is the distribution of a statistic generated by simulation from repeatedly randomly reallocating (reassigning) the response values to treatment groups; the value of the statistic is determined and recorded for each reallocation, and the resulting distribution approximates the sampling distribution of the statistic.
Central limit theorem (CLT) ● (core concept) — The central limit theorem (CLT) states that the sampling distribution of a mean of a random sample has a shape that can be approximated by a normal distribution; the larger the sample is, the better the approximation will be.
Simulating a sampling distribution ● (core concept) — The sampling distribution of a statistic can be simulated by repeatedly generating a large number of random samples from the population assuming known value(s) for the parameter(s); the value of the statistic is determined and recorded for each sample, and the resulting distribution of the sample statistic values approximates the sampling distribution of the statistic.