Unit 4 · Inference for Quantitative Data: Means
● Core concept · ○ Supporting concept
4.1 Sampling Distributions for Sample Means
Sampling distribution of a sample mean ● (core concept) — The distribution of sample-mean values over all possible samples of size n. When sampled values are independent, its mean is μ(x̄) = μ and its standard deviation is σ(x̄) = σ/√n.
Randomization condition (means) ● (core concept) — The data should be collected using a random sample — or, for two-sample procedures, two independent random samples or a randomized experiment. If the data come from an experiment, only this condition is needed, with treatments randomly assigned to experimental units.
10% condition (means) ● (core concept) — When sampling without replacement, the population size N must be at least 10 times the sample size (n ≤ 10%N); for two samples, n1 ≤ 10%N1 and n2 ≤ 10%N2. This condition is unnecessary when the data are from a randomized experiment.
Normality condition (sampling distribution of a sample mean) ● (core concept) — If the population distribution can be modeled by a normal distribution, the sampling distribution of x̄ is normal regardless of sample size. Otherwise it is approximately normal provided n ≥ 30; if the population is extremely skewed, a sample size much larger than 30 may be needed.
4.2 Constructing a Confidence Interval for a Population Mean or Population Mean Difference
t-distribution (Student's t) ● (core concept) — A family of symmetric, bell-shaped distributions with wider tails than the standard normal, identified by their degrees of freedom (based on sample size). With small df they have a narrower peak and fatter tails; as df increases they more closely resemble the standard normal (mean 0, standard deviation 1). They are used for inferences about a population mean μ when σ is unknown and the sample standard deviation s must be used instead.
Degrees of freedom (t) ● (core concept) — The parameter identifying a specific t-distribution, based on sample size: df = n − 1 for one-sample procedures (for matched pairs, n is the number of differences, so df = number of differences − 1). For two-sample procedures the degrees of freedom (found with technology) fall between n1 + n2 − 2 and the smaller of n1 − 1 and n2 − 1.
One-sample t-interval for a population mean ● (core concept) — The confidence interval procedure for the mean of a quantitative variable when σ is unknown: x̄ ± t*(s/√n), where t* is the critical value for the central C% of the t-distribution with n − 1 degrees of freedom.
One-sample t-interval for a population mean difference (matched pairs) ● (core concept) — For a matched pairs design with two dependent samples, calculate the differences between pairs to produce one sample of differences, then use the one-sample t-interval on the mean difference.
Paired vs. independent samples ● (core concept) — In a matched pairs design the two samples are dependent (paired), and differences within each pair are analyzed as a single sample. Two-sample t procedures require two independent samples instead.
Sample data condition (one-sample t procedures) ● (core concept) — One of the following: the population distribution is indicated to be approximately normal, n ≥ 30, or (if n < 30) the sample data distribution is free from strong skewness and outliers. For matched pairs, the number of differences should be ≥ 30, or the differences free from strong skewness and outliers if fewer.
Standard error (sample mean) ● (core concept) — SE(x̄) = s/√n.
Margin of error (t-interval) ● (core concept) — The critical value t* times the standard error: t*(s/√n).
Critical value (t*) ● (core concept) — t* (and −t*) are the values enclosing the middle C% of the appropriate t-distribution.
Stating the parameter in context (mean / mean difference) ● (core concept) — For a confidence interval or test about a population mean or population mean difference, the parameter should reference the population mean (or mean difference), the response variable, and the population in context; for a mean difference, state the order of subtraction.
4.3 Justifying a Claim Based on a Confidence Interval for a Population Mean or Population Mean Difference
Confidence interval interpretation (population mean or mean difference) ● (core concept) — We are C% confident that the interval (a, b) contains the true population mean (or mean difference), where a and b are the lower and upper limits. Because the interval is calculated from a sample, it may or may not contain the parameter.
Relationships among confidence level, margin of error, sample size, and interval width (means) ● (core concept) — For a given sample, increasing the confidence level increases the critical value, the margin of error, and the width of the interval. Increasing the sample size decreases the standard error and narrows the interval (width is approximately proportional to 1/√n).
Justifying a claim based on a confidence interval (population mean or mean difference) ● (core concept) — A confidence interval for a population mean or population mean difference provides an interval of values that may serve as convincing evidence to support a particular claim about the parameter.
4.4 Setting Up a Test for a Population Mean or Population Mean Difference
One-sample t-test for a population mean ● (core concept) — The hypothesis testing procedure for a population mean when the population standard deviation σ is unknown.
One-sample t-test for a population mean difference (matched pairs) ● (core concept) — For a matched pairs design with two dependent samples, calculate the differences between pairs to produce one sample of differences, then use the one-sample t-test on the mean difference.
Null and alternative hypotheses (one-sample t-test for a mean) ● (core concept) — H0: μ = μ0, where μ0 is the null hypothesized value. One-sided alternatives are Ha: μ < μ0 or Ha: μ > μ0; the two-sided alternative is Ha: μ ≠ μ0.
Null and alternative hypotheses (one-sample t-test for a mean difference) ● (core concept) — H0: μd = 0. One-sided alternatives are Ha: μd < 0 or Ha: μd > 0; the two-sided alternative is Ha: μd ≠ 0.
4.5 Carrying Out a Test for a Population Mean or Population Mean Difference
Test statistic (one-sample t-test) ● (core concept) — t = (x̄ − μ0)/(s/√n). The t-statistic has a t-distribution with n − 1 degrees of freedom when the null hypothesis is true, and its p-value is found from the t-distribution using a table or technology.
p-value interpretation (means t-test) ● (core concept) — The p-value is the probability of obtaining a test statistic as extreme or more extreme, in the direction of the alternative hypothesis, assuming the null hypothesis is true (the population mean equals the stated value, or the population means are equal), stated in context.
4.6 Sampling Distributions for the Difference Between Two Sample Means
Sampling distribution of the difference in sample means ● (core concept) — For two independent populations with means μ1, μ2 and standard deviations σ1, σ2, when sampled values are independent, the sampling distribution of x̄1 − x̄2 has mean μ1 − μ2 and standard deviation √(σ1²/n1 + σ2²/n2).
Normality condition (sampling distribution of the difference in sample means) ● (core concept) — The sampling distribution of x̄1 − x̄2 is normal if both population distributions can be modeled by a normal distribution, or approximately normal if n1 ≥ 30 and n2 ≥ 30.
4.7 Constructing a Confidence Interval for the Difference Between Two Population Means
Two-sample t-interval for a difference between population means ● (core concept) — The confidence interval procedure for two independent samples when population standard deviations are unknown: (x̄1 − x̄2) ± t*√(s1²/n1 + s2²/n2), where t* is the critical value for the central C% of the t-distribution with appropriate degrees of freedom, found using technology.
Sample data condition (two-sample t procedures) ● (core concept) — Both samples should have size ≥ 30, or both population distributions are indicated to be approximately normal. If either sample size is less than 30, both sample data distributions should be free from strong skewness and outliers.
Standard error (difference in means) ● (core concept) — SE(x̄1 − x̄2) = √(s1²/n1 + s2²/n2), where s1 and s2 are the sample standard deviations.
Margin of error (two-sample t-interval) ● (core concept) — The critical value t* times the standard error of the difference: t*√(s1²/n1 + s2²/n2).
4.8 Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means
Confidence interval for a difference in means (interpretation and claim) ● (core concept) — We are C% confident the interval (a, b) contains the true difference between the two population means; the computed interval may or may not contain it. If the interval contains 0, there is insufficient evidence of a difference; if it does not contain 0, there is sufficient evidence of a difference.
4.9 Setting Up a Test for the Difference Between Two Population Means
Two-sample t-test for a difference between two population means ● (core concept) — The hypothesis testing procedure for comparing two population means.
Null and alternative hypotheses (two-sample t-test for means) ● (core concept) — H0 states no difference: H0: μ1 − μ2 = 0 or H0: μ1 = μ2. One-sided alternatives are Ha: μ1 < μ2 (or μ1 − μ2 < 0) or Ha: μ1 > μ2 (or μ1 − μ2 > 0); the two-sided alternative is Ha: μ1 ≠ μ2 (or μ1 − μ2 ≠ 0).
4.10 Carrying Out a Test for the Difference Between Two Population Means
Test statistic (two-sample t-test for means) ● (core concept) — t = ((x̄1 − x̄2) − 0)/√(s1²/n1 + s2²/n2). The t-statistic has a t-distribution when the null hypothesis is true; the degrees of freedom (between n1 + n2 − 2 and the smaller of n1 − 1 and n2 − 1) are found using technology.