Part I · Describe
Forty numbers, no story yet
Each dot is one person's answer, one observation. Commute time is a quantitative variable: it has a unit (minutes), and arithmetic on it makes sense.
Scattered like this, the data hide everything interesting. Descriptive statistics summarize what this particular group looks like. They make no claims about anyone outside it.
Describe · Shape
Put every dot on a number line
Slide each dot to its value on a minutes axis and stack the ties. The result is a dot plot, a close cousin of the histogram. Here each column is a 2-minute bin.
Now the distribution has a shape. Most commutes bunch up on the left, and a thin tail stretches right to minutes.
That lopsided shape is called right-skewed. A commute can't drop below zero, but a bad day can run very long. Other shapes you will meet: symmetric, left-skewed, and bimodal (two humps).
Describe · Center
Where is the middle?
The mean adds every value and divides by the count:
Think of the axis as a seesaw: the mean is where the triangle would balance it.
The median is the middle value once you sort. With 40 values it is the average of the 20th and 21st: min.
In right-skewed data the long tail pulls the mean toward it, so the mean usually lands right of the median. Here mean − median = min.
Describe · Outliers
One bad morning
Suppose one commuter's train stalls and their -minute trip becomes 112 minutes.
The mean jumps from to . The median moves from to .
A value far from the rest is an outlier. The median is resistant: it only cares about order, so one extreme value barely nudges it. The mean uses the size of every value, so a single outlier can drag it. With skewed data or outliers, report the median, or both.
Describe · Spread
How far from typical?
Two cities can share an average commute while one is far more predictable, so center alone isn't enough. Each horizontal line is a deviation: one commute minus the mean. Positive ones point right, negative ones left.
Square the deviations, add them, and divide by n − 1 to get the variance. Its square root, the standard deviation, is back in minutes:
Read s as a typical distance from the mean. The shaded band, x̄ ± 1s, holds % of these commuters.
Why n − 1? Deviations are measured from x̄, which was fitted to these same 40 values, so they run slightly small. Dividing by 39 instead of 40 corrects that.
Describe · Box plots
Five numbers in one picture
The five-number summary is the minimum, Q1, median, Q3 and maximum. The quartiles Q1 and Q3 cut off the lowest and highest quarter of the data.
The box spans the middle half. Its width is the interquartile range, IQR = min. A common rule flags anything beyond Q3 + 1.5 × IQR = min (or below Q1 − 1.5 × IQR) as a possible outlier.
Median and IQR are the resistant pair. Mean and standard deviation are the sensitive pair. Pick the pair that suits the shape.
Part II · Infer
The forty are a sample
Our 40 commuters were picked at random from a whole city. In this simulation the population is people (amber), and our sample is blue.
A number describing the population is a parameter: μ for its mean, σ for its standard deviation. The matching number from a sample is a statistic: x̄ and s. Here μ = and x̄ = .
In real life you never see the amber dots, so μ is unknown. Inferential statistics uses the sample to estimate a parameter and to say how far off that estimate is likely to be.
Random selection is what makes this work. A convenient sample, say everyone waiting at one train station, can be biased in ways no formula can fix.
Infer · Sampling distributions
What if we asked again?
A different random 40 would give a different x̄. The chart now draws sample after sample from the population and drops each sample's mean into the histogram at the bottom.
That pile of means is the sampling distribution of x̄. Notice three things:
- It centers on μ. On average, x̄ hits the target.
- It is much narrower than the population. Its spread is the standard error, SE = σ / √n.
- It looks like a bell curve even though commutes are skewed. That is the Central Limit Theorem, and it improves as n grows.
Try n = 5: the means spread wide and keep some of the skew. At n = 100 they crowd tightly around μ.
Infer · Confidence intervals
A range instead of a single guess
We have only one sample, so we estimate the standard error as s / √n = min and build a confidence interval:
t* comes from the t distribution with n − 1 = 39 degrees of freedom. It is slightly larger than 1.96 because s is only an estimate of σ.
The top row is our interval. Each row below applies the same recipe to a fresh sample, and intervals that miss μ turn red. Over many samples, 95% of intervals built this way capture μ.
The 95% describes the method. Any single interval either contains μ or it doesn't; from inside one sample you can't tell which.
Infer · Hypothesis tests
Testing a claim
A transit agency report says the average commute is minutes. Is our sample consistent with that? Write the claim as a null hypothesis and its rival as the alternative:
If H₀ were true, sample means would scatter around like the curve in the chart. Measure how far our x̄ landed, in standard errors:
The p-value is the shaded area: the probability, if H₀ were true, of a sample mean at least this far from μ₀ in either direction. Here p = .
A small p-value says the data would be surprising if H₀ were true. It is not the probability that H₀ is true.
Infer · Errors & power
Two ways to be wrong
We reject H₀ when p < α. Choosing α = 0.05 fixes the dashed decision lines. Any decision rule can fail in two ways:
- Type I error
- Rejecting an H₀ that is true. Its probability is α, the red tails under the null curve.
- Type II error
- Missing a real difference. Its probability is β, the pale blue area under the curve for the true mean.
Because this is a simulation, we know the truth: μ = . With the claim at μ₀ = and n = 40, β ≈ , so the test's power, 1 − β, is about .
Slide the claim toward the true mean: the curves overlap and power falls. Larger samples narrow both curves and raise power.