Describe Penguins: Center and Spread
Mean, median, and measures of variation on Palmer Penguins — all in your browser
Welcome back. Everything on this page runs real R in your browser — same setup as Coding Activity 2. You already pulled a column and computed a mean; today you will summarize a variable with standard tools for center (typical value) and spread (how much values vary).
The first time you click Run Code, your browser downloads R once (a few seconds). After that it is quick.
- Center — where most values sit (
mean(),median()). - Spread — how far values scatter from that center (
sd(),IQR()).
Always say units when you report a number (grams for body mass here).
1. Load body mass
Run this to load Palmer Penguins and store body mass in grams. Two penguins have missing weights (NA); we skip those in every summary below.
2. Center: mean and median
Compute the mean and median of body_mass, each with na.rm = TRUE. Run both lines. (If you skipped the cell above, the setup below runs automatically when you click Run Code here.)
From ?mean and ?median: the first argument x is a numeric vector. By default na.rm = FALSE, so any NA makes the result NA — pass na.rm = TRUE to skip missing values.
Example with a built-in dataset:
mean(mtcars$mpg, na.rm = TRUE)
median(mtcars$mpg, na.rm = TRUE)Apply the same pattern to the object named in the exercise prompt.
mean(body_mass, na.rm = TRUE)
median(body_mass, na.rm = TRUE)Mean ≈ 4,002 g; median ≈ 3,700 g. The mean is pulled upward by a few heavy Gentoo penguins — that is why mean and median differ.
3. Spread: standard deviation
sd() measures typical distance from the mean. Compute the standard deviation of body_mass with na.rm = TRUE.
From ?sd: sd(x, na.rm = FALSE) measures spread around the mean in the same units as x. Like mean(), set na.rm = TRUE when the vector can contain missing values.
Example:
sd(mtcars$hp, na.rm = TRUE)Use sd() on the body-mass vector from this activity.
sd(body_mass, na.rm = TRUE)You should see about 800 g. Standard deviation uses the same units as the data — grams, not grams squared.
4. Variance vs. standard deviation
var() is sd squared, so its units are grams² — awkward for biology papers. Run var() and sd() on body_mass (both with na.rm = TRUE) and notice the difference in scale.
From ?var: variance has the same arguments as sd() — numeric vector x, and na.rm (default FALSE). Variance is in squared units; standard deviation is the square root and is easier to interpret in a sentence.
Compare both on the same column, for example:
var(mtcars$mpg, na.rm = TRUE)
sd(mtcars$mpg, na.rm = TRUE)Run both functions on the penguin body-mass vector here.
var(body_mass, na.rm = TRUE)
sd(body_mass, na.rm = TRUE)Variance ≈ 640,000 g²; sd ≈ 800 g. For a sentence in a lab report, prefer sd in grams.
5. Middle 50%: IQR
The interquartile range (IQR()) spans the middle half of the data — less sensitive to extreme values than sd. Compute it for body_mass.
From ?IQR: IQR(x, na.rm = FALSE, type = 7) returns the distance between the 25th and 75th percentiles (middle 50% of the data). Use na.rm = TRUE when needed.
Example:
IQR(mtcars$wt, na.rm = TRUE)Call it on the vector named in the prompt.
IQR(body_mass, na.rm = TRUE)IQR ≈ 1,200 g — the distance between the 25th and 75th percentiles.
6. Five-number summary
summary() on a numeric vector prints min, quartiles, median, mean, and max in one call. Run it on body_mass.
From ?summary: for a numeric vector, summary() prints min, quartiles, median, mean, and max in one table. No extra arguments are required for a simple numeric vector.
Example:
summary(mtcars$disp)Pass in the body-mass object you created earlier.
summary(body_mass)Use this as a quick sanity check before you write up results.
7. Compare two species
Adelie and Gentoo penguins differ in size. The vectors below are ready — compute mean and sd for each (with na.rm = TRUE). No hypothesis test yet; just describe center and spread.
You already have two numeric vectors in memory. For each vector, call mean(..., na.rm = TRUE) and sd(..., na.rm = TRUE) — four lines total.
Practice the idea on mtcars by splitting on a category, then summarizing each subset:
v4 <- mtcars$mpg[mtcars$cyl == 4]
v8 <- mtcars$mpg[mtcars$cyl == 8]
mean(v4, na.rm = TRUE)
sd(v4, na.rm = TRUE)Repeat for the second group, then do the same for the two species vectors in this exercise.
mean(adelie_mass, na.rm = TRUE)
sd(adelie_mass, na.rm = TRUE)
mean(gentoo_mass, na.rm = TRUE)
sd(gentoo_mass, na.rm = TRUE)Gentoo mean body mass is higher; both groups show substantial spread around their means.
8. Recap
| Function | What it tells you |
|---|---|
mean(..., na.rm = TRUE) |
Average (center) |
median(..., na.rm = TRUE) |
Middle value (robust center) |
sd(..., na.rm = TRUE) |
Typical deviation from the mean (spread, same units) |
var(..., na.rm = TRUE) |
Squared spread (same units as data, squared) |
IQR(..., na.rm = TRUE) |
Spread of the middle 50% |
summary() |
Min, quartiles, median, mean, max |
In Coding Activity 5 you will visualize these distributions with ggplot2.
Keep playing
Try the same summaries on flipper length in millimeters.
Try this with a partner. One person predicts whether flipper length has more or less spread than body mass; the other runs sd() on both. Swap roles.