Describe Penguins: Center and Spread

Mean, median, and measures of variation on Palmer Penguins — all in your browser

Welcome back. Everything on this page runs real R in your browser — same setup as Coding Activity 2. You already pulled a column and computed a mean; today you will summarize a variable with standard tools for center (typical value) and spread (how much values vary).

The first time you click Run Code, your browser downloads R once (a few seconds). After that it is quick.

NoteCenter and spread in one sentence
  • Center — where most values sit (mean(), median()).
  • Spread — how far values scatter from that center (sd(), IQR()).

Always say units when you report a number (grams for body mass here).

1. Load body mass

Run this to load Palmer Penguins and store body mass in grams. Two penguins have missing weights (NA); we skip those in every summary below.

2. Center: mean and median

Compute the mean and median of body_mass, each with na.rm = TRUE. Run both lines. (If you skipped the cell above, the setup below runs automatically when you click Run Code here.)

NoteHint

From ?mean and ?median: the first argument x is a numeric vector. By default na.rm = FALSE, so any NA makes the result NA — pass na.rm = TRUE to skip missing values.

Example with a built-in dataset:

mean(mtcars$mpg, na.rm = TRUE)
median(mtcars$mpg, na.rm = TRUE)

Apply the same pattern to the object named in the exercise prompt.

TipSolution
mean(body_mass, na.rm = TRUE)
median(body_mass, na.rm = TRUE)

Mean ≈ 4,002 g; median ≈ 3,700 g. The mean is pulled upward by a few heavy Gentoo penguins — that is why mean and median differ.

3. Spread: standard deviation

sd() measures typical distance from the mean. Compute the standard deviation of body_mass with na.rm = TRUE.

NoteHint

From ?sd: sd(x, na.rm = FALSE) measures spread around the mean in the same units as x. Like mean(), set na.rm = TRUE when the vector can contain missing values.

Example:

sd(mtcars$hp, na.rm = TRUE)

Use sd() on the body-mass vector from this activity.

TipSolution
sd(body_mass, na.rm = TRUE)

You should see about 800 g. Standard deviation uses the same units as the data — grams, not grams squared.

4. Variance vs. standard deviation

var() is sd squared, so its units are grams² — awkward for biology papers. Run var() and sd() on body_mass (both with na.rm = TRUE) and notice the difference in scale.

NoteHint

From ?var: variance has the same arguments as sd() — numeric vector x, and na.rm (default FALSE). Variance is in squared units; standard deviation is the square root and is easier to interpret in a sentence.

Compare both on the same column, for example:

var(mtcars$mpg, na.rm = TRUE)
sd(mtcars$mpg, na.rm = TRUE)

Run both functions on the penguin body-mass vector here.

TipSolution
var(body_mass, na.rm = TRUE)
sd(body_mass, na.rm = TRUE)

Variance ≈ 640,000 g²; sd ≈ 800 g. For a sentence in a lab report, prefer sd in grams.

5. Middle 50%: IQR

The interquartile range (IQR()) spans the middle half of the data — less sensitive to extreme values than sd. Compute it for body_mass.

NoteHint

From ?IQR: IQR(x, na.rm = FALSE, type = 7) returns the distance between the 25th and 75th percentiles (middle 50% of the data). Use na.rm = TRUE when needed.

Example:

IQR(mtcars$wt, na.rm = TRUE)

Call it on the vector named in the prompt.

TipSolution
IQR(body_mass, na.rm = TRUE)

IQR ≈ 1,200 g — the distance between the 25th and 75th percentiles.

6. Five-number summary

summary() on a numeric vector prints min, quartiles, median, mean, and max in one call. Run it on body_mass.

NoteHint

From ?summary: for a numeric vector, summary() prints min, quartiles, median, mean, and max in one table. No extra arguments are required for a simple numeric vector.

Example:

summary(mtcars$disp)

Pass in the body-mass object you created earlier.

TipSolution
summary(body_mass)

Use this as a quick sanity check before you write up results.

7. Compare two species

Adelie and Gentoo penguins differ in size. The vectors below are ready — compute mean and sd for each (with na.rm = TRUE). No hypothesis test yet; just describe center and spread.

NoteHint

You already have two numeric vectors in memory. For each vector, call mean(..., na.rm = TRUE) and sd(..., na.rm = TRUE) — four lines total.

Practice the idea on mtcars by splitting on a category, then summarizing each subset:

v4 <- mtcars$mpg[mtcars$cyl == 4]
v8 <- mtcars$mpg[mtcars$cyl == 8]
mean(v4, na.rm = TRUE)
sd(v4, na.rm = TRUE)

Repeat for the second group, then do the same for the two species vectors in this exercise.

TipSolution
mean(adelie_mass, na.rm = TRUE)
sd(adelie_mass, na.rm = TRUE)
mean(gentoo_mass, na.rm = TRUE)
sd(gentoo_mass, na.rm = TRUE)

Gentoo mean body mass is higher; both groups show substantial spread around their means.

8. Recap

Function What it tells you
mean(..., na.rm = TRUE) Average (center)
median(..., na.rm = TRUE) Middle value (robust center)
sd(..., na.rm = TRUE) Typical deviation from the mean (spread, same units)
var(..., na.rm = TRUE) Squared spread (same units as data, squared)
IQR(..., na.rm = TRUE) Spread of the middle 50%
summary() Min, quartiles, median, mean, max

In Coding Activity 5 you will visualize these distributions with ggplot2.

Keep playing

Try the same summaries on flipper length in millimeters.

Try this with a partner. One person predicts whether flipper length has more or less spread than body mass; the other runs sd() on both. Swap roles.