Module 2 - Test bank

This bank holds 40 questions: ten in each section.

There are two datasets used in the test bank. The datasets are:

Your module test will contain one question per section. The test will only use one of the datasets, and the dataset will be provided in the test.

For each question, 20% of the mark is for presentation and the other 80% is for the correctness of your script.

You submit a single R script: a header block at the top, a clearly labelled section for each question, and your sentence answers written as comments. The script should run from top to bottom in a project folder that has the csvs in data/. Here is an example of what a test will look like and the expectation for what your script will look like:

A real test looks like this – one question from each section:

Question 1

Five canola fields yielded 52.3, 47.8, 55.1, 44.6, and 50.9 bu/ac, on 160, 320, 240, 130, and 200 acres.

(a) Create a vector yields and a vector acres holding these values.

(b) Create a vector yields_t_ha converting the yields to tonnes per hectare (multiply by 0.0560) in one line.

(c) Compute the mean and standard deviation of yields.

(d) Compute total production in bushels, and the farm’s average yield weighted by acres. In a comment: why does the weighted average differ from the plain mean?

Question 2

The Saskatchewan canola file contains one RM-year average yield per row, measured in bu/ac.

(a) Compute the mean, median, and standard deviation of Yield.

(b) Compute its 90th percentile.

(c) In a comment: the mean sits a little above the median – what does that direction of gap suggest about the shape of canola yields?

Question 3

The full RM file holds all eight crops; lentil yields are recorded in lb/ac.

(a) Filter to the lentil rows, saving the result as lentils. How many rows are there?

(b) Add a yield_kg_ha column (multiply by 1.12), saving back to lentils.

(c) Compute the mean of the new column and of the original column, and check their ratio. In a comment: why is the ratio the conversion factor?

Question 4

Using the full eight-crop RM file.

(a) In one pipeline, compute each RM’s mean Spring Wheat yield across all years, sorted best-first.

(b) Report the top three RMs in a comment.

(c) In a comment: what did each row of the input represent, and what does each row of the result represent?

And here is a script that would earn full marks, including all of the presentation component – with the console session it produces shown beneath it:

# ---
# Title: Module 2 test
# Author: Jordan Field
# Date: 2026-10-15
# Description:
#   Answers to the four test questions, one block per question.
#   Sentence answers are written as comments below each result.
# ---

# Load packages
library(tidyverse)

# ============================================================
# Question 1
# ============================================================

# (a) Yields (bu/ac) and field sizes (acres)
yields <- c(52.3, 47.8, 55.1, 44.6, 50.9)
acres  <- c(160, 320, 240, 130, 200)

# (b) Convert every yield to tonnes per hectare in one step
yields_t_ha <- yields * 0.0560
yields_t_ha

# (c) Centre and spread of the yields
mean(yields)   # 50.14 bu/ac
sd(yields)     # 4.06 bu/ac

# (d) Total production, then the average weighted by acres
total_bu <- sum(yields * acres)
total_bu                 # 52,866 bu
total_bu / sum(acres)    # 50.35 bu/ac
# (d) The weighted average differs from the plain mean because the
# fields differ in size, so each yield should count by its acres.

# ============================================================
# Question 2
# ============================================================

# Read the canola file using a relative path from the project folder
rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")

# (a) Centre and spread of the canola yields
mean(rm_canola$Yield)      # 28.29 bu/ac
median(rm_canola$Yield)    # 26.9 bu/ac
sd(rm_canola$Yield)        # 10.06 bu/ac

# (b) 90th percentile
quantile(rm_canola$Yield, 0.9)   # 42.5 bu/ac

# (c) The mean sits a little above the median, which suggests mild
# right skew: the best RM-years pull the mean up more than the
# worst years pull it down.

# ============================================================
# Question 3
# ============================================================

# Read the full eight-crop file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Keep only the lentil rows
lentils <- filter(rm_yields, Crop == "Lentils")
nrow(lentils)    # 6,338 rows

# (b) Add the converted column (1 lb/ac is about 1.12 kg/ha) and
# overwrite lentils with the version that has the extra column
lentils <- mutate(lentils, yield_kg_ha = Yield * 1.12)

# (c) Means of the new and original columns
mean(lentils$yield_kg_ha)   # 1,353.4 kg/ha
mean(lentils$Yield)         # 1,208.39 lb/ac
# (c) Their ratio is 1.12, the conversion factor -- multiplying every
# value by a constant multiplies the mean by the same constant.

# ============================================================
# Question 4
# ============================================================

# (a) Mean spring wheat yield by RM, best first
rm_yields |>                              # take the data, THEN
  filter(Crop == "Spring Wheat") |>       # keep spring wheat, THEN
  group_by(RM) |>                         # split the rows by RM, THEN
  summarise(mean_yield = mean(Yield)) |>  # one row per RM: its mean, THEN
  arrange(desc(mean_yield))               # best RMs on top

# (b) Top three: RM 369 (47.6), RM 333 (47.2), RM 368 (47.2) bu/ac.
# (c) Each input row was one RM-year observation of spring wheat; each
# result row is one RM, its years collapsed into a single mean.
> # ---
> # Title: Module 2 test
> # Author: Jordan Field
> # Date: 2026-10-15
> # Description:
> #   Answers to the four test questions, one block per question.
> #   Sentence answers are written as comments below each result.
> # ---
> 
> # Load packages
> library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr     1.2.1     ✔ readr     2.2.0
✔ forcats   1.0.1     ✔ stringr   1.6.0
✔ ggplot2   4.0.3     ✔ tibble    3.3.1
✔ lubridate 1.9.5     ✔ tidyr     1.3.2
✔ purrr     1.2.2     
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
> 
> # ============================================================
> # Question 1
> # ============================================================
> 
> # (a) Yields (bu/ac) and field sizes (acres)
> yields <- c(52.3, 47.8, 55.1, 44.6, 50.9)
> acres  <- c(160, 320, 240, 130, 200)
> 
> # (b) Convert every yield to tonnes per hectare in one step
> yields_t_ha <- yields * 0.0560
> yields_t_ha
[1] 2.9288 2.6768 3.0856 2.4976 2.8504
> 
> # (c) Centre and spread of the yields
> mean(yields)   # 50.14 bu/ac
[1] 50.14
> sd(yields)     # 4.06 bu/ac
[1] 4.062388
> 
> # (d) Total production, then the average weighted by acres
> total_bu <- sum(yields * acres)
> total_bu                 # 52,866 bu
[1] 52866
> total_bu / sum(acres)    # 50.35 bu/ac
[1] 50.34857
> # (d) The weighted average differs from the plain mean because the
> # fields differ in size, so each yield should count by its acres.
> 
> # ============================================================
> # Question 2
> # ============================================================
> 
> # Read the canola file using a relative path from the project folder
> rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")
Rows: 10039 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Centre and spread of the canola yields
> mean(rm_canola$Yield)      # 28.29 bu/ac
[1] 28.29239
> median(rm_canola$Yield)    # 26.9 bu/ac
[1] 26.9
> sd(rm_canola$Yield)        # 10.06 bu/ac
[1] 10.06019
> 
> # (b) 90th percentile
> quantile(rm_canola$Yield, 0.9)   # 42.5 bu/ac
 90% 
42.5 
> 
> # (c) The mean sits a little above the median, which suggests mild
> # right skew: the best RM-years pull the mean up more than the
> # worst years pull it down.
> 
> # ============================================================
> # Question 3
> # ============================================================
> 
> # Read the full eight-crop file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Keep only the lentil rows
> lentils <- filter(rm_yields, Crop == "Lentils")
> nrow(lentils)    # 6,338 rows
[1] 6338
> 
> # (b) Add the converted column (1 lb/ac is about 1.12 kg/ha) and
> # overwrite lentils with the version that has the extra column
> lentils <- mutate(lentils, yield_kg_ha = Yield * 1.12)
> 
> # (c) Means of the new and original columns
> mean(lentils$yield_kg_ha)   # 1,353.4 kg/ha
[1] 1353.402
> mean(lentils$Yield)         # 1,208.39 lb/ac
[1] 1208.395
> # (c) Their ratio is 1.12, the conversion factor -- multiplying every
> # value by a constant multiplies the mean by the same constant.
> 
> # ============================================================
> # Question 4
> # ============================================================
> 
> # (a) Mean spring wheat yield by RM, best first
> rm_yields |>                              # take the data, THEN
+   filter(Crop == "Spring Wheat") |>       # keep spring wheat, THEN
+   group_by(RM) |>                         # split the rows by RM, THEN
+   summarise(mean_yield = mean(Yield)) |>  # one row per RM: its mean, THEN
+   arrange(desc(mean_yield))               # best RMs on top
# A tibble: 298 × 2
      RM mean_yield
   <dbl>      <dbl>
 1   369       47.6
 2   333       47.2
 3   368       47.2
 4   271       46.1
 5   303       45.0
 6   427       45.0
 7   404       44.9
 8   496       44.8
 9   493       44.6
10   338       44.5
# ℹ 288 more rows
> 
> # (b) Top three: RM 369 (47.6), RM 333 (47.2), RM 368 (47.2) bu/ac.
> # (c) Each input row was one RM-year observation of spring wheat; each
> # result row is one RM, its years collapsed into a single mean.
  • At least an hour before the test, “Schedule an appointment” in the test for that day.
  • Bring your (charged) laptop to the test.
  • Your test will be accessible at 11:30am.
  • When you get to the room:
  1. Close all applications other than your webbrowser (used to access the test) and VS Code.
  2. Log into zoom and join the breakout room with your name on it.
  3. In Zoom, share your screen (not an application window, but your whole screen) and record your session.
  4. Navigate to Canvas and access your test.
  5. Download the dataset that you need
  6. Complete the test in VS Code.
  7. Upload the completed script to Canvas (while still logged into zoom).
  8. Leave zoom, and upload the screenshot recording to Canvas.
  9. Your done!!

Section 1 — R Basics

Question 1

Five canola fields yielded 52.3, 47.8, 55.1, 44.6, and 50.9 bu/ac, on 160, 320, 240, 130, and 200 acres.

(a) Create a vector yields and a vector acres holding these values.

(b) Using one line of code, create a vector yields_t_ha converting the yields to tonnes per hectare (multiply by 0.0560) in one line.

(c) Compute the mean and standard deviation of yields.

(d) Compute total production in bushels (each field’s yield times its acres, summed), and the farm’s average yield weighted by acres (total bushels divided by total acres).

Answer
# (a) Create the yield and acreage vectors
yields <- c(52.3, 47.8, 55.1, 44.6, 50.9)
acres  <- c(160, 320, 240, 130, 200)

# (b) Convert every yield to tonnes per hectare in one line
yields_t_ha <- yields * 0.0560

# Print the converted yields
yields_t_ha

# (c) Mean and standard deviation of the yields
mean(yields)
sd(yields)

# (d) Total production in bushels
total_bu <- sum(yields * acres)
total_bu

# Average yield weighted by acres
total_bu / sum(acres)

## I find the mean to be 50.14 bu/ac and the sd 4.06.
## Total production is 52,866 bu and the acre-weighted average 50.35 bu/ac.
> # (a) Create the yield and acreage vectors
> yields <- c(52.3, 47.8, 55.1, 44.6, 50.9)
> acres  <- c(160, 320, 240, 130, 200)
> 
> # (b) Convert every yield to tonnes per hectare in one line
> yields_t_ha <- yields * 0.0560
> 
> # Print the converted yields
> yields_t_ha
[1] 2.9288 2.6768 3.0856 2.4976 2.8504
> 
> # (c) Mean and standard deviation of the yields
> mean(yields)
[1] 50.14
> sd(yields)
[1] 4.062388
> 
> # (d) Total production in bushels
> total_bu <- sum(yields * acres)
> total_bu
[1] 52866
> 
> # Average yield weighted by acres
> total_bu / sum(acres)
[1] 50.34857
> 
> ## I find the mean to be 50.14 bu/ac and the sd 4.06.
> ## Total production is 52,866 bu and the acre-weighted average 50.35 bu/ac.

Question 2

Six spring wheat fields yielded 51.2, 48.7, 53.4, 8.2, 50.9, and 49.5 bu/ac; the 8.2 field was hailed out in July.

(a) Save the yields and compute the mean and the median.

(b) In a comment: the two disagree by nearly seven bushels. Which one better describes a typical field on this farm, and why?

(c) Compute the mean of the five unhailed fields (a second vector is fine) and compare it with the median from part a. In a comment: what does the closeness of those two numbers tell you about the median?

Answer
# (a) Yields, including the hailed-out field
yields <- c(51.2, 48.7, 53.4, 8.2, 50.9, 49.5)
mean(yields)
median(yields)


## I find the mean to be 43.65 and the median 50.2.
## (b) The median describes a typical field better: five of the six fields
## sit near 50, and the single 8.2 drags the mean seven bushels below all
## of them while barely touching the median.

# (c) The five unhailed fields
unhailed <- c(51.2, 48.7, 53.4, 50.9, 49.5)
mean(unhailed)

## (c) The unhailed mean is 50.74, almost exactly the full-data median --
## the median was already ignoring the extreme value.
> # (a) Yields, including the hailed-out field
> yields <- c(51.2, 48.7, 53.4, 8.2, 50.9, 49.5)
> mean(yields)
[1] 43.65
> median(yields)
[1] 50.2
> 
> 
> ## I find the mean to be 43.65 and the median 50.2.
> ## (b) The median describes a typical field better: five of the six fields
> ## sit near 50, and the single 8.2 drags the mean seven bushels below all
> ## of them while barely touching the median.
> 
> # (c) The five unhailed fields
> unhailed <- c(51.2, 48.7, 53.4, 50.9, 49.5)
> mean(unhailed)
[1] 50.74
> 
> ## (c) The unhailed mean is 50.74, almost exactly the full-data median --
> ## the median was already ignoring the extreme value.

Question 3

Five barley fields yielded 68.2, 71.5, 64.9, 74.1, and 66.3 bu/ac. Five lentil fields yielded 1450, 1720, 1280, 1610, and 1390 lb/ac.

(a) Compute the mean, the standard deviation, the variance, and the range (maximum minus minimum) of each crop’s yields.

(b) The lentil standard deviation is far larger. In a comment: why can the two standard deviations not be compared directly?

(c) Compute the coefficient of variation (sd divided by mean) for each crop. Which crop’s yields are more variable relative to their own average?

Answer
# (a) Yield vectors for each crop
barley  <- c(68.2, 71.5, 64.9, 74.1, 66.3)
lentils <- c(1450, 1720, 1280, 1610, 1390)

# Mean of each crop
mean(barley)
mean(lentils)

# Standard deviation of each crop
sd(barley)
sd(lentils)

# Variance of each crop
var(barley)
var(lentils)

# Range (maximum minus minimum) of each crop
max(barley) - min(barley)
max(lentils) - min(lentils)

## I find barley mean 69.0, sd 3.77, variance 14.25, range 9.2; lentils
## mean 1490, sd 175.36, variance 30,750, range 440.

## (b) The two standard deviations are in different units (bushels vs pounds
## per acre) and on very different scales, so the raw spreads are not
## comparable.

# (c) Coefficient of variation (sd divided by mean) for each crop
sd(barley) / mean(barley)
sd(lentils) / mean(lentils)

## (c) CV barley 0.055, lentils 0.118 -- lentil yields are more variable
## relative to their own average.
> # (a) Yield vectors for each crop
> barley  <- c(68.2, 71.5, 64.9, 74.1, 66.3)
> lentils <- c(1450, 1720, 1280, 1610, 1390)
> 
> # Mean of each crop
> mean(barley)
[1] 69
> mean(lentils)
[1] 1490
> 
> # Standard deviation of each crop
> sd(barley)
[1] 3.774917
> sd(lentils)
[1] 175.3568
> 
> # Variance of each crop
> var(barley)
[1] 14.25
> var(lentils)
[1] 30750
> 
> # Range (maximum minus minimum) of each crop
> max(barley) - min(barley)
[1] 9.2
> max(lentils) - min(lentils)
[1] 440
> 
> ## I find barley mean 69.0, sd 3.77, variance 14.25, range 9.2; lentils
> ## mean 1490, sd 175.36, variance 30,750, range 440.
> 
> ## (b) The two standard deviations are in different units (bushels vs pounds
> ## per acre) and on very different scales, so the raw spreads are not
> ## comparable.
> 
> # (c) Coefficient of variation (sd divided by mean) for each crop
> sd(barley) / mean(barley)
[1] 0.05470895
> sd(lentils) / mean(lentils)
[1] 0.1176891
> 
> ## (c) CV barley 0.055, lentils 0.118 -- lentil yields are more variable
> ## relative to their own average.

Question 4

Here is a script exactly as saved:

mean(yields)
yields <- c(31.7, 28.4, 33.9, 30.2)

(a) In a comment: what would happen if you ran this script top to bottom in a fresh session, and why?

(b) Rewrite the script in working order, with a header block and a comment on each line.

(c) Compute the standard deviation and median of yields.

Answer
# (a) Run as saved, the script fails on line 1: mean(yields) is asked for
# before yields exists in the fresh session, so R stops with
# R reports that the object yields was not found.

# (b) The working order, with a header
# ---
# This script computes the mean of a yield vector.
# ---

# Create the yield vector first
yields <- c(31.7, 28.4, 33.9, 30.2)

# Then compute its mean
mean(yields)

## I find the mean to be 31.05.

# (c) Standard deviation and median of yields
sd(yields)
median(yields)
> # (a) Run as saved, the script fails on line 1: mean(yields) is asked for
> # before yields exists in the fresh session, so R stops with
> # R reports that the object yields was not found.
> 
> # (b) The working order, with a header
> # ---
> # This script computes the mean of a yield vector.
> # ---
> 
> # Create the yield vector first
> yields <- c(31.7, 28.4, 33.9, 30.2)
> 
> # Then compute its mean
> mean(yields)
[1] 31.05
> 
> ## I find the mean to be 31.05.
> 
> # (c) Standard deviation and median of yields
> sd(yields)
[1] 2.330236
> median(yields)
[1] 30.95

Question 5

Suppose there are four fields:

  • Field 1 had canola, 160 acres, and was irrigated.
  • Field 2 had wheat, 320 acres, and was not irrigated.
  • Field 3 had peas, 240 acres, and was not irrigated.
  • Field 4 had barley, 130 acres, and was irrigated.

(a) Create three vectors of length four, one for the crop grown, one for the seeded acres, and one for whether the field was irrigated.

(b) Combine your three vectors into a data frame called trial and print it in the console.

Answer
# (a) Crop grown: character, so the values are quoted
crop      <- c("Canola", "Wheat", "Peas", "Barley")

# Seeded acres: numeric, no quotes
acres     <- c(160, 320, 240, 130)

# Irrigated or not: logical, no quotes
irrigated <- c(TRUE, FALSE, FALSE, TRUE)

# (b) Combine the three vectors into a data frame
trial <- data.frame(crop = crop,
                    acres = acres,
                    irrigated = irrigated)

# Print the data frame
trial
> # (a) Crop grown: character, so the values are quoted
> crop      <- c("Canola", "Wheat", "Peas", "Barley")
> 
> # Seeded acres: numeric, no quotes
> acres     <- c(160, 320, 240, 130)
> 
> # Irrigated or not: logical, no quotes
> irrigated <- c(TRUE, FALSE, FALSE, TRUE)
> 
> # (b) Combine the three vectors into a data frame
> trial <- data.frame(crop = crop,
+                     acres = acres,
+                     irrigated = irrigated)
> 
> # Print the data frame
> trial
    crop acres irrigated
1 Canola   160      TRUE
2  Wheat   320     FALSE
3   Peas   240     FALSE
4 Barley   130      TRUE

Question 6

The same five canola fields were measured in two years. 2024: 38.2, 41.6, 35.9, 43.1, 39.7 bu/ac. 2025: 45.1, 44.8, 42.3, 49.6, 46.2 bu/ac.

(a) Save the two vectors and compute each field’s change in yield in one line.

(b) Compute the mean change, and the mean percentage change (change over the 2024 yield, times 100).

(c) In a comment: every field improved, but which measure – bushels gained or percent gained – would you use to compare this farm against a lentil farm, and why?

Answer
# (a) The two years, and each field change in one line
y2024 <- c(38.2, 41.6, 35.9, 43.1, 39.7)
y2025 <- c(45.1, 44.8, 42.3, 49.6, 46.2)
change <- y2025 - y2024
change

# (b) Mean change and mean percentage change
mean(change)
mean(change / y2024 * 100)

## I find a mean change of 5.9 bu/ac and a mean percentage change of 15.0%.
## (c) The percentage change: bushels gained are not comparable across crops
## measured on different scales (a lentil farm gains hundreds of lb/ac), but
## percent gained puts both farms on the same footing.
> # (a) The two years, and each field change in one line
> y2024 <- c(38.2, 41.6, 35.9, 43.1, 39.7)
> y2025 <- c(45.1, 44.8, 42.3, 49.6, 46.2)
> change <- y2025 - y2024
> change
[1] 6.9 3.2 6.4 6.5 6.5
> 
> # (b) Mean change and mean percentage change
> mean(change)
[1] 5.9
> mean(change / y2024 * 100)
[1] 15.00729
> 
> ## I find a mean change of 5.9 bu/ac and a mean percentage change of 15.0%.
> ## (c) The percentage change: bushels gained are not comparable across crops
> ## measured on different scales (a lentil farm gains hundreds of lb/ac), but
> ## percent gained puts both farms on the same footing.

Question 7

A farm’s four barley fields: field IDs "B1" to "B4", sizes 155, 240, 160, and 310 acres, yields 66.8, 71.2, 63.5, and 69.9 bu/ac.

(a) Build a data frame barley_fields holding the three variables.

(b) Compute each field’s production (acres times yield) and save it as a vector.

(c) Compute total production, and the farm’s average yield weighted by acres (total production over total acres).

(d) In a comment: compare the weighted average with mean(barley_fields$yield) and say why they differ.

Answer
# (a) The barley fields as a data frame
barley_fields <- data.frame(
  field_id = c("B1", "B2", "B3", "B4"),
  acres    = c(155, 240, 160, 310),
  yield    = c(66.8, 71.2, 63.5, 69.9)
)

# (b) Production of each field
production <- barley_fields$acres * barley_fields$yield
production

# (c) Total production and the acre-weighted average yield
sum(production)
sum(production) / sum(barley_fields$acres)

# (d) The plain mean, for comparison
mean(barley_fields$yield)

## I find total production of 59,271 bu and a weighted average of 68.52.
## (d) The plain mean is 67.85: it counts each field once, while the
## weighted average lets the 310-acre field count for its share of the
## land, and that field yields above average, pulling the weighted figure up.
> # (a) The barley fields as a data frame
> barley_fields <- data.frame(
+   field_id = c("B1", "B2", "B3", "B4"),
+   acres    = c(155, 240, 160, 310),
+   yield    = c(66.8, 71.2, 63.5, 69.9)
+ )
> 
> # (b) Production of each field
> production <- barley_fields$acres * barley_fields$yield
> production
[1] 10354 17088 10160 21669
> 
> # (c) Total production and the acre-weighted average yield
> sum(production)
[1] 59271
> sum(production) / sum(barley_fields$acres)
[1] 68.52139
> 
> # (d) The plain mean, for comparison
> mean(barley_fields$yield)
[1] 67.85
> 
> ## I find total production of 59,271 bu and a weighted average of 68.52.
> ## (d) The plain mean is 67.85: it counts each field once, while the
> ## weighted average lets the 310-acre field count for its share of the
> ## land, and that field yields above average, pulling the weighted figure up.

Question 8

(a) Build a data frame called durum_fields with columns field_id ("D1" to "D4"), region (South, South, Central, North), and yield (28.3, 41.6, 35.2, 44.8).

(b) Report its number of rows, number of columns, and column names using functions.

(c) Using $, compute the mean yield and the range (maximum minus minimum).

Answer
# (a) Build the data frame
durum_fields <- data.frame(
  field_id = c("D1", "D2", "D3", "D4"),
  region = c("South", "South", "Central", "North"),
  yield = c(28.3, 41.6, 35.2, 44.8)
)

# (b) Number of rows
nrow(durum_fields)

# Number of columns
ncol(durum_fields)

# Column names
names(durum_fields)

# (c) Mean yield, using $
mean(durum_fields$yield)

# Range (maximum minus minimum)
max(durum_fields$yield) - min(durum_fields$yield)

## (b) 4 rows, 3 columns; the columns are field_id, region, yield.
## (c) I find the mean to be 37.48 (the console shows 37.475) and the
## range 16.5.
> # (a) Build the data frame
> durum_fields <- data.frame(
+   field_id = c("D1", "D2", "D3", "D4"),
+   region = c("South", "South", "Central", "North"),
+   yield = c(28.3, 41.6, 35.2, 44.8)
+ )
> 
> # (b) Number of rows
> nrow(durum_fields)
[1] 4
> 
> # Number of columns
> ncol(durum_fields)
[1] 3
> 
> # Column names
> names(durum_fields)
[1] "field_id" "region"   "yield"   
> 
> # (c) Mean yield, using $
> mean(durum_fields$yield)
[1] 37.475
> 
> # Range (maximum minus minimum)
> max(durum_fields$yield) - min(durum_fields$yield)
[1] 16.5
> 
> ## (b) 4 rows, 3 columns; the columns are field_id, region, yield.
> ## (c) I find the mean to be 37.48 (the console shows 37.475) and the
> ## range 16.5.

Question 9

Twelve canola fields yielded 24.1, 31.5, 28.8, 35.2, 22.7, 30.4, 27.9, 33.6, 25.5, 29.3, 32.8, and 26.4 bu/ac.

(a) Save the values as a vector and compute the mean.

(b) Compute the 10th, 50th, and 90th percentiles in a single quantile() call.

(c) In a comment: one of your three percentiles should equal the median. Which one, and does it?

Answer
# (a) Canola yields as a vector (bu/ac)
canola <- c(24.1, 31.5, 28.8, 35.2, 22.7, 30.4, 27.9, 33.6, 25.5, 29.3, 32.8, 26.4)

# Mean yield
mean(canola)

# (b) 10th, 50th, and 90th percentiles in one call
quantile(canola, c(0.1, 0.5, 0.9))

# (c) The median, to check against the 50th percentile
median(canola)

## I find the mean to be 29.02.
## (b) 10th 24.24, 50th 29.05, 90th 33.52.
## (c) The 50th percentile is the median by definition; median(canola)
## confirms 29.05.
> # (a) Canola yields as a vector (bu/ac)
> canola <- c(24.1, 31.5, 28.8, 35.2, 22.7, 30.4, 27.9, 33.6, 25.5, 29.3, 32.8, 26.4)
> 
> # Mean yield
> mean(canola)
[1] 29.01667
> 
> # (b) 10th, 50th, and 90th percentiles in one call
> quantile(canola, c(0.1, 0.5, 0.9))
  10%   50%   90% 
24.24 29.05 33.52 
> 
> # (c) The median, to check against the 50th percentile
> median(canola)
[1] 29.05
> 
> ## I find the mean to be 29.02.
> ## (b) 10th 24.24, 50th 29.05, 90th 33.52.
> ## (c) The 50th percentile is the median by definition; median(canola)
> ## confirms 29.05.

Question 10

Four oat fields yielded 84.3, 91.7, 77.2, and 88.5 bu/ac, and oats are selling at $4.15 per bushel.

(a) Save the yields as a vector and the price as a single object.

(b) Compute a revenue-per-acre vector (yield times price) in one line, using the price object rather than retyping the number.

(c) Compute the mean revenue per acre.

(d) In a comment: the price rises to $4.60. What value would you change, and which lines would you rerun to update the results? Why is that easier than typing the price into each calculation?

Answer
# (a) Oat yields (bu/ac) and the price ($/bu)
yields <- c(84.3, 91.7, 77.2, 88.5)
price  <- 4.15

# (b) Revenue per acre, using the price object
revenue <- yields * price

# Print the revenues
revenue

# (c) Mean revenue per acre
mean(revenue)

## (b) The revenues are 349.85, 380.56, 320.38, 367.28 $/ac, rounded to the
## cent (the console shows 349.845 and so on).
## (c) Mean revenue is $354.51 per acre.
## (d) Change price <- 4.60, then rerun the lines below it. The price lives
## in one object, so one edit updates everything -- the same reason a
## spreadsheet keeps a price in one cell.
> # (a) Oat yields (bu/ac) and the price ($/bu)
> yields <- c(84.3, 91.7, 77.2, 88.5)
> price  <- 4.15
> 
> # (b) Revenue per acre, using the price object
> revenue <- yields * price
> 
> # Print the revenues
> revenue
[1] 349.845 380.555 320.380 367.275
> 
> # (c) Mean revenue per acre
> mean(revenue)
[1] 354.5138
> 
> ## (b) The revenues are 349.85, 380.56, 320.38, 367.28 $/ac, rounded to the
> ## cent (the console shows 349.845 and so on).
> ## (c) Mean revenue is $354.51 per acre.
> ## (d) Change price <- 4.60, then rerun the lines below it. The price lives
> ## in one object, so one edit updates everything -- the same reason a
> ## spreadsheet keeps a price in one cell.

Section 2 — Reading and Inspecting Data

Saskatchewan RM Canola

Question 11

This question uses: rm_canola_yields_1990_2025.csv

(a) Run summary(rm_canola). From its output, report the minimum, median, and maximum of Yield, and the first and last Year.

(b) In a comment: are most yields in a plausible range for canola, and do the extremes need investigation? Are the years what the file name promises? Are there missing values?

(c) The minimum is below 2 bu/ac. In a comment: is a canola yield that low necessarily a data error?

Answer
# Load the tidyverse and read the canola file
library(tidyverse)
rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")

# (a) Summary of every column
summary(rm_canola)

## (a) I find Yield has minimum 1.9, median 26.9, and maximum 61.0, and Year runs from 1990 to 2025.
## (b) Most yields are plausible for canola, but the values near 2 bu/ac
## deserve investigation. The years run 1990--2025 and there are no NAs.
## (c) Not necessarily. A severe drought or widespread crop failure could
## produce a very low RM average.
> # Load the tidyverse and read the canola file
> library(tidyverse)
> rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")
Rows: 10039 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Summary of every column
> summary(rm_canola)
      Year            RM               Crop           Yield      
 Min.   :1990   Min.   :  1.0   Length   :10039   Min.   : 1.90  
 1st Qu.:1999   1st Qu.:128.0   N.unique :    1   1st Qu.:21.00  
 Median :2008   Median :253.0   N.blank  :    0   Median :26.90  
 Mean   :2008   Mean   :254.4   Min.nchar:    6   Mean   :28.29  
 3rd Qu.:2017   3rd Qu.:372.0   Max.nchar:    6   3rd Qu.:35.40  
 Max.   :2025   Max.   :622.0                     Max.   :61.00  
        Unit      
 Length   :10039  
 N.unique :    1  
 N.blank  :    0  
 Min.nchar:    5  
 Max.nchar:    5  
                  
> 
> ## (a) I find Yield has minimum 1.9, median 26.9, and maximum 61.0, and Year runs from 1990 to 2025.
> ## (b) Most yields are plausible for canola, but the values near 2 bu/ac
> ## deserve investigation. The years run 1990--2025 and there are no NAs.
> ## (c) Not necessarily. A severe drought or widespread crop failure could
> ## produce a very low RM average.

Question 12

This question uses: rm_canola_yields_1990_2025.csv

(a) Compute the mean, median, standard deviation, and variance of rm_canola$Yield.

(b) Compute the 90th percentile of Yield.

(c) In a comment: what does that difference between the mean and median suggest about the shape of canola yields?

Answer
# Load the tidyverse and read the canola file
library(tidyverse)
rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")

# (a) Centre of the yields
mean(rm_canola$Yield)
median(rm_canola$Yield)

# Spread of the yields
sd(rm_canola$Yield)
var(rm_canola$Yield)

# (b) 90th percentile of the yields
quantile(rm_canola$Yield, probs = 0.9)

## (a) I find a mean of 28.29, a median of 26.9, a standard deviation of 10.06, and a variance of 101.21 -- the sd squared, in squared units.
## (b) The 90th percentile is 42.5 bu/ac.
## (c) A mean above the median suggests mild right skew: the very best RM-years pull the mean up more than the worst years pull it down.
> # Load the tidyverse and read the canola file
> library(tidyverse)
> rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")
Rows: 10039 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Centre of the yields
> mean(rm_canola$Yield)
[1] 28.29239
> median(rm_canola$Yield)
[1] 26.9
> 
> # Spread of the yields
> sd(rm_canola$Yield)
[1] 10.06019
> var(rm_canola$Yield)
[1] 101.2074
> 
> # (b) 90th percentile of the yields
> quantile(rm_canola$Yield, probs = 0.9)
 90% 
42.5 
> 
> ## (a) I find a mean of 28.29, a median of 26.9, a standard deviation of 10.06, and a variance of 101.21 -- the sd squared, in squared units.
> ## (b) The 90th percentile is 42.5 bu/ac.
> ## (c) A mean above the median suggests mild right skew: the very best RM-years pull the mean up more than the worst years pull it down.

Question 13

This question uses: rm_canola_yields_1990_2025.csv

(a) Print the first six rows of rm_canola with head().

(b) Report the number of rows in the data, and compute the first and last year in the data.

(c) There are 295 RMs in Saskatchewan. Divide the row count by the number of years. What does this tell you about the coverage of the dataset?

Answer
# (a) First six rows
head(rm_canola)

# (b) Rows, and the span of years
nrow(rm_canola)
range(rm_canola$Year)

## (b) I find 10,039 rows spanning 1990 to 2025.

# (c) Rows per year
nrow(rm_canola) / 36

## (c) 10,039 / 36 is about 279 rows per year, against the 295 RMs --
## so not every RM reports canola every year; some RM-year combinations
## are absent from the file.
> # (a) First six rows
> head(rm_canola)
# A tibble: 6 × 5
   Year    RM Crop   Yield Unit 
  <dbl> <dbl> <chr>  <dbl> <chr>
1  1990     1 Canola  22   bu/ac
2  1991     1 Canola  24.1 bu/ac
3  1992     1 Canola  15   bu/ac
4  1993     1 Canola  25.3 bu/ac
5  1994     1 Canola  23.1 bu/ac
6  1995     1 Canola  15.9 bu/ac
> 
> # (b) Rows, and the span of years
> nrow(rm_canola)
[1] 10039
> range(rm_canola$Year)
[1] 1990 2025
> 
> ## (b) I find 10,039 rows spanning 1990 to 2025.
> 
> # (c) Rows per year
> nrow(rm_canola) / 36
[1] 278.8611
> 
> ## (c) 10,039 / 36 is about 279 rows per year, against the 295 RMs --
> ## so not every RM reports canola every year; some RM-year combinations
> ## are absent from the file.

Question 14

This question uses: rm_canola_yields_1990_2025.csv

(a) Using $, compute the mean, the standard deviation, and the 10th and 90th percentiles of Yield across the whole file.

(b) In a comment: an RM-year at the 10th percentile grew how much canola? Interpret the number for a grower.

(c) Compute the gap between the 90th and 10th percentiles. In a comment: what does this gap describe that the standard deviation also tries to describe?

Answer
# (a) Mean, sd, and the 10th and 90th percentiles
mean(rm_canola$Yield)
sd(rm_canola$Yield)
quantile(rm_canola$Yield, probs = c(0.10, 0.90))

## I find a mean of 28.29, an sd of 10.06, and percentiles of 16.1 and 42.5.

## (b) An RM-year at the 10th percentile grew 16.1 bu/ac -- nine in ten
## RM-years across the file were at or above that value.

# (c) The 90-10 gap
42.5 - 16.1

## (c) The gap is 26.4 bu/ac. Like the sd it describes how spread out
## yields are, but it reads directly as a bushel distance holding the
## middle 80% of RM-years, and it ignores the extremes entirely.
> # (a) Mean, sd, and the 10th and 90th percentiles
> mean(rm_canola$Yield)
[1] 28.29239
> sd(rm_canola$Yield)
[1] 10.06019
> quantile(rm_canola$Yield, probs = c(0.10, 0.90))
 10%  90% 
16.1 42.5 
> 
> ## I find a mean of 28.29, an sd of 10.06, and percentiles of 16.1 and 42.5.
> 
> ## (b) An RM-year at the 10th percentile grew 16.1 bu/ac -- nine in ten
> ## RM-years across the file were at or above that value.
> 
> # (c) The 90-10 gap
> 42.5 - 16.1
[1] 26.4
> 
> ## (c) The gap is 26.4 bu/ac. Like the sd it describes how spread out
> ## yields are, but it reads directly as a bushel distance holding the
> ## middle 80% of RM-years, and it ignores the extremes entirely.

Question 15

This question uses: rm_canola_yields_1990_2025.csv

(a) Compute the 25th, 50th, and 75th percentiles of Yield in one quantile() call.

(b) Compute the IQR twice: with IQR(), and as the difference between two of your percentiles. Confirm the two match.

(c) In a comment: what fraction of all observations lies between your 25th and 75th percentiles, and how would you describe those observations?

Answer
# Load the tidyverse and read the canola file
library(tidyverse)
rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")

# (a) The three quartiles in one call
quantile(rm_canola$Yield, probs = c(0.25, 0.50, 0.75))

# (b) The IQR, from the function
IQR(rm_canola$Yield)

# The IQR, as the 75th minus the 25th percentile
quantile(rm_canola$Yield, probs = 0.75) - quantile(rm_canola$Yield, probs = 0.25)

## I find quartiles of 21.0, 26.9, and 35.4 bu/ac, and an IQR of 14.4
## both ways.
## (c) Half of all observations lie between the two percentiles. They
## are the middle, typical yields: the lowest quarter and the highest
## quarter of observations sit outside the range.
> # Load the tidyverse and read the canola file
> library(tidyverse)
> rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")
Rows: 10039 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) The three quartiles in one call
> quantile(rm_canola$Yield, probs = c(0.25, 0.50, 0.75))
 25%  50%  75% 
21.0 26.9 35.4 
> 
> # (b) The IQR, from the function
> IQR(rm_canola$Yield)
[1] 14.4
> 
> # The IQR, as the 75th minus the 25th percentile
> quantile(rm_canola$Yield, probs = 0.75) - quantile(rm_canola$Yield, probs = 0.25)
 75% 
14.4 
> 
> ## I find quartiles of 21.0, 26.9, and 35.4 bu/ac, and an IQR of 14.4
> ## both ways.
> ## (c) Half of all observations lie between the two percentiles. They
> ## are the middle, typical yields: the lowest quarter and the highest
> ## quarter of observations sit outside the range.

Manitoba Wheat Variety

Question 16

This question uses: mb_wheat_reported_2020_2025.csv

(a) Run summary(mb_wheat). From its output, report the minimum, median, and maximum of Yield_bu_ac, and the first and last Year.

(b) In a comment: are most yields in a plausible range for spring wheat, and do the extremes need investigation? Are there missing values?

Answer
# Load the tidyverse and read the wheat file
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Summary of every column
summary(mb_wheat)

## (a) I find Yield_bu_ac has minimum 4.5, median 62.2, and maximum 97.8, and Year runs from 2020 to 2025.
## (b) Most yields are plausible for Manitoba spring wheat, but the lowest
## values deserve investigation. There are no missing values in this file.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Summary of every column
> summary(mb_wheat)
      Year         Municipality       Variety         Farms       
 Min.   :2020   Length   :2398   Length   :2398   Min.   :  3.00  
 1st Qu.:2021   N.unique :  96   N.unique :  44   1st Qu.:  4.00  
 Median :2023   N.blank  :   0   N.blank  :   0   Median :  8.00  
 Mean   :2023   Min.nchar:   4   Min.nchar:   5   Mean   : 15.42  
 3rd Qu.:2024   Max.nchar:  28   Max.nchar:  38   3rd Qu.: 21.00  
 Max.   :2025                                     Max.   :148.00  
     Acres        Yield_bu_ac    Reported      
 Min.   :  501   Min.   : 4.50   Mode:logical  
 1st Qu.: 1354   1st Qu.:53.90   TRUE:2398     
 Median : 2941   Median :62.20                 
 Mean   : 6118   Mean   :61.16                 
 3rd Qu.: 7721   3rd Qu.:69.60                 
 Max.   :74604   Max.   :97.80                 
> 
> ## (a) I find Yield_bu_ac has minimum 4.5, median 62.2, and maximum 97.8, and Year runs from 2020 to 2025.
> ## (b) Most yields are plausible for Manitoba spring wheat, but the lowest
> ## values deserve investigation. There are no missing values in this file.

Question 17

This question uses: mb_wheat_reported_2020_2025.csv

(a) Compute the mean, median, standard deviation, and variance of mb_wheat$Yield_bu_ac.

(b) In a comment: what does that difference between the mean and median suggest about the data?

(c) Compute the smallest and largest yields in the data.

Answer
# Load the tidyverse and read the wheat file
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Centre of the yields
mean(mb_wheat$Yield_bu_ac)
median(mb_wheat$Yield_bu_ac)

# Spread of the yields
sd(mb_wheat$Yield_bu_ac)
var(mb_wheat$Yield_bu_ac)

## (a) I find a mean of 61.16, a median of 62.2, a standard deviation of 12.38, and a variance of 153.34.

## (b) A mean below the median suggests mild left skew: unusually low values pull the mean down more than the high values pull it up -- a tail of very poor variety-municipality observations mixed in with the normal ones.

# (c) Smallest and largest yield
range(mb_wheat$Yield_bu_ac)

## (c) range() returns both ends as one vector: the smallest yield is 4.5 and the largest is 97.8 bu/ac.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Centre of the yields
> mean(mb_wheat$Yield_bu_ac)
[1] 61.15617
> median(mb_wheat$Yield_bu_ac)
[1] 62.2
> 
> # Spread of the yields
> sd(mb_wheat$Yield_bu_ac)
[1] 12.38297
> var(mb_wheat$Yield_bu_ac)
[1] 153.338
> 
> ## (a) I find a mean of 61.16, a median of 62.2, a standard deviation of 12.38, and a variance of 153.34.
> 
> ## (b) A mean below the median suggests mild left skew: unusually low values pull the mean down more than the high values pull it up -- a tail of very poor variety-municipality observations mixed in with the normal ones.
> 
> # (c) Smallest and largest yield
> range(mb_wheat$Yield_bu_ac)
[1]  4.5 97.8
> 
> ## (c) range() returns both ends as one vector: the smallest yield is 4.5 and the largest is 97.8 bu/ac.

Question 18

This question uses: mb_wheat_reported_2020_2025.csv

(a) Compute the 25th, 50th, and 75th percentiles of Yield_bu_ac in one quantile() call.

(b) Compute the IQR as the difference between two of your percentiles.

(c) Compare the IQR to the range. Which might give a better sense of the spread of the data?

Answer
# Load the tidyverse and read the wheat file
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) The three quartiles in one call
quantile(mb_wheat$Yield_bu_ac, probs = c(0.25, 0.50, 0.75))

# (b) The IQR, as the 75th minus the 25th percentile
quantile(mb_wheat$Yield_bu_ac, probs = 0.75) - quantile(mb_wheat$Yield_bu_ac, probs = 0.25)

## I find quartiles of 53.9, 62.2, and 69.6 bu/ac, and an IQR of 15.7.

# (c) The range, for comparison
max(mb_wheat$Yield_bu_ac) - min(mb_wheat$Yield_bu_ac)

## (c) The range is 93.3 bu/ac against an IQR of 15.7. The IQR usually
## gives the better sense of spread: the range rests on the single best
## and single worst results, while the IQR is the width of the middle
## half of the data.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) The three quartiles in one call
> quantile(mb_wheat$Yield_bu_ac, probs = c(0.25, 0.50, 0.75))
 25%  50%  75% 
53.9 62.2 69.6 
> 
> # (b) The IQR, as the 75th minus the 25th percentile
> quantile(mb_wheat$Yield_bu_ac, probs = 0.75) - quantile(mb_wheat$Yield_bu_ac, probs = 0.25)
 75% 
15.7 
> 
> ## I find quartiles of 53.9, 62.2, and 69.6 bu/ac, and an IQR of 15.7.
> 
> # (c) The range, for comparison
> max(mb_wheat$Yield_bu_ac) - min(mb_wheat$Yield_bu_ac)
[1] 93.3
> 
> ## (c) The range is 93.3 bu/ac against an IQR of 15.7. The IQR usually
> ## gives the better sense of spread: the range rests on the single best
> ## and single worst results, while the IQR is the width of the middle
> ## half of the data.

Question 19

This question uses: mb_wheat_reported_2020_2025.csv

(a) Compute the 5th and 95th percentiles of Yield_bu_ac in one call.

(b) In a comment: interpret the 95th percentile for a grower.

(c) Compute the gap between the two percentiles. In a comment: what fraction of all results does this span hold?

Answer
# Load the tidyverse and read the wheat file
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) The 5th and 95th percentiles in one call
quantile(mb_wheat$Yield_bu_ac, probs = c(0.05, 0.95))

## (a) I find percentiles of 38.8 and 78.8 bu/ac.

## (b) A result at the 95th percentile yielded 78.8 bu/ac -- only one in
## twenty municipality-variety results beat that.

# (c) The gap between them
quantile(mb_wheat$Yield_bu_ac, probs = 0.95) - quantile(mb_wheat$Yield_bu_ac, probs = 0.05)

## (c) The gap is about 40 bu/ac, and it holds the middle 90% of results.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) The 5th and 95th percentiles in one call
> quantile(mb_wheat$Yield_bu_ac, probs = c(0.05, 0.95))
    5%    95% 
38.785 78.800 
> 
> ## (a) I find percentiles of 38.8 and 78.8 bu/ac.
> 
> ## (b) A result at the 95th percentile yielded 78.8 bu/ac -- only one in
> ## twenty municipality-variety results beat that.
> 
> # (c) The gap between them
> quantile(mb_wheat$Yield_bu_ac, probs = 0.95) - quantile(mb_wheat$Yield_bu_ac, probs = 0.05)
   95% 
40.015 
> 
> ## (c) The gap is about 40 bu/ac, and it holds the middle 90% of results.

Question 20

This question uses: mb_wheat_reported_2020_2025.csv

(a) Compute the range of Yield_bu_ac (largest minus smallest) and the IQR.

(b) In a comment: the two numbers disagree by a lot. What does each one describe?

(c) In a comment: which of the two measures would change if we discovered that the highest value in the data was actually a data error?

Answer
# Load the tidyverse and read the wheat file
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) The range
max(mb_wheat$Yield_bu_ac) - min(mb_wheat$Yield_bu_ac)

# The IQR
IQR(mb_wheat$Yield_bu_ac)

## I find a range of 93.3 bu/ac against an IQR of 15.7.
## (b) The range is the distance between the single best and single worst
## results; the IQR is the width of the middle half.
## (c) Only the range. The 97.8 is the maximum, so removing it moves the
## range directly, while the quartiles sit deep in the middle of 2,398
## results and barely notice.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) The range
> max(mb_wheat$Yield_bu_ac) - min(mb_wheat$Yield_bu_ac)
[1] 93.3
> 
> # The IQR
> IQR(mb_wheat$Yield_bu_ac)
[1] 15.7
> 
> ## I find a range of 93.3 bu/ac against an IQR of 15.7.
> ## (b) The range is the distance between the single best and single worst
> ## results; the IQR is the width of the middle half.
> ## (c) Only the range. The 97.8 is the maximum, so removing it moves the
> ## range directly, while the quartiles sit deep in the middle of 2,398
> ## results and barely notice.

Section 3 — Data Manipulation Functions

Saskatchewan RM Crop Yields

Question 21

This question uses: rm_yields_1990_2025.csv

(a) Filter to the Spring Wheat observations of at least 60 bu/ac (all years) and saving the result. How many rows are in the filtered data frame?

(b) Filter again to keep only the 2023 rows of that result. How many rows are in this second data frame?

(c) From your 2023 data frame, select RM and Yield and sort the data by yield (greatest to least). Which RM tops the list, at what yield?

Answer
# Load the tidyverse
library(tidyverse)

# Read the long yield file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Spring Wheat at 60 bu/ac or better, all years
sw_60 <- filter(rm_yields, Crop == "Spring Wheat" & Yield >= 60)
# Row count of the filtered data frame
nrow(sw_60)

## (a) I find 434 rows at 60 bu/ac or better.

# (b) Keep only the 2023 rows of that result
sw_60_2023 <- filter(sw_60, Year == 2023)
# Row count of the second data frame
nrow(sw_60_2023)

## (b) 53 of them are from 2023.

# (c) Just RM and Yield, sorted best first
sw_best <- select(sw_60_2023, RM, Yield)
arrange(sw_best, desc(Yield))

## (c) RM 187 tops the 2023 list, at 74.7 bu/ac.
> # Load the tidyverse
> library(tidyverse)
> 
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Spring Wheat at 60 bu/ac or better, all years
> sw_60 <- filter(rm_yields, Crop == "Spring Wheat" & Yield >= 60)
> # Row count of the filtered data frame
> nrow(sw_60)
[1] 434
> 
> ## (a) I find 434 rows at 60 bu/ac or better.
> 
> # (b) Keep only the 2023 rows of that result
> sw_60_2023 <- filter(sw_60, Year == 2023)
> # Row count of the second data frame
> nrow(sw_60_2023)
[1] 53
> 
> ## (b) 53 of them are from 2023.
> 
> # (c) Just RM and Yield, sorted best first
> sw_best <- select(sw_60_2023, RM, Yield)
> arrange(sw_best, desc(Yield))
# A tibble: 53 × 2
      RM Yield
   <dbl> <dbl>
 1   187  74.7
 2   494  73  
 3   338  72.8
 4   368  72.4
 5   464  72.2
 6   555  72.2
 7   520  72  
 8   304  71.5
 9   496  70.9
10   588  70.9
# ℹ 43 more rows
> 
> ## (c) RM 187 tops the 2023 list, at 74.7 bu/ac.

Question 22

This question uses: rm_yields_1990_2025.csv

(a) Filter to rows where the crop is Oats or Barley, using |, saving the result. How many rows are in the filtered data frame?

(b) Write the same filter using %in% and confirm the row count matches.

(c) What would filter(rm_yields, Crop == "Oats" & Crop == "Barley") return? In a comment: explain why.

Answer
# Load the tidyverse
library(tidyverse)

# Read the long yield file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Oats or Barley, using the | connector
oats_barley <- filter(rm_yields, Crop == "Oats" | Crop == "Barley")
# Row count of the filtered data frame
nrow(oats_barley)

# (b) The same filter written with %in%
oats_barley2 <- filter(rm_yields, Crop %in% c("Oats", "Barley"))
# The count should match (a)
nrow(oats_barley2)

## I find 19,709 rows both ways.

## (c) It returns zero rows. The & version asks for rows whose crop is
## Oats and Barley at the same time, and no single row can satisfy both
## conditions. "Or" is the right connector, written |.
> # Load the tidyverse
> library(tidyverse)
> 
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Oats or Barley, using the | connector
> oats_barley <- filter(rm_yields, Crop == "Oats" | Crop == "Barley")
> # Row count of the filtered data frame
> nrow(oats_barley)
[1] 19709
> 
> # (b) The same filter written with %in%
> oats_barley2 <- filter(rm_yields, Crop %in% c("Oats", "Barley"))
> # The count should match (a)
> nrow(oats_barley2)
[1] 19709
> 
> ## I find 19,709 rows both ways.
> 
> ## (c) It returns zero rows. The & version asks for rows whose crop is
> ## Oats and Barley at the same time, and no single row can satisfy both
> ## conditions. "Or" is the right connector, written |.

Question 23

This question uses: rm_yields_1990_2025.csv

(a) Filter to Canola in 2024, then in one mutate() call add two columns: yield_t_ha (multiply by 0.0560) and yield_kg_ha (multiply by 56.0). Save the result of each step.

(b) Compute the mean of each new column.

(c) In a comment: without running anything, what is the mean of yield_kg_ha divided by the mean of yield_t_ha, and why?

Answer
# Load the tidyverse
library(tidyverse)

# Read the long yield file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Canola in 2024
canola_2024 <- filter(rm_yields, Crop == "Canola" & Year == 2024)
# Add both converted columns in one mutate() call
canola_2024 <- mutate(canola_2024,
                      yield_t_ha  = Yield * 0.0560,  # tonnes per hectare
                      yield_kg_ha = Yield * 56.0)    # kilograms per hectare

# (b) Mean of each new column
mean(canola_2024$yield_t_ha)
mean(canola_2024$yield_kg_ha)

## (b) I find means of 1.76 t/ha and 1764.11 kg/ha.

## (c) The ratio of the means is 1000: both columns are the same yields scaled by constants, and 56.0 / 0.0560 = 1000 (a tonne is 1000 kg).
> # Load the tidyverse
> library(tidyverse)
> 
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Canola in 2024
> canola_2024 <- filter(rm_yields, Crop == "Canola" & Year == 2024)
> # Add both converted columns in one mutate() call
> canola_2024 <- mutate(canola_2024,
+                       yield_t_ha  = Yield * 0.0560,  # tonnes per hectare
+                       yield_kg_ha = Yield * 56.0)    # kilograms per hectare
> 
> # (b) Mean of each new column
> mean(canola_2024$yield_t_ha)
[1] 1.764115
> mean(canola_2024$yield_kg_ha)
[1] 1764.115
> 
> ## (b) I find means of 1.76 t/ha and 1764.11 kg/ha.
> 
> ## (c) The ratio of the means is 1000: both columns are the same yields scaled by constants, and 56.0 / 0.0560 = 1000 (a tonne is 1000 kg).

Question 24

This question uses: rm_yields_1990_2025.csv

Canola is selling at $14.20 per bushel.

(a) Filter to Canola in 2025 and, in one mutate(), add a revenue_per_ac column (yield times price). Save the result.

(b) Sort by revenue (greatest to least), and report the top RM and its revenue per acre.

(c) Compute the mean revenue per acre.

Answer
# Load the tidyverse and read the data
library(tidyverse)
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# Price object
price <- 14.20

# (a) 2025 canola with revenue per acre
canola_25 <- filter(rm_yields, Crop == "Canola" & Year == 2025)
canola_25 <- mutate(canola_25, revenue_per_ac = Yield * price)

# (b) Best revenue first
arrange(canola_25, desc(revenue_per_ac))

## (b) I find RM 287 on top at $866.20 per acre.

# (c) Mean revenue per acre
mean(canola_25$revenue_per_ac)

## (c) The mean is $623.94 per acre.
> # Load the tidyverse and read the data
> library(tidyverse)
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # Price object
> price <- 14.20
> 
> # (a) 2025 canola with revenue per acre
> canola_25 <- filter(rm_yields, Crop == "Canola" & Year == 2025)
> canola_25 <- mutate(canola_25, revenue_per_ac = Yield * price)
> 
> # (b) Best revenue first
> arrange(canola_25, desc(revenue_per_ac))
# A tibble: 290 × 6
    Year    RM Crop   Yield Unit  revenue_per_ac
   <dbl> <dbl> <chr>  <dbl> <chr>          <dbl>
 1  2025   287 Canola  61   bu/ac           866.
 2  2025   292 Canola  57.3 bu/ac           814.
 3  2025   257 Canola  56.7 bu/ac           805.
 4  2025   259 Canola  56   bu/ac           795.
 5  2025   380 Canola  55.6 bu/ac           790.
 6  2025   128 Canola  55.4 bu/ac           787.
 7  2025   345 Canola  54.9 bu/ac           780.
 8  2025   225 Canola  54.8 bu/ac           778.
 9  2025   276 Canola  54.6 bu/ac           775.
10  2025   260 Canola  54.5 bu/ac           774.
# ℹ 280 more rows
> 
> ## (b) I find RM 287 on top at $866.20 per acre.
> 
> # (c) Mean revenue per acre
> mean(canola_25$revenue_per_ac)
[1] 623.9431
> 
> ## (c) The mean is $623.94 per acre.

Question 25

This question uses: rm_yields_1990_2025.csv

(a) Filter rm_yields to Canola in 2024, saving the result as canola_2024. How many rows are in the filtered data frame?

(b) From canola_2024, select only the RM and Yield columns, saving the result. Report its column names and its number of columns.

(c) In a comment: after (a) and (b), how many rows does rm_yields itself have, and why?

Answer
# Load the tidyverse
library(tidyverse)

# Read the long yield file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Keep only the Canola rows from 2024
canola_2024 <- filter(rm_yields, Crop == "Canola" & Year == 2024)
# Row count of the filtered data frame
nrow(canola_2024)

# (b) Keep only the RM and Yield columns
canola_slim <- select(canola_2024, RM, Yield)
# Column names of the slimmed version
names(canola_slim)
# Number of columns
ncol(canola_slim)

# (c) The original data frame, after (a) and (b)
nrow(rm_yields)

## I find 293 rows of 2024 canola.
## The selected version has the columns RM and Yield, so 2 columns (and still 293 rows).
## rm_yields itself still has 71104 rows: filter() and select() return new data frames and leave the original untouched.
> # Load the tidyverse
> library(tidyverse)
> 
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Keep only the Canola rows from 2024
> canola_2024 <- filter(rm_yields, Crop == "Canola" & Year == 2024)
> # Row count of the filtered data frame
> nrow(canola_2024)
[1] 293
> 
> # (b) Keep only the RM and Yield columns
> canola_slim <- select(canola_2024, RM, Yield)
> # Column names of the slimmed version
> names(canola_slim)
[1] "RM"    "Yield"
> # Number of columns
> ncol(canola_slim)
[1] 2
> 
> # (c) The original data frame, after (a) and (b)
> nrow(rm_yields)
[1] 71104
> 
> ## I find 293 rows of 2024 canola.
> ## The selected version has the columns RM and Yield, so 2 columns (and still 293 rows).
> ## rm_yields itself still has 71104 rows: filter() and select() return new data frames and leave the original untouched.

Manitoba Wheat Variety

Question 26

This question uses: mb_wheat_reported_2020_2025.csv

(a) Filter to rows with at least 10 farms and a yield above 70 bu/ac, saving the result. How many rows are in it?

(b) Sort by yields (greatest to least) and report the top municipality, variety, and year.

(c) In a comment: compare this with filtering on yield alone without the minimum number of farms. Comment on the trustworthiness of these results versus those in part b.

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) At least 10 farms AND yield above 70
solid_high <- filter(mb_wheat, Farms >= 10 & Yield_bu_ac > 70)
nrow(solid_high)

# (b) Best first
arrange(solid_high, desc(Yield_bu_ac))

## I find 273 rows; LOUISE growing SY MANNESS in 2025 tops the list
## at 96.3 bu/ac.

# (c) The same filter on yield alone
high_only <- filter(mb_wheat, Yield_bu_ac > 70)
nrow(high_only)

## (c) Filtering on yield alone keeps 569 rows, and its top result is a
## 97.8 bu/ac average from just 5 farms. An average from a few farms is
## more sensitive to one unusual result, so the part b list is the more
## trustworthy ranking; the farm minimum gives every row a broader base.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) At least 10 farms AND yield above 70
> solid_high <- filter(mb_wheat, Farms >= 10 & Yield_bu_ac > 70)
> nrow(solid_high)
[1] 273
> 
> # (b) Best first
> arrange(solid_high, desc(Yield_bu_ac))
# A tibble: 273 × 7
    Year Municipality      Variety              Farms Acres Yield_bu_ac Reported
   <dbl> <chr>             <chr>                <dbl> <dbl>       <dbl> <lgl>   
 1  2025 LOUISE            SY MANNESS              14  3252        96.3 TRUE    
 2  2025 GRANDVIEW         AAC STARBUCK <SECAN>    10  4456        90.2 TRUE    
 3  2025 ROLAND            SY MANNESS              11  2986        88.6 TRUE    
 4  2025 STE. ANNE         AAC BRANDON (BW 932)    17  2445        87.9 TRUE    
 5  2024 LOUISE            AAC WHEATLAND <SECA…    19  7576        86.9 TRUE    
 6  2025 GRANDVIEW         AAC BRANDON (BW 932)    26 12672        86.4 TRUE    
 7  2024 MORRIS            SY MANNESS              28  5288        84.6 TRUE    
 8  2025 LOUISE            AAC STARBUCK <SECAN>    28 12761        84.4 TRUE    
 9  2024 MACDONALD         SY MANNESS              21  7776        84.1 TRUE    
10  2025 RUSSELL-BINSCARTH AAC WHEATLAND <SECA…    14  7644        84   TRUE    
# ℹ 263 more rows
> 
> ## I find 273 rows; LOUISE growing SY MANNESS in 2025 tops the list
> ## at 96.3 bu/ac.
> 
> # (c) The same filter on yield alone
> high_only <- filter(mb_wheat, Yield_bu_ac > 70)
> nrow(high_only)
[1] 569
> 
> ## (c) Filtering on yield alone keeps 569 rows, and its top result is a
> ## 97.8 bu/ac average from just 5 farms. An average from a few farms is
> ## more sensitive to one unusual result, so the part b list is the more
> ## trustworthy ranking; the farm minimum gives every row a broader base.

Question 27

This question uses: mb_wheat_reported_2020_2025.csv

(a) Filter the data to the 2025 rows, saving the result. How many rows are in the filtered data frame?

(b) From the filtered data, select Municipality, Variety, and Yield_bu_ac, saving the result. Report its column names.

(c) In a comment: how many rows does the original data frame still have, and what happened to the other years’ rows?

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Keep the 2025 rows only
wheat_2025 <- filter(mb_wheat, Year == 2025)
# Row count of the filtered data frame
nrow(wheat_2025)

# (b) Three columns from the 2025 rows
wheat_slim <- select(wheat_2025, Municipality, Variety, Yield_bu_ac)
# Its column names
names(wheat_slim)

# (c) The original data frame, after (a) and (b)
nrow(mb_wheat)

## I find 438 rows for 2025.
## The selected columns are Municipality, Variety, Yield_bu_ac.
## mb_wheat itself still has 2398 rows -- filtering returns a new data frame. The other years' rows are simply absent from the filtered copy; they were not deleted from anything.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Keep the 2025 rows only
> wheat_2025 <- filter(mb_wheat, Year == 2025)
> # Row count of the filtered data frame
> nrow(wheat_2025)
[1] 438
> 
> # (b) Three columns from the 2025 rows
> wheat_slim <- select(wheat_2025, Municipality, Variety, Yield_bu_ac)
> # Its column names
> names(wheat_slim)
[1] "Municipality" "Variety"      "Yield_bu_ac" 
> 
> # (c) The original data frame, after (a) and (b)
> nrow(mb_wheat)
[1] 2398
> 
> ## I find 438 rows for 2025.
> ## The selected columns are Municipality, Variety, Yield_bu_ac.
> ## mb_wheat itself still has 2398 rows -- filtering returns a new data frame. The other years' rows are simply absent from the filtered copy; they were not deleted from anything.

Question 28

This question uses: mb_wheat_reported_2020_2025.csv

(a) Filter the data to the varieties AAC BRANDON (BW 932) or AAC STARBUCK <SECAN> using |, saving the result. How many rows are in the filtered data frame?

(b) Write the same filter using %in% and confirm the count matches.

(c) If in part (a) you used & instead of |, what would the result be? Explain.

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) The two varieties, using the | connector
two_var <- filter(mb_wheat, Variety == "AAC BRANDON (BW 932)" | Variety == "AAC STARBUCK <SECAN>")
# Row count of the filtered data frame
nrow(two_var)

# (b) The same filter written with %in%
two_var2 <- filter(mb_wheat, Variety %in% c("AAC BRANDON (BW 932)", "AAC STARBUCK <SECAN>"))
# The count should match (a)
nrow(two_var2)

## I find 893 rows both ways.
## (c) It would return zero rows. The & version demands both conditions true of the same row, and no single row has a variety that is both names at once. "Or" is the right connector, written |.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) The two varieties, using the | connector
> two_var <- filter(mb_wheat, Variety == "AAC BRANDON (BW 932)" | Variety == "AAC STARBUCK <SECAN>")
> # Row count of the filtered data frame
> nrow(two_var)
[1] 893
> 
> # (b) The same filter written with %in%
> two_var2 <- filter(mb_wheat, Variety %in% c("AAC BRANDON (BW 932)", "AAC STARBUCK <SECAN>"))
> # The count should match (a)
> nrow(two_var2)
[1] 893
> 
> ## I find 893 rows both ways.
> ## (c) It would return zero rows. The & version demands both conditions true of the same row, and no single row has a variety that is both names at once. "Or" is the right connector, written |.

Question 29

This question uses: mb_wheat_reported_2020_2025.csv

(a) Use one mutate() to add acres_per_farm (acres divided by farms), saving the result.

(b) Sort by acres_per_farm, largest first, and report the top municipality and variety.

(c) Compute the mean of acres_per_farm. In a comment, explain what this number means in the context of the data.

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Acres per farm
mb_wheat <- mutate(mb_wheat, acres_per_farm = Acres / Farms)

# (b) Largest first
arrange(mb_wheat, desc(acres_per_farm))

## (b) I find MOUNTAIN growing SY MANNESS on top, at 1,587 acres per farm.

# (c) The average across rows
mean(mb_wheat$acres_per_farm)

## (c) The mean is 370: in a typical municipality-variety-year result,
## each farm growing that variety seeded about 370 acres of it. It
## measures how concentrated a variety is in the hands of its growers --
## a different fact from how much land it covers in total.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Acres per farm
> mb_wheat <- mutate(mb_wheat, acres_per_farm = Acres / Farms)
> 
> # (b) Largest first
> arrange(mb_wheat, desc(acres_per_farm))
# A tibble: 2,398 × 8
    Year Municipality    Variety Farms Acres Yield_bu_ac Reported acres_per_farm
   <dbl> <chr>           <chr>   <dbl> <dbl>       <dbl> <lgl>             <dbl>
 1  2025 MOUNTAIN        SY MAN…     3  4761        77.8 TRUE              1587 
 2  2023 DELORAINE-WINC… AAC EL…     4  5994        82.2 TRUE              1498.
 3  2025 WESTLAKE-GLADS… AAC WH…     3  4483        68.6 TRUE              1494.
 4  2025 NORTH CYPRESS-… AAC WH…     5  7228        76.6 TRUE              1446.
 5  2025 ALONSA          AAC HO…     5  6469        45.1 TRUE              1294.
 6  2025 GLENELLA-LANSD… AAC WH…     3  3504        78.5 TRUE              1168 
 7  2023 MINITONAS-BOWS… AAC BR…     4  4648        66.1 TRUE              1162 
 8  2021 DELORAINE-WINC… AAC EL…     5  5763        62.5 TRUE              1153.
 9  2020 ELLICE-ARCHIE   AAC VI…     4  4606        63.2 TRUE              1152.
10  2025 ROCKWOOD        BOLLES…     4  4363        66.6 TRUE              1091.
# ℹ 2,388 more rows
> 
> ## (b) I find MOUNTAIN growing SY MANNESS on top, at 1,587 acres per farm.
> 
> # (c) The average across rows
> mean(mb_wheat$acres_per_farm)
[1] 370.0478
> 
> ## (c) The mean is 370: in a typical municipality-variety-year result,
> ## each farm growing that variety seeded about 370 acres of it. It
> ## measures how concentrated a variety is in the hands of its growers --
> ## a different fact from how much land it covers in total.

Question 30

This question uses: mb_wheat_reported_2020_2025.csv

(a) Use one mutate() to add total_bu, each variety’s total production in a municipality in bushels (acres times yield), saving the result.

(b) Sort by total_bu, largest first, and report the top municipality and variety.

(c) In a comment: why is the highest-yielding observation not the one with the greatest production?

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Total production: acres times the per-acre yield
mb_wheat <- mutate(mb_wheat, total_bu = Acres * Yield_bu_ac)

# (b) Largest first
arrange(mb_wheat, desc(total_bu))

## (b) I find PORTAGE LA PRAIRIE growing AAC BRANDON (BW 932) in 2020 on
## top, at about 5.1 million bushels.
## (c) Acreage. The top producer yielded a modest 68.8 bu/ac but on
## 74,604 acres, while the highest-yielding result (97.8 bu/ac) sat on
## under 2,000 acres. Total production rewards land as much as yield.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Total production: acres times the per-acre yield
> mb_wheat <- mutate(mb_wheat, total_bu = Acres * Yield_bu_ac)
> 
> # (b) Largest first
> arrange(mb_wheat, desc(total_bu))
# A tibble: 2,398 × 8
    Year Municipality       Variety   Farms  Acres Yield_bu_ac Reported total_bu
   <dbl> <chr>              <chr>     <dbl>  <dbl>       <dbl> <lgl>       <dbl>
 1  2020 PORTAGE LA PRAIRIE AAC BRAN…   122 74604         68.8 TRUE     5132755.
 2  2020 MACDONALD          AAC BRAN…   115 55104.        71.6 TRUE     3945418.
 3  2024 PORTAGE LA PRAIRIE AAC BRAN…    84 51923         67   TRUE     3478841 
 4  2022 MINITONAS-BOWSMAN  AAC VIEW…    74 42055         82.4 TRUE     3465332 
 5  2025 PORTAGE LA PRAIRIE AAC BRAN…    83 45210         75.3 TRUE     3404313 
 6  2023 PORTAGE LA PRAIRIE AAC BRAN…    85 48876         67.5 TRUE     3299130 
 7  2022 PORTAGE LA PRAIRIE AAC BRAN…    86 51489         63.1 TRUE     3248956.
 8  2020 WESTLAKE-GLADSTONE AAC BRAN…    75 49160         65.5 TRUE     3219980 
 9  2022 SWAN VALLEY WEST   AAC VIEW…    71 42788         74.6 TRUE     3191985.
10  2020 WALLACE-WOODWORTH  AAC BRAN…   106 48658         65.2 TRUE     3172502.
# ℹ 2,388 more rows
> 
> ## (b) I find PORTAGE LA PRAIRIE growing AAC BRANDON (BW 932) in 2020 on
> ## top, at about 5.1 million bushels.
> ## (c) Acreage. The top producer yielded a modest 68.8 bu/ac but on
> ## 74,604 acres, while the highest-yielding result (97.8 bu/ac) sat on
> ## under 2,000 acres. Total production rewards land as much as yield.

Section 4 — Pipelines and Grouped Summaries

Saskatchewan RM Crop Yields

Question 31

This question uses: rm_yields_1990_2025.csv

(a) In one pipeline: for 2025 only, compute each crop’s mean yield and observation count, sorted from highest mean to lowest, and save the result as crop_summary_2025.

(b) Write the table to output/crop_summary_2025.csv.

(c) Report the two highest-yielding crops in the sorted table. In a comment: check the Unit column – why is the first-place crop not comparable to the rest, and which bushel crop is the real standout?

Answer
# Load the tidyverse
library(tidyverse)

# Read the long yield file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Each crop in 2025: mean yield and observation count, best first
crop_summary_2025 <- rm_yields |>
  filter(Year == 2025) |>                      # 2025 only
  group_by(Crop) |>                            # one group per crop
  summarise(mean_yield = mean(Yield),          # average yield
            n = n()) |>                        # observation count
  arrange(desc(mean_yield))                    # highest mean first

# Print the sorted table
crop_summary_2025

# (b) Write the table to a csv
write_csv(crop_summary_2025, "output/crop_summary_2025.csv")

## I find Lentils on top at 1897 and Oats second at 97.5.
## (c) The Unit column shows Lentils are recorded in pounds per acre, not
## bushels, so the first row is not comparable to the rest. Among the
## bushel crops, Oats are highest at 97.5 bu/ac.
> # Load the tidyverse
> library(tidyverse)
> 
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Each crop in 2025: mean yield and observation count, best first
> crop_summary_2025 <- rm_yields |>
+   filter(Year == 2025) |>                      # 2025 only
+   group_by(Crop) |>                            # one group per crop
+   summarise(mean_yield = mean(Yield),          # average yield
+             n = n()) |>                        # observation count
+   arrange(desc(mean_yield))                    # highest mean first
> 
> # Print the sorted table
> crop_summary_2025
# A tibble: 8 × 3
  Crop         mean_yield     n
  <chr>             <dbl> <int>
1 Lentils          1897.    218
2 Oats               97.5   199
3 Barley             72.8   288
4 Spring Wheat       53.0   283
5 Durum              48.6   189
6 Canola             43.9   290
7 Peas               43.5   285
8 Flax               30.5   195
> 
> # (b) Write the table to a csv
> write_csv(crop_summary_2025, "output/crop_summary_2025.csv")
> 
> ## I find Lentils on top at 1897 and Oats second at 97.5.
> ## (c) The Unit column shows Lentils are recorded in pounds per acre, not
> ## bushels, so the first row is not comparable to the rest. Among the
> ## bushel crops, Oats are highest at 97.5 bu/ac.

Question 32

This question uses: rm_yields_1990_2025.csv

(a) In one pipeline: filter to Lentils from 2021 through 2025, convert from lb/acre to kg/ha with mutate() (multiply by 1.12), and compute the mean converted yield and n() for each year. Sort from highest mean to lowest and save the result as lentil_summary.

(b) Report the highest and lowest annual means and the counts behind them.

(c) In a comment: in what order do your steps run, and why would putting the summarise() before the mutate() fail?

Answer
# Load the tidyverse
library(tidyverse)

# Read the long yield file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Recent lentils, converted to kg/ha and summarised by year
lentil_summary <- rm_yields |>
  filter(Crop == "Lentils" & Year >= 2021) |>    # recent lentil rows
  mutate(yield_kg_ha = Yield * 1.12) |>          # convert lb/ac to kg/ha
  group_by(Year) |>                               # one group per year
  summarise(mean_kg_ha = mean(yield_kg_ha),       # annual average
            n = n()) |>                           # observations behind it
  arrange(desc(mean_kg_ha))                       # highest mean first

# Print the summary table
lentil_summary

## (b) The highest annual mean is 2025 at 2,124.4 kg/ha from 218
## observations. The lowest is 2021 at 1,057.5 kg/ha from 194.
## (c) The steps filter, convert, group, summarise, and sort in that order.
## Putting summarise() before mutate() would fail because yield_kg_ha would
## not exist when the summary tried to use it.
> # Load the tidyverse
> library(tidyverse)
> 
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Recent lentils, converted to kg/ha and summarised by year
> lentil_summary <- rm_yields |>
+   filter(Crop == "Lentils" & Year >= 2021) |>    # recent lentil rows
+   mutate(yield_kg_ha = Yield * 1.12) |>          # convert lb/ac to kg/ha
+   group_by(Year) |>                               # one group per year
+   summarise(mean_kg_ha = mean(yield_kg_ha),       # annual average
+             n = n()) |>                           # observations behind it
+   arrange(desc(mean_kg_ha))                       # highest mean first
> 
> # Print the summary table
> lentil_summary
# A tibble: 5 × 3
   Year mean_kg_ha     n
  <dbl>      <dbl> <int>
1  2025      2124.   218
2  2024      1478.   213
3  2023      1472.   196
4  2022      1396.   196
5  2021      1058.   194
> 
> ## (b) The highest annual mean is 2025 at 2,124.4 kg/ha from 218
> ## observations. The lowest is 2021 at 1,057.5 kg/ha from 194.
> ## (c) The steps filter, convert, group, summarise, and sort in that order.
> ## Putting summarise() before mutate() would fail because yield_kg_ha would
> ## not exist when the summary tried to use it.

Question 33

This question uses: rm_yields_1990_2025.csv

Canola is selling at $14.20 per bushel.

(a) In one pipeline: filter to Canola observations in 2023-2025, add revenue_per_ac with mutate(), then compute the mean revenue for each year.

(b) Report the three yearly means.

Answer
# Load the tidyverse and read the data
library(tidyverse)
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# Price object
price <- 14.20

# (a) Mean canola revenue per acre by year, 2023-2025
rm_yields |>
  filter(Crop == "Canola" & Year >= 2023) |>      # recent canola rows
  mutate(revenue_per_ac = Yield * price) |>       # dollars per acre
  group_by(Year) |>                               # one group per year
  summarise(mean_revenue = mean(revenue_per_ac))  # average within each year

## (b) I find mean revenues of $481.11 in 2023, $447.33 in 2024, and
## $623.94 in 2025.
> # Load the tidyverse and read the data
> library(tidyverse)
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # Price object
> price <- 14.20
> 
> # (a) Mean canola revenue per acre by year, 2023-2025
> rm_yields |>
+   filter(Crop == "Canola" & Year >= 2023) |>      # recent canola rows
+   mutate(revenue_per_ac = Yield * price) |>       # dollars per acre
+   group_by(Year) |>                               # one group per year
+   summarise(mean_revenue = mean(revenue_per_ac))  # average within each year
# A tibble: 3 × 2
   Year mean_revenue
  <dbl>        <dbl>
1  2023         481.
2  2024         447.
3  2025         624.
> 
> ## (b) I find mean revenues of $481.11 in 2023, $447.33 in 2024, and
> ## $623.94 in 2025.

Question 34

This question uses: rm_yields_1990_2025.csv

(a) In one pipeline, compute the mean canola yield for each year from 2021 to 2025 (filter first, then group).

(b) 2021 was a severe drought year in Saskatchewan. Does the data agree?

Answer
# Load the tidyverse
library(tidyverse)

# Read the long yield file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Mean canola yield for each year from 2021 to 2025
rm_yields |>
  filter(Crop == "Canola" & Year >= 2021) |>   # canola, 2021 onward
  group_by(Year) |>                            # one group per year
  summarise(mean_yield = mean(Yield))          # average within each year

## (b) Yes, the data agrees: the 2021 mean sits roughly a third below the surrounding years.
> # Load the tidyverse
> library(tidyverse)
> 
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Mean canola yield for each year from 2021 to 2025
> rm_yields |>
+   filter(Crop == "Canola" & Year >= 2021) |>   # canola, 2021 onward
+   group_by(Year) |>                            # one group per year
+   summarise(mean_yield = mean(Yield))          # average within each year
# A tibble: 5 × 2
   Year mean_yield
  <dbl>      <dbl>
1  2021       21.9
2  2022       35.4
3  2023       33.9
4  2024       31.5
5  2025       43.9
> 
> ## (b) Yes, the data agrees: the 2021 mean sits roughly a third below the surrounding years.

Question 35

This question uses: rm_yields_1990_2025.csv

(a) In one pipeline: filter to Flax from 2020 through 2025, compute each RM’s mean_yield and number of observations, retain RMs with at least four observations, and sort from highest mean to lowest. Save the result as flax_summary.

(b) Report the top RM, its mean, and its count.

(c) In a comment: why does the minimum-count rule improve the ranking, and why must filter(n >= 4) come after summarise()?

Answer
# Load the tidyverse
library(tidyverse)

# Read the long yield file
rm_yields <- read_csv("data/rm_yields_1990_2025.csv")

# (a) Recent flax means by RM, retaining RMs with at least four reports
flax_summary <- rm_yields |>
  filter(Crop == "Flax" & Year >= 2020) |>    # recent flax rows
  group_by(RM) |>                              # one group per RM
  summarise(mean_yield = mean(Yield),          # average within the RM
            n = n()) |>                        # observations behind it
  filter(n >= 4) |>                            # enough years to compare
  arrange(desc(mean_yield))                    # highest mean first

# Print the summary table
flax_summary

## (b) RM 336 is highest at 34.5 bu/ac from 4 observations.
## (c) The rule removes rankings based on only one or two years. The count
## n is created by summarise(), so it cannot be filtered before that step.
> # Load the tidyverse
> library(tidyverse)
> 
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Recent flax means by RM, retaining RMs with at least four reports
> flax_summary <- rm_yields |>
+   filter(Crop == "Flax" & Year >= 2020) |>    # recent flax rows
+   group_by(RM) |>                              # one group per RM
+   summarise(mean_yield = mean(Yield),          # average within the RM
+             n = n()) |>                        # observations behind it
+   filter(n >= 4) |>                            # enough years to compare
+   arrange(desc(mean_yield))                    # highest mean first
> 
> # Print the summary table
> flax_summary
# A tibble: 191 × 3
      RM mean_yield     n
   <dbl>      <dbl> <int>
 1   336       34.5     4
 2   442       33.1     5
 3    64       32.6     6
 4    63       32.0     5
 5   400       31.4     6
 6   211       31.3     6
 7    61       30.9     6
 8   369       30.8     4
 9   438       30.4     6
10   351       29.8     6
# ℹ 181 more rows
> 
> ## (b) RM 336 is highest at 34.5 bu/ac from 4 observations.
> ## (c) The rule removes rankings based on only one or two years. The count
> ## n is created by summarise(), so it cannot be filtered before that step.

Manitoba Wheat Variety

Question 36

This question uses: mb_wheat_reported_2020_2025.csv

(a) In one pipeline, compute each variety’s total acres (sum(Acres)), sorted greatest to least.

(b) In a comment: totalling acres makes sense, but what would totalling Yield_bu_ac produce, and why is that number meaningless?

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Total acres by variety, largest first
mb_wheat |>
  group_by(Variety) |>                    # one group per variety
  summarise(total_acres = sum(Acres)) |>  # acres are amounts, so they add
  arrange(desc(total_acres))              # biggest variety first


## (b) Totalling Yield_bu_ac would add per-acre rates across rows -- a
## number with no physical meaning, since a rate says how much each acre
## gave, not how much there was. Acres are an amount, so their total is a
## real quantity of land.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Total acres by variety, largest first
> mb_wheat |>
+   group_by(Variety) |>                    # one group per variety
+   summarise(total_acres = sum(Acres)) |>  # acres are amounts, so they add
+   arrange(desc(total_acres))              # biggest variety first
# A tibble: 44 × 2
   Variety                                total_acres
   <chr>                                        <dbl>
 1 AAC BRANDON (BW 932)                      6741112.
 2 AAC STARBUCK <SECAN>                      2503939.
 3 AAC WHEATLAND <SECAN>                     1399093.
 4 AAC VIEWFIELD <FP GENETICS>|BW965| EXP    1114296.
 5 AAC HOCKLEY |BW5044|                       588044.
 6 AAC REDBERRY                               474429.
 7 SY MANNESS                                 389027 
 8 AAC ELIE(BW931)                            356550 
 9 BOLLES <SEED DEPOT>                        273594 
10 AAC HODGE |BW1069|                         176698.
# ℹ 34 more rows
> 
> 
> ## (b) Totalling Yield_bu_ac would add per-acre rates across rows -- a
> ## number with no physical meaning, since a rate says how much each acre
> ## gave, not how much there was. Acres are an amount, so their total is a
> ## real quantity of land.

Question 37

This question uses: mb_wheat_reported_2020_2025.csv

(a) In one pipeline, compute each variety’s mean yield and observation count, sorted by mean yield from greatest to least.

(b) Report the top two varieties, their means, and their counts.

(c) In a comment: Report the count of the top three varieties. Do you find the average yield of each variety trustworthy?

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Each variety: mean yield and observation count, best first
mb_wheat |>
  group_by(Variety) |>                         # one group per variety
  summarise(mean_yield = mean(Yield_bu_ac),    # average yield
            n = n()) |>                        # observation count
  arrange(desc(mean_yield))                    # highest mean first

## (b) I find AAC WESTKING first at 79.2 bu/ac from 17 observations, then CDC GO (BW781) at 77.5 from a single observation.
## (c) The top three counts are 17 (AAC WESTKING), 1 (CDC GO), and
## 2 (AAC SPIKE). I would trust the AAC WESTKING mean: seventeen
## municipality-year results is a reasonable base. The other two rest on
## one or two results, so their averages could be one municipality having
## a good year.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Each variety: mean yield and observation count, best first
> mb_wheat |>
+   group_by(Variety) |>                         # one group per variety
+   summarise(mean_yield = mean(Yield_bu_ac),    # average yield
+             n = n()) |>                        # observation count
+   arrange(desc(mean_yield))                    # highest mean first
# A tibble: 44 × 3
   Variety                 mean_yield     n
   <chr>                        <dbl> <int>
 1 AAC WESTKING                  79.2    17
 2 CDC GO (BW781)                77.5     1
 3 AAC SPIKE                     73.2     2
 4 SY MANNESS                    72.1   108
 5 AAC WHEATLAND <SECAN>         65.8   219
 6 AAC BROADACRES <PROVEN>       65.0    18
 7 AAC HOCKLEY |BW5044|          64.9   177
 8 GLENN                         64.5     4
 9 HARVEST (BW259)               64.4     1
10 AAC CONNERY |PT245|           64.1     2
# ℹ 34 more rows
> 
> ## (b) I find AAC WESTKING first at 79.2 bu/ac from 17 observations, then CDC GO (BW781) at 77.5 from a single observation.
> ## (c) The top three counts are 17 (AAC WESTKING), 1 (CDC GO), and
> ## 2 (AAC SPIKE). I would trust the AAC WESTKING mean: seventeen
> ## municipality-year results is a reasonable base. The other two rest on
> ## one or two results, so their averages could be one municipality having
> ## a good year.

Question 38

This question uses: mb_wheat_reported_2020_2025.csv

(a) In one pipeline: for 2025 only, compute each variety’s mean reported yield and observation count, sorted best-first, and save the result as variety_summary_2025.

(b) Write the table to output/variety_summary_2025.csv.

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Each variety in 2025: mean reported yield and count, best first
variety_summary_2025 <- mb_wheat |>
  filter(Year == 2025) |>                      # 2025 only
  group_by(Variety) |>                         # one group per variety
  summarise(mean_yield = mean(Yield_bu_ac),    # average yield
            n = n()) |>                        # observation count
  arrange(desc(mean_yield))                    # highest mean first

# Print the top of the sorted table
head(variety_summary_2025)

# (b) Write the table to a csv
write_csv(variety_summary_2025, "output/variety_summary_2025.csv")
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Each variety in 2025: mean reported yield and count, best first
> variety_summary_2025 <- mb_wheat |>
+   filter(Year == 2025) |>                      # 2025 only
+   group_by(Variety) |>                         # one group per variety
+   summarise(mean_yield = mean(Yield_bu_ac),    # average yield
+             n = n()) |>                        # observation count
+   arrange(desc(mean_yield))                    # highest mean first
> 
> # Print the top of the sorted table
> head(variety_summary_2025)
# A tibble: 6 × 3
  Variety                                mean_yield     n
  <chr>                                       <dbl> <int>
1 AAC WESTKING                                 79.2    17
2 AAC VIEWFIELD <FP GENETICS>|BW965| EXP       73.5    16
3 AAC SPIKE                                    73.2     2
4 SY MANNESS                                   72.6    62
5 CDC LANDMARK                                 72.3     1
6 AAC WHEATLAND <SECAN>                        71.5    50
> 
> # (b) Write the table to a csv
> write_csv(variety_summary_2025, "output/variety_summary_2025.csv")

Question 39

This question uses: mb_wheat_reported_2020_2025.csv

(a) In one pipeline, compute each municipality’s mean yield and observation count, sorted from highest mean to lowest.

(b) Report the top municipality and how many observations back its mean.

(c) In a comment: a municipality mean pools every variety grown there. What two different things could make one municipality top this table?

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) Mean yield and count by municipality, best first
mb_wheat |>
  group_by(Municipality) |>                 # one group per municipality
  summarise(mean_yield = mean(Yield_bu_ac), # average across its rows
            n = n()) |>                     # rows behind the average
  arrange(desc(mean_yield))                 # best municipality first

## (b) I find LOUISE on top at 72.86 bu/ac, backed by 26 observations.
## (c) Two different things could put a municipality here: genuinely better
## growing conditions, or a mix that happens to lean on high-yielding
## varieties. The mean pools varieties, so the ranking cannot tell the
## land from the lineup.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) Mean yield and count by municipality, best first
> mb_wheat |>
+   group_by(Municipality) |>                 # one group per municipality
+   summarise(mean_yield = mean(Yield_bu_ac), # average across its rows
+             n = n()) |>                     # rows behind the average
+   arrange(desc(mean_yield))                 # best municipality first
# A tibble: 96 × 3
   Municipality      mean_yield     n
   <chr>                  <dbl> <int>
 1 LOUISE                  72.9    26
 2 LAC DU BONNET           71.5    25
 3 ALEXANDER               71.4    17
 4 ELTON                   70.7    23
 5 RHINELAND               70.4    22
 6 THOMPSON                69.6    15
 7 MINITONAS-BOWSMAN       69.5    33
 8 CARTIER                 69.3    19
 9 PEMBINA                 69.2    34
10 EMERSON-FRANKLIN        68.8    16
# ℹ 86 more rows
> 
> ## (b) I find LOUISE on top at 72.86 bu/ac, backed by 26 observations.
> ## (c) Two different things could put a municipality here: genuinely better
> ## growing conditions, or a mix that happens to lean on high-yielding
> ## varieties. The mean pools varieties, so the ranking cannot tell the
> ## land from the lineup.

Question 40

This question uses: mb_wheat_reported_2020_2025.csv

(a) In one pipeline: filter the data to 2023, compute each variety’s mean yield and count, sort best-first, and save the result.

(b) Report the top variety by mean and its count, and the largest count in the table and its variety.

(c) In a comment: are they the same variety? What do the two columns measure differently?

Answer
# Load the tidyverse and read the data
library(tidyverse)
mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")

# (a) 2023 variety means and counts, best first
variety_2023 <- mb_wheat |>
  filter(Year == 2023) |>                   # one year only
  group_by(Variety) |>                      # one group per variety
  summarise(mean_yield = mean(Yield_bu_ac), # average within the variety
            n = n()) |>                     # results behind it
  arrange(desc(mean_yield))                 # best mean first
variety_2023

## (b) I find AC BARRIE (BW 661) with the best 2023 mean, 67.3 bu/ac --
## from a single observation. The largest count is AAC BRANDON (BW 932),
## with 84.
## (c) They are not the same variety. The mean measures how well a variety
## did where it was grown; the count measures how widely it was grown. A
## chart-topping mean on n = 1 is one municipality having a good year.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported

ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
> 
> # (a) 2023 variety means and counts, best first
> variety_2023 <- mb_wheat |>
+   filter(Year == 2023) |>                   # one year only
+   group_by(Variety) |>                      # one group per variety
+   summarise(mean_yield = mean(Yield_bu_ac), # average within the variety
+             n = n()) |>                     # results behind it
+   arrange(desc(mean_yield))                 # best mean first
> variety_2023
# A tibble: 26 × 3
   Variety                                mean_yield     n
   <chr>                                       <dbl> <int>
 1 AC BARRIE (BW 661)                           67.3     1
 2 AAC WHEATLAND <SECAN>                        66.2    44
 3 SY MANNESS                                   65.8    11
 4 AAC STARBUCK <SECAN>                         63.8    76
 5 AAC HOCKLEY |BW5044|                         63.6    51
 6 AAC HODGE |BW1069|                           63.0    30
 7 AAC ELIE(BW931)                              62.9     8
 8 BOLLES <SEED DEPOT>                          61.9    15
 9 AAC VIEWFIELD <FP GENETICS>|BW965| EXP       61.7    22
10 CDC SKRUSH |PT599|                           60.5     1
# ℹ 16 more rows
> 
> ## (b) I find AC BARRIE (BW 661) with the best 2023 mean, 67.3 bu/ac --
> ## from a single observation. The largest count is AAC BRANDON (BW 932),
> ## with 84.
> ## (c) They are not the same variety. The mean measures how well a variety
> ## did where it was grown; the count measures how widely it was grown. A
> ## chart-topping mean on n = 1 is one municipality having a good year.