This bank holds 40 questions: ten in each section.
Section 1 covers R basics: vectors, objects, and scripts
Section 2 covers reading and inspecting data
Section 3 covers data manipulation functions
Section 4 covers pipelines and grouped summaries
There are two datasets used in the test bank. The datasets are:
Saskatchewan RM Crop Yields (1990–2025; canola-only file for Section 2). Average yield for eight crops in every Rural Municipality, in long form. Lentils are recorded in pounds per acre; every other crop is in bushels per acre – the Unit column says which.
Manitoba Wheat Variety (2020–2025). Red spring wheat yield by variety and municipality, one row per municipality-variety-year.
Your module test will contain one question per section. The test will only use one of the datasets, and the dataset will be provided in the test.
For each question, 20% of the mark is for presentation and the other 80% is for the correctness of your script.
You submit a single R script: a header block at the top, a clearly labelled section for each question, and your sentence answers written as comments. The script should run from top to bottom in a project folder that has the csvs in data/. Here is an example of what a test will look like and the expectation for what your script will look like:
NoteA sample test, with a full-marks answer script
A real test looks like this – one question from each section:
Question 1
Five canola fields yielded 52.3, 47.8, 55.1, 44.6, and 50.9 bu/ac, on 160, 320, 240, 130, and 200 acres.
(a) Create a vector yields and a vector acres holding these values.
(b) Create a vector yields_t_ha converting the yields to tonnes per hectare (multiply by 0.0560) in one line.
(c) Compute the mean and standard deviation of yields.
(d) Compute total production in bushels, and the farm’s average yield weighted by acres. In a comment: why does the weighted average differ from the plain mean?
Question 2
The Saskatchewan canola file contains one RM-year average yield per row, measured in bu/ac.
(a) Compute the mean, median, and standard deviation of Yield.
(b) Compute its 90th percentile.
(c) In a comment: the mean sits a little above the median – what does that direction of gap suggest about the shape of canola yields?
Question 3
The full RM file holds all eight crops; lentil yields are recorded in lb/ac.
(a) Filter to the lentil rows, saving the result as lentils. How many rows are there?
(b) Add a yield_kg_ha column (multiply by 1.12), saving back to lentils.
(c) Compute the mean of the new column and of the original column, and check their ratio. In a comment: why is the ratio the conversion factor?
Question 4
Using the full eight-crop RM file.
(a) In one pipeline, compute each RM’s mean Spring Wheat yield across all years, sorted best-first.
(b) Report the top three RMs in a comment.
(c) In a comment: what did each row of the input represent, and what does each row of the result represent?
And here is a script that would earn full marks, including all of the presentation component – with the console session it produces shown beneath it:
# ---# Title: Module 2 test# Author: Jordan Field# Date: 2026-10-15# Description:# Answers to the four test questions, one block per question.# Sentence answers are written as comments below each result.# ---# Load packageslibrary(tidyverse)# ============================================================# Question 1# ============================================================# (a) Yields (bu/ac) and field sizes (acres)yields <-c(52.3, 47.8, 55.1, 44.6, 50.9)acres <-c(160, 320, 240, 130, 200)# (b) Convert every yield to tonnes per hectare in one stepyields_t_ha <- yields *0.0560yields_t_ha# (c) Centre and spread of the yieldsmean(yields) # 50.14 bu/acsd(yields) # 4.06 bu/ac# (d) Total production, then the average weighted by acrestotal_bu <-sum(yields * acres)total_bu # 52,866 butotal_bu /sum(acres) # 50.35 bu/ac# (d) The weighted average differs from the plain mean because the# fields differ in size, so each yield should count by its acres.# ============================================================# Question 2# ============================================================# Read the canola file using a relative path from the project folderrm_canola <-read_csv("data/rm_canola_yields_1990_2025.csv")# (a) Centre and spread of the canola yieldsmean(rm_canola$Yield) # 28.29 bu/acmedian(rm_canola$Yield) # 26.9 bu/acsd(rm_canola$Yield) # 10.06 bu/ac# (b) 90th percentilequantile(rm_canola$Yield, 0.9) # 42.5 bu/ac# (c) The mean sits a little above the median, which suggests mild# right skew: the best RM-years pull the mean up more than the# worst years pull it down.# ============================================================# Question 3# ============================================================# Read the full eight-crop filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Keep only the lentil rowslentils <-filter(rm_yields, Crop =="Lentils")nrow(lentils) # 6,338 rows# (b) Add the converted column (1 lb/ac is about 1.12 kg/ha) and# overwrite lentils with the version that has the extra columnlentils <-mutate(lentils, yield_kg_ha = Yield *1.12)# (c) Means of the new and original columnsmean(lentils$yield_kg_ha) # 1,353.4 kg/hamean(lentils$Yield) # 1,208.39 lb/ac# (c) Their ratio is 1.12, the conversion factor -- multiplying every# value by a constant multiplies the mean by the same constant.# ============================================================# Question 4# ============================================================# (a) Mean spring wheat yield by RM, best firstrm_yields |># take the data, THENfilter(Crop =="Spring Wheat") |># keep spring wheat, THENgroup_by(RM) |># split the rows by RM, THENsummarise(mean_yield =mean(Yield)) |># one row per RM: its mean, THENarrange(desc(mean_yield)) # best RMs on top# (b) Top three: RM 369 (47.6), RM 333 (47.2), RM 368 (47.2) bu/ac.# (c) Each input row was one RM-year observation of spring wheat; each# result row is one RM, its years collapsed into a single mean.
> # ---
> # Title: Module 2 test
> # Author: Jordan Field
> # Date: 2026-10-15
> # Description:
> # Answers to the four test questions, one block per question.
> # Sentence answers are written as comments below each result.
> # ---
>
> # Load packages
> library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr 1.2.1 ✔ readr 2.2.0
✔ forcats 1.0.1 ✔ stringr 1.6.0
✔ ggplot2 4.0.3 ✔ tibble 3.3.1
✔ lubridate 1.9.5 ✔ tidyr 1.3.2
✔ purrr 1.2.2
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag() masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
>
> # ============================================================
> # Question 1
> # ============================================================
>
> # (a) Yields (bu/ac) and field sizes (acres)
> yields <- c(52.3, 47.8, 55.1, 44.6, 50.9)
> acres <- c(160, 320, 240, 130, 200)
>
> # (b) Convert every yield to tonnes per hectare in one step
> yields_t_ha <- yields * 0.0560
> yields_t_ha
[1] 2.9288 2.6768 3.0856 2.4976 2.8504
>
> # (c) Centre and spread of the yields
> mean(yields) # 50.14 bu/ac
[1] 50.14
> sd(yields) # 4.06 bu/ac
[1] 4.062388
>
> # (d) Total production, then the average weighted by acres
> total_bu <- sum(yields * acres)
> total_bu # 52,866 bu
[1] 52866
> total_bu / sum(acres) # 50.35 bu/ac
[1] 50.34857
> # (d) The weighted average differs from the plain mean because the
> # fields differ in size, so each yield should count by its acres.
>
> # ============================================================
> # Question 2
> # ============================================================
>
> # Read the canola file using a relative path from the project folder
> rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")
Rows: 10039 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Centre and spread of the canola yields
> mean(rm_canola$Yield) # 28.29 bu/ac
[1] 28.29239
> median(rm_canola$Yield) # 26.9 bu/ac
[1] 26.9
> sd(rm_canola$Yield) # 10.06 bu/ac
[1] 10.06019
>
> # (b) 90th percentile
> quantile(rm_canola$Yield, 0.9) # 42.5 bu/ac
90%
42.5
>
> # (c) The mean sits a little above the median, which suggests mild
> # right skew: the best RM-years pull the mean up more than the
> # worst years pull it down.
>
> # ============================================================
> # Question 3
> # ============================================================
>
> # Read the full eight-crop file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Keep only the lentil rows
> lentils <- filter(rm_yields, Crop == "Lentils")
> nrow(lentils) # 6,338 rows
[1] 6338
>
> # (b) Add the converted column (1 lb/ac is about 1.12 kg/ha) and
> # overwrite lentils with the version that has the extra column
> lentils <- mutate(lentils, yield_kg_ha = Yield * 1.12)
>
> # (c) Means of the new and original columns
> mean(lentils$yield_kg_ha) # 1,353.4 kg/ha
[1] 1353.402
> mean(lentils$Yield) # 1,208.39 lb/ac
[1] 1208.395
> # (c) Their ratio is 1.12, the conversion factor -- multiplying every
> # value by a constant multiplies the mean by the same constant.
>
> # ============================================================
> # Question 4
> # ============================================================
>
> # (a) Mean spring wheat yield by RM, best first
> rm_yields |> # take the data, THEN
+ filter(Crop == "Spring Wheat") |> # keep spring wheat, THEN
+ group_by(RM) |> # split the rows by RM, THEN
+ summarise(mean_yield = mean(Yield)) |> # one row per RM: its mean, THEN
+ arrange(desc(mean_yield)) # best RMs on top
# A tibble: 298 × 2
RM mean_yield
<dbl> <dbl>
1 369 47.6
2 333 47.2
3 368 47.2
4 271 46.1
5 303 45.0
6 427 45.0
7 404 44.9
8 496 44.8
9 493 44.6
10 338 44.5
# ℹ 288 more rows
>
> # (b) Top three: RM 369 (47.6), RM 333 (47.2), RM 368 (47.2) bu/ac.
> # (c) Each input row was one RM-year observation of spring wheat; each
> # result row is one RM, its years collapsed into a single mean.
NoteProcedure for test
At least an hour before the test, “Schedule an appointment” in the test for that day.
Bring your (charged) laptop to the test.
Your test will be accessible at 11:30am.
When you get to the room:
Close all applications other than your webbrowser (used to access the test) and VS Code.
Log into zoom and join the breakout room with your name on it.
In Zoom, share your screen (not an application window, but your whole screen) and record your session.
Navigate to Canvas and access your test.
Download the dataset that you need
Complete the test in VS Code.
Upload the completed script to Canvas (while still logged into zoom).
Leave zoom, and upload the screenshot recording to Canvas.
Your done!!
Section 1 — R Basics
Question 1
Five canola fields yielded 52.3, 47.8, 55.1, 44.6, and 50.9 bu/ac, on 160, 320, 240, 130, and 200 acres.
(a) Create a vector yields and a vector acres holding these values.
(b) Using one line of code, create a vector yields_t_ha converting the yields to tonnes per hectare (multiply by 0.0560) in one line.
(c) Compute the mean and standard deviation of yields.
(d) Compute total production in bushels (each field’s yield times its acres, summed), and the farm’s average yield weighted by acres (total bushels divided by total acres).
Answer
# (a) Create the yield and acreage vectorsyields <-c(52.3, 47.8, 55.1, 44.6, 50.9)acres <-c(160, 320, 240, 130, 200)# (b) Convert every yield to tonnes per hectare in one lineyields_t_ha <- yields *0.0560# Print the converted yieldsyields_t_ha# (c) Mean and standard deviation of the yieldsmean(yields)sd(yields)# (d) Total production in bushelstotal_bu <-sum(yields * acres)total_bu# Average yield weighted by acrestotal_bu /sum(acres)## I find the mean to be 50.14 bu/ac and the sd 4.06.## Total production is 52,866 bu and the acre-weighted average 50.35 bu/ac.
> # (a) Create the yield and acreage vectors
> yields <- c(52.3, 47.8, 55.1, 44.6, 50.9)
> acres <- c(160, 320, 240, 130, 200)
>
> # (b) Convert every yield to tonnes per hectare in one line
> yields_t_ha <- yields * 0.0560
>
> # Print the converted yields
> yields_t_ha
[1] 2.9288 2.6768 3.0856 2.4976 2.8504
>
> # (c) Mean and standard deviation of the yields
> mean(yields)
[1] 50.14
> sd(yields)
[1] 4.062388
>
> # (d) Total production in bushels
> total_bu <- sum(yields * acres)
> total_bu
[1] 52866
>
> # Average yield weighted by acres
> total_bu / sum(acres)
[1] 50.34857
>
> ## I find the mean to be 50.14 bu/ac and the sd 4.06.
> ## Total production is 52,866 bu and the acre-weighted average 50.35 bu/ac.
Question 2
Six spring wheat fields yielded 51.2, 48.7, 53.4, 8.2, 50.9, and 49.5 bu/ac; the 8.2 field was hailed out in July.
(a) Save the yields and compute the mean and the median.
(b) In a comment: the two disagree by nearly seven bushels. Which one better describes a typical field on this farm, and why?
(c) Compute the mean of the five unhailed fields (a second vector is fine) and compare it with the median from part a. In a comment: what does the closeness of those two numbers tell you about the median?
Answer
# (a) Yields, including the hailed-out fieldyields <-c(51.2, 48.7, 53.4, 8.2, 50.9, 49.5)mean(yields)median(yields)## I find the mean to be 43.65 and the median 50.2.## (b) The median describes a typical field better: five of the six fields## sit near 50, and the single 8.2 drags the mean seven bushels below all## of them while barely touching the median.# (c) The five unhailed fieldsunhailed <-c(51.2, 48.7, 53.4, 50.9, 49.5)mean(unhailed)## (c) The unhailed mean is 50.74, almost exactly the full-data median --## the median was already ignoring the extreme value.
> # (a) Yields, including the hailed-out field
> yields <- c(51.2, 48.7, 53.4, 8.2, 50.9, 49.5)
> mean(yields)
[1] 43.65
> median(yields)
[1] 50.2
>
>
> ## I find the mean to be 43.65 and the median 50.2.
> ## (b) The median describes a typical field better: five of the six fields
> ## sit near 50, and the single 8.2 drags the mean seven bushels below all
> ## of them while barely touching the median.
>
> # (c) The five unhailed fields
> unhailed <- c(51.2, 48.7, 53.4, 50.9, 49.5)
> mean(unhailed)
[1] 50.74
>
> ## (c) The unhailed mean is 50.74, almost exactly the full-data median --
> ## the median was already ignoring the extreme value.
Question 3
Five barley fields yielded 68.2, 71.5, 64.9, 74.1, and 66.3 bu/ac. Five lentil fields yielded 1450, 1720, 1280, 1610, and 1390 lb/ac.
(a) Compute the mean, the standard deviation, the variance, and the range (maximum minus minimum) of each crop’s yields.
(b) The lentil standard deviation is far larger. In a comment: why can the two standard deviations not be compared directly?
(c) Compute the coefficient of variation (sd divided by mean) for each crop. Which crop’s yields are more variable relative to their own average?
Answer
# (a) Yield vectors for each cropbarley <-c(68.2, 71.5, 64.9, 74.1, 66.3)lentils <-c(1450, 1720, 1280, 1610, 1390)# Mean of each cropmean(barley)mean(lentils)# Standard deviation of each cropsd(barley)sd(lentils)# Variance of each cropvar(barley)var(lentils)# Range (maximum minus minimum) of each cropmax(barley) -min(barley)max(lentils) -min(lentils)## I find barley mean 69.0, sd 3.77, variance 14.25, range 9.2; lentils## mean 1490, sd 175.36, variance 30,750, range 440.## (b) The two standard deviations are in different units (bushels vs pounds## per acre) and on very different scales, so the raw spreads are not## comparable.# (c) Coefficient of variation (sd divided by mean) for each cropsd(barley) /mean(barley)sd(lentils) /mean(lentils)## (c) CV barley 0.055, lentils 0.118 -- lentil yields are more variable## relative to their own average.
> # (a) Yield vectors for each crop
> barley <- c(68.2, 71.5, 64.9, 74.1, 66.3)
> lentils <- c(1450, 1720, 1280, 1610, 1390)
>
> # Mean of each crop
> mean(barley)
[1] 69
> mean(lentils)
[1] 1490
>
> # Standard deviation of each crop
> sd(barley)
[1] 3.774917
> sd(lentils)
[1] 175.3568
>
> # Variance of each crop
> var(barley)
[1] 14.25
> var(lentils)
[1] 30750
>
> # Range (maximum minus minimum) of each crop
> max(barley) - min(barley)
[1] 9.2
> max(lentils) - min(lentils)
[1] 440
>
> ## I find barley mean 69.0, sd 3.77, variance 14.25, range 9.2; lentils
> ## mean 1490, sd 175.36, variance 30,750, range 440.
>
> ## (b) The two standard deviations are in different units (bushels vs pounds
> ## per acre) and on very different scales, so the raw spreads are not
> ## comparable.
>
> # (c) Coefficient of variation (sd divided by mean) for each crop
> sd(barley) / mean(barley)
[1] 0.05470895
> sd(lentils) / mean(lentils)
[1] 0.1176891
>
> ## (c) CV barley 0.055, lentils 0.118 -- lentil yields are more variable
> ## relative to their own average.
Question 4
Here is a script exactly as saved:
mean(yields)yields <-c(31.7, 28.4, 33.9, 30.2)
(a) In a comment: what would happen if you ran this script top to bottom in a fresh session, and why?
(b) Rewrite the script in working order, with a header block and a comment on each line.
(c) Compute the standard deviation and median of yields.
Answer
# (a) Run as saved, the script fails on line 1: mean(yields) is asked for# before yields exists in the fresh session, so R stops with# R reports that the object yields was not found.# (b) The working order, with a header# ---# This script computes the mean of a yield vector.# ---# Create the yield vector firstyields <-c(31.7, 28.4, 33.9, 30.2)# Then compute its meanmean(yields)## I find the mean to be 31.05.# (c) Standard deviation and median of yieldssd(yields)median(yields)
> # (a) Run as saved, the script fails on line 1: mean(yields) is asked for
> # before yields exists in the fresh session, so R stops with
> # R reports that the object yields was not found.
>
> # (b) The working order, with a header
> # ---
> # This script computes the mean of a yield vector.
> # ---
>
> # Create the yield vector first
> yields <- c(31.7, 28.4, 33.9, 30.2)
>
> # Then compute its mean
> mean(yields)
[1] 31.05
>
> ## I find the mean to be 31.05.
>
> # (c) Standard deviation and median of yields
> sd(yields)
[1] 2.330236
> median(yields)
[1] 30.95
Question 5
Suppose there are four fields:
Field 1 had canola, 160 acres, and was irrigated.
Field 2 had wheat, 320 acres, and was not irrigated.
Field 3 had peas, 240 acres, and was not irrigated.
Field 4 had barley, 130 acres, and was irrigated.
(a) Create three vectors of length four, one for the crop grown, one for the seeded acres, and one for whether the field was irrigated.
(b) Combine your three vectors into a data frame called trial and print it in the console.
Answer
# (a) Crop grown: character, so the values are quotedcrop <-c("Canola", "Wheat", "Peas", "Barley")# Seeded acres: numeric, no quotesacres <-c(160, 320, 240, 130)# Irrigated or not: logical, no quotesirrigated <-c(TRUE, FALSE, FALSE, TRUE)# (b) Combine the three vectors into a data frametrial <-data.frame(crop = crop,acres = acres,irrigated = irrigated)# Print the data frametrial
> # (a) Crop grown: character, so the values are quoted
> crop <- c("Canola", "Wheat", "Peas", "Barley")
>
> # Seeded acres: numeric, no quotes
> acres <- c(160, 320, 240, 130)
>
> # Irrigated or not: logical, no quotes
> irrigated <- c(TRUE, FALSE, FALSE, TRUE)
>
> # (b) Combine the three vectors into a data frame
> trial <- data.frame(crop = crop,
+ acres = acres,
+ irrigated = irrigated)
>
> # Print the data frame
> trial
crop acres irrigated
1 Canola 160 TRUE
2 Wheat 320 FALSE
3 Peas 240 FALSE
4 Barley 130 TRUE
Question 6
The same five canola fields were measured in two years. 2024: 38.2, 41.6, 35.9, 43.1, 39.7 bu/ac. 2025: 45.1, 44.8, 42.3, 49.6, 46.2 bu/ac.
(a) Save the two vectors and compute each field’s change in yield in one line.
(b) Compute the mean change, and the mean percentage change (change over the 2024 yield, times 100).
(c) In a comment: every field improved, but which measure – bushels gained or percent gained – would you use to compare this farm against a lentil farm, and why?
Answer
# (a) The two years, and each field change in one liney2024 <-c(38.2, 41.6, 35.9, 43.1, 39.7)y2025 <-c(45.1, 44.8, 42.3, 49.6, 46.2)change <- y2025 - y2024change# (b) Mean change and mean percentage changemean(change)mean(change / y2024 *100)## I find a mean change of 5.9 bu/ac and a mean percentage change of 15.0%.## (c) The percentage change: bushels gained are not comparable across crops## measured on different scales (a lentil farm gains hundreds of lb/ac), but## percent gained puts both farms on the same footing.
> # (a) The two years, and each field change in one line
> y2024 <- c(38.2, 41.6, 35.9, 43.1, 39.7)
> y2025 <- c(45.1, 44.8, 42.3, 49.6, 46.2)
> change <- y2025 - y2024
> change
[1] 6.9 3.2 6.4 6.5 6.5
>
> # (b) Mean change and mean percentage change
> mean(change)
[1] 5.9
> mean(change / y2024 * 100)
[1] 15.00729
>
> ## I find a mean change of 5.9 bu/ac and a mean percentage change of 15.0%.
> ## (c) The percentage change: bushels gained are not comparable across crops
> ## measured on different scales (a lentil farm gains hundreds of lb/ac), but
> ## percent gained puts both farms on the same footing.
Question 7
A farm’s four barley fields: field IDs "B1" to "B4", sizes 155, 240, 160, and 310 acres, yields 66.8, 71.2, 63.5, and 69.9 bu/ac.
(a) Build a data frame barley_fields holding the three variables.
(b) Compute each field’s production (acres times yield) and save it as a vector.
(c) Compute total production, and the farm’s average yield weighted by acres (total production over total acres).
(d) In a comment: compare the weighted average with mean(barley_fields$yield) and say why they differ.
Answer
# (a) The barley fields as a data framebarley_fields <-data.frame(field_id =c("B1", "B2", "B3", "B4"),acres =c(155, 240, 160, 310),yield =c(66.8, 71.2, 63.5, 69.9))# (b) Production of each fieldproduction <- barley_fields$acres * barley_fields$yieldproduction# (c) Total production and the acre-weighted average yieldsum(production)sum(production) /sum(barley_fields$acres)# (d) The plain mean, for comparisonmean(barley_fields$yield)## I find total production of 59,271 bu and a weighted average of 68.52.## (d) The plain mean is 67.85: it counts each field once, while the## weighted average lets the 310-acre field count for its share of the## land, and that field yields above average, pulling the weighted figure up.
> # (a) The barley fields as a data frame
> barley_fields <- data.frame(
+ field_id = c("B1", "B2", "B3", "B4"),
+ acres = c(155, 240, 160, 310),
+ yield = c(66.8, 71.2, 63.5, 69.9)
+ )
>
> # (b) Production of each field
> production <- barley_fields$acres * barley_fields$yield
> production
[1] 10354 17088 10160 21669
>
> # (c) Total production and the acre-weighted average yield
> sum(production)
[1] 59271
> sum(production) / sum(barley_fields$acres)
[1] 68.52139
>
> # (d) The plain mean, for comparison
> mean(barley_fields$yield)
[1] 67.85
>
> ## I find total production of 59,271 bu and a weighted average of 68.52.
> ## (d) The plain mean is 67.85: it counts each field once, while the
> ## weighted average lets the 310-acre field count for its share of the
> ## land, and that field yields above average, pulling the weighted figure up.
Question 8
(a) Build a data frame called durum_fields with columns field_id ("D1" to "D4"), region (South, South, Central, North), and yield (28.3, 41.6, 35.2, 44.8).
(b) Report its number of rows, number of columns, and column names using functions.
(c) Using $, compute the mean yield and the range (maximum minus minimum).
Answer
# (a) Build the data framedurum_fields <-data.frame(field_id =c("D1", "D2", "D3", "D4"),region =c("South", "South", "Central", "North"),yield =c(28.3, 41.6, 35.2, 44.8))# (b) Number of rowsnrow(durum_fields)# Number of columnsncol(durum_fields)# Column namesnames(durum_fields)# (c) Mean yield, using $mean(durum_fields$yield)# Range (maximum minus minimum)max(durum_fields$yield) -min(durum_fields$yield)## (b) 4 rows, 3 columns; the columns are field_id, region, yield.## (c) I find the mean to be 37.48 (the console shows 37.475) and the## range 16.5.
> # (a) Build the data frame
> durum_fields <- data.frame(
+ field_id = c("D1", "D2", "D3", "D4"),
+ region = c("South", "South", "Central", "North"),
+ yield = c(28.3, 41.6, 35.2, 44.8)
+ )
>
> # (b) Number of rows
> nrow(durum_fields)
[1] 4
>
> # Number of columns
> ncol(durum_fields)
[1] 3
>
> # Column names
> names(durum_fields)
[1] "field_id" "region" "yield"
>
> # (c) Mean yield, using $
> mean(durum_fields$yield)
[1] 37.475
>
> # Range (maximum minus minimum)
> max(durum_fields$yield) - min(durum_fields$yield)
[1] 16.5
>
> ## (b) 4 rows, 3 columns; the columns are field_id, region, yield.
> ## (c) I find the mean to be 37.48 (the console shows 37.475) and the
> ## range 16.5.
(a) Save the values as a vector and compute the mean.
(b) Compute the 10th, 50th, and 90th percentiles in a single quantile() call.
(c) In a comment: one of your three percentiles should equal the median. Which one, and does it?
Answer
# (a) Canola yields as a vector (bu/ac)canola <-c(24.1, 31.5, 28.8, 35.2, 22.7, 30.4, 27.9, 33.6, 25.5, 29.3, 32.8, 26.4)# Mean yieldmean(canola)# (b) 10th, 50th, and 90th percentiles in one callquantile(canola, c(0.1, 0.5, 0.9))# (c) The median, to check against the 50th percentilemedian(canola)## I find the mean to be 29.02.## (b) 10th 24.24, 50th 29.05, 90th 33.52.## (c) The 50th percentile is the median by definition; median(canola)## confirms 29.05.
> # (a) Canola yields as a vector (bu/ac)
> canola <- c(24.1, 31.5, 28.8, 35.2, 22.7, 30.4, 27.9, 33.6, 25.5, 29.3, 32.8, 26.4)
>
> # Mean yield
> mean(canola)
[1] 29.01667
>
> # (b) 10th, 50th, and 90th percentiles in one call
> quantile(canola, c(0.1, 0.5, 0.9))
10% 50% 90%
24.24 29.05 33.52
>
> # (c) The median, to check against the 50th percentile
> median(canola)
[1] 29.05
>
> ## I find the mean to be 29.02.
> ## (b) 10th 24.24, 50th 29.05, 90th 33.52.
> ## (c) The 50th percentile is the median by definition; median(canola)
> ## confirms 29.05.
Question 10
Four oat fields yielded 84.3, 91.7, 77.2, and 88.5 bu/ac, and oats are selling at $4.15 per bushel.
(a) Save the yields as a vector and the price as a single object.
(b) Compute a revenue-per-acre vector (yield times price) in one line, using the price object rather than retyping the number.
(c) Compute the mean revenue per acre.
(d) In a comment: the price rises to $4.60. What value would you change, and which lines would you rerun to update the results? Why is that easier than typing the price into each calculation?
Answer
# (a) Oat yields (bu/ac) and the price ($/bu)yields <-c(84.3, 91.7, 77.2, 88.5)price <-4.15# (b) Revenue per acre, using the price objectrevenue <- yields * price# Print the revenuesrevenue# (c) Mean revenue per acremean(revenue)## (b) The revenues are 349.85, 380.56, 320.38, 367.28 $/ac, rounded to the## cent (the console shows 349.845 and so on).## (c) Mean revenue is $354.51 per acre.## (d) Change price <- 4.60, then rerun the lines below it. The price lives## in one object, so one edit updates everything -- the same reason a## spreadsheet keeps a price in one cell.
> # (a) Oat yields (bu/ac) and the price ($/bu)
> yields <- c(84.3, 91.7, 77.2, 88.5)
> price <- 4.15
>
> # (b) Revenue per acre, using the price object
> revenue <- yields * price
>
> # Print the revenues
> revenue
[1] 349.845 380.555 320.380 367.275
>
> # (c) Mean revenue per acre
> mean(revenue)
[1] 354.5138
>
> ## (b) The revenues are 349.85, 380.56, 320.38, 367.28 $/ac, rounded to the
> ## cent (the console shows 349.845 and so on).
> ## (c) Mean revenue is $354.51 per acre.
> ## (d) Change price <- 4.60, then rerun the lines below it. The price lives
> ## in one object, so one edit updates everything -- the same reason a
> ## spreadsheet keeps a price in one cell.
(a) Run summary(rm_canola). From its output, report the minimum, median, and maximum of Yield, and the first and last Year.
(b) In a comment: are most yields in a plausible range for canola, and do the extremes need investigation? Are the years what the file name promises? Are there missing values?
(c) The minimum is below 2 bu/ac. In a comment: is a canola yield that low necessarily a data error?
Answer
# Load the tidyverse and read the canola filelibrary(tidyverse)rm_canola <-read_csv("data/rm_canola_yields_1990_2025.csv")# (a) Summary of every columnsummary(rm_canola)## (a) I find Yield has minimum 1.9, median 26.9, and maximum 61.0, and Year runs from 1990 to 2025.## (b) Most yields are plausible for canola, but the values near 2 bu/ac## deserve investigation. The years run 1990--2025 and there are no NAs.## (c) Not necessarily. A severe drought or widespread crop failure could## produce a very low RM average.
> # Load the tidyverse and read the canola file
> library(tidyverse)
> rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")
Rows: 10039 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Summary of every column
> summary(rm_canola)
Year RM Crop Yield
Min. :1990 Min. : 1.0 Length :10039 Min. : 1.90
1st Qu.:1999 1st Qu.:128.0 N.unique : 1 1st Qu.:21.00
Median :2008 Median :253.0 N.blank : 0 Median :26.90
Mean :2008 Mean :254.4 Min.nchar: 6 Mean :28.29
3rd Qu.:2017 3rd Qu.:372.0 Max.nchar: 6 3rd Qu.:35.40
Max. :2025 Max. :622.0 Max. :61.00
Unit
Length :10039
N.unique : 1
N.blank : 0
Min.nchar: 5
Max.nchar: 5
>
> ## (a) I find Yield has minimum 1.9, median 26.9, and maximum 61.0, and Year runs from 1990 to 2025.
> ## (b) Most yields are plausible for canola, but the values near 2 bu/ac
> ## deserve investigation. The years run 1990--2025 and there are no NAs.
> ## (c) Not necessarily. A severe drought or widespread crop failure could
> ## produce a very low RM average.
(a) Compute the mean, median, standard deviation, and variance of rm_canola$Yield.
(b) Compute the 90th percentile of Yield.
(c) In a comment: what does that difference between the mean and median suggest about the shape of canola yields?
Answer
# Load the tidyverse and read the canola filelibrary(tidyverse)rm_canola <-read_csv("data/rm_canola_yields_1990_2025.csv")# (a) Centre of the yieldsmean(rm_canola$Yield)median(rm_canola$Yield)# Spread of the yieldssd(rm_canola$Yield)var(rm_canola$Yield)# (b) 90th percentile of the yieldsquantile(rm_canola$Yield, probs =0.9)## (a) I find a mean of 28.29, a median of 26.9, a standard deviation of 10.06, and a variance of 101.21 -- the sd squared, in squared units.## (b) The 90th percentile is 42.5 bu/ac.## (c) A mean above the median suggests mild right skew: the very best RM-years pull the mean up more than the worst years pull it down.
> # Load the tidyverse and read the canola file
> library(tidyverse)
> rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")
Rows: 10039 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Centre of the yields
> mean(rm_canola$Yield)
[1] 28.29239
> median(rm_canola$Yield)
[1] 26.9
>
> # Spread of the yields
> sd(rm_canola$Yield)
[1] 10.06019
> var(rm_canola$Yield)
[1] 101.2074
>
> # (b) 90th percentile of the yields
> quantile(rm_canola$Yield, probs = 0.9)
90%
42.5
>
> ## (a) I find a mean of 28.29, a median of 26.9, a standard deviation of 10.06, and a variance of 101.21 -- the sd squared, in squared units.
> ## (b) The 90th percentile is 42.5 bu/ac.
> ## (c) A mean above the median suggests mild right skew: the very best RM-years pull the mean up more than the worst years pull it down.
(a) Print the first six rows of rm_canola with head().
(b) Report the number of rows in the data, and compute the first and last year in the data.
(c) There are 295 RMs in Saskatchewan. Divide the row count by the number of years. What does this tell you about the coverage of the dataset?
Answer
# (a) First six rowshead(rm_canola)# (b) Rows, and the span of yearsnrow(rm_canola)range(rm_canola$Year)## (b) I find 10,039 rows spanning 1990 to 2025.# (c) Rows per yearnrow(rm_canola) /36## (c) 10,039 / 36 is about 279 rows per year, against the 295 RMs --## so not every RM reports canola every year; some RM-year combinations## are absent from the file.
> # (a) First six rows
> head(rm_canola)
# A tibble: 6 × 5
Year RM Crop Yield Unit
<dbl> <dbl> <chr> <dbl> <chr>
1 1990 1 Canola 22 bu/ac
2 1991 1 Canola 24.1 bu/ac
3 1992 1 Canola 15 bu/ac
4 1993 1 Canola 25.3 bu/ac
5 1994 1 Canola 23.1 bu/ac
6 1995 1 Canola 15.9 bu/ac
>
> # (b) Rows, and the span of years
> nrow(rm_canola)
[1] 10039
> range(rm_canola$Year)
[1] 1990 2025
>
> ## (b) I find 10,039 rows spanning 1990 to 2025.
>
> # (c) Rows per year
> nrow(rm_canola) / 36
[1] 278.8611
>
> ## (c) 10,039 / 36 is about 279 rows per year, against the 295 RMs --
> ## so not every RM reports canola every year; some RM-year combinations
> ## are absent from the file.
(a) Using $, compute the mean, the standard deviation, and the 10th and 90th percentiles of Yield across the whole file.
(b) In a comment: an RM-year at the 10th percentile grew how much canola? Interpret the number for a grower.
(c) Compute the gap between the 90th and 10th percentiles. In a comment: what does this gap describe that the standard deviation also tries to describe?
Answer
# (a) Mean, sd, and the 10th and 90th percentilesmean(rm_canola$Yield)sd(rm_canola$Yield)quantile(rm_canola$Yield, probs =c(0.10, 0.90))## I find a mean of 28.29, an sd of 10.06, and percentiles of 16.1 and 42.5.## (b) An RM-year at the 10th percentile grew 16.1 bu/ac -- nine in ten## RM-years across the file were at or above that value.# (c) The 90-10 gap42.5-16.1## (c) The gap is 26.4 bu/ac. Like the sd it describes how spread out## yields are, but it reads directly as a bushel distance holding the## middle 80% of RM-years, and it ignores the extremes entirely.
> # (a) Mean, sd, and the 10th and 90th percentiles
> mean(rm_canola$Yield)
[1] 28.29239
> sd(rm_canola$Yield)
[1] 10.06019
> quantile(rm_canola$Yield, probs = c(0.10, 0.90))
10% 90%
16.1 42.5
>
> ## I find a mean of 28.29, an sd of 10.06, and percentiles of 16.1 and 42.5.
>
> ## (b) An RM-year at the 10th percentile grew 16.1 bu/ac -- nine in ten
> ## RM-years across the file were at or above that value.
>
> # (c) The 90-10 gap
> 42.5 - 16.1
[1] 26.4
>
> ## (c) The gap is 26.4 bu/ac. Like the sd it describes how spread out
> ## yields are, but it reads directly as a bushel distance holding the
> ## middle 80% of RM-years, and it ignores the extremes entirely.
(a) Compute the 25th, 50th, and 75th percentiles of Yield in one quantile() call.
(b) Compute the IQR twice: with IQR(), and as the difference between two of your percentiles. Confirm the two match.
(c) In a comment: what fraction of all observations lies between your 25th and 75th percentiles, and how would you describe those observations?
Answer
# Load the tidyverse and read the canola filelibrary(tidyverse)rm_canola <-read_csv("data/rm_canola_yields_1990_2025.csv")# (a) The three quartiles in one callquantile(rm_canola$Yield, probs =c(0.25, 0.50, 0.75))# (b) The IQR, from the functionIQR(rm_canola$Yield)# The IQR, as the 75th minus the 25th percentilequantile(rm_canola$Yield, probs =0.75) -quantile(rm_canola$Yield, probs =0.25)## I find quartiles of 21.0, 26.9, and 35.4 bu/ac, and an IQR of 14.4## both ways.## (c) Half of all observations lie between the two percentiles. They## are the middle, typical yields: the lowest quarter and the highest## quarter of observations sit outside the range.
> # Load the tidyverse and read the canola file
> library(tidyverse)
> rm_canola <- read_csv("data/rm_canola_yields_1990_2025.csv")
Rows: 10039 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) The three quartiles in one call
> quantile(rm_canola$Yield, probs = c(0.25, 0.50, 0.75))
25% 50% 75%
21.0 26.9 35.4
>
> # (b) The IQR, from the function
> IQR(rm_canola$Yield)
[1] 14.4
>
> # The IQR, as the 75th minus the 25th percentile
> quantile(rm_canola$Yield, probs = 0.75) - quantile(rm_canola$Yield, probs = 0.25)
75%
14.4
>
> ## I find quartiles of 21.0, 26.9, and 35.4 bu/ac, and an IQR of 14.4
> ## both ways.
> ## (c) Half of all observations lie between the two percentiles. They
> ## are the middle, typical yields: the lowest quarter and the highest
> ## quarter of observations sit outside the range.
(a) Run summary(mb_wheat). From its output, report the minimum, median, and maximum of Yield_bu_ac, and the first and last Year.
(b) In a comment: are most yields in a plausible range for spring wheat, and do the extremes need investigation? Are there missing values?
Answer
# Load the tidyverse and read the wheat filelibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Summary of every columnsummary(mb_wheat)## (a) I find Yield_bu_ac has minimum 4.5, median 62.2, and maximum 97.8, and Year runs from 2020 to 2025.## (b) Most yields are plausible for Manitoba spring wheat, but the lowest## values deserve investigation. There are no missing values in this file.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Summary of every column
> summary(mb_wheat)
Year Municipality Variety Farms
Min. :2020 Length :2398 Length :2398 Min. : 3.00
1st Qu.:2021 N.unique : 96 N.unique : 44 1st Qu.: 4.00
Median :2023 N.blank : 0 N.blank : 0 Median : 8.00
Mean :2023 Min.nchar: 4 Min.nchar: 5 Mean : 15.42
3rd Qu.:2024 Max.nchar: 28 Max.nchar: 38 3rd Qu.: 21.00
Max. :2025 Max. :148.00
Acres Yield_bu_ac Reported
Min. : 501 Min. : 4.50 Mode:logical
1st Qu.: 1354 1st Qu.:53.90 TRUE:2398
Median : 2941 Median :62.20
Mean : 6118 Mean :61.16
3rd Qu.: 7721 3rd Qu.:69.60
Max. :74604 Max. :97.80
>
> ## (a) I find Yield_bu_ac has minimum 4.5, median 62.2, and maximum 97.8, and Year runs from 2020 to 2025.
> ## (b) Most yields are plausible for Manitoba spring wheat, but the lowest
> ## values deserve investigation. There are no missing values in this file.
(a) Compute the mean, median, standard deviation, and variance of mb_wheat$Yield_bu_ac.
(b) In a comment: what does that difference between the mean and median suggest about the data?
(c) Compute the smallest and largest yields in the data.
Answer
# Load the tidyverse and read the wheat filelibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Centre of the yieldsmean(mb_wheat$Yield_bu_ac)median(mb_wheat$Yield_bu_ac)# Spread of the yieldssd(mb_wheat$Yield_bu_ac)var(mb_wheat$Yield_bu_ac)## (a) I find a mean of 61.16, a median of 62.2, a standard deviation of 12.38, and a variance of 153.34.## (b) A mean below the median suggests mild left skew: unusually low values pull the mean down more than the high values pull it up -- a tail of very poor variety-municipality observations mixed in with the normal ones.# (c) Smallest and largest yieldrange(mb_wheat$Yield_bu_ac)## (c) range() returns both ends as one vector: the smallest yield is 4.5 and the largest is 97.8 bu/ac.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Centre of the yields
> mean(mb_wheat$Yield_bu_ac)
[1] 61.15617
> median(mb_wheat$Yield_bu_ac)
[1] 62.2
>
> # Spread of the yields
> sd(mb_wheat$Yield_bu_ac)
[1] 12.38297
> var(mb_wheat$Yield_bu_ac)
[1] 153.338
>
> ## (a) I find a mean of 61.16, a median of 62.2, a standard deviation of 12.38, and a variance of 153.34.
>
> ## (b) A mean below the median suggests mild left skew: unusually low values pull the mean down more than the high values pull it up -- a tail of very poor variety-municipality observations mixed in with the normal ones.
>
> # (c) Smallest and largest yield
> range(mb_wheat$Yield_bu_ac)
[1] 4.5 97.8
>
> ## (c) range() returns both ends as one vector: the smallest yield is 4.5 and the largest is 97.8 bu/ac.
(a) Compute the 25th, 50th, and 75th percentiles of Yield_bu_ac in one quantile() call.
(b) Compute the IQR as the difference between two of your percentiles.
(c) Compare the IQR to the range. Which might give a better sense of the spread of the data?
Answer
# Load the tidyverse and read the wheat filelibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) The three quartiles in one callquantile(mb_wheat$Yield_bu_ac, probs =c(0.25, 0.50, 0.75))# (b) The IQR, as the 75th minus the 25th percentilequantile(mb_wheat$Yield_bu_ac, probs =0.75) -quantile(mb_wheat$Yield_bu_ac, probs =0.25)## I find quartiles of 53.9, 62.2, and 69.6 bu/ac, and an IQR of 15.7.# (c) The range, for comparisonmax(mb_wheat$Yield_bu_ac) -min(mb_wheat$Yield_bu_ac)## (c) The range is 93.3 bu/ac against an IQR of 15.7. The IQR usually## gives the better sense of spread: the range rests on the single best## and single worst results, while the IQR is the width of the middle## half of the data.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) The three quartiles in one call
> quantile(mb_wheat$Yield_bu_ac, probs = c(0.25, 0.50, 0.75))
25% 50% 75%
53.9 62.2 69.6
>
> # (b) The IQR, as the 75th minus the 25th percentile
> quantile(mb_wheat$Yield_bu_ac, probs = 0.75) - quantile(mb_wheat$Yield_bu_ac, probs = 0.25)
75%
15.7
>
> ## I find quartiles of 53.9, 62.2, and 69.6 bu/ac, and an IQR of 15.7.
>
> # (c) The range, for comparison
> max(mb_wheat$Yield_bu_ac) - min(mb_wheat$Yield_bu_ac)
[1] 93.3
>
> ## (c) The range is 93.3 bu/ac against an IQR of 15.7. The IQR usually
> ## gives the better sense of spread: the range rests on the single best
> ## and single worst results, while the IQR is the width of the middle
> ## half of the data.
(a) Compute the 5th and 95th percentiles of Yield_bu_ac in one call.
(b) In a comment: interpret the 95th percentile for a grower.
(c) Compute the gap between the two percentiles. In a comment: what fraction of all results does this span hold?
Answer
# Load the tidyverse and read the wheat filelibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) The 5th and 95th percentiles in one callquantile(mb_wheat$Yield_bu_ac, probs =c(0.05, 0.95))## (a) I find percentiles of 38.8 and 78.8 bu/ac.## (b) A result at the 95th percentile yielded 78.8 bu/ac -- only one in## twenty municipality-variety results beat that.# (c) The gap between themquantile(mb_wheat$Yield_bu_ac, probs =0.95) -quantile(mb_wheat$Yield_bu_ac, probs =0.05)## (c) The gap is about 40 bu/ac, and it holds the middle 90% of results.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) The 5th and 95th percentiles in one call
> quantile(mb_wheat$Yield_bu_ac, probs = c(0.05, 0.95))
5% 95%
38.785 78.800
>
> ## (a) I find percentiles of 38.8 and 78.8 bu/ac.
>
> ## (b) A result at the 95th percentile yielded 78.8 bu/ac -- only one in
> ## twenty municipality-variety results beat that.
>
> # (c) The gap between them
> quantile(mb_wheat$Yield_bu_ac, probs = 0.95) - quantile(mb_wheat$Yield_bu_ac, probs = 0.05)
95%
40.015
>
> ## (c) The gap is about 40 bu/ac, and it holds the middle 90% of results.
(a) Compute the range of Yield_bu_ac (largest minus smallest) and the IQR.
(b) In a comment: the two numbers disagree by a lot. What does each one describe?
(c) In a comment: which of the two measures would change if we discovered that the highest value in the data was actually a data error?
Answer
# Load the tidyverse and read the wheat filelibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) The rangemax(mb_wheat$Yield_bu_ac) -min(mb_wheat$Yield_bu_ac)# The IQRIQR(mb_wheat$Yield_bu_ac)## I find a range of 93.3 bu/ac against an IQR of 15.7.## (b) The range is the distance between the single best and single worst## results; the IQR is the width of the middle half.## (c) Only the range. The 97.8 is the maximum, so removing it moves the## range directly, while the quartiles sit deep in the middle of 2,398## results and barely notice.
> # Load the tidyverse and read the wheat file
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) The range
> max(mb_wheat$Yield_bu_ac) - min(mb_wheat$Yield_bu_ac)
[1] 93.3
>
> # The IQR
> IQR(mb_wheat$Yield_bu_ac)
[1] 15.7
>
> ## I find a range of 93.3 bu/ac against an IQR of 15.7.
> ## (b) The range is the distance between the single best and single worst
> ## results; the IQR is the width of the middle half.
> ## (c) Only the range. The 97.8 is the maximum, so removing it moves the
> ## range directly, while the quartiles sit deep in the middle of 2,398
> ## results and barely notice.
(a) Filter to the Spring Wheat observations of at least 60 bu/ac (all years) and saving the result. How many rows are in the filtered data frame?
(b) Filter again to keep only the 2023 rows of that result. How many rows are in this second data frame?
(c) From your 2023 data frame, select RM and Yield and sort the data by yield (greatest to least). Which RM tops the list, at what yield?
Answer
# Load the tidyverselibrary(tidyverse)# Read the long yield filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Spring Wheat at 60 bu/ac or better, all yearssw_60 <-filter(rm_yields, Crop =="Spring Wheat"& Yield >=60)# Row count of the filtered data framenrow(sw_60)## (a) I find 434 rows at 60 bu/ac or better.# (b) Keep only the 2023 rows of that resultsw_60_2023 <-filter(sw_60, Year ==2023)# Row count of the second data framenrow(sw_60_2023)## (b) 53 of them are from 2023.# (c) Just RM and Yield, sorted best firstsw_best <-select(sw_60_2023, RM, Yield)arrange(sw_best, desc(Yield))## (c) RM 187 tops the 2023 list, at 74.7 bu/ac.
> # Load the tidyverse
> library(tidyverse)
>
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Spring Wheat at 60 bu/ac or better, all years
> sw_60 <- filter(rm_yields, Crop == "Spring Wheat" & Yield >= 60)
> # Row count of the filtered data frame
> nrow(sw_60)
[1] 434
>
> ## (a) I find 434 rows at 60 bu/ac or better.
>
> # (b) Keep only the 2023 rows of that result
> sw_60_2023 <- filter(sw_60, Year == 2023)
> # Row count of the second data frame
> nrow(sw_60_2023)
[1] 53
>
> ## (b) 53 of them are from 2023.
>
> # (c) Just RM and Yield, sorted best first
> sw_best <- select(sw_60_2023, RM, Yield)
> arrange(sw_best, desc(Yield))
# A tibble: 53 × 2
RM Yield
<dbl> <dbl>
1 187 74.7
2 494 73
3 338 72.8
4 368 72.4
5 464 72.2
6 555 72.2
7 520 72
8 304 71.5
9 496 70.9
10 588 70.9
# ℹ 43 more rows
>
> ## (c) RM 187 tops the 2023 list, at 74.7 bu/ac.
(a) Filter to rows where the crop is Oats or Barley, using |, saving the result. How many rows are in the filtered data frame?
(b) Write the same filter using %in% and confirm the row count matches.
(c) What would filter(rm_yields, Crop == "Oats" & Crop == "Barley") return? In a comment: explain why.
Answer
# Load the tidyverselibrary(tidyverse)# Read the long yield filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Oats or Barley, using the | connectoroats_barley <-filter(rm_yields, Crop =="Oats"| Crop =="Barley")# Row count of the filtered data framenrow(oats_barley)# (b) The same filter written with %in%oats_barley2 <-filter(rm_yields, Crop %in%c("Oats", "Barley"))# The count should match (a)nrow(oats_barley2)## I find 19,709 rows both ways.## (c) It returns zero rows. The & version asks for rows whose crop is## Oats and Barley at the same time, and no single row can satisfy both## conditions. "Or" is the right connector, written |.
> # Load the tidyverse
> library(tidyverse)
>
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Oats or Barley, using the | connector
> oats_barley <- filter(rm_yields, Crop == "Oats" | Crop == "Barley")
> # Row count of the filtered data frame
> nrow(oats_barley)
[1] 19709
>
> # (b) The same filter written with %in%
> oats_barley2 <- filter(rm_yields, Crop %in% c("Oats", "Barley"))
> # The count should match (a)
> nrow(oats_barley2)
[1] 19709
>
> ## I find 19,709 rows both ways.
>
> ## (c) It returns zero rows. The & version asks for rows whose crop is
> ## Oats and Barley at the same time, and no single row can satisfy both
> ## conditions. "Or" is the right connector, written |.
(a) Filter to Canola in 2024, then in one mutate() call add two columns: yield_t_ha (multiply by 0.0560) and yield_kg_ha (multiply by 56.0). Save the result of each step.
(b) Compute the mean of each new column.
(c) In a comment: without running anything, what is the mean of yield_kg_ha divided by the mean of yield_t_ha, and why?
Answer
# Load the tidyverselibrary(tidyverse)# Read the long yield filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Canola in 2024canola_2024 <-filter(rm_yields, Crop =="Canola"& Year ==2024)# Add both converted columns in one mutate() callcanola_2024 <-mutate(canola_2024,yield_t_ha = Yield *0.0560, # tonnes per hectareyield_kg_ha = Yield *56.0) # kilograms per hectare# (b) Mean of each new columnmean(canola_2024$yield_t_ha)mean(canola_2024$yield_kg_ha)## (b) I find means of 1.76 t/ha and 1764.11 kg/ha.## (c) The ratio of the means is 1000: both columns are the same yields scaled by constants, and 56.0 / 0.0560 = 1000 (a tonne is 1000 kg).
> # Load the tidyverse
> library(tidyverse)
>
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Canola in 2024
> canola_2024 <- filter(rm_yields, Crop == "Canola" & Year == 2024)
> # Add both converted columns in one mutate() call
> canola_2024 <- mutate(canola_2024,
+ yield_t_ha = Yield * 0.0560, # tonnes per hectare
+ yield_kg_ha = Yield * 56.0) # kilograms per hectare
>
> # (b) Mean of each new column
> mean(canola_2024$yield_t_ha)
[1] 1.764115
> mean(canola_2024$yield_kg_ha)
[1] 1764.115
>
> ## (b) I find means of 1.76 t/ha and 1764.11 kg/ha.
>
> ## (c) The ratio of the means is 1000: both columns are the same yields scaled by constants, and 56.0 / 0.0560 = 1000 (a tonne is 1000 kg).
(a) Filter to Canola in 2025 and, in one mutate(), add a revenue_per_ac column (yield times price). Save the result.
(b) Sort by revenue (greatest to least), and report the top RM and its revenue per acre.
(c) Compute the mean revenue per acre.
Answer
# Load the tidyverse and read the datalibrary(tidyverse)rm_yields <-read_csv("data/rm_yields_1990_2025.csv")# Price objectprice <-14.20# (a) 2025 canola with revenue per acrecanola_25 <-filter(rm_yields, Crop =="Canola"& Year ==2025)canola_25 <-mutate(canola_25, revenue_per_ac = Yield * price)# (b) Best revenue firstarrange(canola_25, desc(revenue_per_ac))## (b) I find RM 287 on top at $866.20 per acre.# (c) Mean revenue per acremean(canola_25$revenue_per_ac)## (c) The mean is $623.94 per acre.
> # Load the tidyverse and read the data
> library(tidyverse)
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # Price object
> price <- 14.20
>
> # (a) 2025 canola with revenue per acre
> canola_25 <- filter(rm_yields, Crop == "Canola" & Year == 2025)
> canola_25 <- mutate(canola_25, revenue_per_ac = Yield * price)
>
> # (b) Best revenue first
> arrange(canola_25, desc(revenue_per_ac))
# A tibble: 290 × 6
Year RM Crop Yield Unit revenue_per_ac
<dbl> <dbl> <chr> <dbl> <chr> <dbl>
1 2025 287 Canola 61 bu/ac 866.
2 2025 292 Canola 57.3 bu/ac 814.
3 2025 257 Canola 56.7 bu/ac 805.
4 2025 259 Canola 56 bu/ac 795.
5 2025 380 Canola 55.6 bu/ac 790.
6 2025 128 Canola 55.4 bu/ac 787.
7 2025 345 Canola 54.9 bu/ac 780.
8 2025 225 Canola 54.8 bu/ac 778.
9 2025 276 Canola 54.6 bu/ac 775.
10 2025 260 Canola 54.5 bu/ac 774.
# ℹ 280 more rows
>
> ## (b) I find RM 287 on top at $866.20 per acre.
>
> # (c) Mean revenue per acre
> mean(canola_25$revenue_per_ac)
[1] 623.9431
>
> ## (c) The mean is $623.94 per acre.
(a) Filter rm_yields to Canola in 2024, saving the result as canola_2024. How many rows are in the filtered data frame?
(b) From canola_2024, select only the RM and Yield columns, saving the result. Report its column names and its number of columns.
(c) In a comment: after (a) and (b), how many rows does rm_yields itself have, and why?
Answer
# Load the tidyverselibrary(tidyverse)# Read the long yield filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Keep only the Canola rows from 2024canola_2024 <-filter(rm_yields, Crop =="Canola"& Year ==2024)# Row count of the filtered data framenrow(canola_2024)# (b) Keep only the RM and Yield columnscanola_slim <-select(canola_2024, RM, Yield)# Column names of the slimmed versionnames(canola_slim)# Number of columnsncol(canola_slim)# (c) The original data frame, after (a) and (b)nrow(rm_yields)## I find 293 rows of 2024 canola.## The selected version has the columns RM and Yield, so 2 columns (and still 293 rows).## rm_yields itself still has 71104 rows: filter() and select() return new data frames and leave the original untouched.
> # Load the tidyverse
> library(tidyverse)
>
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Keep only the Canola rows from 2024
> canola_2024 <- filter(rm_yields, Crop == "Canola" & Year == 2024)
> # Row count of the filtered data frame
> nrow(canola_2024)
[1] 293
>
> # (b) Keep only the RM and Yield columns
> canola_slim <- select(canola_2024, RM, Yield)
> # Column names of the slimmed version
> names(canola_slim)
[1] "RM" "Yield"
> # Number of columns
> ncol(canola_slim)
[1] 2
>
> # (c) The original data frame, after (a) and (b)
> nrow(rm_yields)
[1] 71104
>
> ## I find 293 rows of 2024 canola.
> ## The selected version has the columns RM and Yield, so 2 columns (and still 293 rows).
> ## rm_yields itself still has 71104 rows: filter() and select() return new data frames and leave the original untouched.
(a) Filter to rows with at least 10 farms and a yield above 70 bu/ac, saving the result. How many rows are in it?
(b) Sort by yields (greatest to least) and report the top municipality, variety, and year.
(c) In a comment: compare this with filtering on yield alone without the minimum number of farms. Comment on the trustworthiness of these results versus those in part b.
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) At least 10 farms AND yield above 70solid_high <-filter(mb_wheat, Farms >=10& Yield_bu_ac >70)nrow(solid_high)# (b) Best firstarrange(solid_high, desc(Yield_bu_ac))## I find 273 rows; LOUISE growing SY MANNESS in 2025 tops the list## at 96.3 bu/ac.# (c) The same filter on yield alonehigh_only <-filter(mb_wheat, Yield_bu_ac >70)nrow(high_only)## (c) Filtering on yield alone keeps 569 rows, and its top result is a## 97.8 bu/ac average from just 5 farms. An average from a few farms is## more sensitive to one unusual result, so the part b list is the more## trustworthy ranking; the farm minimum gives every row a broader base.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) At least 10 farms AND yield above 70
> solid_high <- filter(mb_wheat, Farms >= 10 & Yield_bu_ac > 70)
> nrow(solid_high)
[1] 273
>
> # (b) Best first
> arrange(solid_high, desc(Yield_bu_ac))
# A tibble: 273 × 7
Year Municipality Variety Farms Acres Yield_bu_ac Reported
<dbl> <chr> <chr> <dbl> <dbl> <dbl> <lgl>
1 2025 LOUISE SY MANNESS 14 3252 96.3 TRUE
2 2025 GRANDVIEW AAC STARBUCK <SECAN> 10 4456 90.2 TRUE
3 2025 ROLAND SY MANNESS 11 2986 88.6 TRUE
4 2025 STE. ANNE AAC BRANDON (BW 932) 17 2445 87.9 TRUE
5 2024 LOUISE AAC WHEATLAND <SECA… 19 7576 86.9 TRUE
6 2025 GRANDVIEW AAC BRANDON (BW 932) 26 12672 86.4 TRUE
7 2024 MORRIS SY MANNESS 28 5288 84.6 TRUE
8 2025 LOUISE AAC STARBUCK <SECAN> 28 12761 84.4 TRUE
9 2024 MACDONALD SY MANNESS 21 7776 84.1 TRUE
10 2025 RUSSELL-BINSCARTH AAC WHEATLAND <SECA… 14 7644 84 TRUE
# ℹ 263 more rows
>
> ## I find 273 rows; LOUISE growing SY MANNESS in 2025 tops the list
> ## at 96.3 bu/ac.
>
> # (c) The same filter on yield alone
> high_only <- filter(mb_wheat, Yield_bu_ac > 70)
> nrow(high_only)
[1] 569
>
> ## (c) Filtering on yield alone keeps 569 rows, and its top result is a
> ## 97.8 bu/ac average from just 5 farms. An average from a few farms is
> ## more sensitive to one unusual result, so the part b list is the more
> ## trustworthy ranking; the farm minimum gives every row a broader base.
(a) Filter the data to the 2025 rows, saving the result. How many rows are in the filtered data frame?
(b) From the filtered data, select Municipality, Variety, and Yield_bu_ac, saving the result. Report its column names.
(c) In a comment: how many rows does the original data frame still have, and what happened to the other years’ rows?
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Keep the 2025 rows onlywheat_2025 <-filter(mb_wheat, Year ==2025)# Row count of the filtered data framenrow(wheat_2025)# (b) Three columns from the 2025 rowswheat_slim <-select(wheat_2025, Municipality, Variety, Yield_bu_ac)# Its column namesnames(wheat_slim)# (c) The original data frame, after (a) and (b)nrow(mb_wheat)## I find 438 rows for 2025.## The selected columns are Municipality, Variety, Yield_bu_ac.## mb_wheat itself still has 2398 rows -- filtering returns a new data frame. The other years' rows are simply absent from the filtered copy; they were not deleted from anything.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Keep the 2025 rows only
> wheat_2025 <- filter(mb_wheat, Year == 2025)
> # Row count of the filtered data frame
> nrow(wheat_2025)
[1] 438
>
> # (b) Three columns from the 2025 rows
> wheat_slim <- select(wheat_2025, Municipality, Variety, Yield_bu_ac)
> # Its column names
> names(wheat_slim)
[1] "Municipality" "Variety" "Yield_bu_ac"
>
> # (c) The original data frame, after (a) and (b)
> nrow(mb_wheat)
[1] 2398
>
> ## I find 438 rows for 2025.
> ## The selected columns are Municipality, Variety, Yield_bu_ac.
> ## mb_wheat itself still has 2398 rows -- filtering returns a new data frame. The other years' rows are simply absent from the filtered copy; they were not deleted from anything.
(a) Filter the data to the varieties AAC BRANDON (BW 932)orAAC STARBUCK <SECAN> using |, saving the result. How many rows are in the filtered data frame?
(b) Write the same filter using %in% and confirm the count matches.
(c) If in part (a) you used & instead of |, what would the result be? Explain.
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) The two varieties, using the | connectortwo_var <-filter(mb_wheat, Variety =="AAC BRANDON (BW 932)"| Variety =="AAC STARBUCK <SECAN>")# Row count of the filtered data framenrow(two_var)# (b) The same filter written with %in%two_var2 <-filter(mb_wheat, Variety %in%c("AAC BRANDON (BW 932)", "AAC STARBUCK <SECAN>"))# The count should match (a)nrow(two_var2)## I find 893 rows both ways.## (c) It would return zero rows. The & version demands both conditions true of the same row, and no single row has a variety that is both names at once. "Or" is the right connector, written |.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) The two varieties, using the | connector
> two_var <- filter(mb_wheat, Variety == "AAC BRANDON (BW 932)" | Variety == "AAC STARBUCK <SECAN>")
> # Row count of the filtered data frame
> nrow(two_var)
[1] 893
>
> # (b) The same filter written with %in%
> two_var2 <- filter(mb_wheat, Variety %in% c("AAC BRANDON (BW 932)", "AAC STARBUCK <SECAN>"))
> # The count should match (a)
> nrow(two_var2)
[1] 893
>
> ## I find 893 rows both ways.
> ## (c) It would return zero rows. The & version demands both conditions true of the same row, and no single row has a variety that is both names at once. "Or" is the right connector, written |.
(a) Use one mutate() to add acres_per_farm (acres divided by farms), saving the result.
(b) Sort by acres_per_farm, largest first, and report the top municipality and variety.
(c) Compute the mean of acres_per_farm. In a comment, explain what this number means in the context of the data.
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Acres per farmmb_wheat <-mutate(mb_wheat, acres_per_farm = Acres / Farms)# (b) Largest firstarrange(mb_wheat, desc(acres_per_farm))## (b) I find MOUNTAIN growing SY MANNESS on top, at 1,587 acres per farm.# (c) The average across rowsmean(mb_wheat$acres_per_farm)## (c) The mean is 370: in a typical municipality-variety-year result,## each farm growing that variety seeded about 370 acres of it. It## measures how concentrated a variety is in the hands of its growers --## a different fact from how much land it covers in total.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Acres per farm
> mb_wheat <- mutate(mb_wheat, acres_per_farm = Acres / Farms)
>
> # (b) Largest first
> arrange(mb_wheat, desc(acres_per_farm))
# A tibble: 2,398 × 8
Year Municipality Variety Farms Acres Yield_bu_ac Reported acres_per_farm
<dbl> <chr> <chr> <dbl> <dbl> <dbl> <lgl> <dbl>
1 2025 MOUNTAIN SY MAN… 3 4761 77.8 TRUE 1587
2 2023 DELORAINE-WINC… AAC EL… 4 5994 82.2 TRUE 1498.
3 2025 WESTLAKE-GLADS… AAC WH… 3 4483 68.6 TRUE 1494.
4 2025 NORTH CYPRESS-… AAC WH… 5 7228 76.6 TRUE 1446.
5 2025 ALONSA AAC HO… 5 6469 45.1 TRUE 1294.
6 2025 GLENELLA-LANSD… AAC WH… 3 3504 78.5 TRUE 1168
7 2023 MINITONAS-BOWS… AAC BR… 4 4648 66.1 TRUE 1162
8 2021 DELORAINE-WINC… AAC EL… 5 5763 62.5 TRUE 1153.
9 2020 ELLICE-ARCHIE AAC VI… 4 4606 63.2 TRUE 1152.
10 2025 ROCKWOOD BOLLES… 4 4363 66.6 TRUE 1091.
# ℹ 2,388 more rows
>
> ## (b) I find MOUNTAIN growing SY MANNESS on top, at 1,587 acres per farm.
>
> # (c) The average across rows
> mean(mb_wheat$acres_per_farm)
[1] 370.0478
>
> ## (c) The mean is 370: in a typical municipality-variety-year result,
> ## each farm growing that variety seeded about 370 acres of it. It
> ## measures how concentrated a variety is in the hands of its growers --
> ## a different fact from how much land it covers in total.
(a) Use one mutate() to add total_bu, each variety’s total production in a municipality in bushels (acres times yield), saving the result.
(b) Sort by total_bu, largest first, and report the top municipality and variety.
(c) In a comment: why is the highest-yielding observation not the one with the greatest production?
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Total production: acres times the per-acre yieldmb_wheat <-mutate(mb_wheat, total_bu = Acres * Yield_bu_ac)# (b) Largest firstarrange(mb_wheat, desc(total_bu))## (b) I find PORTAGE LA PRAIRIE growing AAC BRANDON (BW 932) in 2020 on## top, at about 5.1 million bushels.## (c) Acreage. The top producer yielded a modest 68.8 bu/ac but on## 74,604 acres, while the highest-yielding result (97.8 bu/ac) sat on## under 2,000 acres. Total production rewards land as much as yield.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Total production: acres times the per-acre yield
> mb_wheat <- mutate(mb_wheat, total_bu = Acres * Yield_bu_ac)
>
> # (b) Largest first
> arrange(mb_wheat, desc(total_bu))
# A tibble: 2,398 × 8
Year Municipality Variety Farms Acres Yield_bu_ac Reported total_bu
<dbl> <chr> <chr> <dbl> <dbl> <dbl> <lgl> <dbl>
1 2020 PORTAGE LA PRAIRIE AAC BRAN… 122 74604 68.8 TRUE 5132755.
2 2020 MACDONALD AAC BRAN… 115 55104. 71.6 TRUE 3945418.
3 2024 PORTAGE LA PRAIRIE AAC BRAN… 84 51923 67 TRUE 3478841
4 2022 MINITONAS-BOWSMAN AAC VIEW… 74 42055 82.4 TRUE 3465332
5 2025 PORTAGE LA PRAIRIE AAC BRAN… 83 45210 75.3 TRUE 3404313
6 2023 PORTAGE LA PRAIRIE AAC BRAN… 85 48876 67.5 TRUE 3299130
7 2022 PORTAGE LA PRAIRIE AAC BRAN… 86 51489 63.1 TRUE 3248956.
8 2020 WESTLAKE-GLADSTONE AAC BRAN… 75 49160 65.5 TRUE 3219980
9 2022 SWAN VALLEY WEST AAC VIEW… 71 42788 74.6 TRUE 3191985.
10 2020 WALLACE-WOODWORTH AAC BRAN… 106 48658 65.2 TRUE 3172502.
# ℹ 2,388 more rows
>
> ## (b) I find PORTAGE LA PRAIRIE growing AAC BRANDON (BW 932) in 2020 on
> ## top, at about 5.1 million bushels.
> ## (c) Acreage. The top producer yielded a modest 68.8 bu/ac but on
> ## 74,604 acres, while the highest-yielding result (97.8 bu/ac) sat on
> ## under 2,000 acres. Total production rewards land as much as yield.
(a) In one pipeline: for 2025 only, compute each crop’s mean yield and observation count, sorted from highest mean to lowest, and save the result as crop_summary_2025.
(b) Write the table to output/crop_summary_2025.csv.
(c) Report the two highest-yielding crops in the sorted table. In a comment: check the Unit column – why is the first-place crop not comparable to the rest, and which bushel crop is the real standout?
Answer
# Load the tidyverselibrary(tidyverse)# Read the long yield filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Each crop in 2025: mean yield and observation count, best firstcrop_summary_2025 <- rm_yields |>filter(Year ==2025) |># 2025 onlygroup_by(Crop) |># one group per cropsummarise(mean_yield =mean(Yield), # average yieldn =n()) |># observation countarrange(desc(mean_yield)) # highest mean first# Print the sorted tablecrop_summary_2025# (b) Write the table to a csvwrite_csv(crop_summary_2025, "output/crop_summary_2025.csv")## I find Lentils on top at 1897 and Oats second at 97.5.## (c) The Unit column shows Lentils are recorded in pounds per acre, not## bushels, so the first row is not comparable to the rest. Among the## bushel crops, Oats are highest at 97.5 bu/ac.
> # Load the tidyverse
> library(tidyverse)
>
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Each crop in 2025: mean yield and observation count, best first
> crop_summary_2025 <- rm_yields |>
+ filter(Year == 2025) |> # 2025 only
+ group_by(Crop) |> # one group per crop
+ summarise(mean_yield = mean(Yield), # average yield
+ n = n()) |> # observation count
+ arrange(desc(mean_yield)) # highest mean first
>
> # Print the sorted table
> crop_summary_2025
# A tibble: 8 × 3
Crop mean_yield n
<chr> <dbl> <int>
1 Lentils 1897. 218
2 Oats 97.5 199
3 Barley 72.8 288
4 Spring Wheat 53.0 283
5 Durum 48.6 189
6 Canola 43.9 290
7 Peas 43.5 285
8 Flax 30.5 195
>
> # (b) Write the table to a csv
> write_csv(crop_summary_2025, "output/crop_summary_2025.csv")
>
> ## I find Lentils on top at 1897 and Oats second at 97.5.
> ## (c) The Unit column shows Lentils are recorded in pounds per acre, not
> ## bushels, so the first row is not comparable to the rest. Among the
> ## bushel crops, Oats are highest at 97.5 bu/ac.
(a) In one pipeline: filter to Lentils from 2021 through 2025, convert from lb/acre to kg/ha with mutate() (multiply by 1.12), and compute the mean converted yield and n() for each year. Sort from highest mean to lowest and save the result as lentil_summary.
(b) Report the highest and lowest annual means and the counts behind them.
(c) In a comment: in what order do your steps run, and why would putting the summarise() before the mutate() fail?
Answer
# Load the tidyverselibrary(tidyverse)# Read the long yield filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Recent lentils, converted to kg/ha and summarised by yearlentil_summary <- rm_yields |>filter(Crop =="Lentils"& Year >=2021) |># recent lentil rowsmutate(yield_kg_ha = Yield *1.12) |># convert lb/ac to kg/hagroup_by(Year) |># one group per yearsummarise(mean_kg_ha =mean(yield_kg_ha), # annual averagen =n()) |># observations behind itarrange(desc(mean_kg_ha)) # highest mean first# Print the summary tablelentil_summary## (b) The highest annual mean is 2025 at 2,124.4 kg/ha from 218## observations. The lowest is 2021 at 1,057.5 kg/ha from 194.## (c) The steps filter, convert, group, summarise, and sort in that order.## Putting summarise() before mutate() would fail because yield_kg_ha would## not exist when the summary tried to use it.
> # Load the tidyverse
> library(tidyverse)
>
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Recent lentils, converted to kg/ha and summarised by year
> lentil_summary <- rm_yields |>
+ filter(Crop == "Lentils" & Year >= 2021) |> # recent lentil rows
+ mutate(yield_kg_ha = Yield * 1.12) |> # convert lb/ac to kg/ha
+ group_by(Year) |> # one group per year
+ summarise(mean_kg_ha = mean(yield_kg_ha), # annual average
+ n = n()) |> # observations behind it
+ arrange(desc(mean_kg_ha)) # highest mean first
>
> # Print the summary table
> lentil_summary
# A tibble: 5 × 3
Year mean_kg_ha n
<dbl> <dbl> <int>
1 2025 2124. 218
2 2024 1478. 213
3 2023 1472. 196
4 2022 1396. 196
5 2021 1058. 194
>
> ## (b) The highest annual mean is 2025 at 2,124.4 kg/ha from 218
> ## observations. The lowest is 2021 at 1,057.5 kg/ha from 194.
> ## (c) The steps filter, convert, group, summarise, and sort in that order.
> ## Putting summarise() before mutate() would fail because yield_kg_ha would
> ## not exist when the summary tried to use it.
(a) In one pipeline: filter to Canola observations in 2023-2025, add revenue_per_ac with mutate(), then compute the mean revenue for each year.
(b) Report the three yearly means.
Answer
# Load the tidyverse and read the datalibrary(tidyverse)rm_yields <-read_csv("data/rm_yields_1990_2025.csv")# Price objectprice <-14.20# (a) Mean canola revenue per acre by year, 2023-2025rm_yields |>filter(Crop =="Canola"& Year >=2023) |># recent canola rowsmutate(revenue_per_ac = Yield * price) |># dollars per acregroup_by(Year) |># one group per yearsummarise(mean_revenue =mean(revenue_per_ac)) # average within each year## (b) I find mean revenues of $481.11 in 2023, $447.33 in 2024, and## $623.94 in 2025.
> # Load the tidyverse and read the data
> library(tidyverse)
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # Price object
> price <- 14.20
>
> # (a) Mean canola revenue per acre by year, 2023-2025
> rm_yields |>
+ filter(Crop == "Canola" & Year >= 2023) |> # recent canola rows
+ mutate(revenue_per_ac = Yield * price) |> # dollars per acre
+ group_by(Year) |> # one group per year
+ summarise(mean_revenue = mean(revenue_per_ac)) # average within each year
# A tibble: 3 × 2
Year mean_revenue
<dbl> <dbl>
1 2023 481.
2 2024 447.
3 2025 624.
>
> ## (b) I find mean revenues of $481.11 in 2023, $447.33 in 2024, and
> ## $623.94 in 2025.
(a) In one pipeline, compute the mean canola yield for each year from 2021 to 2025 (filter first, then group).
(b) 2021 was a severe drought year in Saskatchewan. Does the data agree?
Answer
# Load the tidyverselibrary(tidyverse)# Read the long yield filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Mean canola yield for each year from 2021 to 2025rm_yields |>filter(Crop =="Canola"& Year >=2021) |># canola, 2021 onwardgroup_by(Year) |># one group per yearsummarise(mean_yield =mean(Yield)) # average within each year## (b) Yes, the data agrees: the 2021 mean sits roughly a third below the surrounding years.
> # Load the tidyverse
> library(tidyverse)
>
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Mean canola yield for each year from 2021 to 2025
> rm_yields |>
+ filter(Crop == "Canola" & Year >= 2021) |> # canola, 2021 onward
+ group_by(Year) |> # one group per year
+ summarise(mean_yield = mean(Yield)) # average within each year
# A tibble: 5 × 2
Year mean_yield
<dbl> <dbl>
1 2021 21.9
2 2022 35.4
3 2023 33.9
4 2024 31.5
5 2025 43.9
>
> ## (b) Yes, the data agrees: the 2021 mean sits roughly a third below the surrounding years.
(a) In one pipeline: filter to Flax from 2020 through 2025, compute each RM’s mean_yield and number of observations, retain RMs with at least four observations, and sort from highest mean to lowest. Save the result as flax_summary.
(b) Report the top RM, its mean, and its count.
(c) In a comment: why does the minimum-count rule improve the ranking, and why must filter(n >= 4) come after summarise()?
Answer
# Load the tidyverselibrary(tidyverse)# Read the long yield filerm_yields <-read_csv("data/rm_yields_1990_2025.csv")# (a) Recent flax means by RM, retaining RMs with at least four reportsflax_summary <- rm_yields |>filter(Crop =="Flax"& Year >=2020) |># recent flax rowsgroup_by(RM) |># one group per RMsummarise(mean_yield =mean(Yield), # average within the RMn =n()) |># observations behind itfilter(n >=4) |># enough years to comparearrange(desc(mean_yield)) # highest mean first# Print the summary tableflax_summary## (b) RM 336 is highest at 34.5 bu/ac from 4 observations.## (c) The rule removes rankings based on only one or two years. The count## n is created by summarise(), so it cannot be filtered before that step.
> # Load the tidyverse
> library(tidyverse)
>
> # Read the long yield file
> rm_yields <- read_csv("data/rm_yields_1990_2025.csv")
Rows: 71104 Columns: 5
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Crop, Unit
dbl (3): Year, RM, Yield
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Recent flax means by RM, retaining RMs with at least four reports
> flax_summary <- rm_yields |>
+ filter(Crop == "Flax" & Year >= 2020) |> # recent flax rows
+ group_by(RM) |> # one group per RM
+ summarise(mean_yield = mean(Yield), # average within the RM
+ n = n()) |> # observations behind it
+ filter(n >= 4) |> # enough years to compare
+ arrange(desc(mean_yield)) # highest mean first
>
> # Print the summary table
> flax_summary
# A tibble: 191 × 3
RM mean_yield n
<dbl> <dbl> <int>
1 336 34.5 4
2 442 33.1 5
3 64 32.6 6
4 63 32.0 5
5 400 31.4 6
6 211 31.3 6
7 61 30.9 6
8 369 30.8 4
9 438 30.4 6
10 351 29.8 6
# ℹ 181 more rows
>
> ## (b) RM 336 is highest at 34.5 bu/ac from 4 observations.
> ## (c) The rule removes rankings based on only one or two years. The count
> ## n is created by summarise(), so it cannot be filtered before that step.
(a) In one pipeline, compute each variety’s total acres (sum(Acres)), sorted greatest to least.
(b) In a comment: totalling acres makes sense, but what would totalling Yield_bu_ac produce, and why is that number meaningless?
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Total acres by variety, largest firstmb_wheat |>group_by(Variety) |># one group per varietysummarise(total_acres =sum(Acres)) |># acres are amounts, so they addarrange(desc(total_acres)) # biggest variety first## (b) Totalling Yield_bu_ac would add per-acre rates across rows -- a## number with no physical meaning, since a rate says how much each acre## gave, not how much there was. Acres are an amount, so their total is a## real quantity of land.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Total acres by variety, largest first
> mb_wheat |>
+ group_by(Variety) |> # one group per variety
+ summarise(total_acres = sum(Acres)) |> # acres are amounts, so they add
+ arrange(desc(total_acres)) # biggest variety first
# A tibble: 44 × 2
Variety total_acres
<chr> <dbl>
1 AAC BRANDON (BW 932) 6741112.
2 AAC STARBUCK <SECAN> 2503939.
3 AAC WHEATLAND <SECAN> 1399093.
4 AAC VIEWFIELD <FP GENETICS>|BW965| EXP 1114296.
5 AAC HOCKLEY |BW5044| 588044.
6 AAC REDBERRY 474429.
7 SY MANNESS 389027
8 AAC ELIE(BW931) 356550
9 BOLLES <SEED DEPOT> 273594
10 AAC HODGE |BW1069| 176698.
# ℹ 34 more rows
>
>
> ## (b) Totalling Yield_bu_ac would add per-acre rates across rows -- a
> ## number with no physical meaning, since a rate says how much each acre
> ## gave, not how much there was. Acres are an amount, so their total is a
> ## real quantity of land.
(a) In one pipeline, compute each variety’s mean yield and observation count, sorted by mean yield from greatest to least.
(b) Report the top two varieties, their means, and their counts.
(c) In a comment: Report the count of the top three varieties. Do you find the average yield of each variety trustworthy?
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Each variety: mean yield and observation count, best firstmb_wheat |>group_by(Variety) |># one group per varietysummarise(mean_yield =mean(Yield_bu_ac), # average yieldn =n()) |># observation countarrange(desc(mean_yield)) # highest mean first## (b) I find AAC WESTKING first at 79.2 bu/ac from 17 observations, then CDC GO (BW781) at 77.5 from a single observation.## (c) The top three counts are 17 (AAC WESTKING), 1 (CDC GO), and## 2 (AAC SPIKE). I would trust the AAC WESTKING mean: seventeen## municipality-year results is a reasonable base. The other two rest on## one or two results, so their averages could be one municipality having## a good year.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Each variety: mean yield and observation count, best first
> mb_wheat |>
+ group_by(Variety) |> # one group per variety
+ summarise(mean_yield = mean(Yield_bu_ac), # average yield
+ n = n()) |> # observation count
+ arrange(desc(mean_yield)) # highest mean first
# A tibble: 44 × 3
Variety mean_yield n
<chr> <dbl> <int>
1 AAC WESTKING 79.2 17
2 CDC GO (BW781) 77.5 1
3 AAC SPIKE 73.2 2
4 SY MANNESS 72.1 108
5 AAC WHEATLAND <SECAN> 65.8 219
6 AAC BROADACRES <PROVEN> 65.0 18
7 AAC HOCKLEY |BW5044| 64.9 177
8 GLENN 64.5 4
9 HARVEST (BW259) 64.4 1
10 AAC CONNERY |PT245| 64.1 2
# ℹ 34 more rows
>
> ## (b) I find AAC WESTKING first at 79.2 bu/ac from 17 observations, then CDC GO (BW781) at 77.5 from a single observation.
> ## (c) The top three counts are 17 (AAC WESTKING), 1 (CDC GO), and
> ## 2 (AAC SPIKE). I would trust the AAC WESTKING mean: seventeen
> ## municipality-year results is a reasonable base. The other two rest on
> ## one or two results, so their averages could be one municipality having
> ## a good year.
(a) In one pipeline: for 2025 only, compute each variety’s mean reported yield and observation count, sorted best-first, and save the result as variety_summary_2025.
(b) Write the table to output/variety_summary_2025.csv.
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Each variety in 2025: mean reported yield and count, best firstvariety_summary_2025 <- mb_wheat |>filter(Year ==2025) |># 2025 onlygroup_by(Variety) |># one group per varietysummarise(mean_yield =mean(Yield_bu_ac), # average yieldn =n()) |># observation countarrange(desc(mean_yield)) # highest mean first# Print the top of the sorted tablehead(variety_summary_2025)# (b) Write the table to a csvwrite_csv(variety_summary_2025, "output/variety_summary_2025.csv")
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Each variety in 2025: mean reported yield and count, best first
> variety_summary_2025 <- mb_wheat |>
+ filter(Year == 2025) |> # 2025 only
+ group_by(Variety) |> # one group per variety
+ summarise(mean_yield = mean(Yield_bu_ac), # average yield
+ n = n()) |> # observation count
+ arrange(desc(mean_yield)) # highest mean first
>
> # Print the top of the sorted table
> head(variety_summary_2025)
# A tibble: 6 × 3
Variety mean_yield n
<chr> <dbl> <int>
1 AAC WESTKING 79.2 17
2 AAC VIEWFIELD <FP GENETICS>|BW965| EXP 73.5 16
3 AAC SPIKE 73.2 2
4 SY MANNESS 72.6 62
5 CDC LANDMARK 72.3 1
6 AAC WHEATLAND <SECAN> 71.5 50
>
> # (b) Write the table to a csv
> write_csv(variety_summary_2025, "output/variety_summary_2025.csv")
(a) In one pipeline, compute each municipality’s mean yield and observation count, sorted from highest mean to lowest.
(b) Report the top municipality and how many observations back its mean.
(c) In a comment: a municipality mean pools every variety grown there. What two different things could make one municipality top this table?
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) Mean yield and count by municipality, best firstmb_wheat |>group_by(Municipality) |># one group per municipalitysummarise(mean_yield =mean(Yield_bu_ac), # average across its rowsn =n()) |># rows behind the averagearrange(desc(mean_yield)) # best municipality first## (b) I find LOUISE on top at 72.86 bu/ac, backed by 26 observations.## (c) Two different things could put a municipality here: genuinely better## growing conditions, or a mix that happens to lean on high-yielding## varieties. The mean pools varieties, so the ranking cannot tell the## land from the lineup.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) Mean yield and count by municipality, best first
> mb_wheat |>
+ group_by(Municipality) |> # one group per municipality
+ summarise(mean_yield = mean(Yield_bu_ac), # average across its rows
+ n = n()) |> # rows behind the average
+ arrange(desc(mean_yield)) # best municipality first
# A tibble: 96 × 3
Municipality mean_yield n
<chr> <dbl> <int>
1 LOUISE 72.9 26
2 LAC DU BONNET 71.5 25
3 ALEXANDER 71.4 17
4 ELTON 70.7 23
5 RHINELAND 70.4 22
6 THOMPSON 69.6 15
7 MINITONAS-BOWSMAN 69.5 33
8 CARTIER 69.3 19
9 PEMBINA 69.2 34
10 EMERSON-FRANKLIN 68.8 16
# ℹ 86 more rows
>
> ## (b) I find LOUISE on top at 72.86 bu/ac, backed by 26 observations.
> ## (c) Two different things could put a municipality here: genuinely better
> ## growing conditions, or a mix that happens to lean on high-yielding
> ## varieties. The mean pools varieties, so the ranking cannot tell the
> ## land from the lineup.
(a) In one pipeline: filter the data to 2023, compute each variety’s mean yield and count, sort best-first, and save the result.
(b) Report the top variety by mean and its count, and the largest count in the table and its variety.
(c) In a comment: are they the same variety? What do the two columns measure differently?
Answer
# Load the tidyverse and read the datalibrary(tidyverse)mb_wheat <-read_csv("data/mb_wheat_reported_2020_2025.csv")# (a) 2023 variety means and counts, best firstvariety_2023 <- mb_wheat |>filter(Year ==2023) |># one year onlygroup_by(Variety) |># one group per varietysummarise(mean_yield =mean(Yield_bu_ac), # average within the varietyn =n()) |># results behind itarrange(desc(mean_yield)) # best mean firstvariety_2023## (b) I find AC BARRIE (BW 661) with the best 2023 mean, 67.3 bu/ac --## from a single observation. The largest count is AAC BRANDON (BW 932),## with 84.## (c) They are not the same variety. The mean measures how well a variety## did where it was grown; the count measures how widely it was grown. A## chart-topping mean on n = 1 is one municipality having a good year.
> # Load the tidyverse and read the data
> library(tidyverse)
> mb_wheat <- read_csv("data/mb_wheat_reported_2020_2025.csv")
Rows: 2398 Columns: 7
── Column specification ────────────────────────────────────────────────────────
Delimiter: ","
chr (2): Municipality, Variety
dbl (4): Year, Farms, Acres, Yield_bu_ac
lgl (1): Reported
ℹ Use `spec()` to retrieve the full column specification for this data.
ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.
>
> # (a) 2023 variety means and counts, best first
> variety_2023 <- mb_wheat |>
+ filter(Year == 2023) |> # one year only
+ group_by(Variety) |> # one group per variety
+ summarise(mean_yield = mean(Yield_bu_ac), # average within the variety
+ n = n()) |> # results behind it
+ arrange(desc(mean_yield)) # best mean first
> variety_2023
# A tibble: 26 × 3
Variety mean_yield n
<chr> <dbl> <int>
1 AC BARRIE (BW 661) 67.3 1
2 AAC WHEATLAND <SECAN> 66.2 44
3 SY MANNESS 65.8 11
4 AAC STARBUCK <SECAN> 63.8 76
5 AAC HOCKLEY |BW5044| 63.6 51
6 AAC HODGE |BW1069| 63.0 30
7 AAC ELIE(BW931) 62.9 8
8 BOLLES <SEED DEPOT> 61.9 15
9 AAC VIEWFIELD <FP GENETICS>|BW965| EXP 61.7 22
10 CDC SKRUSH |PT599| 60.5 1
# ℹ 16 more rows
>
> ## (b) I find AC BARRIE (BW 661) with the best 2023 mean, 67.3 bu/ac --
> ## from a single observation. The largest count is AAC BRANDON (BW 932),
> ## with 84.
> ## (c) They are not the same variety. The mean measures how well a variety
> ## did where it was grown; the count measures how widely it was grown. A
> ## chart-topping mean on n = 1 is one municipality having a good year.