Module 2 lab: getting started in R

Work through this at your own pace. Each step ends with a check – something you should be able to see on your screen before moving on. If a check fails, that is the point to put your hand up rather than pushing on.

The whole thing should take about an hour. It covers the same ground as 4  Getting Started with R and 5  Loading Data and Packages in R, so those chapters are the place to look if you want more detail on any step.

Step 1: Install R and Positron

Two separate installs, and the order matters. R is the language – the thing that actually does the arithmetic. Positron is the program you write R in: an editor, a place to see your files, and a window onto R. Installing Positron does not install R, which is why you do both.

  1. Install R from https://cran.rstudio.com. Pick the version for your operating system and accept the defaults. There is nothing to configure.
  2. Install Positron from https://positron.posit.co.
  3. Open Positron.

Positron looks for R on your computer when it starts. If R is not installed yet, it opens fine but has nothing to run your code with, which is the usual cause of a missing > prompt.

Check: Positron opens and there is a pane at the bottom with a > prompt in it. If Positron opens but there is no >, look for an R interpreter selector in the top right and choose the R you just installed. If your R install finished after Positron was already open, quit Positron and reopen it so it looks again.

Step 2: Set up a folder

Everything for this course goes in one folder, so that your data and your code live together.

The reason is that R has to be told where files are. If your script is on the Desktop and your data is in Downloads, every line that reads a file has to spell out the full path from the top of your hard drive, and that path is different on every computer. Keep them in one folder and you can write short paths like data/barley_trial.csv that work anywhere. Step 7 is where this pays off.

Make a folder called AREC_261 somewhere you will find it again – Documents is fine, and if you use OneDrive, put it there so it is backed up. Inside it make two subfolders:

AREC_261/
  code/
  data/

Scripts go in code/, files you downloaded go in data/. Later you will want an output/ folder as well for things your scripts produce.

In Positron, use File > Open Folder and open AREC_261. Open the folder itself, not a file inside it. This sets the folder R treats as home, so everything below is relative to it.

The Explorer pane is also the easiest way to add subfolders: right-click in it, or use the new-folder button at the top of the pane.

Check: the Explorer pane on the left shows AREC_261 with code and data inside it. If the pane shows some other folder name at the top, you have the wrong folder open – File > Open Folder again.

Step 3: The console and the script

These are the two places you can type R, and the difference matters.

The console is the pane at the bottom with the > prompt. You type a command, press Enter, and R runs it immediately. Nothing is saved. Close Positron and everything you typed there is gone.

A script is a text file of R commands that you save. Nothing runs when you type it. You write the commands, save the file, and then send them to the console to run – all at once, or a line at a time.

The distinction people find confusing at first: the script is where the code lives, the console is where it runs. Even when you run a script, the commands are executed in the console, and that is where the output appears.

Two consequences follow from that, and they are the source of most early confusion.

The first is that typing something into a script has no effect at all until you send it. A student who writes acres <- 320 in a script, does not run it, and then asks for acres in the console gets Error: object 'acres' not found. The line exists as text in a file; R has never seen it.

The second is that the console remembers things your script does not contain. If you create an object in the console and then use it in your script, the script works for you today and fails for anyone else, because the line that made the object was never written down. This is why the rule is to build up the script, not the console. Anything you want to keep, put in the script and run it from there.

Use the console for one-off things: a quick calculation, checking what a function does, looking at a variable, installing a package. Use a script for anything you will want to run again, which is almost everything.

If the console ever fills with output you no longer want to look at, Ctrl+L clears it. That clears the display only – your objects are still there.

Check: nothing to do here. Read it and carry on.

Step 4: Type commands in the console

Click in the console and type each of these, pressing Enter after each one.

4 + 5
4 * 5
4 + 5 * 2
(4 + 5) * 2

The last two give different answers – 14 and 18. R follows the usual order of operations, so brackets are the way to force the addition first.

Now save an answer as an object:

acres <- 320
yield <- 42.5
acres * yield

An object is a name you give to a value so you can use it again without retyping it. The <- assigns: it takes what is on the right and stores it under the name on the left. After acres <- 320, R remembers acres until you close it, and you can use acres anywhere you would have written 320.

Notice that acres <- 320 prints nothing. Assigning is silent. To see what is in an object, type its name on its own line. That is why the third line above is just acres * yield – no assignment, so R prints the result instead of storing it.

The shortcut for <- is Alt+- (Windows) or Option+- (Mac), which also puts the spaces in for you.

Check: acres * yield prints [1] 13600. The [1] is not part of the answer – it just labels the first value on the line.

If instead you get Error: object 'acres' not found, R does not have an object by that name. Usually the assignment line was never run, or the name was typed differently the second time. Names are case-sensitive, so Acres and acres are two different objects.

So far each object has held a single number. A vector holds several values of the same type, and you build one with c():

yields <- c(72, 81, 68, 77, 74)
yields

The c stands for combine. Without it, yields <- 72, 81, 68 is not valid R at all.

R will do arithmetic on the whole vector at once, which is the point of them. You do not write a loop over the five fields; you write the calculation once and R applies it to every element:

yields * 0.0218
mean(yields)
max(yields)
length(yields)

Check: mean(yields) prints [1] 74.4 and length(yields) prints [1] 5.

Comparisons work the same way, giving one answer per element rather than one answer overall:

yields > 75
sum(yields > 75)

Check: the first prints FALSE TRUE FALSE TRUE FALSE. The second prints [1] 2 – sum() counts the TRUEs, which is a handy way of asking how many.

Text goes in quotation marks, numbers do not:

varieties <- c("CDC Copeland", "AAC Synergy", "CDC Austenson")
varieties

A vector holds one type of thing. If you mix them, R does not complain – it converts everything to text instead. Try it:

c(72, 81, "68")

That prints "72" "81" "68", with quotes around all three, and mean() on it returns NA with a warning that the argument is not numeric. A stray quotation mark around one number in a long c() will do this to the whole vector, and the only visible sign is the quote marks in the output.

A larger version of the same mistake is to quote the whole thing:

yields <- "c(72, 81, 68, 77, 74)"
yields
length(yields)
mean(yields)

yields prints as [1] "c(72, 81, 68, 77, 74)" – one piece of text that happens to look like R code. length(yields) is 1, not 5, and mean(yields) gives NA with argument is not numeric or logical: returning NA. R never ran the c(); the quotes told it to store those characters. sd() on the same object gives NA with a different warning, NAs introduced by coercion.

The length is the fast way to tell. Five numbers should give length() of 5, and anything that returns 1 when you expected several is a string.

Practice

Five canola fields yielded 52.3, 47.8, 55.1, 44.6 and 50.9 bu/ac, on 160, 320, 240, 130 and 200 acres. Save each as a vector. Compute the mean and standard deviation of the yields, then the total production in bushels and the farm’s average yield weighted by acres. Why is the weighted average not the same as the plain mean?

# Yields in bu/ac and field sizes in acres
yields <- c(52.3, 47.8, 55.1, 44.6, 50.9)
acres  <- c(160, 320, 240, 130, 200)

# Centre and spread of the yields
mean(yields)
sd(yields)

# Total bushels, then the acre-weighted average yield
total_bu <- sum(yields * acres)
total_bu
total_bu / sum(acres)

The mean is 50.14 bu/ac and the standard deviation 4.06. Total production is 52866 bushels, and dividing by the 1050 total acres gives 50.35 bu/ac.

The plain mean counts each field once whatever its size. The weighted average lets the 320-acre field count for its share of the land, and that field yielded above average, so the weighted figure comes out slightly higher.

Step 5: Write a script and run it

In Positron, File > New File > R Script. Save it into your code folder as module2_lab2.R. The .R ending is what tells Positron to treat the file as R code and colour it accordingly; save it as .txt and you lose that.

Type this into the script, including the header:

# ---
# Title: Module 2 lab, second run
# Author: Your Name
# Date: 2026-09-22
# Description:
#   Practising the console, scripts, packages and reading data.
# ---

# Field size and yield
acres <- 320
yield <- 42.5

# Total production in bushels
production <- acres * yield
production

Nothing has happened yet. Typing code in a script does not run it.

To run it: put your cursor on a line and press Cmd+Enter (Mac) or Ctrl+Enter (Windows). That sends the line to the console and moves down. Hold the keys and step through the whole script. Watch the console as you go – each line you send appears there, followed by whatever it produced.

Two other ways to run, for when you want more than a line at a time. Select several lines and the same shortcut sends all of them. To run the whole file from the top, use Cmd+Shift+Enter (Mac) or Ctrl+Shift+Enter (Windows), or the Run button at the top right of the editor.

Check: the console shows [1] 13600, and your script is saved in code/ with the header at the top.

An unsaved file shows a dot or a mark beside its name in the tab. Cmd+S or Ctrl+S saves.

Anything after a # is a comment – R ignores it. You can comment out a whole block by selecting it and pressing Cmd+Shift+C or Ctrl+Shift+C, which is useful for temporarily switching off a few lines without deleting them. The header block and the short notes above each step are for whoever reads your code later, which is usually you.

Now add a vector to the script and run it:

# Yields from five fields, in bushels per acre
field_yields <- c(72, 81, 68, 77, 74)

# The average across the five
mean(field_yields)

Practice

Add to your script: a vector of the acres of those same five fields – 160, 320, 240, 80 and 120 – then work out the total production in bushels across all five. Comment each step.

# Acres in each of the five fields
field_acres <- c(160, 320, 240, 80, 120)

# Production in each field, then the farm total
field_production <- field_yields * field_acres
sum(field_production)

field_yields * field_acres multiplies the two vectors position by position, giving five numbers, and sum() adds them. The total is 68800 bushels.

Step 6: Packages

R on its own is fairly small. A package is a bundle of extra functions somebody else has written. Installing one downloads it to your computer; loading one makes it available in your current session. Two steps, and they happen at different rates: you install once per computer, and you load once per session.

The tidyverse is not one package but a collection of them. readr reads files, dplyr manipulates data, ggplot2 makes graphs. Installing the tidyverse gets you all of them, and library(tidyverse) loads the main ones together.

Install it. This takes a few minutes and prints a lot of output. Type it in the console, not the script:

install.packages("tidyverse")

Some of that output appears in red even when nothing has gone wrong. What matters is whether it finishes and gives you the > prompt back, not whether the text was red.

If you installed it last week you can skip this. You install a package once. You load it every session. Put this line in your script, under the header:

# Load packages
library(tidyverse)

Run that line. You will see a message listing the packages it attached and some conflicts – that message is normal and is not an error. The conflicts line is telling you that two loaded packages both have a function of the same name, and which one wins.

Note the quotes. They go around the name when you install (install.packages("tidyverse")) but not when you load (library(tidyverse)).

Check: library(tidyverse) runs and the message mentions dplyr, ggplot2 and readr.

The reason install.packages() goes in the console and library() goes in the script: you do not want to reinstall the tidyverse every time you run your code, but you do need to load it every time.

Two errors come out of this step, and they mean different things.

Error in library(tidyverse) : there is no package called 'tidyverse' means it is not installed – the install either failed or was never run. Run install.packages("tidyverse") again and read the end of the output.

Error in read_csv(...) : could not find function "read_csv" means it is installed but not loaded in this session. Run library(tidyverse). You will hit this one every time you restart Positron and start partway down your script, which is exactly why the library() line belongs at the top of the file rather than in the console.

Step 7: Read in data

Download barley_trial.csv and save it into the data folder inside AREC_261. It should appear under data in the Explorer pane. If it does not, it went somewhere else – most likely Downloads.

Add these lines to your script and run them:

# Read the barley trial data
barley <- read_csv("data/barley_trial.csv")

# Look at it
barley
glimpse(barley)

read_csv() comes from the tidyverse, which is why Step 6 had to come first.

The assignment matters as much as the reading. read_csv("data/barley_trial.csv") on its own prints the data and forgets it. With barley <- in front, the data is stored under the name barley and everything after this point can use it.

The three lines do different jobs. read_csv() prints a note about what it found – the number of rows and columns, and the type it guessed for each column. barley prints the first ten rows so you can see actual values. glimpse() turns the table on its side: one line per column, with its type and the start of its values, which is the fastest way to check that a column of numbers was read as numbers rather than text.

The path "data/barley_trial.csv" is relative – it means “the data folder inside the folder I have open.” It works because you opened AREC_261 in Step 2. The alternative is an absolute path like /Users/yourname/Documents/AREC_261/data/barley_trial.csv, which works only on your machine. Relative paths are why you can hand the whole folder to someone else and have their code run.

The forward slash separates folder from file. On Windows, paths are normally shown with backslashes, but a backslash means something special inside R text. Write forward slashes on Windows too.

Check: glimpse(barley) prints 120 rows and 5 columns: field_id, fertilizer_kg_ha, rainfall_mm, variety, yield_bu_acre. field_id and variety should be <chr>, the other three <dbl>.

If you get Error: 'data/barley_trial.csv' does not exist in current working directory, R looked where you told it and found nothing. Note that the error message ends with the folder it looked in – read that first, because it tells you which of these it is:

  • The folder in the message is not AREC_261. You have the wrong folder open in Positron. File > Open Folder on AREC_261 and run the line again.
  • The folder is right but the file is not in data/. It is probably still in Downloads.
  • Both look right. Then the name is different from what you typed. Windows sometimes saves it as barley_trial.csv.txt, and capitalisation counts. Click the file in the Explorer pane and compare the name character by character with the one in your quotes.

getwd() prints the folder R is currently working from, if you want to check it directly.

Describing a column

glimpse() tells you the shape of the table. To get numbers out of a single column you need the $, which pulls one column out as a vector:

# One column, as a vector of 120 numbers
barley$yield_bu_acre

# Centre
mean(barley$yield_bu_acre)
median(barley$yield_bu_acre)

# Spread
sd(barley$yield_bu_acre)
var(barley$yield_bu_acre)
range(barley$yield_bu_acre)

Check: the mean is 78.96, the median 78.9, the standard deviation 8.22, and range() prints both ends at once, 55.3 98.4.

The $ is what turns a table into something mean() can work with. mean(barley) does not work, because a data frame of five columns has no single average, and mean(yield_bu_acre) does not work either – that is Step 8’s problem, further down.

quantile() gives you cut points other than the middle. Ask for several at once by passing a vector of probabilities:

# The quartiles
quantile(barley$yield_bu_acre, probs = c(0.25, 0.50, 0.75))

# The middle half, two ways
IQR(barley$yield_bu_acre)
quantile(barley$yield_bu_acre, probs = 0.75) - quantile(barley$yield_bu_acre, probs = 0.25)

Check: the quartiles are 73.0, 78.9 and 84.075, and both IQR lines give 11.075.

summary() does most of that in one call, for every column at once:

summary(barley)

For yield_bu_acre it prints a minimum of 55.30, first quartile 73.00, median 78.90, mean 78.96, third quartile 84.08 and maximum 98.40. That is the quickest first look at a file you have not seen before: implausible minimums and maximums show up immediately, and so do missing values, which summary() reports as an extra NA's line under any column that has them.

Practice

Compute the 90th percentile of yield_bu_acre, and the gap between the 90th and the 10th. What does that gap describe, and how does it differ from the range?

# The 10th and 90th percentiles in one call
quantile(barley$yield_bu_acre, probs = c(0.10, 0.90))

# The range, for comparison
range(barley$yield_bu_acre)

The percentiles are 68.28 and 89.64, a gap of 21.36 bu/ac. That gap holds the middle 80% of the fields. The range is 98.4 minus 55.3, or 43.1 bu/ac, and it rests entirely on the single best and single worst field. Both describe spread; the percentile gap ignores the extremes, and the range is made of nothing else.

Step 8: The tidyverse functions

Four functions do most of the work in R. Each one takes a data frame and gives you back a data frame. They split the job up by what they change: filter() picks rows, select() picks columns, arrange() reorders rows, and mutate() adds columns. Everything else in this module is those four in combination.

The first argument is always the data frame, and the column names after it are written bare – no quotes, no barley$. Inside these functions R already knows which table you mean, so yield_bu_acre is enough.

Add these to your script, running each as you go. Read the output before moving to the next one.

# Keep only the rows where yield was above 80 bu/acre
filter(barley, yield_bu_acre > 80)

# Keep only the CDC Copeland fields
filter(barley, variety == "CDC Copeland")

# Sort from highest yield to lowest
arrange(barley, desc(yield_bu_acre))

# Keep only three of the columns
select(barley, field_id, variety, yield_bu_acre)

# Add a column: yield converted to tonnes per hectare
mutate(barley, yield_t_ha = yield_bu_acre * 0.0218)

Barley yields higher than canola, so the cutoff in that first filter() is higher too. At 80 bu/acre it splits the trial roughly in half. Ask for anything above 60 and you get 119 of the 120 fields back, which tells you nothing.

Note the double == in filter(). A single = means something else in R, and this catches almost everybody at least once. filter(barley, variety = "AAC Synergy") gives an error saying it detected a named input, with a second line guessing that you meant ==. Take the guess.

The other two filter() mistakes look similar but say different things. Get the column name wrong and the error ends object 'Variety' not found – capitalisation counts, and glimpse() is the quick way to see the real names. Forget the quotes around the text value and you get unexpected symbol, because R tried to read CDC Copeland as two object names. Text values need quotes; column names do not.

arrange() sorts smallest first by default, which is why desc() is there. Give it two columns and it sorts by the first, then uses the second to break ties:

# Sort by variety, and within each variety by yield
arrange(barley, variety, desc(yield_bu_acre))

Check: the top three rows are all AAC Synergy – B060 at 98.4, B082 at 98.0, B102 at 95.7 – because A sorts before C, and within AAC Synergy the highest yield comes first.

select() also works the other way round. Put a minus in front of a column to drop it rather than keep it, which is easier when you want most of them:

# Everything except the two input columns
select(barley, -fertilizer_kg_ha, -rainfall_mm)

mutate() takes as many new columns as you want, separated by commas, and the new column goes on the far right:

# Two conversions at once
mutate(barley,
       yield_t_ha = yield_bu_acre * 0.0218,
       fert_lb_acre = fertilizer_kg_ha * 0.892)

Here the single = is correct. Inside mutate() you are naming a new column, not testing whether two things match.

Check: the first filter() returns 55 rows, the CDC Copeland one returns 39, and the top row after arrange(barley, desc(yield_bu_acre)) is field B060 at 98.4 bu/acre.

None of these changed barley. They printed a result and threw it away. Run barley again after all of that and it still has its 120 rows and its 5 original columns – no filtering, no sorting, no yield_t_ha. This is the single most common source of “but I already did that” in the first few weeks.

To keep a result you have to assign it:

# Now the result is saved
top_fields <- filter(barley, yield_bu_acre > 80)
nrow(top_fields)

Check: nrow(top_fields) prints [1] 55.

top_fields is a new data frame sitting beside barley, not a change to it. You could also write barley <- filter(barley, yield_bu_acre > 80), which replaces barley with the 55 rows – and then the other 65 are gone until you read the file again. Assign to a new name unless you have a reason not to.

Practice

Keep only the CDC Copeland fields that yielded more than 80 bu/acre, and save the result as an object called good_copeland. How many are there?

# CDC Copeland fields that beat 80 bu/acre
good_copeland <- filter(barley, variety == "CDC Copeland", yield_bu_acre > 80)
nrow(good_copeland)

Two conditions separated by a comma means both must hold. There are 17 such fields.

Swap the comma for a | and you get either condition instead of both – every CDC Copeland field plus every field over 80 whatever its variety, which is 77 rows. The comma is the one you want almost always.

Two conditions on the same column

A comma will not get you two varieties, because no row is two varieties at once:

# Zero rows: no field is both
filter(barley, variety == "AAC Synergy", variety == "CDC Austenson")

# Either one
filter(barley, variety == "AAC Synergy" | variety == "CDC Austenson")

# The same thing, shorter
filter(barley, variety %in% c("AAC Synergy", "CDC Austenson"))

Check: the first returns 0 rows, and the other two both return 81.

The zero-row result is the one to watch for, because it is not an error. R does exactly what you asked and there is nothing in the output to say the question was wrong. %in% is the readable form once you have more than two values to match.

The column name on its own

The bare column names in Step 8 work only inside a dplyr function. Outside one, R has no idea which table you mean:

mean(yield_bu_acre)
quantile(yield_bu_acre, 0.75)

Both give Error: object 'yield_bu_acre' not found. R is not being obtuse – yield_bu_acre is not an object in your session. It is the name of a column inside barley, and nothing outside filter(), mutate() and friends knows to look there. Use barley$yield_bu_acre instead.

The same error appears if you drop the data frame from a dplyr call:

filter(variety == "CDC Copeland")
summarise(mean_yield = mean(yield_bu_acre))

Those give Error: object 'variety' not found and Error: object 'yield_bu_acre' not found. The first argument was missing, so filter() treated variety == "CDC Copeland" as the data frame and tried to evaluate it in the session, where variety does not exist. The message names the column rather than the missing data frame, which is why it is easy to misread as a spelling problem.

Quoting the data frame instead gives a different message:

filter("barley", variety == "CDC Copeland")

Error: no applicable method for 'filter' applied to an object of class "character". R is saying it does not know how to filter a piece of text. Data frame names go bare, like column names; only data values take quotes.

Quoting a column name is worse, because it runs:

filter(barley, "variety" == "CDC Copeland")

Zero rows, no error. You compared the word "variety" to the word "CDC Copeland", which is FALSE for every row, so every row was dropped. The tell is # A tibble: 0 × 5 where you expected 39 rows.

Practice

Keep only the AAC Synergy fields, saving the result as synergy. How many rows? Then add a yield_t_ha column to synergy (multiply the yield by 0.0218), saving it back to synergy. Compute the mean of the new column and of the original, and check their ratio. How many rows does barley have after all this?

# Keep the AAC Synergy fields
synergy <- filter(barley, variety == "AAC Synergy")
nrow(synergy)

# Add the converted column, saving back over synergy
synergy <- mutate(synergy, yield_t_ha = yield_bu_acre * 0.0218)

# Means of the new and original columns
mean(synergy$yield_t_ha)
mean(synergy$yield_bu_acre)

# The original is untouched
nrow(barley)

There are 50 AAC Synergy fields. The mean of the converted column is 1.788 t/ha against 82.036 bu/ac, a ratio of exactly 0.0218 – multiplying every value by a constant multiplies the mean by the same constant.

barley still has 120 rows. filter() and mutate() return new data frames; they do not reach back and change the one you gave them.

The synergy <- mutate(synergy, ...) line is the part students leave off. Without the assignment the new column is printed and discarded, and mean(synergy$yield_t_ha) on the next line prints [1] NA with two warnings: Unknown or uninitialised column: yield_t_ha and argument is not numeric or logical: returning NA. That is not an error, so the script keeps running and the NA goes into whatever comes next.

Step 9: The pipe

Most real work needs several of those functions in a row. You could nest them, but it reads inside-out and gets unreadable fast:

# Don't write this
head(arrange(filter(barley, variety == "AAC Synergy"), desc(yield_bu_acre)), 3)

The pipe |> takes whatever is on its left and passes it as the first argument to the function on its right. The same thing, written in the order it happens:

# The AAC Synergy fields, best yield first, top three
barley |>
  filter(variety == "AAC Synergy") |>
  arrange(desc(yield_bu_acre)) |>
  select(field_id, yield_bu_acre) |>
  head(3)

Read |> as “and then”. Start with barley, and then keep the AAC Synergy rows, and then sort them, and then keep two columns, and then show the first three.

The two versions do exactly the same thing. Notice what happened to the arguments: filter(barley, variety == "AAC Synergy") becomes barley |> filter(variety == "AAC Synergy"). The data frame moved out in front of the pipe, so the function no longer needs it. Everything that stays inside the brackets is the part that is specific to that step.

Each step receives the result of the one above it, not the original barley. By the time head(3) runs, it is looking at a table of 50 rows cut to two columns and sorted, not at the file. This is also why order matters: select() before arrange(desc(yield_bu_acre)) is fine here because yield survives the select, but drop the yield column first and the sort has nothing to sort on.

The keyboard shortcut for |> is Cmd+Shift+M (Mac) or Ctrl+Shift+M (Windows).

When you break a pipeline over several lines, the |> goes at the end of the line, not the start of the next one. R reads a line ending in |> as unfinished and keeps going; a line ending in ) looks complete, so it runs, and the next line starting with |> is a syntax error. If your console shows a + instead of a > after running a pipeline, R is still waiting for you to finish it – press Escape and look for a missing bracket.

Check: the top three are B060 at 98.4, B082 at 98.0, and B102 at 95.7.

Build the pipeline one line at a time. Write the filter(), run it, look at the output, then add the arrange(). If you write all five lines and then run it, and something is wrong, you have no idea which line did it. Highlighting the first two lines and running just those – without the trailing |> – is the quickest way to see where a long pipeline goes wrong.

Practice

In one pipeline: take the fields that got less than 200 mm of rain, sort them by yield with the best first, keep the field, variety and yield columns, and show the top five.

# The driest fields, best yield first
barley |>
  filter(rainfall_mm < 200) |>
  arrange(desc(yield_bu_acre)) |>
  select(field_id, variety, yield_bu_acre) |>
  head(5)

There are 31 fields under 200 mm. The best of them is B107 at 88.2 bu/acre, then B027 at 87.8.

Step 10: group_by and summarise

This is the step that turns a table of rows into an answer, and it is worth going slowly.

summarise() collapses many rows into one:

# One number for the whole dataset
barley |>
  summarise(mean_yield = mean(yield_bu_acre))

Check: one row, 79.0 bu/acre, the average across all 120 fields.

One number for the whole farm is rarely what you want. You want it per variety, per year, per field – and that is what group_by() does. It does not change the data and it prints almost nothing. What it does is attach a label to the table saying “from here on, treat these as separate piles.”

# Marking the data does nothing visible
barley |>
  group_by(variety)

Check: the same 120 rows you started with. The only difference is the second line of the header: # Groups: variety [3]. R is telling you there are three piles waiting.

Now run the same summarise() against the grouped table:

# One row per variety
barley |>
  group_by(variety) |>
  summarise(
    n = n(),
    mean_yield = mean(yield_bu_acre)
  )

Check: three rows – AAC Synergy 50 fields averaging 82.0, CDC Austenson 31 averaging 74.2, CDC Copeland 39 averaging 78.8.

The same summarise() gave one row before and three rows now. That is the whole idea: summarise() always collapses each pile to one row, and group_by() decides what the piles are.

n() counts the rows in each pile, and it is worth including every time. A mean over 50 fields and a mean over 8 are not equally trustworthy, and the count is the cheapest way to see which you have.

You can ask for as many summaries as you want, and each becomes a column:

# Five numbers per variety
barley |>
  group_by(variety) |>
  summarise(
    n = n(),
    mean_yield = mean(yield_bu_acre),
    sd_yield = sd(yield_bu_acre),
    min_yield = min(yield_bu_acre),
    max_yield = max(yield_bu_acre)
  )

Check: AAC Synergy has the highest mean at 82.0 and the widest range, 64.1 to 98.4.

Grouping by something that is not a column yet

You are not limited to the columns in the file. Make the grouping column first with mutate(), then group by it. Here the fields are split into dry and wet on either side of 235 mm:

# Split the fields into two rainfall bands, then compare them
barley |>
  mutate(rain_band = if_else(rainfall_mm < 235, "dry", "wet")) |>
  group_by(rain_band) |>
  summarise(
    n = n(),
    mean_yield = mean(yield_bu_acre)
  )

Check: 60 fields in each band. The dry half averages 76.5 bu/acre, the wet half 81.4.

That is a five-bushel difference from rainfall alone, which is the sort of comparison this pattern is for.

Grouping by two things at once

Give group_by() two columns and you get one row per combination:

# Variety and rainfall band together
barley |>
  mutate(rain_band = if_else(rainfall_mm < 235, "dry", "wet")) |>
  group_by(variety, rain_band) |>
  summarise(
    n = n(),
    mean_yield = mean(yield_bu_acre),
    .groups = "drop"
  )

Check: six rows, three varieties times two bands. AAC Synergy runs 78.3 dry and 85.8 wet; CDC Austenson 72.7 and 75.4; CDC Copeland 77.0 and 81.0. Every variety does better in the wet half, and AAC Synergy gains the most.

.groups = "drop" tells R to forget the grouping afterwards. Without it, the result stays grouped by variety and the next thing you do to it behaves unexpectedly. If you ever get a result you cannot explain after a summarise(), check whether the table is still grouped.

The shortcut

Counting rows per group is so common it has its own function:

# These two do the same thing
barley |> group_by(variety) |> summarise(n = n())
barley |> count(variety)

Check: both give AAC Synergy 50, CDC Austenson 31, CDC Copeland 39.

Practice

On the fields that got more than 160 kg/ha of fertilizer, which variety had the highest average yield? You need a filter(), a group_by() and a summarise() in one pipeline, and sorting the result makes the answer easy to read.

# Average yield by variety, on the heavily fertilized fields only
barley |>
  filter(fertilizer_kg_ha > 160) |>
  group_by(variety) |>
  summarise(
    n = n(),
    mean_yield = mean(yield_bu_acre)
  ) |>
  arrange(desc(mean_yield))

AAC Synergy comes out highest at 85.4 bu/acre, but over only 8 fields, against 11 for CDC Copeland at 83.2. Eight fields is not many to draw a conclusion from, and the gap is small.

Note the order of the pipeline. filter() comes before group_by(), so the piles are built from the heavily fertilized fields only. Put the filter() after the summarise() and you would be filtering the three summary rows instead, which is a different question.

Practice

Which single field beat its own variety’s average by the most? This one needs group_by() with mutate() rather than summarise().

# How far each field sits from its variety's mean
barley |>
  group_by(variety) |>
  mutate(vs_variety_mean = yield_bu_acre - mean(yield_bu_acre)) |>
  ungroup() |>
  arrange(desc(vs_variety_mean)) |>
  select(field_id, variety, yield_bu_acre, vs_variety_mean) |>
  head(3)

B093, a CDC Austenson field, at 17.1 bu/acre above its variety average.

The difference from everything above: summarise() collapses each pile to one row, mutate() keeps all 120 rows and computes the group’s mean for each one. So mean(yield_bu_acre) here is the mean of that row’s variety, not of the whole file. ungroup() afterwards for the same reason as .groups = "drop".

Two ways a grouped pipeline goes wrong

Both of these run without an error, which is what makes them worth recognising.

The first is summarise() where the question needed a group_by() in front of it:

# Asked for the mean by variety, wrote this
barley |>
  summarise(mean_yield = mean(yield_bu_acre))

That gives a one-row table with mean_yield of 79.0. It is a correct average of all 120 fields, and it is not the answer to a question about varieties. One row where you expected three is the whole signal. Count the rows in the result against the number of groups you asked for, every time.

The second is a pipeline that stops at group_by():

# Nothing is actually summarised
barley |>
  group_by(variety)

You get back all 120 rows with # Groups: variety [3] in the header. group_by() only labels the piles; something has to collapse them. A result the same size as the data you started with means no summarise() ran.

Save the filter

A filter whose result is never assigned is the version of this that produces a plausible wrong number rather than an obvious one:

# The filter runs, prints, and is thrown away
filter(barley, rainfall_mm > 250)

# So this is the mean of all 120 fields, not the wet ones
mean(barley$yield_bu_acre)

That prints 78.96, which looks like a perfectly reasonable answer. The wet fields actually average 81.71 over 52 rows:

wet <- filter(barley, rainfall_mm > 250)
nrow(wet)
mean(wet$yield_bu_acre)

Check: 52 rows and a mean of 81.71.

The second version of this mistake is subtler: the filter is saved, but the line after it goes back to barley out of habit. Whenever a step narrows the data, the next line has to name the narrowed object.

Practice

On the fields that got more than 250 mm of rain, compute each variety’s mean yield and the number of fields behind it, sorted best-first. Save the table, then write it to output/wet_variety_summary.csv. Which variety tops the list, and would you trust the bottom row as much as the top?

# Variety means on the wettest fields only
wet_summary <- barley |>
  filter(rainfall_mm > 250) |>
  group_by(variety) |>
  summarise(
    n = n(),
    mean_yield = mean(yield_bu_acre)
  ) |>
  arrange(desc(mean_yield))

wet_summary

# Save the table for a report
write_csv(wet_summary, "output/wet_variety_summary.csv")

AAC Synergy tops the list at 86.8 bu/ac over 22 fields, then CDC Copeland at 80.7 over 16, then CDC Austenson at 74.9 over 14.

The counts are close enough here that all three means rest on a similar base, so the ranking is about as trustworthy as this file allows. That is the point of carrying n() – when the bottom row turns out to be two fields, the ordering is telling you about two fields.

You need an output folder inside AREC_261 before the write_csv() line will run. Without it the error is Cannot open file for writing: followed by the path you asked for. R will not create the folder for you.

Practice

Barley is selling at $5.75 per bushel. In one pipeline, add a revenue-per-acre column and compute each variety’s mean revenue, best first. Put the price in an object rather than typing the number into the calculation.

# The price lives in one place
price <- 5.75

# Mean revenue per acre by variety
barley |>
  mutate(revenue_per_ac = yield_bu_acre * price) |>
  group_by(variety) |>
  summarise(
    n = n(),
    mean_revenue = mean(revenue_per_ac)
  ) |>
  arrange(desc(mean_revenue))

AAC Synergy comes out at $471.71 per acre, CDC Copeland at $453.14, CDC Austenson at $426.61.

The order is the same as the yield ranking in Step 10, and it has to be: every yield was multiplied by the same number, so nothing can change places. The ranking would only move if the varieties sold at different prices.

The price object is there so that a price change is one edit. Type 5.75 into the mutate() and you have to find it again next time.

Checking your own work

Nothing above errors when it goes wrong. A filter that dropped the wrong rows, a mean taken over the whole file, a column that was never saved – all of them print something that looks like an answer. So the checking has to be deliberate.

Read the row count after every filter. nrow() is one line and it tells you whether the filter did what you meant. Going from 120 to 52 is a filter working; going to 120 means the condition matched everything, and going to 0 means it matched nothing, which is usually a spelling or capitalisation problem in a text value rather than an empty dataset.

Check that a mean sits between the minimum and the maximum. It is a crude test and it catches real mistakes – a mean above every value in the column usually means the wrong column got summarised, and an NA means a stale or non-numeric column crept in.

Print the object rather than assuming. After synergy <- filter(...), type synergy and look at it. The Variables pane will also tell you how many rows and columns an object has, which is the fastest way to catch an assignment that never ran.

When a pipeline gives an answer you cannot explain, run it a line at a time. Highlight the first step, run it, look at the result, then add the next. The point where the data stops looking the way you expect is where the problem is, and finding that takes a few seconds rather than the several minutes staring at all five lines costs.

Then restart R and run the whole script from the top. Use Session > Restart R, or quit and reopen Positron, so that every object is gone and the script has to rebuild them all. This is the check that matters most, because the console remembers things your script does not contain. A script that only works because you defined an object by hand twenty minutes ago will fail for the TA, and your test is marked on a script that has to run clean from top to bottom in a fresh session.

If you finish early

The best use of any time left is the test bank: Module 2 test bank. Your test draws one question from each section of it, so work through examples from every section – R basics, reading and inspecting data, data manipulation functions, and pipelines with grouped summaries – rather than doing five of the kind you already find easy. Every question has a worked answer behind a dropdown.

Doing them here rather than at home is the point of the lab: this is the one time you can get an answer to “why did mine come out different?” while the file is still open in front of you.

Two other things worth trying if you want a break from the bank.

Break things on purpose, so you recognise them later. Run each of these and read what comes back:

filter(barley, variety = "AAC Synergy")
filter(barley, Variety == "AAC Synergy")
mean(yield_bu_acre)
filter(barley, variety == "aac synergy")
barley |> group_by(variety) |> summarise(mean_yield = mean(yield_bu_accre))

The first uses one = where it needs two, and the message guesses the fix. The second has the wrong capitalisation in a column name and ends object 'Variety' not found. The third refers to a column without saying which data frame it is in.

The last two are the more interesting pair. The fourth has the right column and the wrong capitalisation in the value, so it runs and gives # A tibble: 0 × 5 – zero rows, no error, no complaint. The fifth misspells a column inside summarise(), and the error names the column and then adds a line telling you which group it failed in. Both of these are what a real mistake looks like: R either says nothing at all, or says something precise that you have to read past the first line to use.

Before you leave

Save your script.

The whole of Module 2 is in that one file: a header, library(tidyverse), reading a csv, and a few pipelines. That is the shape of most R work you will do in this course.