Lab: getting started in R
Work through this at your own pace. Each step ends with a check – something you should be able to see on your screen before moving on. If a check fails, that is the point to put your hand up rather than pushing on.
I have lost my voice today, so I will be coming around the room rather than talking you through it. Write your question down or point at the screen and I will read it.
The whole thing should take about an hour. It covers the same ground as 4 Getting Started with R and 5 Loading Data and Packages in R, so those chapters are the place to look if you want more detail on any step.
Step 1: Install R and Positron
Two separate installs, and the order matters. R is the language. Positron is the program you write R in.
- Install R from https://cran.rstudio.com. Pick the version for your operating system and accept the defaults.
- Install Positron from https://positron.posit.co.
- Open Positron.
Check: Positron opens and there is a pane at the bottom with a > prompt in it. If Positron opens but there is no >, look for an R interpreter selector in the top right and choose the R you just installed.
Step 2: Set up a folder
Everything for this course goes in one folder, so that your data and your code live together.
Make a folder called AREC_261 somewhere you will find it again – Documents is fine, and if you use OneDrive, put it there so it is backed up. Inside it make two subfolders:
AREC_261/
code/
data/
In Positron, use File > Open Folder and open AREC_261. You should see the folder appear in the Explorer pane on the left.
Check: the Explorer pane on the left shows AREC_261 with code and data inside it.
Step 3: The console and the script
These are the two places you can type R, and the difference matters.
The console is the pane at the bottom with the > prompt. You type a command, press Enter, and R runs it immediately. Nothing is saved. Close Positron and everything you typed there is gone.
A script is a text file of R commands that you save. Nothing runs when you type it. You write the commands, save the file, and then send them to the console to run – all at once, or a line at a time.
The distinction people find confusing at first: the script is where the code lives, the console is where it runs. Even when you run a script, the commands are executed in the console, and that is where the output appears.
Use the console for one-off things: a quick calculation, checking what a function does, looking at a variable. Use a script for anything you will want to run again, which is almost everything.
Check: nothing to do here. Read it and carry on.
Step 4: Type commands in the console
Click in the console and type each of these, pressing Enter after each one.
4 + 5
4 * 5
4 + 5 * 2
(4 + 5) * 2The last two give different answers. R follows the usual order of operations.
Now save an answer as an object:
acres <- 320
yield <- 42.5
acres * yieldThe <- assigns. After acres <- 320, R remembers acres until you close it.
Check: acres * yield prints [1] 13600. The [1] is not part of the answer – it just labels the first value on the line.
So far each object has held a single number. A vector holds several values of the same type, and you build one with c():
yields <- c(48, 52, 47, 55, 50)
yieldsR will do arithmetic on the whole vector at once, which is the point of them:
yields * 0.056
mean(yields)
max(yields)
length(yields)Check: mean(yields) prints [1] 50.4 and length(yields) prints [1] 5.
Text goes in quotation marks, numbers do not:
varieties <- c("InVigor", "DEKALB", "Clearfield")
varietiesStep 5: Write a script and run it
In Positron, File > New File > R Script. Save it into your code folder as module2_lab.R.
Type this into the script, including the header:
# ---
# Title: Module 2 lab
# Author: Your Name
# Date: 2026-09-15
# Description:
# Practising the console, scripts, packages and reading data.
# ---
# Field size and yield
acres <- 320
yield <- 42.5
# Total production in bushels
production <- acres * yield
productionNothing has happened yet. Typing code in a script does not run it.
To run it: put your cursor on a line and press Cmd+Enter (Mac) or Ctrl+Enter (Windows). That sends the line to the console and moves down. Hold the keys and step through the whole script.
Check: the console shows [1] 13600, and your script is saved in code/ with the header at the top.
Anything after a # is a comment – R ignores it. The header block and the short notes above each step are for whoever reads your code later, which is usually you.
Now add a vector to the script and run it:
# Yields from five fields, in bushels per acre
field_yields <- c(48, 52, 47, 55, 50)
# The average across the five
mean(field_yields)Practice
Add to your script: a vector of the acres of those same five fields – 160, 320, 240, 80 and 120 – then work out the total production in bushels across all five. Comment each step.
# Acres in each of the five fields
field_acres <- c(160, 320, 240, 80, 120)
# Production in each field, then the farm total
field_production <- field_yields * field_acres
sum(field_production)field_yields * field_acres multiplies the two vectors position by position, giving five numbers, and sum() adds them. The total is 46000 bushels.
Step 6: Packages
R on its own is fairly small. A package is a bundle of extra functions somebody else has written. Installing one downloads it to your computer; loading one makes it available in your current session.
Install the tidyverse. This takes a few minutes and prints a lot of output. Type it in the console, not the script:
install.packages("tidyverse")You install a package once. You load it every session. Put this line in your script, under the header:
# Load packages
library(tidyverse)Run that line. You will see a message listing the packages it attached and some conflicts – that message is normal and is not an error.
Check: library(tidyverse) runs and the message mentions dplyr, ggplot2 and readr.
The reason install.packages() goes in the console and library() goes in the script: you do not want to reinstall the tidyverse every time you run your code, but you do need to load it every time.
Step 7: Read in data
Download canola_trial.csv and save it into the data folder inside AREC_261.
Add these lines to your script and run them:
# Read the canola trial data
canola <- read_csv("data/canola_trial.csv")
# Look at it
canola
glimpse(canola)read_csv() comes from the tidyverse, which is why Step 6 had to come first.
The path "data/canola_trial.csv" is relative – it means “the data folder inside the folder I have open.” It works because you opened AREC_261 in Step 2. This is why the folder setup was worth the minute it took.
Check: glimpse(canola) prints 120 rows and 5 columns: field_id, fertilizer_kg_ha, rainfall_mm, variety, yield_bu_acre.
If you get Error: 'data/canola_trial.csv' does not exist in current working directory, then one of three things is wrong: the file is not in data/, it is named something else (Windows sometimes adds .txt), or you have a different folder open in Positron.
Step 8: The tidyverse functions
Four functions do most of the work in R. Each one takes a data frame and gives you back a data frame.
Add these to your script, running each as you go. Read the output before moving to the next one.
# Keep only the rows where yield was above 55 bu/acre
filter(canola, yield_bu_acre > 55)
# Keep only the InVigor fields
filter(canola, variety == "InVigor")
# Sort from highest yield to lowest
arrange(canola, desc(yield_bu_acre))
# Keep only three of the columns
select(canola, field_id, variety, yield_bu_acre)
# Add a column: yield converted to tonnes per hectare
mutate(canola, yield_t_ha = yield_bu_acre * 0.056)Note the double == in filter(). A single = means something else in R, and this catches almost everybody at least once.
Check: the first filter() returns 47 rows, the InVigor one returns 42, and the top row after arrange() is field C112 at 70.4 bu/acre.
None of these changed canola. They printed a result and threw it away. To keep a result you have to assign it:
# Now the result is saved
top_fields <- filter(canola, yield_bu_acre > 55)
nrow(top_fields)Check: nrow(top_fields) prints [1] 47.
Practice
Keep only the Clearfield fields that yielded more than 55 bu/acre, and save the result as an object called good_clearfield. How many are there?
# Clearfield fields that beat 55 bu/acre
good_clearfield <- filter(canola, variety == "Clearfield", yield_bu_acre > 55)
nrow(good_clearfield)Two conditions separated by a comma means both must hold. There are 13 such fields.
Step 9: The pipe
Most real work needs several of those functions in a row. You could nest them, but it reads inside-out and gets unreadable fast:
# Don't write this
head(arrange(filter(canola, variety == "InVigor"), desc(yield_bu_acre)), 3)The pipe |> takes whatever is on its left and passes it as the first argument to the function on its right. The same thing, written in the order it happens:
# The InVigor fields, best yield first, top three
canola |>
filter(variety == "InVigor") |>
arrange(desc(yield_bu_acre)) |>
select(field_id, yield_bu_acre) |>
head(3)Read |> as “and then”. Start with canola, and then keep the InVigor rows, and then sort them, and then keep two columns, and then show the first three.
The keyboard shortcut for |> is Cmd+Shift+M (Mac) or Ctrl+Shift+M (Windows).
Check: the top three are C112 at 70.4, C045 at 62.8, and C033 at 62.5.
Build the pipeline one line at a time. Write the filter(), run it, look at the output, then add the arrange(). If you write all five lines and then run it, and something is wrong, you have no idea which line did it.
Practice
In one pipeline: take the fields that got less than 200 mm of rain, sort them by yield with the best first, keep the field, variety and yield columns, and show the top five.
# The driest fields, best yield first
canola |>
filter(rainfall_mm < 200) |>
arrange(desc(yield_bu_acre)) |>
select(field_id, variety, yield_bu_acre) |>
head(5)The best of the dry fields is C001 at 56.7 bu/acre, tied with C118.
Step 10: group_by and summarise
summarise() collapses many rows into one:
# One number for the whole dataset
canola |>
summarise(mean_yield = mean(yield_bu_acre))That is the average across all 120 fields. More often you want it per group, which is what group_by() is for:
# One row per variety
canola |>
group_by(variety) |>
summarise(
n = n(),
mean_yield = mean(yield_bu_acre)
)group_by() on its own does nothing visible. It marks the data so that the next summarise() runs separately within each group.
Check: three rows – Clearfield 39 fields averaging 53.5, DEKALB 39 averaging 52.2, InVigor 42 averaging 54.5.
n() counts the rows in each group. You can add as many summaries as you want:
canola |>
group_by(variety) |>
summarise(
n = n(),
mean_yield = mean(yield_bu_acre),
max_yield = max(yield_bu_acre),
mean_rain = mean(rainfall_mm)
)Practice
On the fields that got more than 140 kg/ha of fertilizer, which variety had the highest average yield? You need a filter(), a group_by() and a summarise() in one pipeline, and sorting the result makes the answer easy to read.
# Average yield by variety, on the heavily fertilized fields only
canola |>
filter(fertilizer_kg_ha > 140) |>
group_by(variety) |>
summarise(
n = n(),
mean_yield = mean(yield_bu_acre)
) |>
arrange(desc(mean_yield))DEKALB comes out highest at 59.6 bu/acre, but over only 8 fields, against 13 for InVigor at 59.4. Eight fields is not many to draw a conclusion from, and the gap is small.
If you finish early
Two things worth trying.
Write a summary table out to a file, which is how you get a result into a report:
# Mean yield by variety, saved to the output folder
variety_summary <- canola |>
group_by(variety) |>
summarise(mean_yield = mean(yield_bu_acre))
write_csv(variety_summary, "output/variety_summary.csv")You will need an output folder inside AREC_261 first.
Then break something on purpose, so you recognise it later. Run each of these and read the error:
filter(canola, variety = "InVigor")
filter(canola, Variety == "InVigor")
mean(yield_bu_acre)The first uses one = where it needs two. The second has the wrong capitalisation. The third refers to a column without saying which data frame it is in. All three are mistakes you will make this term, and knowing what they look like saves a lot of time.
Before you leave
Save your script.
The whole of Module 2 is in that one file: a header, library(tidyverse), reading a csv, and a few pipelines. That is the shape of most R work you will do in this course.