13 AI Tools
AI tools are becoming increasingly central to the work of data analysts. Much of the workflow in the future will involve using AI tools for routine tasks, which allows you to focus on higher-level thinking. Many of you have already used AI tools such as Copilot to describe error messages, or re-write code. In this chapter we will go much further – using AI to write code based on our instructions (vibe coding) and even having AI agents run and verify code for us. It is mind-blowing to me that three years ago none of this was possible. We are living in a time of great change – and while there are concerns about the pace of AI, it is also a time of great opportunity and excitement.
The parts of data analysis that used to take the longest – and, as you are learning, are the most frustrating – are remembering syntax, looking up a function, debugging a typo and reshaping a file into the form you need. Those are the parts AI handles best. What is left is deciding what questions you want to ask, choosing the methods, and judging whether the answer is right.
That last part is where the caution comes in. The same tool that writes a correct script will also produce code that runs cleanly and answers the wrong question. So the practical question is whether you can check the result. A task is safer to hand to an AI when an error will be obvious or easy to test. It is riskier when a wrong answer looks plausible.
13.1 What a Language Model Produces
A large language model is trained to predict and generate text from the context it receives. That training includes natural language and a great deal of code. Additional training makes the model better at following instructions and producing useful answers.
The important word here is generate. The model does not retrieve a correct answer from a database. It generates a response that fits your request and the context available to it. This is why it can explain an R error very well, but also invent a function name that sounds as though it should exist.
The model’s wording does not measure its certainty. A correct function name and an invented one can arrive in the same confident tone. People often use hesitation as a clue that another person is unsure; generated text does not provide that clue reliably. Judge an answer from its code, sources and results rather than from how confident it sounds.
13.2 Three Ways to Use AI for Code
Broadly, there are three ways to use AI for help with data analysis.
A chatbot. You are likely familiar with chatting with Claude, ChatGPT or Gemini through a web page or app. Chatbots are useful for asking questions, brainstorming and learning about a new method. You can paste in an error message, attach a screenshot or upload a script and ask what went wrong. The chatbot only knows what you give it in that conversation, along with any files or saved project instructions it can access.
An in-editor assistant. This type of assistant lives inside another program. It can see the open script, workbook or other selected project context. It offers completions while you type and answers questions without as much copying between windows. GitHub Copilot in Positron and Copilot in Excel are examples. The convenience is useful, but it also makes suggestions very easy to accept without reading them.
A coding agent. An agent can inspect several files, edit them, run code, read the resulting errors and try again. You give it a task rather than asking for one line of code. This makes an agent much more useful for work that crosses several scripts or files. It also gives the agent more opportunities to make changes you did not intend. Your job shifts from typing each line to defining the task, setting limits and reviewing what changed.
Up to this point, you have mostly used AI through a chatbot. There is also a course chatbot that has access to this textbook and should have more context about the examples and conventions used here. In this module we will begin using AI inside the software where we do the analysis.
13.3 Current AI products
There are a handful of companies that provide most of the AI tools that are currently being used. While this is a dynamic space, I’ve tried to list the current products that are available below to help you get a sense of what is available to you.
Claude is Anthropic’s general chatbot. It accepts documents, spreadsheets, images and code, and conversations can be organized into projects with shared instructions and files. Claude Code is Anthropic’s coding agent. It works with a folder or repository and can inspect files, edit them and run commands. A subscription is needed to access this coding agent.
ChatGPT is OpenAI’s general chatbot. It can work with uploaded files, run data analysis and use tools such as web search. Codex is OpenAI’s coding agent. It can work with a repository, edit files, run commands and show the changes for review.
Gemini is Google’s general chatbot. It is closely connected to Google products such as Drive, Docs and Sheets, and it can accept code files and repositories as context. Google also provides Antigravity, a coding agent built on Gemini, which we use in Chapter 15.
Microsoft Copilot is Microsoft’s name for a family of AI assistants. There is a Copilot chatbot, which we use in Chapter 14. However, perhaps more useful are the tools that are embedded in Microsoft products like Word and Excel. The Excel copilot tool can create formulas, summaries and charts using the workbook that is already open.
GitHub Copilot is a coding assistant that works in editors such as Positron, as well as on GitHub itself. It began mainly as an inline code-completion tool, but it now also includes chat and agent modes that can edit several files and run code. Note that Copilot relies on models from Anthropic, OpenAI, and others. We do not use it in this course, but it is free for students through GitHub Education and is worth knowing about if you keep working in Positron.
Cursor is a code editor built around AI assistance. It looks and behaves much like VS Code, but chat, code completion and coding agents are central parts of the editor. It can use models from several AI companies.
Open-source and open-weight models can be downloaded and run on your own computer or on a server you control. This can provide more control over the model and where the data go, although running a capable model may require substantial computing power. Typically, most individual users prefer to use one of the big AI companies rather than these open source products. However, companies that are heavy users have started using free, open source products for mundane tasks, while giving more complex tasks to frontier models from companies like Anthropic and OpenAI.
Features, access and product names change quickly, so check the current documentation before deciding which one to use.
13.4 Eight principles for data analysis with AI
You are still responsible. AI can perform much of the work, but you remain responsible for the data, methods, results, and conclusions. You should understand the analysis thoroughly and be able to explain and defend it.
Use AI to explore, but verify before relying on it. AI is especially valuable for inexpensive exploration. You may have an interesting idea that would take days to code, but an AI agent can investigate it quickly. The result can help you decide whether the idea is worth pursuing. Other agents can help identify obvious problems, but agreement between agents is not proof. If you decide to pursue the analysis, you must verify the data, code, methods, and results yourself.
Protect your data. Do not provide confidential, personal, proprietary, or sensitive data to an AI system unless that system has been approved for such data. When possible, use de-identified data, synthetic examples, or a description of the data structure.
Manage the context. AI needs the relevant files, definitions, background, and project conventions to do good work. Too little context encourages guessing, but too much irrelevant context can also reduce performance. Maintain a current
README.md, data dictionary, or project instruction file that records the information the AI needs. Update it as the project changes. You can also ask your AI agent to continually maintain this and other files that provide them with context in their next session.Define the task and the desired output. Be explicit about what you want the AI to produce, what constraints it must follow, which files it may use or change, and how it should determine whether the task is complete.
Make AI show and explain its work. Ask the AI to describe its approach, assumptions, data transformations, exclusions, and analytical choices. Require inspectable code and supporting sources where appropriate. An explanation supports understanding, but it does not replace verification.
Use multiple agents and models strategically. Do not rely entirely on one model or one conversation. Use different agents for different roles: one might develop an analysis, another audit it, and another attempt an independent implementation. Give reviewers specific tasks rather than simply asking whether the first answer is correct.
Document the analysis and your use of AI. Preserve the raw data, code, prompts or instructions that materially affected the analysis, and important analytical decisions. Disclose meaningful AI assistance and ensure that someone else could reproduce the final results without relying on the original AI conversation.
13.5 Common Failure Modes
Invented functions and packages
A model may suggest a plausible function such as summarise_by() or read_excel_sheet() even when it does not exist. R stops with an error, so this is usually easy to find. Search the package’s official reference before installing a package or changing the script around an unfamiliar function.
In Excel, an invented or unavailable function produces #NAME?. The same error can mean that a real Microsoft 365 function is unavailable in an older Excel version, so check the function’s Microsoft support page and its version list.
Assumed column names
If the model has not seen the data, it may write code for columns named yield, crop and year. The file may use Yield, Crop_Name and Harvest_Year. Supply glimpse() output or a short description of the columns before asking for code. For Excel, supply the exact header row, two or three made-up example rows, the sheet and range, and whether the range is an Excel Table.
Out-of-date syntax
Packages change. An old function or argument may have been replaced since the model’s training data was collected. For example, gather() still appears in older code, while current tidyr uses pivot_longer(). “Superseded” means an older function still works but is no longer recommended; “deprecated” normally warns; “defunct” errors. Package documentation is the check.
Numbers produced without a reproducible calculation
Do not ask a chatbot to calculate a reported statistic from pasted rows and then trust the number. Ask for code, run it in R and keep the calculation in the script. In Excel, every reported number should come from a visible formula rather than a typed constant; use FORMULATEXT when the calculation needs to be audited.
Assumptions from another place
Answers about agriculture may default to US crops, agencies and geography. Saskatchewan data may use rural municipalities rather than counties, SCIC definitions rather than a US crop-insurance program, and bushels per acre rather than another yield unit.
Figure 13.1 shows the same 2025 canola yield in three units. A conversion error can change the numerical scale without causing an R error.
The failures differ in how easy they are to notice. An invented function stops the script. A wrong unit or incomplete join may not.
13.6 Tasks That Are Easy to Check
AI assistance is useful when the result can be tested quickly:
- Explain an R error message and identify the line it refers to.
- Suggest the name of a function, followed by a check of its documentation.
- Draft repetitive syntax such as a
case_when()block. - Explain an unfamiliar line of code.
- Produce a first draft of a plot whose labels, groups and values can be inspected.
- Explain an Excel error such as
#SPILL!,#VALUE!or#REF!, followed by checking the referenced cells. - Draft a
SUMIFSorXLOOKUP, followed by filtering the source table and checking matching rows.
In each case, R or the user provides a separate check.
13.7 Tasks Requiring More Judgment
Some tasks require information the model does not have:
- Choosing the analysis without a clear question.
- Deciding whether the observations represent the population of interest.
- Judging whether an agricultural result is plausible in its stated units and location.
- Choosing which caveat matters to a producer, client or policy maker.
- Supplying references that have not been checked against the original source.
AI can help list possibilities, but the final decision has to be justified from the question, data and subject matter.
13.8 Data That Should Not Be Sent
Prompts and code sent to a hosted AI service leave the computer. What is stored, for how long, and whether it is used to improve a service depends on the provider, account and organization settings.
The datasets in this course are public or synthetic. A producer’s field records, financial statements, contact information, employee data and material covered by a confidentiality agreement are different. Do not paste someone else’s confidential data into a consumer AI tool. Describe the column names, types and problem without supplying the values, or use a service approved by the organization that owns the data.
An in-editor assistant may use the current file or other project context. Check what context has been selected and close or exclude files that should not be sent.
Copilot in Excel sends Table content to Microsoft’s service; it is not a local calculation. An organization-managed Microsoft 365 account may operate within that organization’s data protections, while a consumer chatbot account does not inherit them. Follow the data owner’s policy and use headers plus made-up rows when approval is uncertain.
13.9 Course Policy
AI tools may be used for module practice and the course project. They may not be used on tests.
Tests check whether you can read the data, choose an analysis and recognize an implausible answer without the tool doing those jobs for you. For submitted work, you must be able to explain every line of code and account for every reported number.
A 2026 working paper reports an observational comparison involving 26,811 Chinese secondary students who self-selected into using generative AI (Strömberg et al. 2026). After adoption, homework scores rose and completion times fell, while exam scores fell.
The paper argues that students substituted the tool for practice. That is a plausible interpretation, but its subgroup comparison compares students who changed how they did their homework in different ways; it is not a clean experimental test of the mechanism. The result therefore does not imply that every use of AI reduces learning.
Other courses have their own rules. Check the syllabus for each course before using AI on assessed work.
13.10 The Rest of This Module
The next chapter uses a chatbot, which is the simplest way to work and the one most of you have already tried. The chapter after that moves to a coding agent, which reads your files and runs your code rather than waiting for you to paste things in. We then apply the agent to graphing, because a figure makes many coding errors easy to see. The final chapter, on Excel, is optional. This module practises AI assistance mainly in Positron because scripts make the generated work visible; the same checking habits apply to formulas and charts produced in Excel.
- 3Blue1Brown, But what is a GPT? gives a visual explanation of language-model prediction.
- Simon Willison, Not all AI-assisted programming is vibe coding develops the explanation standard used above.


