10 Graphing best practices
A good chart tells a story. It draws the reader’s attention to a pattern, comparison, or relationship in the data. Charts do this by turning numbers into visual comparisons of position, length, area, or colour. When those comparisons make the story in the data easier to see, a chart is more useful than a table of numbers.
The figures in this chapter are illustrations rather than exercises: they are drawn from published charts and from small examples generated in the chapter’s own code, so there are no data files to download. The chapters that follow use the Saskatchewan variety data and name their files at the top.
One of the best charts I’ve seen1 tells the story of Napoleon’s invasion and subsequent retreat from Russia in 1812 (Figure 10.1).
The width of the band represents the size of the army, its position follows the route of advance and retreat, and the line below records temperature and dates during the retreat. The tan band begins with 422,000 men crossing into Russia. The black band returns with 10,000. Think of all the information that is embodied in this graph – the size of the army, where they travelled, their dates of travel (on their retreat), and temperature. A reader could never take in this type of information in a set of data tables.
Figure 10.2 was published by the Georgia Department of Public Health in May 2020. It appears to show COVID-19 cases falling steadily in the five counties with the most confirmed cases.
The dates along the horizontal axis are out of order. They were sorted by case count, so the decline is built into the chart rather than present in the data. The labels contain the information needed to catch the problem, but only a reader who checks every date will find it.
These examples set the two jobs of a chart. It should make the comparison easy to see, and it should represent the data honestly.
10.1 General principles
Decide whether a chart actually helps
First, you need to decide whether you actually need a graph to get the point across to your audience. Sometimes tables, or just text, do a better job. For example, if I wanted to compare US and Canadian wheat production, I could do so through Figure 10.3. Or I could simply write: In 2024 the United States produced 53.7 million tonnes of wheat and Canada produced 34.9 million. Turning this text into a graph only distracts the reader.
However, if we wanted to give the reader a sense of Canada’s place among the world’s leading producers of wheat, then Figure 10.4 would do that better than a table or text.
Check the scale and order
Charts can mislead while still plotting correct numbers. Figure 10.2 misled audiences by changing the order of the dates. There are many other ways to mislead. A lot of this comes down to altering the way data is presented relative to the audience’s expectations. In Figure 10.2 we expect dates to be ordered chronologically.
One easy way to mislead is to start your vertical axis at a number other than zero. We generally expect that the relative size of bars on a bar chart reflects the relative differences in the values being plotted. But when the vertical axis starts at a number other than zero this isn’t true.
Consider the two panels in Figure 10.5. Both show the same five malting barley varieties, with yields from 103 to 114 bushels per acre. In panel A, the vertical axis starts at 100. This makes the bar for CDC Copeland appear to be about one-fifth the size of the bar for CDC Churchill. In panel B, we see that the bar for Copeland is 90% the size of the bar for Churchill, reflecting that Copeland’s yields were roughly 90% of Churchill’s in the trials.
Excel chooses axis bounds automatically and may produce panel A by default when the values are close together. For a column chart, right-click the value axis, choose Format Axis, and set the Minimum bound to 0. The zero baseline displays the size difference honestly, but it does not establish that small differences between trial averages are meaningful.
Line charts are different. A line shows position rather than length, so its vertical axis does not always need to begin at zero. The displayed range should still be wide enough that ordinary movement does not look like a crisis.
Use descriptive labels
Every chart should tell the reader what they are looking at. At a minimum, this usually means clear labels for the horizontal and vertical axes, along with tick-mark labels that make the scale easy to interpret. Labels should also make the units clear. If a chart shows yield, for example, the reader should not have to guess whether it is measured in bushels per acre, tonnes per hectare, or something else.
There is some art to labelling a chart. Too little information leaves the reader guessing, while too much text can make the chart feel cluttered. Include the information the reader needs to interpret the figure, but avoid labels that simply repeat information already obvious from the chart.
When a figure will appear in a report, it generally looks best to leave the chart title blank in your statistical software and give the figure a caption in the document itself. For example, if you create a chart in R and insert it into a Word report, you can leave the title blank in R and use Word’s captioning function instead. This keeps figure titles consistent with the formatting and numbering used throughout the report.
Remove clutter
Gridlines, borders, markers and decoration all compete with the data. Keep the elements that help the comparison and remove the rest. Some formatting that I would suggest (some of this is personal preference):
- A minimal number of horizontal gridlines, spaced at intervals of 1, 2, 5, 10, 20, 25, 50 or 100 (not 3, 7 or 12).
- No vertical gridlines, except for scatter plots.
- When dealing with large numbers, put the vertical axis in thousands or millions.
Three-dimensional effects distort length and area. Patterned fills were useful when charts had to be printed in black and white, but on a screen they make the bars harder to read.
Label lines directly
A legend makes the reader match each line to a separate key. When there is room, put the series name at the end of the line instead. Admittedly, both of the panels in Figure 10.8 are a bit busy. But at least in panel B, you can easily identify which line corresponds to a particular variety.
Legends are still useful when direct labels would overlap, or when the same symbols appear in several panels.
Annotate what needs explanation
A short note within the chart can explain an unusual point or change, or draw the reader’s attention to an important part of the story in the data (as in panel B of Figure 10.9). Use notes sparingly, however: too many can clutter the chart and obscure the very patterns they are meant to highlight (as in panel A of Figure 10.9).
Adding more information to charts
While most charts are two dimensional, there are ways of bringing in more information, through colours, shapes and sizes. Figure 10.1 is a great example of artfully condensing a considerable amount of information into a single plot. Another example is Figure 10.10, where the point size represents spring wheat seeded acres in addition to wheat and barley yield, and the point colour marks each risk zone’s soil zone.
However, one should be wary of trying to cram too much information into a single plot. A particular no-no is to use two different vertical axes on the same plot, as in Figure 10.11. This is a poor choice for a number of reasons, notably that the audience has to work to understand which line corresponds to which axis. It is almost always better to plot these as two separate charts.
Be intentional with colour
Colour can distinguish groups, show an ordered scale or draw attention to one observation. For example, in Figure 10.4 we highlighted Canada in the plot, as I was assuming our audience would be Canadian. If we were presenting on the possible impacts on world wheat markets of Russia’s invasion of Ukraine, then we might want to highlight Russia and Ukraine. In Figure 10.10, we denoted different soil zones with different colours.
Choose colours that remain distinguishable for readers with colour-vision deficiencies and when printed in grey. It is also best practice to distinguish series by colour and some other means, where possible. For example, if you have three lines on a graph, you can use colours, direct labels, and dashes to differentiate them. Again, much depends on your particular data. In Figure 10.13 the lines on the graph are all very distinct, so direct labels are enough to allow the audience to distinguish each line across the whole graph. If the lines intersected each other more, we might want to add different dashing to the lines.
10.2 Types of charts
Bar charts: compare categories
A bar chart compares a value across categories such as yield by crop, acres by region or revenue by customer. Sort bars by value unless the categories have a natural order (as in the example of Figure 10.2, where bars should have been ordered chronologically). Start the value axis at zero. Note that bar charts often look better when horizontal, particularly where there are a large number of categories.
Line charts: show change over time
A line chart shows movement over a continuous scale, usually time. The connection between points implies that the order and interval have meaning. Most commonly they are used with a measure of time on the horizontal axis. Do not connect unordered categories such as Barley, Canola and Oats with a line.
Scatter plots: compare two numeric variables
A scatter plot places one numeric variable on each axis and one observation at each point.
The points rise from left to right in this 2025 cross-section. Both crops in a zone shared the same growing season, so the pattern says little by itself about whether zones that are good for wheat are generally good for barley. Module 10 develops the numerical tools for describing relationships like this.
Stacked and grouped bars: compare composition
A stacked bar shows a total and its components.
One issue with stacked bar charts is that only the bottom segment and the total of each bar can be easily compared. In Figure 10.15 it is difficult to compare the seeded acreage of canola over time.
Grouped bars give each component a zero baseline, but they become crowded as the number of groups grows, making it hard to compare different crops within one year and crops across time.
An alternative is plotting the crops in separate plots, as in Figure 10.17.
Pie charts: show a simple share
A pie chart shows parts of a total. Most authorities on charts really dislike pie charts. One reason is that angles and areas are harder to compare than lengths, so a bar chart is usually a better way of visualizing differences. Labels on pie charts are also somewhat hard to get right, particularly when pie slices are small. If a pie chart is used, keep the number of slices small and label them directly with percentages.
Ranges, trend lines and reference lines
Charts often are used to show averages. However, one might also want to give the audience some sense of the variation in the data. This can be done through shading around the line. The shaded area might represent the minimum and maximum, or particular percentiles. The definition of the shaded area should be clearly stated in the figure caption.
In Figure 10.19 we use shading to indicate the 10th and 90th percentiles of the risk-zone yields.
A fitted line summarizes the trend in a scatter plot, as in Figure 10.20. I tend to advocate against trend lines – if a trend is not so obvious that it can be seen in the raw data, why include the line?
A reference line can show an average, a target, a break-even value or last year’s result, as in Figure 10.21. Include one when the comparison against it is the point of the chart: yields against the five-year average, prices against break-even. If the reader is not meant to compare each bar against that value, the line is just another element competing for attention, and the information may belong in the caption or the surrounding text instead.
- Edward Tufte, The Visual Display of Quantitative Information – the classic on chart design, including the Minard chart that opens this chapter.
- Cole Nussbaumer Knaflic, Storytelling with Data – a practical book on choosing and decluttering charts, available through the USask Library. Her blog continues the book with short worked examples.
- William S. Cleveland, Visualizing Data – a more technical treatment of how to plot data so that patterns can actually be seen; the source of many of the principles behind good statistical graphics.
I first saw this chart in The Visual Display of Quantitative Information by Edward Tufte, who said “it may well be the best statistical graphic ever drawn” (Tufte 2001, 40).↩︎




























