About me

Data visualisation specialist, mainly using R, Python, and D3.


Background in statistics, operational research, and data science.


Author of several R packages (mainly for visualisation).


Co-author of Royal Statistical Society’s Best Practices for Data Visualisation guidance.

Grid of R package hex logos

What makes data science successful?

Not just:

  • A good model
  • Good code
  • A beautiful visualisation
  • A reproducible analysis

Can someone else understand it, use it, question it, and build on it?

Code is communication

We often think:

code → computer


But actually:

code → computer
code → future me
code → colleague
code → collaborator
code → user

A piece of code has an audience

Who might encounter the code you write?

  • You, tomorrow
  • You, six months from now
  • A colleague
  • A collaborator
  • A maintainer
  • A student
  • Someone who finds it on GitHub

Write for the people who will have to understand it.

It works!

Consider this perfectly functional code:

x <- read.csv("data.csv")

x$x2 <- x$x * 100

x <- x[x$x2 > 50, ]

plot(x$x, x$y)

It works.

But what does it mean?

The problem isn’t the code

The problem is context. Someone else needs to know:

  • What is this data?
  • Why multiply by 100?
  • Why 50?
  • What does x represent?
  • Why this plot?
  • What decision is it supporting?

Technical reproducibility isn’t the same as understanding.

Your audience doesn’t care about your code

They probably don’t care whether you used R or Python. And that’s okay.


They might care about:

  • What happened?
  • Why did it happen?
  • What does this mean?
  • How certain are we?
  • What should we do next?

Code → communication

A data analysis might begin as:

data
  ↓
code
  ↓
plot

 

But the person making the decision sees:

question
  ↓
evidence
  ↓
understanding
  ↓
decision

Data visualisation is an interface

Analysis

  • Data
  • Code
  • Statistics
  • Models

Audience

  • Questions
  • Decisions
  • Actions
  • Curiosity


A chart isn’t the end of an analysis. It’s the beginning of a conversation.

Different questions need different answers

A technical collaborator might ask:

“What’s the uncertainty around this estimate?”

A manager might ask:

“Is this difference meaningful?”

A policymaker might ask:

“What should we do?”

A member of the public might ask:

“What does this mean for me?”

Line charts with one highlighted

Area charts styled as jam jars

A chart can be technically correct…

…and still fail.

Maybe:

  • It’s not interesting enough for people to pay attention
  • The title doesn’t explain the point
  • The annotation assumes too much knowledge
  • The colours aren’t meaningful
  • The uncertainty is hidden
  • The important comparison isn’t obvious

Correctness is necessary, but not sufficient.

Translation isn’t dumbing things down

Good communication doesn’t mean removing complexity.

It means:

  • Choosing what matters
  • Providing context
  • Explaining unfamiliar concepts
  • Making uncertainty visible
  • Connecting evidence to the question

Simple communication can sit on top of complicated analysis.

Collaboration starts with translation

When building for non-technical users, we constantly translate.

Why did productivity change last month?

I’ve built a regression model. Here are the estimates, residuals, and confidence intervals.

Our models show that people who did X increased their productivity slightly, but this may just have been down to chance in who was surveyed. Our estimates aren’t precise enough to be sure.

Should we do anything differently next month?

Speaking the same language

If I asked you to document this code:

x <- read.csv("data.csv")

x$x2 <- x$x * 100

x <- x[x$x2 > 50, ]

plot(x$x, x$y)


You probably assume I mean “add some code comments”.

Reproducibility != communication

We usually talk about reproducibile analysis as:

“Can someone run my code?”


And documentation as:

“Can someone understand my code?”


But there’s another question:

Can someone understand what I was trying to do?

“How do I get started?”

  • README — the 2-minute “what is this and how do I use it?”
  • Quarto / Markdown pages
  • Examples / tutorials
  • Issue templates / contribution guides
  • Feedback form

The non-technical collaborator is not the audience

They’re part of the team. They bring:

  • Domain knowledge
  • Context
  • Different questions
  • Different priorities
  • Knowledge of the eventual users

Sometimes the person who can’t explain your code is the person who best understands whether your answer is useful.

Collaboration changes the work

When we show our work to someone else:

  • They ask questions
  • They spot assumptions
  • They find bugs
  • They suggest different interpretations
  • They identify missing context

Feedback isn’t just quality control. It changes what we build.

The beginner is a collaborator too

Beginners:

  • Ask “obvious” questions
  • Find confusing documentation
  • Challenge assumptions
  • See things differently

Make contribution cheap


Collaboration doesn’t have to mean:

“Please clone this repository, create a branch and submit a pull request.”

Make contribution cheap

Contribution can mean conversations, comments in Word documents, post-it notes:

  • “This doesn’t make sense.”
  • “Have you considered…?”
  • “Here’s an example.”
  • “I spotted a typo.”
  • “Our users would need…”
  • “Can we see this another way?”

All of these are contributions.

Poverty data gaps explorer

Airtable form screenshot

Programming in public

Sharing unfinished work with a community can feel uncomfortable.

But publishing code, examples, visualisations and ideas creates opportunities for:

  • Feedback
  • Questions
  • Collaborations
  • Teaching
  • New ideas
  • New people

The R community gets this right

Think about the things that make R more than a programming language:

  • Open source packages and documentation
  • Blogs, and tutorials
  • User groups, meetups, and conferences
  • Communities like R-Ladies, rainbowR, and DSLC
  • TidyTuesday

The technology matters. But the people around the technology matter more.

TidyTuesday

One dataset. Thousands of possibilities.

People share:

  • Code
  • Visualisations
  • Ideas
  • Questions
  • Techniques

The community learns by seeing what everyone else did.

TidyTuesday logo

A TidyTuesday Community

TidyTuesday is a fantastic example of growing a community by making contribution easier:

  • Data in CSV format.
  • R/Python/Julia packages for loading data.
  • R package for making GitHub pull requests with new datasets. You don’t need to know Git.
  • Friendly spaces (e.g. Slack) to ask questions and get help.

Don’t ask everyone to become an expert before they’re allowed through the door.