Data sets
This page lists every data set used in the book, in alphabetical order by file name. Click a file name to download it directly. Most are .Rdata files that can be loaded with load(); a couple are plain .csv files used to illustrate reading data from text files.
| File | Description |
|---|---|
afl24.Rdata |
Attendance figures for every AFL game played between 1987 and 2010, used to illustrate confidence intervals (10 Estimating unknown quantities from a sample). |
aflsmall.Rdata |
Winning margins for AFL games in the 2010 season, used throughout the introduction to descriptive statistics and graphics (5 Descriptive statistics, 6 Drawing graphs). |
aflsmall2.Rdata |
Winning margins for AFL games in every season from 1987 to 2010, used to compare many years’ worth of data with boxplots (6 Drawing graphs). |
agpp.Rdata |
Survey responses on an attitude before and after an intervention, used to illustrate McNemar’s test (12 Categorical data analysis). |
anscombesquartet.Rdata |
Anscombe’s Quartet: four small \(X\)/\(Y\) data sets that share identical summary statistics but look very different when plotted (5 Descriptive statistics). |
awesome.Rdata |
Fictitious “test of awesomeness” scores for two independent groups, used to introduce the two-sample Wilcoxon (Mann-Whitney) test (13 The t-test). |
awesome2.Rdata |
The same awesomeness scores as awesome.Rdata, but stored as two separate vectors rather than one data frame (13 The t-test). |
berkeley.Rdata |
The 1973 UC Berkeley graduate admissions data by department and gender, used to illustrate Simpson’s paradox (1 Why do we learn statistics?). |
booksales.csv |
A small table of monthly book sales figures, used as a first example of reading a CSV file into R (4 R mechanics, 7 Pragmatic matters). |
booksales.Rdata |
The same book sales data as booksales.csv, stored as an .Rdata file to illustrate load() (4 R mechanics). |
booksales2.csv |
The book sales data re-saved with unusual delimiters, quoting, and a missing-value code, used to teach the more obscure arguments of read.csv() (7 Pragmatic matters). |
cakes.Rdata |
A small matrix of quality ratings for different cake recipes across repeated tastings, used to illustrate transposing a matrix (7 Pragmatic matters). |
chapek9.Rdata |
Fictitious survey data on whether humans and robots prefer puppies, flowers, or data files, used for chi-square tests of independence (12 Categorical data analysis, 17 Bayesian statistics). |
chico.Rdata |
Paired early- and late-semester test scores for Dr Chico’s class, used to introduce the paired-samples \(t\)-test (13 The t-test, 17 Bayesian statistics). |
clinicaltrial.Rdata |
A fictitious clinical trial comparing drug and therapy combinations on mood, used throughout the one-way and factorial ANOVA chapters (5 Descriptive statistics, 14 Comparing several means (one-way ANOVA), 16 Factorial ANOVA, 17 Bayesian statistics). |
coffee.Rdata |
A fictitious, unbalanced experiment on how milk and sugar affect coffee-fuelled “babbling,” used for factorial ANOVA (16 Factorial ANOVA). |
effort.Rdata |
Fictitious data relating hours worked to grade received for 10 students, used to illustrate correlation (5 Descriptive statistics). |
happy.Rdata |
Before-and-after happiness scores for students taking a statistics class, used to introduce the one-sample Wilcoxon signed-rank test (13 The t-test). |
harpo.Rdata |
Grades for Dr Harpo’s class, split by tutorial group, used to introduce the independent-samples \(t\)-test (13 The t-test, 17 Bayesian statistics). |
likert.Rdata |
Raw Likert-scale survey responses, used to illustrate factors and frequency tables (7 Pragmatic matters). |
nightgarden.Rdata |
Transcribed dialogue (speaker and utterance) from the TV show In the Night Garden, used throughout the data handling and scripting chapters (7 Pragmatic matters, 8 Basic programming). |
nightgarden2.Rdata |
A variant In the Night Garden data frame, used to illustrate subsetting data frames with square brackets (7 Pragmatic matters). |
parenthood.Rdata |
100 days of sleep and grumpiness measurements, used throughout the correlation and regression chapters (5 Descriptive statistics, 6 Drawing graphs, 15 Linear regression, 17 Bayesian statistics). |
parenthood2.Rdata |
The same sleep and grumpiness data as parenthood.Rdata, but with some values deleted, used to illustrate handling missing data in correlations (5 Descriptive statistics). |
randomness.Rdata |
The card suits chosen by 200 people asked to pick “randomly,” used for chi-square goodness-of-fit and independence tests (12 Categorical data analysis). |
repeated.Rdata |
A repeated-measures working-memory/reaction-time experiment in wide form, used to teach reshaping data between wide and long form (7 Pragmatic matters). |
rtfm.Rdata |
A small \(2 \times 2\) factorial data set on class attendance, textbook reading, and grades, used to show that ANOVA is a special case of the linear model (16 Factorial ANOVA). |
salem.Rdata |
A small, sparse contingency table from a fictitious witch-trial study, used to illustrate Fisher’s exact test (12 Categorical data analysis). |
work.Rdata |
Hours worked, tasks completed, pay, and weekday, a mixed numeric/categorical data set used to illustrate the correlate() function (5 Descriptive statistics). |
zeppo.Rdata |
Grades for 20 psychology students in Dr Zeppo’s class, used to introduce the one-sample \(z\)-test (13 The t-test). |