Data sets

This page lists every data set used in the book, in alphabetical order by file name. Click a file name to download it directly. Most are .Rdata files that can be loaded with load(); a couple are plain .csv files used to illustrate reading data from text files.

File Description
afl24.Rdata Attendance figures for every AFL game played between 1987 and 2010, used to illustrate confidence intervals (10  Estimating unknown quantities from a sample).
aflsmall.Rdata Winning margins for AFL games in the 2010 season, used throughout the introduction to descriptive statistics and graphics (5  Descriptive statistics, 6  Drawing graphs).
aflsmall2.Rdata Winning margins for AFL games in every season from 1987 to 2010, used to compare many years’ worth of data with boxplots (6  Drawing graphs).
agpp.Rdata Survey responses on an attitude before and after an intervention, used to illustrate McNemar’s test (12  Categorical data analysis).
anscombesquartet.Rdata Anscombe’s Quartet: four small \(X\)/\(Y\) data sets that share identical summary statistics but look very different when plotted (5  Descriptive statistics).
awesome.Rdata Fictitious “test of awesomeness” scores for two independent groups, used to introduce the two-sample Wilcoxon (Mann-Whitney) test (13  The t-test).
awesome2.Rdata The same awesomeness scores as awesome.Rdata, but stored as two separate vectors rather than one data frame (13  The t-test).
berkeley.Rdata The 1973 UC Berkeley graduate admissions data by department and gender, used to illustrate Simpson’s paradox (1  Why do we learn statistics?).
booksales.csv A small table of monthly book sales figures, used as a first example of reading a CSV file into R (4  R mechanics, 7  Pragmatic matters).
booksales.Rdata The same book sales data as booksales.csv, stored as an .Rdata file to illustrate load() (4  R mechanics).
booksales2.csv The book sales data re-saved with unusual delimiters, quoting, and a missing-value code, used to teach the more obscure arguments of read.csv() (7  Pragmatic matters).
cakes.Rdata A small matrix of quality ratings for different cake recipes across repeated tastings, used to illustrate transposing a matrix (7  Pragmatic matters).
chapek9.Rdata Fictitious survey data on whether humans and robots prefer puppies, flowers, or data files, used for chi-square tests of independence (12  Categorical data analysis, 17  Bayesian statistics).
chico.Rdata Paired early- and late-semester test scores for Dr Chico’s class, used to introduce the paired-samples \(t\)-test (13  The t-test, 17  Bayesian statistics).
clinicaltrial.Rdata A fictitious clinical trial comparing drug and therapy combinations on mood, used throughout the one-way and factorial ANOVA chapters (5  Descriptive statistics, 14  Comparing several means (one-way ANOVA), 16  Factorial ANOVA, 17  Bayesian statistics).
coffee.Rdata A fictitious, unbalanced experiment on how milk and sugar affect coffee-fuelled “babbling,” used for factorial ANOVA (16  Factorial ANOVA).
effort.Rdata Fictitious data relating hours worked to grade received for 10 students, used to illustrate correlation (5  Descriptive statistics).
happy.Rdata Before-and-after happiness scores for students taking a statistics class, used to introduce the one-sample Wilcoxon signed-rank test (13  The t-test).
harpo.Rdata Grades for Dr Harpo’s class, split by tutorial group, used to introduce the independent-samples \(t\)-test (13  The t-test, 17  Bayesian statistics).
likert.Rdata Raw Likert-scale survey responses, used to illustrate factors and frequency tables (7  Pragmatic matters).
nightgarden.Rdata Transcribed dialogue (speaker and utterance) from the TV show In the Night Garden, used throughout the data handling and scripting chapters (7  Pragmatic matters, 8  Basic programming).
nightgarden2.Rdata A variant In the Night Garden data frame, used to illustrate subsetting data frames with square brackets (7  Pragmatic matters).
parenthood.Rdata 100 days of sleep and grumpiness measurements, used throughout the correlation and regression chapters (5  Descriptive statistics, 6  Drawing graphs, 15  Linear regression, 17  Bayesian statistics).
parenthood2.Rdata The same sleep and grumpiness data as parenthood.Rdata, but with some values deleted, used to illustrate handling missing data in correlations (5  Descriptive statistics).
randomness.Rdata The card suits chosen by 200 people asked to pick “randomly,” used for chi-square goodness-of-fit and independence tests (12  Categorical data analysis).
repeated.Rdata A repeated-measures working-memory/reaction-time experiment in wide form, used to teach reshaping data between wide and long form (7  Pragmatic matters).
rtfm.Rdata A small \(2 \times 2\) factorial data set on class attendance, textbook reading, and grades, used to show that ANOVA is a special case of the linear model (16  Factorial ANOVA).
salem.Rdata A small, sparse contingency table from a fictitious witch-trial study, used to illustrate Fisher’s exact test (12  Categorical data analysis).
work.Rdata Hours worked, tasks completed, pay, and weekday, a mixed numeric/categorical data set used to illustrate the correlate() function (5  Descriptive statistics).
zeppo.Rdata Grades for 20 psychology students in Dr Zeppo’s class, used to introduce the one-sample \(z\)-test (13  The t-test).