Adair, G. (1984). The hawthorne effect: A reconsideration of the
methodological artifact. Journal of Applied Psychology,
69, 334–345.
Agresti, A. (1996). An introduction to categorical data
analysis. Wiley.
Agresti, A. (2002). Categorical data analysis (2nd ed.). Wiley.
Akaike, H. (1974). A new look at the statistical model identification.
IEEE Transactions on Automatic Control, 19, 716–723.
Amrhein, V., Greenland, S., & McShane, B. (2019). Retire statistical
significance. Nature, 567, 305–307.
Anscombe, F. J. (1973). Graphs in statistical analysis. American
Statistician, 27, 17–21.
Bengtsson, H. (2022).
R.matlab: Read and write MAT files and call
MATLAB from within r.
https://github.com/HenrikBengtsson/R.matlab
Benjamin, D. J., Berger, J. O., Johannesson, M., Nosek, B. A.,
Wagenmakers, E.-J., Berk, R., Bollen, K. A., et al. (2018). Redefine
statistical significance. Nature Human Behaviour,
2(1), 6–10.
Benjamini, Y., & Hochberg, Y. (1995). Controlling the false
discovery rate: A practical and powerful approach to multiple testing.
Journal of the Royal Statistical Society, Series B,
57(1), 289–300.
Bickel, P. J., Hammel, E. A., & O’Connell, J. W. (1975). Sex bias in
graduate admissions: Data from Berkeley. Science,
187, 398–404.
Box, G. E. P. (1976). Science and statistics. Journal of the
American Statistical Association, 71, 791–799.
Box, J. F. (1987). Guinness, gosset, fisher, and small samples.
Statistical Science, 2, 45–52.
Braun, J., & Murdoch, D. J. (2007). A first course in
statistical programming with R. Cambridge University
Press Cambridge.
Brown, M. B., & Forsythe, A. B. (1974). Robust tests for the
equality of variances. Journal of the American Statistical
Association, 69, 364–367.
Bürkner, P.-C. (2017). Brms: An R package for
Bayesian multilevel models using Stan.
Journal of Statistical Software, 80(1), 1–28.
Campbell, D. T., & Stanley, J. C. (1963). Experimental and
quasi-experimental designs for research. Houghton Mifflin.
Champely, S. (2020).
Pwr: Basic functions for power analysis.
https://github.com/heliosdrm/pwr
Cochran, W. G. (1954). The χ2 test of goodness of
fit. The Annals of Mathematical Statistics, 23,
315–345.
Cohen, J. (1988). Statistical power analysis for the behavioral
sciences (2nd ed.). Lawrence Erlbaum.
Cook, R. D., & Weisberg, S. (1983). Diagnostics for
heteroscedasticity in regression. Biometrika, 70,
1–10.
Cramér, H. (1946). Mathematical methods of statistics.
Princeton University Press.
Cumming, G. (2014). The new statistics: Why and how.
Psychological
Science,
25(1), 7–29.
https://doi.org/10.1177/0956797613504966
Devezer, B., Navarro, D. J., Vandekerckhove, J., & Buzbas, E. O.
(2021). The case for formal methodology in scientific reform.
Royal
Society Open Science,
8(3), 200805.
https://doi.org/10.1098/rsos.200805
Dunn, O. J. (1961). Multiple comparisons among means. Journal of the
American Statistical Association, 56, 52–64.
Ellis, P. D. (2010). The essential guide to effect sizes:
Statistical power, meta-analysis, and the interpretation of research
results. Cambridge University Press.
Ellman, M. (2002). Soviet repression statistics: Some comments.
Europe-Asia Studies, 54(7), 1151–1172.
Evans, J. St. B. T., Barston, J. L., & Pollard, P. (1983). On the
conflict between logic and belief in syllogistic reasoning. Memory
and Cognition, 11, 295–306.
Evans, M., Hastings, N., & Peacock, B. (2011). Statistical
distributions (3rd ed). Wiley.
Fisher, R. A. (1922a). On the interpretation of χ2 from contingency
tables, and the calculation of p. Journal of the Royal
Statistical Society, 84, 87–94.
Fisher, R. A. (1922b). On the mathematical foundation of theoretical
statistics. Philosophical Transactions of the Royal Society A,
222, 309–368.
Fisher, R. A. (1925). Statistical methods for research workers.
Oliver; Boyd.
Fox, J., & Weisberg, S. (2011). An R companion to
applied regression (2nd ed.). Sage.
Fox, J., Weisberg, S., & Price, B. (2026).
Car: Companion to
applied regression.
https://github.com/bprice2652/car_repo
Fox, J., Weisberg, S., Price, B., Friendly, M., & Hong, J. (2026).
Effects: Effect displays for linear, generalized linear, and other
models.
https://cran.r-project.org/package=effects
Friendly, M. (2011).
HistData: Data sets from the history of
statistics and data visualization.
http://CRAN.R-project.org/package=HistData
Friendly, M., Dray, S., Li, P., & Bellhouse, D. (2025).
HistData: Data sets from the history of statistics and data
visualization.
https://friendly.github.io/HistData/
Gelman, A., & Stern, H. (2006). The difference between
“significant” and “not significant” is not
itself statistically significant. The American Statistician,
60, 328–331.
Genz, A., Bretz, F., Miwa, T., Mi, X., & Hothorn, T. (2026).
Mvtnorm: Multivariate normal and t distributions.
https://codeberg.org/thothorn/mvtnorm
Gunel, E., & Dickey, J. (1974). Bayes factors for independence in
contingency tables. Biometrika, 545–557.
Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements
of statistical learning: Data mining, inference and prediction (2nd
ed). Springer.
Hays, W. L. (1994). Statistics (5th ed.). Harcourt Brace.
Hedges, L. V. (1981). Distribution theory for glass’s estimator of
effect size and related estimators. Journal of Educational
Statistics, 6, 107–128.
Hedges, L. V., & Olkin, I. (1985). Statistical methods for
meta-analysis. Academic Press.
Hoekstra, R., Morey, R. D., Rouder, J. N., & Wagenmakers, E.-J.
(2014). Robust misinterpretation of confidence intervals.
Psychonomic Bulletin & Review,
21(5), 1157–1164.
https://doi.org/10.3758/s13423-013-0572-3
Hogg, R. V., McKean, J. V., & Craig, A. T. (2005). Introduction
to mathematical statistics (6th ed.). Pearson.
Holm, S. (1979). A simple sequentially rejective multiple test
procedure. Scandinavian Journal of Statistics, 6,
65–70.
Hothersall, D. (2004). History of psychology. McGraw-Hill.
Hothorn, T., Zeileis, A., Farebrother, R. W., & Cummins, C. (2022).
Lmtest: Testing linear regression models.
Hsu, J. C. (1996). Multiple comparisons: Theory and methods.
Chapman; Hall.
Ioannidis, J. P. A. (2005). Why most published research findings are
false. PLoS Med, 2(8), 697–701.
Jeffreys, H. (1961). The theory of probability (3rd ed.).
Oxford.
John, L. K., Loewenstein, G., & Prelec, D. (2012). Measuring the
prevalence of questionable research practices with incentives for truth
telling. Psychological Science, 23(5), 524–532.
Johnson, V. E. (2013). Revised standards for statistical evidence.
Proceedings of the National Academy of Sciences, (48),
19313–19317.
Kahneman, D., & Tversky, A. (1973). On the psychology of prediction.
Psychological Review, 80, 237–251.
Kass, R. E., & Raftery, A. E. (1995). Bayes factors. Journal of
the American Statistical Association, 90, 773–795.
Keynes, J. M. (1923). A tract on monetary reform. Macmillan;
Company.
Kruschke, J. K. (2015). Doing Bayesian data analysis: A
tutorial with R, JAGS, and
Stan (2nd ed.). Academic Press.
Kruskal, W. H., & Wallis, W. A. (1952). Use of ranks in
one-criterion variance analysis. Journal of the American Statistical
Association, 47, 583–621.
Kühberger, A., Fritz, A., & Scherndl, T. (2014). Publication bias in
psychology: A diagnosis based on the correlation between effect size and
sample size. Public Library of Science One, 9, 1–8.
Lakens, D. (2013). Calculating and reporting effect sizes to facilitate
cumulative science: A practical primer for
t-tests and ANOVAs.
Frontiers in
Psychology,
4, 863.
https://doi.org/10.3389/fpsyg.2013.00863
Lakens, D., & Caldwell, A. R. (2021). Simulation-based power
analysis for factorial analysis of variance designs.
Advances in
Methods and Practices in Psychological Science,
4(1).
https://doi.org/10.1177/2515245920951503
Langsrud, Ø. (2003). ANOVA for unbalanced data: Use type II instead of
type III sums of squares.
Statistics and Computing,
13(2), 163–167.
https://doi.org/10.1023/A:1023260610025
Larntz, K. (1978). Small-sample comparisons of exact levels for
chi-squared goodness-of-fit statistics. Journal of the American
Statistical Association, 73, 253–263.
Lee, M. D., & Wagenmakers, E.-J. (2014). Bayesian cognitive
modeling: A practical course. Cambridge University Press.
Lehmann, E. L. (2011). Fisher, Neyman, and the creation
of classical statistics. Springer.
Levene, H. (1960). Robust tests for equality of variances. In I. O. et
al (Ed.), Contributions to probability and statistics: Essays in
honor of harold hotelling (pp. 278–292). Stanford University Press.
Ligges, U., Maechler, M., & Schnackenberg, S. (2026).
scatterplot3d: 3D scatter plot.
Long, J. S., & Ervin, L. H. (2000). Using heteroscedasticity
consistent standard errors in thee linear regression model. The
American Statistician, 54, 217–224.
Matloff, N., & Matloff, N. S. (2011). The art of R
programming: A tour of statistical software design. No Starch
Press.
McGrath, R. E., & Meyer, G. J. (2006). When effect sizes disagree:
The case of r and d. Psychological Methods,
11, 386–401.
McNemar, Q. (1947). Note on the sampling error of the difference between
correlated proportions or percentages. Psychometrika,
12, 153–157.
Meehl, P. H. (1967). Theory testing in psychology and physics: A
methodological paradox. Philosophy of Science, 34,
103–115.
Morales, M., R Development Core Team, with code developed by the, R-help
listserv community, with general advice from the, & Duncan Murdoch.,
especially. (2020). Sciplot: Scientific graphing functions for
factorial designs.
Morey, R. D., & Davis-Stober, C. P. (2026). On the poor statistical
properties of the p-curve meta-analytic procedure.
Journal of the
American Statistical Association,
121(553), 741–753.
https://doi.org/10.1080/01621459.2025.2544397
Morey, R. D., Hoekstra, R., Rouder, J. N., Lee, M. D., &
Wagenmakers, E.-J. (2016). The fallacy of placing confidence in
confidence intervals.
Psychonomic Bulletin & Review,
23(1), 103–123.
https://doi.org/10.3758/s13423-015-0947-8
Morey, R. D., & Rouder, J. N. (2015).
BayesFactor: Computation
of bayes factors for common designs.
http://CRAN.R-project.org/package=BayesFactor
Morey, R. D., & Rouder, J. N. (2026).
BayesFactor: Computation
of bayes factors for common designs.
https://richarddmorey.github.io/BayesFactor/
Munafò, M. R., Nosek, B. A., Bishop, D. V. M., Button, K. S., Chambers,
C. D., Percie du Sert, N., Simonsohn, U., Wagenmakers, E.-J., Ware, J.
J., & Ioannidis, J. P. A. (2017). A manifesto for reproducible
science. Nature Human Behaviour, 1(1), 0021.
Navarro, D. (2026).
Lsr: Companion to "learning statistics with
r".
https://github.com/djnavarro/lsr
Navarro, D. J. (2019). Between the devil and the deep blue sea: Tensions
between scientific judgement and statistical model selection.
Computational Brain & Behavior, 2, 28–34.
Nosek, B. A., Alter, G., Banks, G. C., Borsboom, D., Bowman, S. D.,
Breckler, S. J., Buck, S., Chambers, C. D., Chin, G., Christensen, G.,
et al. (2015). Promoting an open research culture. Science,
348(6242), 1422–1425.
Open Science Collaboration. (2015). Estimating the reproducibility of
psychological science. Science, 349(6251), aac4716.
Pearson, K. (1900). On the criterion that a given system of deviations
from the probable in the case of a correlated system of variables is
such that it can be reasonably supposed to have arisen from random
sampling. Philosophical Magazine, 50, 157–175.
Pfungst, O. (1911). Clever hans (the horse of mr. Von osten): A
contribution to experimental animal and human psychology (C. L.
Rahn, Tran.). Henry Holt.
R Core Team. (2026a).
Foreign: Read data stored by minitab, s, SAS,
SPSS, stata, systat, weka, dBase, ... https://svn.r-project.org/R-packages/trunk/foreign/
R Core Team. (2026b).
R: A language and environment for statistical
computing. R Foundation for Statistical Computing.
https://doi.org/10.32614/R.manuals
Revelle, W. (2026).
Psych: Procedures for psychological,
psychometric, and personality research.
https://personality-project.org/r/psych/
Rosenthal, R. (1966). Experimenter effects in behavioral
research. Appleton.
Rouder, J. N., Speckman, P. L., Sun, D., Morey, R. D., & Iverson, G.
(2009). Bayesian t-tests for accepting and rejecting the null
hypothesis. Psychonomic Bulletin & Review, 16,
225–237.
Sahai, H., & Ageel, M. I. (2000). The analysis of variance:
Fixed, random and mixed models. Birkhauser.
Schoot, R. van de, Depaoli, S., King, R., Kramer, B., Märtens, K.,
Tadesse, M. G., Vannucci, M., Gelman, A., Veen, D., Willemsen, J., &
Yau, C. (2021). Bayesian statistics and modelling. Nature Reviews
Methods Primers, 1, 1.
Shaffer, J. P. (1995). Multiple hypothesis testing. Annual Review of
Psychology, 46, 561–584.
Shapiro, S. S., & Wilk, M. B. (1965). An analysis of variance test
for normality (complete samples). Biometrika, 52,
591–611.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011).
False-positive psychology: Undisclosed flexibility in data collection
and analysis allows presenting anything as significant.
Psychological Science, 22(11), 1359–1366.
Simonsohn, U., Nelson, L. D., & Simmons, J. P. (2014). P-curve: A
key to the file-drawer.
Journal of Experimental Psychology:
General,
143(2), 534–547.
https://doi.org/10.1037/a0033242
Sokal, R. R., & Rohlf, F. J. (1994). Biometry: The principles
and practice of statistics in biological research (3rd ed.).
Freeman.
Spector, P. (2008). Data manipulation with R.
Springer.
Steegen, S., Tuerlinckx, F., Gelman, A., & Vanpaemel, W. (2016).
Increasing transparency through a multiverse analysis. Perspectives
on Psychological Science, 11(5), 702–712.
Stevens, S. S. (1946). On the theory of scales of measurement.
Science, 103, 677–680.
Stigler, S. M. (1986). The history of statistics. Harvard
University Press.
Stroebe, W., Postmes, T., & Spears, R. (2012). Scientific misconduct
and the myth of self-correction in science. Perspectives on
Psychological Science, 7(6), 670–688.
Student, A. (1908). The probable error of a mean. Biometrika,
6, 1–2.
Teetor, P. (2011). R cookbook. O’Reilly.
Warnes, G. R., Bolker, B., Bonebakker, L., Gentleman, R., Huber, W.,
Liaw, A., Lumley, T., Maechler, M., Magnusson, A., Moeller, S.,
Schwartz, M., Venables, B., & Galili, T. (2025).
Gplots: Various
r programming tools for plotting data.
https://github.com/talgalili/gplots
Warnes, G. R., Gorjanc, G., Magnusson, A., Andronic, L., Rogers, J.,
MacQueen, D., & Korosec, A. (2024).
Gdata: Various r programming
tools for data manipulation.
https://github.com/r-gregmisc/gdata
Wasserstein, R. L., & Lazar, N. A. (2016). The ASA’s statement on
p-values: Context, process, and purpose. The American
Statistician, 70(2), 129–133.
Wasserstein, R. L., Schirm, A. L., & Lazar, N. A. (2019). Moving to
a world beyond "p < 0.05". The American Statistician,
73(sup1), 1–19.
Welch, B. L. (1947). The generalization of
“Student’s” problem when several different
population variances are involved. Biometrika, 34,
28–35.
Welch, B. L. (1951). On the comparison of several mean values: An
alternative approach. Biometrika, 38, 330–336.
White, H. (1980). A heteroskedasticity-consistent covariance matrix
estimator and a direct test for heteroskedasticity.
Econometrika, 48, 817–838.
Whittingham, M. J., Stephens, P. A., Bradbury, R. B., & Freckleton,
R. P. (2006). Why do we still use stepwise modelling in ecology and
behaviour?
Journal of Animal Ecology,
75(5),
1182–1189.
https://doi.org/10.1111/j.1365-2656.2006.01141.x
Wickham, H. (2007). Reshaping data with the reshape package. Journal
of Statistical Software, 21.
Wickham, H. (2014). Tidy data.
Journal of Statistical Software,
59(10), 1–23.
https://doi.org/10.18637/jss.v059.i10
Wickham, H. (2016).
ggplot2: Elegant graphics for data analysis
(2nd ed.). Springer.
https://doi.org/10.1007/978-3-319-24277-4
Wickham, H., Averick, M., Bryan, J., Chang, W., McGowan, L. D.,
François, R., Grolemund, G., Hayes, A., Henry, L., Hester, J., Kuhn, M.,
Pedersen, T. L., Miller, E., Bache, S. M., Müller, K., Ooms, J.,
Robinson, D., Seidel, D. P., Spinu, V., … Yutani, H. (2019). Welcome to
the tidyverse.
Journal of Open Source Software,
4(43),
1686.
https://doi.org/10.21105/joss.01686
Wickham, H., & Grolemund, G. (2017). R for data science: Import,
tidy, transform, visualize, and model data. O’Reilly Media.
Wilkinson, L., Wills, D., Rope, D., Norton, A., & Dubbs, R. (2006).
The grammar of graphics. Springer.
Xie, Y. (2025).
Knitr: A general-purpose package for dynamic report
generation in r.
https://yihui.org/knitr/
Yates, F. (1934). Contingency tables involving small numbers and the
χ2 test.
Supplement to the Journal of the Royal Statistical Society,
1, 217–235.