When John List sets out to provide an A-to-Z compendium on designing and using experiments, the answer is almost surely that you should want to read at least some of it. His new Experimental Economics textbook is 20 chapters that cover many of the issues that come up in doing both lab and field experiments.
What is covered and what did I like?
The book is divided into 5 parts: 1) the basics of experimental methods; 2) designing experiments; 3) dealing with complications like spillovers, non-random attrition, incomplete compliance, and compromised randomizations; 4) building scientific knowledge, which covers replication and scaling; and 5) the ethical and practical sides of economic experiments.
Several nice themes and approaches are used throughout the book:
· He covers both lab experiments as well as field experiments, which leads to coverage of issues like within-subject designs that are more common in lab experiments, as well as some examples where theory rather than another treatment provides the counterfactual. He argues that “stealth” within-subject designs have a lot to offer field experimentalists and are underused. They are used frequently by tech companies for example, where you may be exposed to many different marketing messages, platform design tweaks, etc. One reason for underuse in development is that many of our interventions take a lot of time to run, and then a long time to see outcomes materialize – but they can really boost power when treatment and outcome can all happen quickly.
· Each chapter uses “running examples” taken from different lab and field experiments in the literature that he uses to discuss and provide specific examples of what the chapter is about. Our readers will be pleased with over a dozen of these coming from development work.
· From the start, he makes clear to think about designing not just for internal validity, but also for two aspects of external validity: what he refers to as both horizontal scaling (will what worked here work elsewhere) and vertical scaling (will something that worked in one school in LA work if scaled up to 1,000 schools?). This builds up to a whole chapter (chapter 16) devoted to generalizability and scaling, where he provides a framework for thinking through these issues.
· The book covers the standard core material well, but also more recent advances such as the use of causal forests for treatment effect heterogeneity (chapter 7), using statistical surrogates for long-term effects (chapter 9), and issues like false positives on underpowered studies. There is nice discussion of moderators (chapter 7) and mediators (chapter 8), and an in-depth chapter on attrition.
My favorite part of the book, and the one I would most recommend to students, is part 5, which covers ethics (chapter 17), pre-treatment administrative responsibilities like IRBs, trial registries, PAPs, and data use agreements (DUAs) (chapter 18), and how to write up a paper (chapter 20).
On ethics and IRBs, he provides a nuanced discussion of a couple of cases that have received discussion in development: Bertrand et al. (2007) on driver’s licenses in India, and Coville et al. (2020) on interventions to get households to pay for water in Kenya. A few specific points I noted when reading these chapters were that perceptions of how ethical participants think the study is can affect who selects into participating in a study (p 572); that we should think of ethics in a dynamic sense, and that failing to learn today about what programs work and at what benefit-cost may not be fair to future generations (p.586), and a couple of words of advice to IRBs – that despite their rules to the contrary “we consider money to be a benefit to subjects” (p592) and that while requiring elaborate consent forms may protect the IRB’s research institution, it may run counter to the purpose of protecting the subjects (p. 646).
His advice on how to write up an experimental paper in Chapter 20 is likely to be very useful for many doing their first few experiments. He discusses his BEC approach, which is not a New York bagel order, but “backdrop, evidence, conclusions”, which he gives as a framework at the paragraph, section, and paper level. He discusses how to start writing a paper (from the inside-out, starting with the standardized design and methods sections), when to release a paper to the outside world, and provides a 21-item checklist for reporting experimental results.
A few specific notes I made for myself of pages I’m likely to refer back to:
P130 He gives an example from one of his own early experiments on bidding behavior for sports cards to illustrate a case where the standard deviation in the treatment group was twice that of the control, so that he should have allocated relatively more sample to the treatment group – it is very rare that we assume the variances will be that different, but he argues previous work would have suggested that here.
P194 Discussion of what to do with matched pair designs when you have attrition – he suggests a reweighting approach.
P225-227 – discussion of the issues around separating extensive and intensive margin effects of treatment, and a bounding approach under a monotonicity assumption.
P365 – notes that he does not advise using whether attrition rates are similar in treatment and control as a viable attrition test – because differential attrition can still maintain internal validity, and equal attrition rates may coincide with its violation.
P424-426 – discusses ways to identify the characteristics of compliers when you have imperfect compliance, and how to bound the ATE
What’s missing, and where do I disagree a little
When someone has written a massive book, it always seems churlish to point out what is left out, especially without saying what you would cut from the existing big book to find space for these other topics (ok, I’ll nominate parts of chapter 19 on optimal use of incentives in the lab as less important, as well as a really long list of further reading that could be put online), but here’s what I’d like to have seen a bit more on:
1) The practical messiness of doing experiments, especially with small samples
For a development economist doing field experiments and reading this book, there is a lot to learn, but also a feeling that the book is missing a lot of the messiness and uncertainties of working with governments and policy experiments. We also don’t get behind-the-scenes details from List’s work with companies like Lyft and Walmart to hear how they thought through some of the trade-offs and issues. For example: (i) while power calculations are discussed, there is a lot less emphasis on approaches to maximizing power from small samples than I would cover in a development class, and rules like the inverse-square rule of take-up are not discussed; (ii) how to choose which and how many outcomes to look at, and how studies may be powered for less interesting but closer in the causal chain outcomes; (iii) when does it make sense to use machine-learning approaches with the sample sizes found in many policy experiments, and warnings about the sensitivity to putting in too many variables; (iv) how to think about registrations, PAPs and the like when it is unclear what the intervention is an what outcomes you’ll be able to measure until after it has been partly underway, etc.
Rachel Glennester and Kudzai Takavarasha's book Running Randomized Evaluations has a lot more of the practical details of working with policymakers and is a nice complement here. There are four online “how to” guides to come out in mid-July, which should cover some more of the practical details.
2) More granular discussion of some of the tricky issues researchers face with IRBs and DUAs.
E.g. What amount of payment is ethical to pay people to take part in your study before it becomes coercive? How should we go about doing experiments designed to enforce the law that will be harmful to those breaking it? If a company requests the right not to be named in a DUA, how do you ensure you can describe enough about the company and intervention in the research still, without them complaining that they are indirectly named, etc.
3) What is not covered that would have been good to cover.
It would have been nice to see discussion of adaptive experiments, and of Bayesian experiment methods which incorporate informative priors. The discussion on multiple hypothesis testing only covers FWER, and not FDRs, nor cases where you may not want to adjust at all. Issues with doing experiments on tech platforms with transaction level data and many different potential stopping points would also be something where I’d love to hear more from the Walmart/Lyft experiences.
4) Where I differ/disagree a little:
a. Chapter 9 has a discussion of when to use the difference-in-means (post) estimator versus the difference-in-differences estimator – but typically Ancova will be better than either. The book doesn’t discuss Ancova or the advantages of using it.
b. Lexicographic preferences for internal validity. Following standard discussion of internal and external validity, he notes that a “lack of internal validity is literally a showstopper” and is a pre-requisite for external validity. I’m very sympathetic to this, and indeed it is the goal of experiments to get clean internal validity. In Chapter 19 he discusses from the viewpoint of where to put your budget how much to go for improving internal validity over generalizability. But I’m also sympathetic to the view that sometimes there may be a bias-variance trade-off, and if treatment effects are heterogeneous, then the MSE of a biased estimate from the relevant population may sometimes be less than from extrapolating an internal valid estimate from another population. The big problem is that typically it is really hard to know how much bias is in non-experimental estimates. I have argued before that not everything needs a power calculation or an experiment - but many of the problems we address do and this book is an excellent guide to then help do so.
Join the Conversation