Published on Development Impact

Measuring What Matters: Some Guidelines for Index Creation - Guest Post by Bruce Wydick

This page in:
Measuring What Matters: Some Guidelines for Index Creation  - Guest Post by Bruce Wydick

For years, I used indices in my development research with a level of confidence that, in retrospect, may not have been entirely earned. I knew running impact estimations on indices increased statistical power and reduced over-testing.  I knew how to create the normal indices used by other development researchers and which indices referees expected to see. But if someone in a seminar had asked why I chose that particular index, I probably would have responded with a long explanation that sounded thoughtful but was mostly a sophisticated way of saying, “That’s what everyone else does.”

In development economics, index construction often seems to be passed down like a family recipe—handed from advisor to advisee, replicated from paper to paper, rarely re-examined from first principles.  We create wealth indices, dwelling indices, empowerment indices, mental health indices, aspiration indices, social capital indices, human flourishing indices, and dozens more.

This is because index creation serves a variety of purposes for development researchers: First, by combining multiple related outcomes into an index, we obtain summary measures of impact that serve as useful punchlines in abstracts.  Because these summary measures aggregate individual outcomes that may show borderline significance in themselves, they typically improve statistical power.  Moreover, testing a single aggregated outcome reduces the need for a multiple-hypothesis-testing correction.  In pre-analysis plans, impact results on aggregated indices are often cases where we can justify using standard p-values over adjusted q-values.

Increasingly, development researchers are also using indices to measure abstract constructs. The rise of behavioral development economics has led to an explosion of research on abstract constructs such as aspirations, agency, hope, trust, grit, and social cohesion, all typically measured through indices. The figure below illustrates the remarkable growth of abstract constructs over the last quarter century used in paper titles published in the top five economics and development economics journals:

Image

Yet surprisingly little attention is paid to the basic question: What exactly makes an index a good index?  Compared to the oceans of ink (digital and real) that we have spilled on causal identification, surprisingly little has been spilled on index construction.  Yet if we measure a construct poorly, even the most beautifully identified causal estimate may be estimating impacts on the wrong thing.  This oversight is costly: Misguided index construction is one of the most underappreciated sources of bias in empirical development research.

In a new methods paper on index construction, I try to address a number of unsettled questions on indices that routinely plague many of us who work in applied development research:

Question #1: What are the different types and categories of indices?

So glad you asked.  One primary index distinction is between construct indices and welfare indices. Construct indices measure a unified construct such as agency, aspirations, or food insecurity. Welfare indices aggregate outcomes into a summary measure of well-being, such as a dwelling quality index or the UNDP's Human Development Index. This distinction reflects disciplinary origins. Psychology evolved as a measurement science because it studies difficult-to-observe constructs such as anxiety and depression. Economics, in contrast, historically dealt with the prices and quantities of widgets (pretty easy to measure as they roll off the conveyor belt) and occasionally on combining observable outcomes into welfare-relevant summaries.

A second distinction is between reflective and formative indices. The difference lies in the direction of causality. In a reflective index, a latent trait causes the observed indicators: an aggression disorder, for example, may be reflected in violent behavior. Unusual levels of kicking and punching therefore represent a possible component of such an index. In a formative index, the indicators themselves create the construct, as in the HDI or a nutrition index. Many construct indices are reflective and many welfare indices are formative, but there are important exceptions. Financial literacy is a construct index typically created formatively, while food insecurity is a reflective index that manifests itself through worry about food, reduced food quality, and missed meals but can also serve as a measure of basic human welfare.

A third distinction is between statistically weighted and conceptually weighted indices. Statistically weighted indices assign weights from the data, as in the Anderson (2008) index, principal components, and factor analysis. Conceptually weighted indices assign weights based on theory or normative judgment, often equally across components, as in the HDI, simple scale indices, and Kling et al. (2007) index, with unequal weights also possible.  When using index construction to increase statistical power or simply as an aggregation tool, explanation of an index can be easier with conceptually weighted indices, particularly if they assign equal weights to variables.  However, interpretation may be easier with certain statistically weighted indices, such as factor analysis, which are designed to more precisely measure the underlying construct.  Here we can legitimately say that “the intervention increased agency by 0.50σ” rather than more truthfully but more awkwardly saying “the intervention increased an equally weighted index of variables intended to measure agency by 0.50σ”.

Question #2: What types of variables should I use in an index?

This depends on what type of index you want to create. The paper presents an arguable hierarchy of data that economists favor in creating formative welfare indices.  Objective outcomes (grade completion, household income, freedom from disease) top this pyramid.  Economists want to measure behavior under constraints.

But what about reflective indices?  While formative indices are created from constrained behavior, psychology likes to measure unconstrained behavior and feelings.  For years psychologists have incorporated reported characteristics and hypothetical responses that free participants from the same behavioral constraints that are essential to recover economic welfare in formative indices. While constraints in traditional development economics inform development outcomes, in psychological measurement they can appear as statistical noise obscuring identification of an underlying construct.

Image

Question #3: Is there a case where “less is more” in index construction?

Yes, less indeed may be more.  The paper points to consideration of parsimonious indices, where it is shown in a proposition that there is a point at which adding marginal variables to a reflective index actually reduces its reliability, its ability to hone in on a particular latent construct.  Added to this is the problem that too many survey questions increase participant fatigue and reduce data quality. (For example, in an experiment to identify the costs of survey fatigue on data quality, Jeong et al. (2023) find that an extra hour of surveying increased the probability of a respondent skipping a question by 10%–64% and reduced reported consumption by 25%.)  The paper introduces two simple ML-based index-pruners, which foster parsimony in formative indices (based on the LASSO) and reflective indices (based on the Elastic Net).

Question #4: OK, so what index should I use?

Now we are getting to the heart of index decisions.  The paper provides an “index decision tree” to create a guide for index choice. 

There are numerous past mistakes that many development economists (including this researcher) have made in index choice, but the paper highlights one particularly common one: the use of Anderson and Kling methods, best suited to formative and hybrid index construction, to construct reflective indices.  Confirmatory factor analysis (CFA), or perhaps principal components (with small n), is the right tool for the reflective index construction job.   The first proposition in the paper demonstrates that creation of a reflective index with a vector of variable weights incongruent to those comprising the latent factor will (nearly) always underestimate the treatment effect on a latent construct. 

To illustrate this, consider measuring child aspirations, where your index contains three variables, aspired educational level, aspired profession, and grit.  Correlation is strong between the first two, but their correlation is weaker with grit.  In an aspirations index, CFA will assign larger index weights to the two strongly correlated variables and a smaller weight on the weakly correlated variable. But by virtue of its GLS properties, the Anderson index will place a relatively bigger weight on the weakly correlated variable because it contains more independent information.  Now it makes sense to assume the aspired education and profession variables are more correlated with the underlying aspirations construct than the relatively weakly correlated grit variable.  This is where CFA shines—in latent construct measurement. But note how by weighting up the grit variable, use of the Anderson index is going to yield a “gritty” aspirations construct measure relative to CFA, which is more likely to reflect latent aspirations.  This creates attenuation bias in aspirations estimates.

Image

Unfortunately, this is not the end of it.  A primary motive for using indices is to increase statistical power.  This is listed as a possible path to the use of an Anderson index in the reflective index branch of decision tree.  But it unlikely to be the optimal path. One is less likely to reject no-effect nulls with individual variables than with an aggregated index. However, statistical power is obtained not just through smaller standard errors, but through a better signal-to-noise ratio.  And even though the Anderson index is good at beating down noise, it loses to CFA in signal clarity.   This sets the stage for an interesting index horserace between strong psychometric signal power and optimal econometric noise reduction. 

Who wins this race?   Using data from three small RCTs on a microenterprise intervention, the index methods paper runs this race.  What it finds is that while Anderson indices consistently display lower standard errors than CFA, its inherent attenuation bias consistently leads it to underestimate the impact of entrepreneurial aspirations relative to CFA.  The result is that better measurement more often wins out over lower standard errors: Using a CFA index for the dependent variable results in more rejected nulls than the Anderson index.  So when running tests on reflective indices, default to CFA over Anderson or Kling indices—you’ll get your measurement cake and likely get to eat your favorite statistical-power cake too.

Proper index construction matters.  And with the proliferation of behavioral research in the development field, it has become more and more important to measure index-based outcomes correctly.  The hope for this methods paper is that it helps foster an ongoing conversation on best practices in index construction.  But one of its key conclusions is that when we measure well, statistical power will often take care of itself.

Bruce Wydick is Professor of Economics at the University of San Francisco, Adjunct Professor at the University of California at Davis, and Distinguished Research Affiliate at the University of Notre Dame. 


Join the Conversation

The content of this field is kept private and will not be shown publicly
Remaining characters: 1000