Showing posts with label quantitative work. Show all posts
Showing posts with label quantitative work. Show all posts

Tuesday, May 12, 2015

ELF Must Die!

My elf lament:  I have written about this before, but need to revisit because the problem is still with us.  Which problem is that?  ELF?  No, not Santa's helpers but Ethnic Fractionalization Indexes.  Huh?  In the study of civil war and other topics, scholars frequently say "ethnicity may matter, so let's see if it does by tossing in this indicator."  Ethnic fractionalization refers to how diverse a society is by essentially measuring how likely is it for a person to bump into a person from another ethnic group. 

Scholars of ethnic conflict do not argue that more diversity means more conflict, but that the existence of some diversity means that ethnic conflict is possible.  Not so much ethnic conflict in North Korea, for example.  But if we care about ethnic politics, then we still have to think about demography--how ethnicity is structured in the political system.  Rather than focusing on more diversity--that there are many, many ethnic groups, we can and should think about whether they are polarized or concentrated.

Polarization: whether the political system revolves around a few large groups?  We might expect more ethnic conflict in such cases than when there are so many groups that no one dominates.  If ethnic conflict is driven by fears of domination (see Donald Horowitz and not just my take on the Ethnic Security Dilemma), then an ethnically polarized country is more likely to have ethnic conflict.

Concentration: are the groups intermixed or are they concentrated?  While the earliest versions of the ethnic security dilemma argued that intermixing created vulnerability and thus conflict, the statistical findings show the opposite.  Concentrated ethnic groups engage in more violence precisely because they are not deterred by their vulnerability.

Of course, the real answer to all of this is that ethnicity affects politics via institutions (probably not an accident this is my most cited work) so one might want to consider how ethnic groups are represented or excluded or empowered or repressed via electoral laws, federalism, and the like.  I used to be in that business and hope to do some more work in that area in the near future.

Anyhow, I keep seeing manuscripts that I review for journals as well as other stuff that keeps using ELF when the author wants to include ethnicity in their model somehow.  But again, the problem is this: what does diversity mean?  Dropping an indicator of diversity into a model does not "account for" or "take seriously" ethnicity.  It is absolutely the least one can do.  And if one gets results, what do those results mean?  If there is generally no theory of ethnic conflict included in the discussion, just a few lines justifying the need to think about ethnicity, then what would a significant correlation mean?

So, my advice to scholars and students is this: if you want to think about how ethnic politics might affect the dependent variable that you care about, then ... THINK about ethnic politics. 


Tuesday, April 1, 2014

The Status of Stats

I am not known for being a statistics whiz.  I have published quantitative work, but I am seen, rightly so, as more comfortable with qualitative work, comparing apples and oranges.  Still, I had the gumption to offer advice on twitter about data today.  What and why?

GDELT was a new dataset that seemed to promise heaps of utility to those who wanted to study event data--which are counts of particular events of interest and handy for analyzing events over time.  It came under fire recently for a variety of reasons.  I did not use the dataset nor do I work with event data, so I am not in any position to judge the dataset itself.

However, I do have experience of working with data that has been criticized.  The Minorities at Risk Project was an effort originally to assess which ethnic groups might be at risk of violence.  The collection selected the groups that were either mobilized or already facing discrimination.  As a result, it was not so good for questions related to why groups become mobilized or face discrimination since those that are not those things were left out.  For the questions I tended to ask, it was less problematic--which groups at risk tend to get more or less international support (least problematic), which groups at risk were more likely to be secessionist or irredentist (a bit problematic), which institutions are associated with more or less ethnic conflict (more problematic).

Once the dataset was criticized by some big names, pop.  It got harder to publish stuff as reviewers scoffed at any findings emanating from the dataset.  The good news for MAR fans is that this led to an NSF project that funded efforts to address the selection bias problem.  The first piece addressing the new dataset has recently been accepted.  The second piece is in the works, and now I am back in the business of pondering the relationships between institutions and ethnic conflict (the delays are my fault now for being distracted by other projects).

Anyhow, the relevance of my experience is this: GDELT is now tarnished, which means it will be harder for stuff to get published as reviewers will be harder to convince.  The peer review process depends on convincing reviewers of the importance of the question, the soundness of the research design, the quality of the data, the interpretation of the findings, and so on.  Given my experience, I expect that using GDELT will be risky if you want publications in the near term.  Over time, the problems might be fixed or might not be that bad.  But for now, its reputation is lousy.  No, this post is not going to do the dirty work of making its reputation bad.  That much has already been achieved.  I am just making it clear that it does not matter if one believes the data to be spiffy or not, but what lies in the minds of reviewers.

So, user beware.  Here be dragons.