Politics

Who Builds the Index

On rankings, proxies, and the politics of measurement.

Global indices keep producing the same hierarchy. Different teams, different decades, same leaderboard. The rankings claim neutrality, but it is always Nordic countries at the top, Global South countries at the bottom, and the societies that design the measures scoring well on them. So, what if the problem is not the index but who gets to build it?

A four-panel comic titled Who Builds the Index? The Hidden Bias of Global Rankings, captioned: global rankings are not neutral; they favor Western models, ignoring diverse social realities. Panel one, rankings are mirrors, not windows: a suited man studies himself in an ornate full-length mirror; a caption reads they measure resemblance to the North Atlantic model, not actual social effectiveness. Panel two, the universal adapter burns local systems: a blue icon labelled Danish-style formal contracts faces a golden emblem labelled Nigeria’s communal trust (esusu) across a lightning split; a caption reads Nigeria’s esusu is invisible to surveys that only recognize Danish-style formal contracts. Panel three, perception scores carry high-stakes consequences: $74.5 billion erupts in a comic blast; a caption reads subjective elements in sovereign ratings may have cost African countries up to $74.5 billion in capital. Panel four, transition toward measurement pluralism: a woman swings a hammer labelled Afrobarometer and AfCRA tools at a padlock labelled Northern-designed metrics while a man braces beside her; a caption reads Global South institutions like Afrobarometer and AfCRA are building tools to challenge the monopoly of Northern-designed metrics.
Illustration generated with Google Gemini.

In Nigeria, people extend credit, share resources, and honor commitments over long periods without written contracts. The circles that do it have a name in every major Nigerian language: esusu, ajo, isusu, adashe. In Denmark, the standing advice for even a parent-to-child loan is a signed debt instrument, proof for the tax authority that the money is a loan and not a gift. Denmark sits near the top of global trust surveys. Nigeria sits near the bottom. The difference is not the presence of trust. It is what the survey recognizes as trust.

Cross-national rankings of trust, governance, and institutional quality keep telling the same story. Nordic countries at the top. The broader North Atlantic world as the unspoken standard. Much of Africa, the Middle East, Latin America, and South and Southeast Asia near the bottom. The consistency is usually taken as proof that the differences are real. There is another possibility. The indices share assumptions deep enough that they no longer register as assumptions. They select similar proxies, reward similar institutional forms, and treat one particular way of organizing collective life as the default. They do not so much discover a global ranking as redraw the world in the image of the societies that designed the measuring tools. The problem is not measurement itself. It is the quiet license of one region’s standards to travel as universal while realities organized differently are filed as noise.

An indicator is always a simplification. It turns something complex into a number that can be compared across countries and years. Simplification is inevitable. The question is which features are kept, which are discarded, and who decides. Those choices carry normative weight even when presented as methodological hygiene.

The dominant global indicators are produced overwhelmingly by institutions in the North Atlantic world. The Worldwide Governance Indicators were developed by a research team at the World Bank. The Corruption Perceptions Index is run from Berlin. None were adopted by treaty. No government voted to be measured by them. They became “global” through uptake, citation, and institutional power rather than consent. Their authority rests on three pillars: institutional location (a World Bank metric enters lending documents with institutional weight), citation (once a score appears in academic papers it is treated as data), and consequences (when numbers shape aid, investment, and diplomacy, they become hard to ignore). Many of these “global” facts are aggregations of expert opinion and perception surveys. The Corruption Perceptions Index does not measure corruption; it measures how selected experts and executives perceive it. An accurate title exposes the operation in ten seconds: not the Corruption Perceptions Index but what selected experts and executives told a Berlin organization. The data survives the retitle. The jurisdiction does not. The industry has run the operation on itself once. In 2014 the Fund for Peace renamed its Failed States Index the Fragile States Index, its director conceding that the name itself had become the problem. The methodology stayed. The label fell.

Yet the system’s defenders have a point that any serious critique must meet. A sovereign-debt analyst in London or New York must allocate risk across dozens of countries by quarter’s end. That person cannot read fifty ethnographies on reputation-based exchange or customary mediation. The algorithm demands a single integer. In this sense the indices function like a universal power adapter: they force every country’s unique energy grid into one rigid prong shape so that capital can flow at all. The adapter burns out local systems, but without it the global financial current has nowhere to plug in. Lenders know the indicators are crude. Many know they file large cultural differences under the heading of “noise.” Still, the fiction is treated as necessary. It makes the spreadsheet run.

Acknowledging this mechanical reality does not excuse the damage. Efficiency for the lender is purchased at the price of accuracy and autonomy for the measured society. Some standardization may be unavoidable; a global financial system cannot operate on fifty separate grammars of trust. The problem is not that an adapter exists. It is that this particular adapter is treated as the only possible standard, and its design forecloses the question of whether other adapters could be built.

The generalized trust question asks: “Generally speaking, would you say that most people can be trusted, or that you need to be very careful in dealing with people?” Aggregate scores track closely with state capacity. In high-capacity environments, a transaction with a stranger is low-risk because formal institutions stand behind it. Where formal enforcement is weak, people may trust intensely within kinship, religious, or commercial networks, yet answer with caution because they do not trust strangers or state institutions. The trust is real. The instrument does not see it. The survey asks about “most people,” which in many contexts means strangers. Caution around strangers without formal legal recourse is rational. That is not a deficit of trust. It is a deficit of formal enforcement, and a blind spot of the instrument. Denmark’s environment makes trust in strangers cheap; Nigeria’s makes it expensive. The difference is not moral quality. It is institutional infrastructure and what the measurement system is built to register.

Governance rankings repeat the pattern. Their operational definitions privilege formal, rule-based, bureaucratic governance of the kind found in high-capacity liberal states. Rule of law is measured through contract enforcement and judicial independence. Government effectiveness is measured through bureaucratic performance. A society that resolves disputes through customary authority or local mediation may be stable and effective at delivering public goods, yet score poorly on dimensions that recognize only formal state structures. The index does not ask whether disputes are resolved fairly. It asks whether they are resolved through formal courts. The technical problems (non-comparable sources, highly correlated dimensions, thin data for many African countries) are symptoms of a benchmark that is rarely stated: the North Atlantic regulatory state. Because the indices are calibrated to recognize this model, countries that do not resemble it are scored as lacking governance rather than as practicing a different kind. The resulting hierarchy looks like a discovery of the data. In significant part it is an artifact of how the data were constructed.

None of this means measurement is impossible. A better trust measure would ask about specific relationships and would distinguish forms of trust. A better governance index would ask whether disputes are resolved fairly and whether public goods reach people, regardless of the administrative channel. Such measures would still require normative choices, but they would make those choices visible and would register institutional forms the current indices render illegible.

Misdescription is only part of the problem. The rankings also act on the world they claim to measure. The Millennium Challenge Corporation screens candidate countries through a scorecard of policy indicators. One of them, Control of Corruption, is drawn from the Worldwide Governance Indicators and has served as what the agency itself calls a hard hurdle: fail that single number and the country fails the whole scorecard. A perception composite decides in one line whether a country may compete for American development money. Washington’s own development analysts have objected that the indicators were never designed for pass-fail decisions about aid. The gate uses them anyway. The intervention that follows a low score can weaken the capacity it claims to strengthen. Local institutions are sidelined. The next round of rankings records another low score. The indicator helped create the condition it described.

The deeper damage is domestic. Labels of corruption, poor governance, or low trust enter local political discourse. Reformers cite them. Governments deflect with them. Citizens sometimes internalize them as collective shortcoming. This is not always passive. Local activists can strategically weaponize the scores as leverage against unaccountable leaders even while knowing the metrics are flawed. Both dynamics can operate at once. Strategic use does not cancel the quieter damage. A reformer seizes on a low score to push sweeping anti-corruption laws. To satisfy the metric’s definition of good governance, the campaign may inadvertently advocate dismantling highly functional informal kinship and reputation networks that look like “nepotism” to a distant auditor. The society begins teaching itself to a test designed by people who do not understand it. A community with dense mutual accountability starts to see itself as “low-trust” simply because the survey only registers trust mediated by formal legal institutions. The ranking does not merely misdescribe. Over time it can rewrite how a society understands itself.

An indicator is a gauge when the number is read by the scored office’s own principal: a ministry tracking its clinics, a board reading its auditor. It is a verdict when the number travels outward to third parties (lenders, donors, editors, visa desks) while the scored party has no channel back into the definition, the weights, or the result. The statistics can be identical. The circuits run in opposite directions. Every instrument named here runs the second circuit. The position is older than the indices. The colonial state was the first external scorer: it counted populations, certified which authorities were legitimate, and filed the verdicts in a language the scored could not contest. Independence changed the letterhead. The seat of the scorer did not move.

Complete refusal has limited practical reach. Countries are already enrolled through lending, risk assessment, and aid. Contestation (treating the scores as claims rather than facts, demanding transparency about what is measured and against what standard) is more durable. Yet contestation alone is insufficient. If you do not build your own tools, you remain permanently downstream of someone else’s definition of progress.

The goal is not to replace one bias with another. It is to build alternative data institutions open about their own normative commitments. Afrobarometer, the Ibrahim Index of African Governance, and the African Governance Report series represent real attempts. They are imperfect, but they exist. Presenting measurement pluralism as a moral ideal is not enough. Running continent-wide surveys requires tens of millions of dollars, capital that frequently comes from the same Northern institutions that sustain the dominant models. Financial dependency can compromise independence. By Afrobarometer’s own account, its interviews already feed the Worldwide Governance Indicators and other Northern products: African counts travel north as inputs; the verdicts travel south as scores. Even a well-designed Southern index must still penetrate actual risk models on Wall Street. The live test is running now. The African Union has chartered its own credit rating agency, AfCRA, private-sector-led precisely to pre-empt the capture charge, with first ratings on local-currency debt promised for 2026, after a UNDP study estimated that subjective elements in sovereign ratings may have cost African countries up to $74.5 billion. Its obstacle is not method. Regulation binds institutional investors to the three incumbent agencies, so an African rating carries little force until the incumbents’ system recognizes it. Measurement pluralism therefore risks creating a secondary hierarchy rather than genuine equality of measurement. Idealism without a map of the friction is not a solution; it is a postponement.

Credibility does not come from claiming neutrality while enforcing a local standard as universal. It comes from making normative choices visible and contestable. Every index decides what to count, what to exclude, and what to treat as success. The problem is not the existence of multiple standards. The problem is the pretense that one local standard speaks for the world, and the use of institutional power to enforce that pretense.

Comments

Comments are read before they appear. Disagreement is welcome; abuse is not.