"Missing Evidence: Tracking Academic Data Use around the World"
Executive Summary
This paper introduces a novel method to track data usage in academic research across 216 countries over a period of 20 years, aiming to uncover patterns and gaps in data-driven research. The core findings highlight:
-
Global Data Usage Patterns: High-income countries receive almost 50% of all data-driven research papers, despite accounting for just 15% of the global population. Conversely, low-income countries, comprising 10% of the world population, contribute only about 5% of data-driven research.
-
Correlation with Economic Indicators: Data-driven research is significantly linked to Gross Domestic Product per Capita and population size, accounting for around 75% of the variation across countries.
-
Boosting Data Supply vs. Demand: The paper emphasizes the importance of enhancing statistical capacity, with a particular focus on the availability of geospatial data at the first administrative level, population censuses within the last decade, and multiple labor force surveys over the past 10 years, each contributing to a 1.1%, 0.3%, and 0.4% increase in data use, respectively.
-
Classification of Countries: Countries are categorized into four groups—Deserts, Savannas, Grasslands, and Forests—based on their ability to increase their data supply or demand. Deserts are characterized by low data demand, indicating potential for increasing both data supply and demand.
Methodology and Findings
The study utilizes a large dataset of English-language academic research papers sourced from the Semantic Scholar Open Research Corpus (s2orc), digitizing millions of papers globally and providing access through APIs. By training a natural language model to predict the use of data in research articles, the researchers were able to estimate the amount of data-driven research per country, independent of the location of the researchers involved.
Key findings include:
- Country-specific Data Usage: The model's predictions showed a strong correlation (0.99) with manual coding of articles, validating its accuracy.
- Economic Correlation: The quantity of data-driven research is closely tied to GDP per capita and population size, suggesting economic factors significantly influence research focus.
- Geospatial and Demographic Data Impact: Enhanced availability of geospatial data at the first administrative level increases data usage by 1.1%, population censuses by 0.3%, and multiple labor force surveys by 0.4%.
Policy Implications
The paper advocates for strategies aimed at increasing both the supply and demand of data in countries. For countries identified as 'Deserts'—those with low data demand—the focus should shift towards increasing data accessibility and literacy. Conversely, for countries already rich in data products but receiving limited research attention, efforts should concentrate on enhancing data demand, potentially through initiatives that make existing data more accessible to researchers.
Conclusion
By illuminating the disparities in data-driven research across countries and identifying key drivers of data usage, this study provides a roadmap for policymakers and researchers looking to optimize the impact of data on public policy and societal improvements. The proposed classification framework offers a practical tool for prioritizing interventions based on the unique challenges faced by different nations.