About the research
Beyond the
correlation.
A systematic, privacy-preserving way to measure and explain uneven population coverage in aggregated mobile-phone application data.
The study
A framework for assessing bias when individual characteristics are unavailable.
Mobile-phone application data can provide frequent, geographically detailed signals of population activity. But access to digital technology and participation in particular platforms are unequal. These inequalities can make aggregate digital counts a selective view of the population.
Carmen Cabrera and Francisco Rowe compare four digital-trace datasets with the 2021 Census. Their framework first measures population coverage bias by local authority, then uses explainable machine learning to identify the demographic, socioeconomic and geographic contexts associated with that bias.
Strong geographic proportionality is useful evidence. It is not sufficient evidence of representativeness.
The approach works with aggregated area-level counts. It does not require—and cannot recover—individual identities or individual demographic profiles.
Research metrics
The study at a glance.
What was compared
The four sources are Twitter/X, Meta, Multi-app1 and Multi-app2. They differ in how signals are generated, combined and filtered, so they should not be treated as interchangeable samples of one target population.
What was modelled
The analysis compares a conventional proportional baseline with a flexible explainable machine-learning framework. It assesses local departures from the baseline and tests how coverage bias relates to contextual features, including nonlinear relationships.
Meta baseline used in the main story
Across 331 local authorities, the Pearson correlation between census population and Meta’s active-account estimate is 0.913. The fitted through-origin proportional model corresponds to 8.095 estimates per 100 residents. The middle 90% of observed local rates is 4.609 to 12.308.
Article attention
Follow the conversation around the published article.
Online attention is kept separate from the research metrics above. It describes discussion around the published article; it is not a measure of research quality, citation impact, readership or representativeness.
The wider conclusion
Do not assume one correction will transfer.
The opening scrollytelling act uses Meta to show that a strong count correlation can coexist with substantial variation in local coverage rates. Its second act compares the same places across all four sources and then shows the nonlinear area-level relationships identified by the accepted models.
The sources are not same-period, same-unit samples. Twitter/X uses inferred monthly home locations of active accounts and Meta uses average nighttime active-account estimates for March 2021; Multi-app1 uses inferred home locations of qualifying devices from the first week of April; Multi-app2 uses inferred home locations from preprocessed multi-application GPS data for November. The fitted index normalises each source’s scale but does not make the numerators or dates interchangeable.
- The four datasets have different coverage profiles. The contextual features associated with coverage bias—and the shape of those relationships—vary by source.
- Their local observed rates differ. An area can sit above one source’s fitted proportional rate and below another’s; neither position is a representativeness score.
- The relationships are often nonlinear. Three illustrative crops from the accepted model figure show a curve with reversal for Meta and population density, an S-shape for Twitter/X and local share aged 20–29, and a threshold for Multi-app1 and local share with Level 4 qualifications. They are feature-level examples, not source-wide signatures.
- A universal adjustment should not be assumed to work. Any correction should be tailored to the data-generating process and observed bias structure, then independently validated for that source.
Area-level bias fingerprints
Four display groups span five model domains.
The accepted analysis covers demographic, socioeconomic, resource-accessibility, mobility and geographic features. The story combines mobility and geography in one display group, matching the organisation of the accepted model results.
Explore the four-source radial atlas and nonlinear examples in Act II
How the website radial atlas relates to the accepted analysis
The atlas derives new website profiles from four pinned released R1 main-model feature-importance files (random holdout, no lagged covariates; fb_tts for Meta). It does not reproduce the accepted radial figure, whose archived build uses a different Twitter/Meta lineage. The paper is the source for published figures; the pinned R1 files are the source for this atlas.
Radius reports relative mean absolute SHAP importance within each source. It does not show direction, causality, population shares or group-specific inclusion rates. The features describe local authorities, not the people using each dataset.
Coverage is a quantity. Representativeness is a relationship to the target population. More of the first does not guarantee the second.
The fitted-rate index and the SHAP panels answer related but different questions. The first compares local identifiers per resident with each source’s fitted proportional rate; the second models how the paper’s coverage-bias outcome is associated with area context.
Every point in the nonlinear panels is a local authority, not a person or population group. The panels cannot identify which individuals are included or missing, do not establish causes, and should be compared by shape rather than magnitude.
See the cross-source comparison, radial atlas and SHAP examples in the main story. Model diagnostics are available in the paper and reproducibility materials; the site's exact rankings come from the pinned R1 files.
Interpretation
What the analysis can—and cannot—say.
It can show
- how aggregate digital counts compare with census populations across areas;
- where local coverage differs from a fitted proportional rate;
- whether coverage bias varies systematically with area-level context; and
- which features are most useful for explaining those differences.
It cannot show
- which identifiable individuals are absent or duplicated;
- an individual person’s probability of appearing in the data;
- causal effects of contextual features on digital participation; or
- that a place with more active-account estimates is inherently better represented.
The public-facing map reports the local rate minus the fitted rate, so positive values mean “more than fitted” and negative values mean “fewer than fitted.” This intuitive display has the opposite sign to the paper’s residual-bias convention; magnitudes and substantive conclusions are unchanged.
Paper and resources
Open analysis materials.
Article: Cabrera C, Rowe F. 2026. Making hidden biases visible in population location data from mobile phones. Royal Society Open Science 13: 251703. DOI: 10.1098/rsos.251703.