About the research
Beyond assessing correlation
A systematic, privacy-preserving way to measure and explain uneven population coverage and representativeness in aggregated mobile-phone application data.
The study
A framework for assessing bias when individual characteristics are unavailable
Mobile-phone application data can provide frequent, geographically detailed signals of population activity. But access to digital technology and participation in particular platforms are unequal. These inequalities can make aggregate digital counts a selective view of the population.
Carmen Cabrera and Francisco Rowe compare four digital-trace datasets with the 2021 Census. Their framework first measures population coverage bias by local authority, then uses explainable machine learning to identify the demographic, socioeconomic, resource accessibility, mobility and geographic contexts associated with that bias.
Strong geographic proportionality is useful evidence. It is not sufficient evidence of representativeness.
The proposed approach in the research works with aggregated area-level counts. It does not require—and cannot recover—individual identities or individual demographic profiles.
Research metrics
The study at a glance
What was compared
The four sources are Twitter/X, Meta, Multi-app1 and Multi-app2. They differ in how signals are generated, combined and filtered, so they should not be treated as interchangeable samples of one target population.
What was modelled
The analysis compares a conventional proportional baseline with a flexible explainable machine-learning framework. It assesses local departures from the baseline and tests how coverage bias relates to contextual features, including nonlinear relationships.
Meta baseline used in the main story
Across 331 local authorities, the Pearson correlation between census population and Meta’s active-user estimate is 0.913. The fitted through-origin proportional model corresponds to 8.095 estimates per 100 residents. The middle 90% of observed local rates is 4.609 to 12.308.
Article attention
Follow the conversation around the published article
Online attention is kept separate from the research metrics above. It describes discussion around the published article; it is not a measure of research quality, citation impact, readership or representativeness.
The wider conclusion
Do not assume one correction will transfer
The opening scrollytelling uses Meta to show that a strong count correlation can coexist with substantial variation in local coverage rates. It also compares the same places across four digital-trace data sources and then shows the nonlinear area-level relationships identified by the accepted models.
The data used: Twitter/X uses inferred monthly home locations of active accounts and Meta uses active-user population estimates for March 2021; Multi-app1 uses inferred home locations of qualifying devices from the first week of April; Multi-app2 uses inferred home locations from preprocessed multi-application GPS data for November. The fitted index normalises each source’s scale.
- The four datasets have different coverage profiles. The contextual features associated with coverage bias—and the shape of those relationships—vary by source.
- Their local observed rates differ. An area can be under-represented in one dataset, but over-represented in another.
- The relationships are often nonlinear. Three illustrative examples show a curve with reversal for Meta and population density, an S-shape for Twitter/X and local share aged 20–29, and a threshold for Multi-app1 and local share with Level 4 qualifications.
- A universal adjustment should not be assumed to work similarly when used to adjust all digital-trace sources. Any correction should be tailored to the data-generating process and observed bias structure, then independently validated for that source.
Area-level bias fingerprints
Four digital-trace sources are assessed across five key population attributes
The accepted analysis covers demographic, socioeconomic, resource-accessibility, mobility and geographic features.
Explore the four-source radial atlas and nonlinear examples
How the website radial atlas relates to the accepted analysis
The atlas derives new website profiles from four pinned released R1 main-model feature-importance files (random holdout, no lagged covariates; fb_tts for Meta). It does not reproduce the accepted radial figure, whose archived build uses a different Twitter/Meta lineage. The paper is the source for published figures; the pinned R1 files are the source for this atlas.
Radius reports relative mean absolute SHAP importance within each source. It does not show direction, causality, population shares or group-specific inclusion rates. The features describe local authorities, not the people using each dataset.
Coverage represents numbers (i.e. a quantity). Representativeness is a relationship to the target population. More of the first does not guarantee the second.
The reported coverage rate and SHAP visualisations answer related but different questions. The first compares local identifiers per resident with each digital-trace data source’s model coverage rate. The second models how that coverage bias rate is associated with area-level contextual attributes of the local populations.
Individual points in the nonlinear panels represent local authorities. They do not represent individual people or population groups. The panels cannot identify which individuals are included or missing, do not establish causes, and should be compared by shape rather than magnitude.
See the cross-source comparison, radial atlas and SHAP examples in the main story. Model diagnostics are available in the paper and reproducibility materials.
Interpretation
What the analysis can—and cannot—say
It can show
- how aggregate digital counts compare with census populations across areas;
- where local coverage differs from a fitted proportional rate;
- whether coverage bias varies systematically with contextual attributes of local populations; and
- which features are most useful for explaining those differences.
It cannot show
- which identifiable individuals are absent or duplicated;
- an individual person’s probability of appearing in the data;
- causal effects of contextual features on digital participation; or
- that a place with more active-user estimates is inherently better represented.
The public-facing map reports the local rate minus the fitted rate, so positive values mean “more than fitted or over-represented” and negative values mean “fewer than fitted or under-represented.”
Paper and resources
Open analysis materials
Article: Cabrera C, Rowe F. 2026. Making hidden biases visible in population location data from mobile phones. Royal Society Open Science 13: 251703. DOI: 10.1098/rsos.251703.