About the research

Beyond the
correlation.

A systematic, privacy-preserving way to measure and explain uneven population coverage in aggregated mobile-phone application data.

The study

A framework for assessing bias when individual characteristics are unavailable.

Mobile-phone application data can provide frequent, geographically detailed signals of population activity. But access to digital technology and participation in particular platforms are unequal. These inequalities can make aggregate digital counts a selective view of the population.

Carmen Cabrera and Francisco Rowe compare four digital-trace datasets with the 2021 Census. Their framework first measures population coverage bias by local authority, then uses explainable machine learning to identify the demographic, socioeconomic and geographic contexts associated with that bias.

Strong geographic proportionality is useful evidence. It is not sufficient evidence of representativeness.

The approach works with aggregated area-level counts. It does not require—and cannot recover—individual identities or individual demographic profiles.

Research metrics

The study at a glance.

4Digital-trace datasetsTwo single-application and two multi-application sources
331Local authoritiesThe analytical sample across England and Wales
2021Census benchmarkArea-level resident population counts
5Groups of contextual featuresDemographic, socioeconomic, accessibility, mobility and geographic

What was compared

The four sources are Twitter/X, Meta, Multi-app1 and Multi-app2. They differ in how signals are generated, combined and filtered, so they should not be treated as interchangeable samples of one target population.

What was modelled

The analysis compares a conventional proportional baseline with a flexible explainable machine-learning framework. It assesses local departures from the baseline and tests how coverage bias relates to contextual features, including nonlinear relationships.

Meta baseline used in the main story

Across 331 local authorities, the Pearson correlation between census population and Meta’s active-account estimate is 0.913. The fitted through-origin proportional model corresponds to 8.095 estimates per 100 residents. The middle 90% of observed local rates is 4.609 to 12.308.

Article attention

Follow the conversation around the published article.

Read the published article

Online attention is kept separate from the research metrics above. It describes discussion around the published article; it is not a measure of research quality, citation impact, readership or representativeness.

The wider conclusion

Do not assume one correction will transfer.

The opening scrollytelling act uses Meta to show that a strong count correlation can coexist with substantial variation in local coverage rates. Its second act compares the same places across all four sources and then shows the nonlinear area-level relationships identified by the accepted models.

The sources are not same-period, same-unit samples. Twitter/X uses inferred monthly home locations of active accounts and Meta uses average nighttime active-account estimates for March 2021; Multi-app1 uses inferred home locations of qualifying devices from the first week of April; Multi-app2 uses inferred home locations from preprocessed multi-application GPS data for November. The fitted index normalises each source’s scale but does not make the numerators or dates interchangeable.

  1. The four datasets have different coverage profiles. The contextual features associated with coverage bias—and the shape of those relationships—vary by source.
  2. Their local observed rates differ. An area can sit above one source’s fitted proportional rate and below another’s; neither position is a representativeness score.
  3. The relationships are often nonlinear. Three illustrative crops from the accepted model figure show a curve with reversal for Meta and population density, an S-shape for Twitter/X and local share aged 20–29, and a threshold for Multi-app1 and local share with Level 4 qualifications. They are feature-level examples, not source-wide signatures.
  4. A universal adjustment should not be assumed to work. Any correction should be tailored to the data-generating process and observed bias structure, then independently validated for that source.

Area-level bias fingerprints

Four display groups span five model domains.

The accepted analysis covers demographic, socioeconomic, resource-accessibility, mobility and geographic features. The story combines mobility and geography in one display group, matching the organisation of the accepted model results.

Explore the four-source radial atlas and nonlinear examples in Act II

How the website radial atlas relates to the accepted analysis

The atlas derives new website profiles from four pinned released R1 main-model feature-importance files (random holdout, no lagged covariates; fb_tts for Meta). It does not reproduce the accepted radial figure, whose archived build uses a different Twitter/Meta lineage. The paper is the source for published figures; the pinned R1 files are the source for this atlas.

Radius reports relative mean absolute SHAP importance within each source. It does not show direction, causality, population shares or group-specific inclusion rates. The features describe local authorities, not the people using each dataset.

Coverage is a quantity. Representativeness is a relationship to the target population. More of the first does not guarantee the second.

The fitted-rate index and the SHAP panels answer related but different questions. The first compares local identifiers per resident with each source’s fitted proportional rate; the second models how the paper’s coverage-bias outcome is associated with area context.

Every point in the nonlinear panels is a local authority, not a person or population group. The panels cannot identify which individuals are included or missing, do not establish causes, and should be compared by shape rather than magnitude.

See the cross-source comparison, radial atlas and SHAP examples in the main story. Model diagnostics are available in the paper and reproducibility materials; the site's exact rankings come from the pinned R1 files.

Interpretation

What the analysis can—and cannot—say.

It can show

  • how aggregate digital counts compare with census populations across areas;
  • where local coverage differs from a fitted proportional rate;
  • whether coverage bias varies systematically with area-level context; and
  • which features are most useful for explaining those differences.

It cannot show

  • which identifiable individuals are absent or duplicated;
  • an individual person’s probability of appearing in the data;
  • causal effects of contextual features on digital participation; or
  • that a place with more active-account estimates is inherently better represented.

The public-facing map reports the local rate minus the fitted rate, so positive values mean “more than fitted” and negative values mean “fewer than fitted.” This intuitive display has the opposite sign to the paper’s residual-bias convention; magnitudes and substantive conclusions are unchanged.

Paper and resources

Open analysis materials.

Article: Cabrera C, Rowe F. 2026. Making hidden biases visible in population location data from mobile phones. Royal Society Open Science 13: 251703. DOI: 10.1098/rsos.251703.