Media brief

The dots line up.
The coverage varies.

A concise guide to the finding, evidence and language to use when reporting the research.

Key finding

A strong population correlation does not establish representativeness.

Digital-trace counts from more populated local authorities tend to be larger. That relationship is important, but the published study shows that local population coverage can still vary systematically—and often nonlinearly—with demographic, socioeconomic and geographic context.

Correlation is a first diagnostic, not a certificate of representativeness.

The larger conclusion is source-specific: the same local authority can fall on opposite sides of different sources’ fitted rates, while coverage bias is associated with different area contexts in nonlinear ways. Any adjustment should be designed and independently validated for that source; one correction should not be assumed to transfer to another.

Key numbers

The evidence in two acts.

  • 331 local authorities in England and Wales.
  • r = .91 Pearson correlation between census population and Meta’s average nighttime active-account estimate.
  • 8.09 per 100 residents fitted proportional Meta rate.
  • 4.61–12.31 per 100 middle 90% of local Meta rates.
  • r = .80–.95 raw-count Pearson correlations across the four released source series.
  • 300 of 331 authorities sit above their source-specific fitted rate in at least one dataset and below it in another.

Illustrative pair: Watford and North East Derbyshire each had about 102,000 census residents, but their Meta estimates corresponded to 2.37 and 15.89 active-account estimates per 100 residents—a 6.7-fold difference.

The published models also show three illustrative nonlinear forms: an S-shape for Twitter/X and local share aged 20–29, a curve with reversal for Meta and population density, and a threshold for Multi-app1 and local share with Level 4 qualifications. These are local-authority-level model associations, not demographic profiles of observed users.

Reporting language

Precise words preserve the finding.

Please use

  • “Meta active-account estimates” rather than “Meta users” or “people”;
  • “local population coverage” or “estimates per 100 census residents”;
  • “above or below that source’s fitted proportional rate” for the cross-source comparison;
  • “area-level model association” for the nonlinear panels;
  • “source-specific measures from different 2021 snapshots” when describing all four datasets together;
  • “associated with contextual features” rather than causal language; and
  • “England and Wales” for the 331-area public story.

Please avoid

  • claiming that the analysis identifies which people are missing;
  • reporting 300 of 331, or 91%, as a share of residents or people represented;
  • treating the four sources as same-period, same-unit samples—their numerators and observation dates differ;
  • describing a high-coverage area as necessarily more representative;
  • inferring that Twitter/X represents young adults, Meta represents urban people or Multi-app1 represents people with qualifications;
  • treating active-account estimates as unique resident counts; or
  • implying that one correction can be applied to every digital source.

Contact and assets

For interviews and graphics.

Authors: Carmen Cabrera and Francisco Rowe, Geographic Data Science Lab, University of Liverpool.

Email: c.cabrera@liverpool.ac.uk · fcorowe@liverpool.ac.uk

Published article: Making hidden biases visible in population location data from mobile phones, Royal Society Open Science 13: 251703.