DEBIAS data story · Published in Royal Society Open Science

Who is missing
from the map?

The dots line up. The local coverage still varies.

Mobile-app counts can closely track census population while producing very different observed records per resident across places. A high correlation is a useful first check—not proof of representativeness.

Aggregate active-account estimates—not identified people. This comparison can show differences in aggregate coverage. It cannot identify missing people; the estimates are not verified unique-resident counts or individual inclusion probabilities.

Meta active-account example · ≈102,000 residents each · illustrative extreme from 331 areas

Watford

102,246 residents

2.37 estimates per 100

Meta estimate: 2,419

6.7×

different
local rates

North East
Derbyshire

102,001 residents

15.89 estimates per 100

Meta estimate: 16,204

Meta average nighttime active-account estimates · 2021 Census · England and Wales

2 / 4 · Counts

Read from the start ↑
Census population and Meta active-account estimates A scatter plot of 331 local authorities. The relationship is strong, but local rates vary. Loading the evidence…

Larger places generally produce larger counts. Pearson r = .91 across 331 local authorities.

2 · Counts

The counts still line up.

Across all 331 areas, census population and Meta’s active-account estimate have a Pearson correlation of r = .91. Larger places generally generate more records. This is where many checks stop.

Correlation shows that the counts move together. It does not show that every place is observed at the same rate.

Scatter plot showing a strong positive relationship between census population and Meta active-account estimates; Pearson r equals 0.91.
The standard check looks strong.

3 · Rates

Put every place on the same scale.

Per 100 census residents, all 331 local rates spread around a fitted Meta rate of 8.09. The middle 90% spans 4.61–12.31 estimates per 100 residents—a 2.7× range.

The pair is an extreme. The full distribution shows that local variation is ordinary, not exceptional.

Strip plot showing 331 local Meta coverage rates around a fitted rate of 8.09 estimates per 100 residents.
Dividing by population reveals the spread.

4 · Map

Put the differences back in place.

The map shows each area’s local rate relative to the fitted Meta rate. Teal means fewer estimated accounts per resident than fitted; coral means more.

More accounts do not mean better representation. The pattern is local—not a simple regional divide.

Map showing local authorities with fewer or more Meta active-account estimates per resident than the fitted rate.
Local departure from the fitted Meta rate.

Act II · Different sources, different patterns

The same place looks different through different data.

The Meta map is one lens. Across four datasets, the same area can land on opposite sides of its source’s fitted population rate—even though every source has a strong count correlation with census population.

Each source has its own fitted baseline. The comparison diagnoses relative observed-identifier coverage; it is not a representativeness score or a percentage of residents included. The four measures and observation dates are not interchangeable.

What differs between the four sources?

Twitter/X uses inferred monthly home locations of active accounts and Meta uses average nighttime active-account estimates for March 2021. Multi-app1 uses inferred home locations of qualifying devices from the first week of April; Multi-app2 uses inferred home locations from preprocessed multi-application GPS data for November. The fitted index normalises scale; it does not make the sources equivalent.

5 / 7 · Four sources

Act II ↑
Watford and North East Derbyshire across four digital sources The two local authorities change sides of their source-specific fitted rates when the data source changes. Loading the four-source comparison…

The pair changes position when the source changes. One times marks each source’s own fitted proportional rate.

5 · Four sources

The same pair switches sides.

Watford is below Meta’s fitted rate but above the other three source baselines. North East Derbyshire is above Meta and Multi-app2, but below Twitter/X and Multi-app1.

The places have not changed. The source has.

Relative to each source’s fitted proportional rate
SourceWatfordNorth East Derbyshire
Twitter/X1.09× 9% above0.49× 51% below
Meta0.29× 71% below1.96× 96% above
Multi-app11.44× 44% above0.83× 17% below
Multi-app21.27× 27% above1.21× 21% above
Values compare local observed identifiers per resident with a separate fitted rate for each source. They are not percentages of residents represented.

6 · 331 areas

91% change position when the source changes.

Using the paper’s fitted proportional baseline, 300 of 331 local authorities sit above their source-specific fitted rate in at least one dataset and below it in another.

The pair is memorable. The England-and-Wales pattern shows it is not an edge case.

Calculated from the released LAD-level data. The 91% describes cross-source position relative to fitted rates—not the percentage of people represented.

7 · Why one fix fails

The area-level signals change with the source.

The local characteristics that help explain coverage bias change from source to source—across demography, socioeconomic conditions, resource accessibility, mobility and geography. These describe places, not users.

Website re-render of the accepted-revision main-model outputs—not a reproduction of the paper's radial figure.

Selected source: Twitter/X

Choose a source. Labels mark area characteristics above 0.5 on that source's relative-importance scale. The cutoff is not a significance test or a measure of how biased a source is.

A

Demographic context

Site-styled radial fingerprint of relative demographic feature importance for the Twitter/X coverage-bias model.
Ten demographic characteristics of local areas.
B

Socioeconomic context

Site-styled radial fingerprint of relative socioeconomic feature importance for the Twitter/X coverage-bias model.
Eleven socioeconomic characteristics of local areas.
C

Resource accessibility

Site-styled radial fingerprint of relative resource-access feature importance for the Twitter/X coverage-bias model.
Four Census household proxies—not direct measures of internet or device access.
D

Mobility & geography

Site-styled radial fingerprint of relative mobility and geographic feature importance for the Twitter/X coverage-bias model.
Two model domains combined in one display group.

1 of 4 · Demographic context

Area characteristics—not user attributes. Radius is relative mean absolute SHAP importance within each source: 1 marks that source's highest-scoring model feature and 0 its lowest, not “no effect.” Use the switcher to compare which features rank relatively high within each source. Do not compare polygon area or raw mean absolute SHAP magnitudes across sources; the models are fitted separately. Importance does not show direction, causality, population shares or group-specific inclusion rates.

View the complete feature-importance values

Swipe horizontally to see raw values and ranks.

Twitter/X accepted-revision main-model feature importance
ContextArea characteristicRaw mean absolute SHAPRelative importance (0–1)Source rank
DemographicPopulation aged 0–90.000026917940211838550.14085612355112
DemographicPopulation aged 10–190.00006496456005508030.3399463712883
DemographicPopulation aged 60–690020
DemographicHouseholds with six or more people0.000038056687923887260.1991429319616
DemographicFemale population0020
DemographicPopulation aged 30–390020
DemographicPopulation aged 50–590.0000123323123064081940.06453248993917
DemographicPopulation aged 20–290.0001911023783221658711
DemographicPopulation aged 40–490020
DemographicPopulation aged 70 and over0020
SocioeconomicRoutine occupations0020
SocioeconomicPopulation with Level 4 qualifications or above0.0000272739892629349780.14271925604711
SocioeconomicPopulation with no qualifications0020
SocioeconomicNever worked and long-term unemployed0.0000095914208288930870.05018996054918
SocioeconomicSmall employers and own-account workers0.000057709544869241020.3019823477655
SocioeconomicFull-time students0.0000304191670377052940.1591773336629
SocioeconomicLower supervisory and technical occupations0.000030050393062360260.15724761421710
SocioeconomicSemi-routine occupations0020
SocioeconomicIntermediate occupations0.0000191706599074994140.10031617646915
SocioeconomicHigher managerial, administrative and professional occupations0.0000349534579072496840.182904358467
SocioeconomicLower managerial, administrative and professional occupations0.0000599202623225605540.3135505839784
Resource accessibilityHouseholds without central heating0.0000206645782189140.10813354810313
Resource accessibilityHouseholds without a car or van0.000128998710644499750.6750240984812
Resource accessibilityHouseholds owned0.0000199365674817549330.10432401551914
Resource accessibilityHouseholds not deprived in any dimension0.0000089924980823960380.04705591924819
Mobility & geographyPopulation density0.000033675244725125830.1762157280348
Mobility & geographyRural population0020
Mobility & geographyPopulation born outside the UK0020
Mobility & geographyEmployed residents working mainly at or from home0.0000126882717663450060.06639515362316
Mobility & geographyPopulation resident in the UK for less than two years0020

What stands out changes. So does the shape of the relationship.

The relationships bend, reverse and jump.

The radial fingerprints rank relative feature importance. These three accepted empirical examples show how individual area-level relationships behave—and why a straight-line correction can miss the local pattern.

Accepted empirical example CURVED REVERSAL U-like · Geographic context

MetaPopulation density

Accepted-figure SHAP crop showing a curved, reversing area-level association between population density and modelled Meta coverage bias.
The association changes direction across the density range.X standardised · Y SHAP contribution
Accepted empirical example S-SHAPE Demographic context

Twitter/XArea share aged 20–29

Accepted-figure SHAP crop showing an S-shaped area-level association between local share aged 20 to 29 and modelled Twitter/X coverage bias.
Low, middle and high area values follow different phases.X standardised · Y SHAP contribution
Accepted empirical example THRESHOLD Socioeconomic context

Multi-app1Area share with Level 4 qualifications

Accepted-figure SHAP crop showing a threshold-shaped area-level association between local share with Level 4 qualifications and modelled Multi-app1 coverage bias.
The association changes abruptly around a step.X standardised · Y SHAP contribution

How to read the two model views: the radial fingerprints show relative mean absolute SHAP importance within a source; the dependence plots show the shape and direction of individual associations with predicted coverage bias. Each point in the dependence plots is one local authority, not a person or group. These are feature-level examples, not universal source signatures, and neither view identifies inclusion rates or causal effects.

See the methods, provenance and interpretation boundaries

What this changes

Representativeness must be diagnosed source by source.

Different datasets place the same areas on different sides of their fitted rates, and their biases relate to local characteristics in different, often nonlinear ways. A correction validated for one source should not be assumed to transfer to another.

Measure before you infer. Validate before you adjust.

See why different datasets need source-specific diagnosis and validation

Explore your area

How does your local authority compare?

See its observed Meta rate, its distance from the fitted benchmark and a precise interpretation note.

Find a local authority

For journalists

Use the finding without losing the caveat.

Download the paired-area graphic, publication-ready figures and concise reporting notes.

Open the media kit

About the research

A framework for diagnosing aggregate coverage bias.

Carmen Cabrera and Francisco Rowe of the University of Liverpool’s Geographic Data Science Lab compared four mobile-app datasets with the 2021 Census across 331 local authority areas in England and Wales. Their published Royal Society Open Science study, Making hidden biases visible in population location data from mobile phones, combines proportional diagnostics and explainable machine learning to examine where aggregate coverage differs and which area characteristics are associated with that variation.

The analysis diagnoses area-level patterns; it cannot identify missing individuals.

Read the methods, limitations and wider findings

Research metrics

Study at a glance

4
mobile-app datasets
331
local authority areas
2021
Census benchmark
r = .80–.95
raw-count Pearson correlations

Calculated from the released local-authority data. The correlations describe how counts move with population; they are not representativeness scores.

Article attention

Follow the conversation around the article.

Online attention is kept separate from the study’s research metrics. It is not a measure of research quality, readership or representativeness.

Read the published article

Sources and notes

Analysis uses released local-authority counts, Meta’s average nighttime active-account estimates for March 2021, 2021 Census resident populations and official 2021 local-authority boundaries. The public map displays local rate minus fitted rate; this has the opposite sign to the paper’s residual-bias convention.