Watford
102,246 residents
2.37 estimates per 100
Meta estimate: 2,419
DEBIAS data story · Published in Royal Society Open Science
Aggregate active-account estimates—not identified people. This comparison can show differences in aggregate coverage. It cannot identify missing people; the estimates are not verified unique-resident counts or individual inclusion probabilities.
Meta active-account example · ≈102,000 residents each · illustrative extreme from 331 areas
102,246 residents
2.37 estimates per 100
Meta estimate: 2,419
different
local rates
102,001 residents
15.89 estimates per 100
Meta estimate: 16,204
Meta average nighttime active-account estimates · 2021 Census · England and Wales
2 / 4 · Counts
Read from the start ↑Larger places generally produce larger counts. Pearson r = .91 across 331 local authorities.
2 · Counts
Across all 331 areas, census population and Meta’s active-account estimate have a Pearson correlation of r = .91. Larger places generally generate more records. This is where many checks stop.
Correlation shows that the counts move together. It does not show that every place is observed at the same rate.
3 · Rates
Per 100 census residents, all 331 local rates spread around a fitted Meta rate of 8.09. The middle 90% spans 4.61–12.31 estimates per 100 residents—a 2.7× range.
The pair is an extreme. The full distribution shows that local variation is ordinary, not exceptional.
4 · Map
The map shows each area’s local rate relative to the fitted Meta rate. Teal means fewer estimated accounts per resident than fitted; coral means more.
More accounts do not mean better representation. The pattern is local—not a simple regional divide.
Act II · Different sources, different patterns
The Meta map is one lens. Across four datasets, the same area can land on opposite sides of its source’s fitted population rate—even though every source has a strong count correlation with census population.
Each source has its own fitted baseline. The comparison diagnoses relative observed-identifier coverage; it is not a representativeness score or a percentage of residents included. The four measures and observation dates are not interchangeable.
Twitter/X uses inferred monthly home locations of active accounts and Meta uses average nighttime active-account estimates for March 2021. Multi-app1 uses inferred home locations of qualifying devices from the first week of April; Multi-app2 uses inferred home locations from preprocessed multi-application GPS data for November. The fitted index normalises scale; it does not make the sources equivalent.
5 / 7 · Four sources
Act II ↑The pair changes position when the source changes. One times marks each source’s own fitted proportional rate.
5 · Four sources
Watford is below Meta’s fitted rate but above the other three source baselines. North East Derbyshire is above Meta and Multi-app2, but below Twitter/X and Multi-app1.
The places have not changed. The source has.
| Source | Watford | North East Derbyshire |
|---|---|---|
| Twitter/X | 1.09× 9% above | 0.49× 51% below |
| Meta | 0.29× 71% below | 1.96× 96% above |
| Multi-app1 | 1.44× 44% above | 0.83× 17% below |
| Multi-app2 | 1.27× 27% above | 1.21× 21% above |
6 · 331 areas
Using the paper’s fitted proportional baseline, 300 of 331 local authorities sit above their source-specific fitted rate in at least one dataset and below it in another.
The pair is memorable. The England-and-Wales pattern shows it is not an edge case.
7 · Why one fix fails
The local characteristics that help explain coverage bias change from source to source—across demography, socioeconomic conditions, resource accessibility, mobility and geography. These describe places, not users.
Website re-render of the accepted-revision main-model outputs—not a reproduction of the paper's radial figure.
Selected source: Twitter/X
Choose a source. Labels mark area characteristics above 0.5 on that source's relative-importance scale. The cutoff is not a significance test or a measure of how biased a source is.
1 of 4 · Demographic context
Area characteristics—not user attributes. Radius is relative mean absolute SHAP importance within each source: 1 marks that source's highest-scoring model feature and 0 its lowest, not “no effect.” Use the switcher to compare which features rank relatively high within each source. Do not compare polygon area or raw mean absolute SHAP magnitudes across sources; the models are fitted separately. Importance does not show direction, causality, population shares or group-specific inclusion rates.
Swipe horizontally to see raw values and ranks.
| Context | Area characteristic | Raw mean absolute SHAP | Relative importance (0–1) | Source rank |
|---|---|---|---|---|
| Demographic | Population aged 0–9 | 0.00002691794021183855 | 0.140856123551 | 12 |
| Demographic | Population aged 10–19 | 0.0000649645600550803 | 0.339946371288 | 3 |
| Demographic | Population aged 60–69 | 0 | 0 | 20 |
| Demographic | Households with six or more people | 0.00003805668792388726 | 0.199142931961 | 6 |
| Demographic | Female population | 0 | 0 | 20 |
| Demographic | Population aged 30–39 | 0 | 0 | 20 |
| Demographic | Population aged 50–59 | 0.000012332312306408194 | 0.064532489939 | 17 |
| Demographic | Population aged 20–29 | 0.00019110237832216587 | 1 | 1 |
| Demographic | Population aged 40–49 | 0 | 0 | 20 |
| Demographic | Population aged 70 and over | 0 | 0 | 20 |
| Socioeconomic | Routine occupations | 0 | 0 | 20 |
| Socioeconomic | Population with Level 4 qualifications or above | 0.000027273989262934978 | 0.142719256047 | 11 |
| Socioeconomic | Population with no qualifications | 0 | 0 | 20 |
| Socioeconomic | Never worked and long-term unemployed | 0.000009591420828893087 | 0.050189960549 | 18 |
| Socioeconomic | Small employers and own-account workers | 0.00005770954486924102 | 0.301982347765 | 5 |
| Socioeconomic | Full-time students | 0.000030419167037705294 | 0.159177333662 | 9 |
| Socioeconomic | Lower supervisory and technical occupations | 0.00003005039306236026 | 0.157247614217 | 10 |
| Socioeconomic | Semi-routine occupations | 0 | 0 | 20 |
| Socioeconomic | Intermediate occupations | 0.000019170659907499414 | 0.100316176469 | 15 |
| Socioeconomic | Higher managerial, administrative and professional occupations | 0.000034953457907249684 | 0.18290435846 | 7 |
| Socioeconomic | Lower managerial, administrative and professional occupations | 0.000059920262322560554 | 0.313550583978 | 4 |
| Resource accessibility | Households without central heating | 0.000020664578218914 | 0.108133548103 | 13 |
| Resource accessibility | Households without a car or van | 0.00012899871064449975 | 0.675024098481 | 2 |
| Resource accessibility | Households owned | 0.000019936567481754933 | 0.104324015519 | 14 |
| Resource accessibility | Households not deprived in any dimension | 0.000008992498082396038 | 0.047055919248 | 19 |
| Mobility & geography | Population density | 0.00003367524472512583 | 0.176215728034 | 8 |
| Mobility & geography | Rural population | 0 | 0 | 20 |
| Mobility & geography | Population born outside the UK | 0 | 0 | 20 |
| Mobility & geography | Employed residents working mainly at or from home | 0.000012688271766345006 | 0.066395153623 | 16 |
| Mobility & geography | Population resident in the UK for less than two years | 0 | 0 | 20 |
What stands out changes. So does the shape of the relationship.
The radial fingerprints rank relative feature importance. These three accepted empirical examples show how individual area-level relationships behave—and why a straight-line correction can miss the local pattern.
MetaPopulation density
Twitter/XArea share aged 20–29
Multi-app1Area share with Level 4 qualifications
How to read the two model views: the radial fingerprints show relative mean absolute SHAP importance within a source; the dependence plots show the shape and direction of individual associations with predicted coverage bias. Each point in the dependence plots is one local authority, not a person or group. These are feature-level examples, not universal source signatures, and neither view identifies inclusion rates or causal effects.
See the methods, provenance and interpretation boundariesWhat this changes
Different datasets place the same areas on different sides of their fitted rates, and their biases relate to local characteristics in different, often nonlinear ways. A correction validated for one source should not be assumed to transfer to another.
Measure before you infer. Validate before you adjust.
See why different datasets need source-specific diagnosis and validationExplore your area
See its observed Meta rate, its distance from the fitted benchmark and a precise interpretation note.
Find a local authorityFor journalists
Download the paired-area graphic, publication-ready figures and concise reporting notes.
Open the media kitAbout the research
Carmen Cabrera and Francisco Rowe of the University of Liverpool’s Geographic Data Science Lab compared four mobile-app datasets with the 2021 Census across 331 local authority areas in England and Wales. Their published Royal Society Open Science study, Making hidden biases visible in population location data from mobile phones, combines proportional diagnostics and explainable machine learning to examine where aggregate coverage differs and which area characteristics are associated with that variation.
The analysis diagnoses area-level patterns; it cannot identify missing individuals.
Read the methods, limitations and wider findingsResearch metrics
Calculated from the released local-authority data. The correlations describe how counts move with population; they are not representativeness scores.
Article attention
Online attention is kept separate from the study’s research metrics. It is not a measure of research quality, readership or representativeness.
Analysis uses released local-authority counts, Meta’s average nighttime active-account estimates for March 2021, 2021 Census resident populations and official 2021 local-authority boundaries. The public map displays local rate minus fitted rate; this has the opposite sign to the paper’s residual-bias convention.