DEBIAS data story · Published in Royal Society Open Science

Who is missing
from the map?

Rapidly unfolding social challenges such as natural disasters, conflicts and epidemics require timely and granular data to understand population changes and appropriately inform planning and decision-making.

However, existing data platforms, including censuses and surveys, are often slow and expensive. Passively generated data via digital technology, including mobile phone apps, have demonstrated great utility in supporting responses to many urgent social challenges, such as the COVID-19 pandemic, the war in Ukraine and flood-induced disaster management in Pakistan. They can provide temporally and geographically granular population data in near-real time.

Yet, digital-trace data can be biased and provide a distorted representation of local populations. Generally, the correlation between the geographical distributions of resident and active-user populations is assessed, but this does not offer an assessment of bias or representativeness.

High correlation may exist. But local coverage may still vary and local populations remain poorly represented.

Digital-trace counts provide a useful signal but they do not represent the local populations. Comparing active-user counts from Meta illustrates wide variation in coverage for places with similar population counts.

1 / 8 · Pair

Read from the start ↑

Watford

102,246 residents

Meta active-user estimate: 2,419

2.37 accounts per 100

6.7×

different local
population coverage rates

North East
Derbyshire

102,001 residents

Meta active-user estimate: 16,204

15.89 accounts per 100

1 / 8 · Pair

Comparing two places

Places can have similar resident population counts but have widely different coverage levels in digital-trace data. Using data from Meta, we illustrate how population coverage varies across UK local authorities. Comparing Watford and North East Derbyshire illustrate how two places with similar resident population counts can have a 6.7 time difference in population coverage.

Watford and North East Derbyshire have similar resident populations but Meta coverage rates of 2.37 and 15.89 accounts per 100 residents.

2 / 8 · Counts

Population counts still line up.

Across local authorities, census populations and Meta’s active-user population estimates have a Pearson correlation of r = .91. More populated places generally have more accounts. Yet, this high correlation does not quantify limited coverage or representation.

Local variations in coverage rates reveal which populations are under- or over-represented in digital-trace data.

Scatter plot showing a strong positive relationship between census population and Meta active-user estimates; Pearson r equals 0.91.

3 / 8 · Population coverage rates

Putting places on the same scale

Expressing Meta coverage as user accounts per 100 census residents, local variations become clearer. They vary widely around Meta’s fitted rate of 8.09 per 100 residents across local authorities.

The middle 90% of areas spans 4.61 to 12.31 accounts per 100 residents, equivalent to a 2.7x range. Such level of variation is typical in digital-trace data from various sources.

Strip plot showing 331 local Meta coverage rates around a fitted rate of 8.09 estimates per 100 residents.
Dividing by population reveals the spread.

4 / 8 · Map

Mapping population coverage rates

Mapping coverage rates relative to the fitted Meta rate reveals under- (darker teal colours) and over-represented (darker coral colours) places.

Coverage rates vary systematically but differently with demographic, socioeconomic, accessibility, mobility and geographical features of local areas. Analysing these relationships helps more precisely identify which population groups may be under- or overrepresented in digital-trace data.

Map showing local authorities with fewer or more Meta active-user estimates per resident than the fitted rate.
Difference between local and fitted rate · estimates per 100

Different data sources, different patterns

The same place is portrayed differently through different data.

Meta is one data source and offers one representation of local populations. But there are more sources. Each may display a high correlation score but offers its own distorted representation of local populations. We investigate four sources.

Each data source has its own fitted coverage rate. Still, the same area can land on opposite sides of a fitted coverage rate depending on the digital-trace data source used to represent their local populations.

What differs between the four sources?

Twitter/X uses inferred monthly home locations of active accounts and Meta uses active-user population estimates for March 2021. Multi-app1 uses inferred home locations of qualifying devices from the first week of April; Multi-app2 uses inferred home locations from preprocessed multi-application GPS data for November. The fitted index normalises scale; it does not make the sources equivalent.

5 / 8 · Sources

Sources intro ↑
Watford and North East Derbyshire across four digital sources The two local authorities change sides of their source-specific fitted rates when the data source changes. Loading the four-source comparison…

Each source portrays the same place differently.

5 / 8 · Sources

Population coverage for the same area can vary widely across data sources

Focusing on Watford and North East Derbyshire shows large differences in coverage by data source. For Watford, the fitted coverage rate is below for Meta, but above for Twitter/X, Multi-app1 and Multi-app2. For North East Derbyshire, it is above for Meta and Multi-app2, but below for Twitter/X and Multi-app1.

The places have not changed. The source has.

Relative to each source’s fitted population coverage rate
SourceWatfordNorth East Derbyshire
Twitter/X1.09× 9% above0.49× 51% below
Meta0.29× 71% below1.96× 96% above
Multi-app11.44× 44% above0.83× 17% below
Multi-app21.27× 27% above1.21× 21% above

91% of local authorities change position

300 of 331 local authority areas display higher or lower coverage rates than the fitted or expected rate in at least one data source. This indicates that patterns of under- or over-representation across areas are the norm, rather than an exception.

300 of 331 authorities change sides across the four source-specific fitted rates from being over- to under-represented or vice versa.

6 / 8 · Understand the influence of local attributes

Understanding which populations are under- or over-represented.

Coverage rates vary systematically with local population attributes. By analysing this variation, we can identify which population groups are likely under- or over-represented in digital-trace data. We investigate differences across our four data sources: Twitter/X, Meta, Multi-App1 and Multi-App2.

We use a machine-learning based model (XGBoost) to analyse which place-based attributes are the most important predictors of low or high population coverage bias. These attributes identify the population groups which are under- or over-represented in a particular digital-trace dataset.

Selected source: Twitter/X

Choose a digital-trace data source. The radial plots below identify the most important place-based attributes associated with variability in population coverage rates. For example, the share of the population aged 20–29 and share of households without a car are the most important predictors of low population coverage rates.

We find that the associations between local attributes and population coverage rates can flatten, reverse direction or change shape across the range of observed records. Machine learning is employed to capture these non-linear relationships.

A

Demographic context

Site-styled radial fingerprint of relative demographic feature importance for the Twitter/X coverage-bias model.
Ten demographic characteristics of local areas.
B

Socioeconomic context

Site-styled radial fingerprint of relative socioeconomic feature importance for the Twitter/X coverage-bias model.
Eleven socioeconomic characteristics of local areas.
C

Resource accessibility

Site-styled radial fingerprint of relative resource-access feature importance for the Twitter/X coverage-bias model.
Four Census household proxies—not direct measures of internet or device access.
D

Mobility & geography

Site-styled radial fingerprint of relative mobility and geographic feature importance for the Twitter/X coverage-bias model.
Two model domains combined in one display group.

1 of 4 · Demographic context

View the complete feature-importance values

Swipe horizontally to see raw values and ranks.

Twitter/X accepted-revision main-model feature importance
ContextArea characteristicRaw mean absolute SHAPRelative importance (0–1)Source rank
DemographicPopulation aged 0–90.000026917940211838550.14085612355112
DemographicPopulation aged 10–190.00006496456005508030.3399463712883
DemographicPopulation aged 60–690020
DemographicHouseholds with six or more people0.000038056687923887260.1991429319616
DemographicFemale population0020
DemographicPopulation aged 30–390020
DemographicPopulation aged 50–590.0000123323123064081940.06453248993917
DemographicPopulation aged 20–290.0001911023783221658711
DemographicPopulation aged 40–490020
DemographicPopulation aged 70 and over0020
SocioeconomicRoutine occupations0020
SocioeconomicPopulation with Level 4 qualifications or above0.0000272739892629349780.14271925604711
SocioeconomicPopulation with no qualifications0020
SocioeconomicNever worked and long-term unemployed0.0000095914208288930870.05018996054918
SocioeconomicSmall employers and own-account workers0.000057709544869241020.3019823477655
SocioeconomicFull-time students0.0000304191670377052940.1591773336629
SocioeconomicLower supervisory and technical occupations0.000030050393062360260.15724761421710
SocioeconomicSemi-routine occupations0020
SocioeconomicIntermediate occupations0.0000191706599074994140.10031617646915
SocioeconomicHigher managerial, administrative and professional occupations0.0000349534579072496840.182904358467
SocioeconomicLower managerial, administrative and professional occupations0.0000599202623225605540.3135505839784
Resource accessibilityHouseholds without central heating0.0000206645782189140.10813354810313
Resource accessibilityHouseholds without a car or van0.000128998710644499750.6750240984812
Resource accessibilityHouseholds owned0.0000199365674817549330.10432401551914
Resource accessibilityHouseholds not deprived in any dimension0.0000089924980823960380.04705591924819
Mobility & geographyPopulation density0.000033675244725125830.1762157280348
Mobility & geographyRural population0020
Mobility & geographyPopulation born outside the UK0020
Mobility & geographyEmployed residents working mainly at or from home0.0000126882717663450060.06639515362316
Mobility & geographyPopulation resident in the UK for less than two years0020

Local attributes do not relate to coverage in a simple, linear way

The illustrative examples below show the shape of the relationship between key place-based attributes and population coverage rates for different data sources. Radial plots indicate which place-based attributes are important predictors of low coverage rates, but they do not tell us how these attributes influence coverage. For example, the radial plots show that the percentage of people aged 20–29 is an important predictor of low coverage rates in Twitter/X. The middle plot below shows how this relationship unfolds: as the percentage of people aged 20–29 increases, population coverage in Twitter/X data also tends to increase; that is, poor coverage becomes smaller. Here, the y-axis represents the inverse of population coverage rates.

We find that the associations between local attributes and population coverage rates can flatten, reverse direction or change shape across the range of observed records. Machine learning is employed to capture these non-linear relationships.

Curved Mobility & geography

MetaPopulation density

Accepted-figure SHAP crop showing a curved, reversing area-level association between population density and modelled Meta coverage bias.
S-Shape Demographic

Twitter/XArea share aged 20–29

Accepted-figure SHAP crop showing an S-shaped area-level association between local share aged 20 to 29 and modelled Twitter/X coverage bias.
Threshold Socioeconomic

Multi-app1Area share with Level 4 qualifications

Accepted-figure SHAP crop showing a threshold-shaped area-level association between local share with Level 4 qualifications and modelled Multi-app1 coverage bias.
See the methods, interpretation and assumptions

7 / 8 · Conclusions

Digital-trace population estimates are valuable, but their representativeness must be assessed and adjusted.

Across four digital-trace data sources, we find substantial local differences in population coverage. These differences vary across places and data sources, and they are associated with demographic, socioeconomic, resource accessibility, mobility and geographical attributes of local populations. These variations reflect biases resulting from the under- and over-representation of certain population groups in the digital platforms used to collect the data. Assessing the representation of digital-trace data is therefore required to define appropriate ways of adjusting them and using them for reliable population inference.

Assess before you infer.

Read the methods, limitations and wider findings

8 / 8 · Learn more

Learn more

Explore

How does your local authority compare?

See its observed Meta rate, its distance from the fitted benchmark and a precise interpretation note.

Find a local authority

About the research

A framework for assessing population coverage bias

Carmen Cabrera and Francisco Rowe of the University of Liverpool’s Geographic Data Science Lab compared four digital-trace datasets with the 2021 Census across 331 local authority areas in England and Wales. Their article published in Royal Society Open Science, Making hidden biases visible in population location data from mobile phones, examines how population coverage varies and which area-level characteristics are associated with that variation.

The analysis focuses on assessing area-level patterns while protecting the anonymity of individual users of digital technologies.

Read the methods, limitations and wider findings

Study at a glance

Key study figures

4
mobile-app datasets
331
local authority areas
2021
Census benchmark
91%
change position across data sources

Article attention

Follow the conversation around the article.

The Altmetric badge summarises online attention to the article, including mentions in news outlets, blogs and social media.

Read the published article

Sources and notes