Skip to content
Ethics

Face Recognition Accuracy Across Demographics: What the Research Shows

Facial recognition accuracy varies by race, sex and age. NIST's 2019 test of 189 algorithms found many made 10 to 100 times more false matches on West African, East African and East Asian faces than on Eastern European faces. The main cause is unbalanced training data; the best algorithms show much smaller gaps, but none is perfectly even.

Wendy WeiAugust 9, 2026 · 5 min read
A world globe turned to show Africa and Asia
Image: Wendy Cope, CC BY 2.0, via Wikimedia Commons (resized)

The Research Evidence

Systematic evaluation of commercial and academic face recognition systems reveals consistent accuracy disparities across demographic groups. The 2019 NIST Face Recognition Vendor Test (FRVT), the most comprehensive evaluation to date, found that many algorithms produced 10 to 100 times more false positives (wrongly matching two different people) for West African, East African and East Asian faces than for Eastern European faces. False non-match rates (failing to match the same person) also varied between groups.

Gender disparities are also documented. Multiple independent studies have found higher error rates for female faces than male faces in many commercial systems, and compounded disparities for the intersection of gender and ethnicity, particularly for darker-skinned women. These are not findings confined to low-quality systems; they appear in some of the highest-performing models in their respective evaluations.

Why Disparities Occur

The root causes are primarily in training data. Systems trained on imbalanced datasets, which describe the majority of systems until very recently, receive more within-group comparison examples from overrepresented groups. This provides stronger gradient signal for those groups, producing more discriminative, better-calibrated representations. Underrepresented groups receive weaker training signal, producing less precise embeddings that are harder to separate.

The problem compounds at deployment. If most of the celebrity database entries are from particular demographic groups, users from underrepresented groups have fewer closely matched options, reducing the quality of top matches even when the embedding is accurate.

What Ollie Does

Ollie's network was trained on MS1MV2, a large dataset of 85,742 people photographed in many conditions. The celebrity database is built from the most-viewed living people on Wikipedia, so it spans many countries, ages and backgrounds, giving people from all backgrounds a rich set of potential matches.

No system fully eliminates accuracy disparities, the research field continues to develop better approaches. Transparency about the issue and ongoing monitoring are the appropriate responses. If you think your results are consistently poor, try a few different photos in soft, even light; if that doesn't help, the contact address in the privacy policy reaches the person who runs Ollie.

Frequently Asked Questions

Is face recognition equally accurate for all racial groups?

Not in most systems. Research consistently finds higher error rates for darker-skinned faces and women in many commercial and academic face recognition systems, primarily due to training data imbalance.

What dataset was Ollie trained on?

MS1MV2, a research dataset of about 5.8 million photos of 85,742 people. Every person in the LFW benchmark was removed from it before training, so the model's 98.5% LFW score is measured on people it never saw.

Try it yourself

Find your celebrity lookalike

Upload a photo and see which celebrities you look most like. Free to try, and your photo is never stored.

Find my celebrity look alike

Related Articles