Skip to content
Machine Learning

VGGFace2 Explained: The Dataset That Shaped Modern Face Recognition

VGGFace2 is a face recognition dataset from Oxford's Visual Geometry Group, released in 2018, with 3.31 million images of 9,131 people, about 362 per person. It was built for variety in pose, age, lighting and ethnicity. Ollie's model was not trained on VGGFace2; it was trained on MS1MV2.

Liam BradleyJune 18, 2026 · 4 min read
Archive binders full of photo slides and negatives
Image: Silvia Florindi, CC BY 4.0, via Wikimedia Commons (resized)

What Is VGGFace2?

VGGFace2 is a large-scale face recognition dataset developed by the Visual Geometry Group at the University of Oxford, released in 2018. It contains 3.31 million images of 9,131 subjects (celebrities and public figures), with an average of 362.6 images per subject. Subjects were selected to span a wide range of ages, ethnicities, and professions.

What distinguishes VGGFace2 from earlier datasets is its emphasis on variation. Each subject is represented across wide variation in pose (0°–65° yaw), age (across decades for many subjects), ethnicity, and image quality conditions. This variation is what makes models trained on it robust to real-world photo conditions.

How It Compares to Earlier Datasets

The original VGGFace dataset (2015) contained 2.6 million images of 2,622 subjects, impressive at the time but limited in identity diversity. The MS-Celeb dataset (2016) contained 10 million images but had widely reported label noise and demographic imbalance. VGGFace2 addressed these issues with more careful collection methodology and deliberate demographic balancing.

LFW (Labeled Faces in the Wild), the classic benchmark, contains only 13,000 images of 5,749 subjects, far too small for training but useful for evaluation. VGGFace2 occupies the training-focused end of the data spectrum: large enough to train powerful representations, diverse enough to generalise broadly.

Frequently Asked Questions

What is VGGFace2?

VGGFace2 is a large-scale face recognition dataset containing 3.31 million images of 9,131 subjects, developed by Oxford's Visual Geometry Group. It emphasises variation in pose, age, and ethnicity for robust model training.

Was Ollie trained on VGGFace2?

No. Ollie's network was trained from scratch on MS1MV2, about 5.8 million photos of 85,742 people. VGGFace2 is covered here because it shaped how modern face datasets are built.

Try it yourself

Find your celebrity lookalike

Upload a photo and see which celebrities you look most like. Free to try, and your photo is never stored.

Find my celebrity look alike

Related Articles