LFW: The Classic Standard
Labeled Faces in the Wild (LFW) was introduced in 2007 and became the standard benchmark for a decade. It contains 13,233 face images of 5,749 people sourced from the internet, paired for verification testing (same person / different person). On this benchmark, human-level performance was approximately 97.5%. Modern deep learning systems achieve 99.8%+,essentially solving the benchmark.
LFW's limitations are widely acknowledged: it has relatively well-lit, roughly frontal photos; it is skewed toward white Western males; and it is too easy for current systems. Its scores no longer differentiate between high-performing methods.
IARPA Janus: Harder Benchmarks
IARPA's Janus benchmarks (IJB-A, IJB-B, IJB-C) were designed to be substantially harder than LFW. IJB-C contains 11,779 subjects including both still images and video frames, with deliberate inclusion of difficult pose, illumination, and expression conditions. NIST's FRVT (Face Recognition Vendor Test) is the most comprehensive independent evaluation, testing commercial and research systems on very large datasets including millions of photos.
Performance on IJB-C at the standard verification threshold (False Match Rate = 0.01%) benchmarks typically around 95–97% for top systems, significantly harder than LFW. NIST FRVT remains the gold standard for evaluating real-world face recognition system performance.
