Illustration for: Foundation Models Come to Biometrics: What GPT-Scale AI Means for Identity Verification

The arrival of large language models fundamentally changed how people think about AI. Instead of training a specialist model for each task from scratch, a single enormous model trained on vast quantities of data could be adapted to perform a remarkable range of tasks with minimal additional training. This paradigm shift is now arriving in biometrics — and it has significant implications for identity verification and fraud detection.

What is a foundation model?

A foundation model is a large neural network trained on a very large and diverse dataset, producing rich general-purpose representations that can be adapted to downstream tasks. In the vision domain, models such as DINOv2 (Meta) and OpenCLIP have been trained on hundreds of millions of images. The key property that makes foundation models attractive for biometric applications is generalisation: a foundation model may generalise more robustly to novel scenarios — including attack types that did not exist when the model was trained.

Foundation models for presentation attack detection

Research from EINSTEIN consortium partners at NTNU and Hochschule Darmstadt has explored using DINOv2 and OpenCLIP as the backbone for iris presentation attack detection. The results are striking: fine-tuning these foundation models with a small neural network head achieves state-of-the-art performance, outperforming purpose-built deep learning approaches — and doing so with substantially less task-specific training data.

Vision Language Models for attack detection

An even more radical approach uses Vision Language Models (VLMs) for biometric attack detection through in-context learning — providing the model with a small number of labelled examples in its context window and asking it to classify new inputs based on that context alone, without any fine-tuning. Research from EINSTEIN partners has demonstrated that this framework achieves competitive performance for both physical presentation attacks and digital morphing attack detection.

Foundation models are not a panacea. Their very generality makes their behaviour harder to characterise and audit — a significant challenge for the explainability requirements of the EU AI Act. But they represent a genuinely new capability for biometric systems: the ability to generalise to attack scenarios that were not anticipated at training time.

© 2026 EINSTEIN Consortium. EINSTEIN is funded by the European Union’s Horizon Europe programme (GA No. 101121280) and by UKRI (IFS 10093453). Views expressed are those of the authors only. www.einstein-horizon.eu