FaceForensics++: Learning to Detect Manipulated Facial Images
- Document
- 25 January 2019
- Event
- 25 January 2019
- Retrieved
- 16 September 2026
The shot
On 25 January 2019, researchers from the Technical University of Munich, the University Federico II of Naples and the University of Erlangen-Nuremberg posted 'FaceForensics++: Learning to Detect Manipulated Facial Images' to arXiv, revising it through a third version on 26 August 2019. The paper, by Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies and Matthias Nießner, introduces a dataset and automated benchmark built from four facial-manipulation methods: DeepFakes, Face2Face, FaceSwap and NeuralTextures.
What the documents show
The paper states its dataset construction directly: it draws on 1,000 original video sequences — the authors' GitHub repository specifies these come from 977 YouTube source videos, each showing 'a trackable mostly frontal face without occlusions' — and applies the four methods 'at random compression level and size' to produce 'a database of over 1.8 million manipulated images,' described as 'an order of magnitude larger than comparable, publicly available, forgery datasets' at the time. That is the authors' own comparison to prior academic datasets, not a claim about datasets published afterward. The repository documents that a fifth method, FaceShifter, was added in July 2020, and that it separately incorporates the Deep Fake Detection Dataset released by Google and Jigsaw — a later addition, not part of the original paper's construction.
The workflow
The paper proposes its benchmark as a shared evaluation protocol: a hidden test set, hosted on the authors' own university server, lets a detection method be scored on manipulated images its developers did not train on, which it frames as necessary to standardize comparisons across approaches. Obtaining the full dataset requires an access-form request, a gatekeeping step the repository describes as part of distribution rather than an open download. This is distinct from the separately organized Deepfake Detection Challenge dataset released the following year; a citation should specify which of the two a detector was tested against, since the paper does not merge them.
What the tool does not change
Editorially: FaceForensics++ is a research dataset for training and testing detection software, not a live monitoring service, and its scope is limited to facial manipulation — it does not address manipulated audio, full-body synthesis or non-facial scene manipulation. The paper frames the underlying problem as a loss of trust in digital content, not a claim that its benchmark alone resolves detection; a detector scoring well on its hidden test set has been evaluated against the four 2019-era manipulation methods, not against techniques developed since.
- Was a detector's reported accuracy measured on the FaceForensics++ hidden test set, the separately organized Deepfake Detection Challenge data, or another collection entirely?
- Does the manipulation method a detector was tested against — DeepFakes, Face2Face, FaceSwap, NeuralTextures, or the later-added FaceShifter — match the technique actually in question?
- How old is the cited detection result relative to newer generative methods the 2019 dataset's four manipulation types do not include?
FaceForensics++ set an early, well-documented bar for facial-manipulation detection research, but it is a 2019-era instrument with a specific, named scope. A claim built on it should say which manipulation methods and which test set were actually used.
Sources & reading trail
The paper's own statement of authors' institutions, dataset construction, the four manipulation methods, and its scale claim relative to prior datasets.
Source published: 25 January 2019 · Retrieved: 16 September 2026
The authors' own repository specifying the 977 YouTube source videos, the July 2020 addition of FaceShifter, and the separate incorporated Deep Fake Detection Dataset.
Source published: Not established · Retrieved: 16 September 2026
Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.