FNR: A Similarity and Transformer-Based Approach to Detect Multi-Modal Fake News in Social Media
Social Network Analysis and Mining, 2023Cited by 37
arXiv preprint, 13 June 2026
The rapid advancement of generative AI models is leading to more realistic deepfake media, encompassing the manipulation of audio, video, or both. Numerous studies have yielded promising intra-domain results; however, these models frequently exhibit decreased efficacy when faced with data from dissimilar domains. We propose EAV-DFD, a generalised deep ensemble audio-visual model combined with a domain adaptation mechanism using a teacher-student framework, to improve performance across unseen domains. Using FakeAVCeleb as the primary domain and DFDC, Deepfake_TIMIT and PolyGlotFake as unseen domains, the framework improves AUC by 4.09%, 17.94% and 0.5% respectively while training the student model on only a small portion of the target data, and can interpret which modality has been manipulated.
Deepfake detectors that score well on the dataset they were trained on tend to collapse on media from anywhere else. EAV-DFD combines an ensemble over audio, visual and cross-modal signals with a teacher-student domain adaptation step, so a small sample from a new domain is enough to recover much of the lost accuracy.
The ensemble structure has a second benefit: the model can report which modality carries the manipulation, not just that something is fake.
@misc{abolhasani2026teacher,
title = {{Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection}},
author = {Elham Abolhasani and Maryam Ramezani and Hamid R. Rabiee},
year = {2026},
eprint = {2606.15117},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.15117}
}Social Network Analysis and Mining, 2023Cited by 37
arXiv preprint, 2025Preprint
IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining, 2019Cited by 2