Evaluating Publicly Accessible Deepfake Detection Systems: Establishing Practical Reliability Benchmarks for Real World Image-Based Deepfakes
Main Article Content
Abstract
Deepfake technology has rapidly advanced in recent years, raising significant concerns about misinformation, identity manipulation, and the misuse of synthetic media. Although many deepfake detection systems reported in academic literature demonstrate strong benchmark accuracy, relatively little research focuses on the practical reliability and accessibility of publicly usable detection tools for ordinary users. This study addresses that gap through a practical evaluation of publicly accessible deepfake detection systems under benchmark and real-world conditions.
A preliminary accessibility survey of publicly referenced deepfake detection platforms was first conducted. Several systems were found to be restricted, unstable, enterprise-focused, video-only, or unavailable for consistent public experimentation. As a result, Hive Moderation was selected as the primary detector for evaluation. The methodology included evaluating benchmark image samples from FaceForensics++ and Celeb-DF, comparing detector performance with human participants, and testing real-world deepfake images collected from online sources.
The results demonstrated a noticeable performance decline on more realistic, real-world deepfakes. Human participants and the evaluated detector exhibited varying performance across benchmark datasets, while both showed reduced accuracy on Celeb-DF and real-world samples. In addition, several incorrect classifications occurred with extremely high confidence scores, suggesting inconsistency in prediction certainty and practical reliability.
Article Details
Section
COPYRIGHT
Submission of a manuscript implies: that the work described has not been published before, that it is not under consideration for publication elsewhere; that if and when the manuscript is accepted for publication, the authors agree to automatic transfer of the copyright to the publisher.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work
- The journal allows the author(s) to retain publishing rights without restrictions.
- The journal allows the author(s) to hold the copyright without restrictions.
References
[1] T. Wang, X. Liao, K. P. Chow, X. Lin, Y. Wang. Deepfake detection: A comprehensive survey from the reliability perspective. ACM Computing Surveys. Vol. 57, no. 3, pg. 1–35, 2024. https://doi.org/10.1145/3699932.
[2] M. Rettinger, B. Beaumont, N.-A. Le-Khac, H.-H. Nguyen-Le. How effective are publicly accessible deepfake detection tools? a comparative evaluation of open-source and free-to-use platforms. arXiv. arXiv:2603.04456, 2026.
[3]Y. Lu, T. Ebrahimi. Assessment framework for deepfake detection in real-world situations. EURASIP Journal on Image and Video Processing. Vol. 2024, no. 1, pg. 6, 2024. https://doi.org/10.1186/s13640-024-00691-0.
[4] N. M. Müller, P. Czempin, F. Dieckmann, A. Froghyar, K. Böttinger. Does audio deepfake detection generalize? arXiv preprint arXiv:2203.16263, 2022.
[5] M. Groh, Z. Epstein, C. Firestone, R. Picard. Deepfake detection by human crowds, machines, and machine-informed crowds. Proceedings of the National Academy of Sciences. Vol. 119, no. 1, pg. e2110013119, 2022. https://doi.org/10.1073/pnas.2110013119.
[6] B. C. Soundarya, H. L. Gururaj. Deepfake detection: critical review of state-of-the-art approaches and future perspectives. Discover Applied Sciences. Vol. 8, 2026, https://doi.org/10.1007/s42452-025-08174-9.
[7] V. N. Tran, S. G. Kwon, S. H. Lee, H. S. Le, K. R. Kwon. Generalization of forgery detection with meta deepfake detection model. IEEE Access. Vol. 11, pg. 535–546, 2022. https://doi.org/10.1109/ACCESS.2022.3233969 .
[8] Y. Lai, H. Wang, J. Yang, X. Kang, B. Li, L. Shen, Z. Yu. GM-DF: Generalized multi-scenario deepfake detection. Proceedings of the 33rd ACM International Conference on Multimedia. pg. 4300–4309, 2025, DOI: 10.1145/3746027.3755386.