A Hybrid Explainable Artificial Intelligence Framework for Deepfake Attribution and Cross-Platform Disinformation Campaign Tracking
Abstract
The rapid advancement of Generative Artificial Intelligence (GenAI) has accelerated the creation and dissemination of highly realistic deepfakes, posing significant cybersecurity threats through identity fraud, misinformation, political manipulation, social engineering, and coordinated cross-platform disinformation campaigns. Although existing deep learning-based deepfake detection systems achieve high classification accuracy, most are limited to binary detection (real versus fake) and provide little support for source attribution, explainability, cross-platform campaign analysis, or digital forensics. These limitations reduce their effectiveness in cyber threat intelligence and forensic investigations. This study aimed to develop a Hybrid Explainable Artificial Intelligence (XAI) Framework for Deepfake Attribution and Cross-Platform Disinformation Campaign Tracking. The objectives were to develop a multimodal attribution framework for images, videos, audio, text, and metadata; integrate Explainable Artificial Intelligence techniques for transparent decision-making; implement graph-based cross-platform disinformation tracking; incorporate cyber threat intelligence and digital provenance; and evaluate the framework against existing approaches. A Design Science Research (DSR) methodology was adopted. The framework was implemented using Vision Transformers (ViT), Temporal Transformers, Wav2Vec 2.0, RoBERTa, Random Forest, XGBoost, and Multilayer Perceptron (MLP) within a weighted ensemble architecture. SHAP, LIME, and Grad-CAM were employed for explainability, while Graph Attention Networks (GATs) and Temporal Graph Neural Networks (TGNNs) enabled campaign tracking. Evaluation was conducted using FaceForensics++, Celeb-DF, DFDC, FakeAVCeleb, ASVspoof, Fakeddit, LIAR, CoAID, Twitter/X, and PHEME datasets. The proposed framework achieved 92.1% accuracy, 91.2% precision, 90.9% recall, 91.0% F1-score, and 95.9% ROC-AUC. It further attained 90.8% attribution accuracy, 98.3% Top-5 attribution precision, 94.2% explanation fidelity, 93.4% community detection accuracy, and an end-to-end inference latency of 63.8 ms, demonstrating near real-time performance. The study concludes that integrating multimodal attribution, Explainable AI, graph analytics, cyber threat intelligence, and digital forensics significantly enhances deepfake attribution and coordinated disinformation tracking beyond conventional detection systems. The proposed framework provides an effective solution for cybersecurity operations, digital forensics, law enforcement, social media monitoring, and national security applications.
Downloads
References
Abbas, F., & Taeihagh, A. (2024). Unmasking deepfakes: A systematic review of deepfake detection and generation techniques using artificial intelligence. Expert Systems with Applications, 252, 124260. https://doi.org/10.1016/j.eswa.2024.124260
Alshahrani, A., Alghamdi, W., Alsubai, S., Aljameel, S. S., Alharbi, A., & Alshamrani, A. (2024). Explainable artificial intelligence: A survey of needs, techniques, applications, and future direction. Neurocomputing, 599, 128111. https://doi.org/10.1016/j.neucom.2024.128111
Altuncu, E., Franqueira, V. N. L., & Li, S. (2024). Deepfake: Definitions, performance metrics and standards, datasets, and a meta-review. Frontiers in Big Data, 7, 1400024. https://doi.org/10.3389/fdata.2024.1400024
Fernández Gambín, Á., Yazidi, A., Vasilakos, A., Haugerud, H., & Djenouri, Y. (2024). Deepfakes: Current and future trends. Artificial Intelligence Review, 57(3), 64. https://doi.org/10.1007/s10462-023-10679-x
Juefei-Xu, F., Wang, R., Huang, Y., Guo, Q., Ma, L., & Liu, Y. (2021). Countering malicious DeepFakes: Survey, battleground, and horizon (arXiv Preprint No. arXiv:2103.00218). arXiv. https://doi.org/10.48550/arXiv.2103.00218
Khan, A. A., Laghari, A. A., Inam, S. A., Ullah, S., Shahzad, M., & Syed, D. (2025). A survey on multimedia-enabled deepfake detection: State-of-the-art tools and techniques, emerging trends, current challenges & limitations, and future directions. Discover Computing, 28, Article 48. https://doi.org/10.1007/s10791-025-09550-0
Khan, S. A., & Dang-Nguyen, D.-T. (2024). Deepfake detection: Analyzing model generalization across architectures, datasets, and pre-training paradigms. IEEE Access, 12, 1880–1908. https://doi.org/10.1109/ACCESS.2023.3348450
Khanjani, B., Watson, S., & Shah, M. (2023). A survey of deepfake detection using deep learning. Neurocomputing.
Khoo, B., Phan, R. C.-W., & Lim, C.-H. (2022). Deepfake attribution: On the source identification of artificially generated images. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 12(3), e1438. https://doi.org/10.1002/widm.1438
Li, H., Chen, J., Wang, B., García, S., & Camacho, D. (2026). Explainable Artificial Intelligence for deepfake detection: Pipeline, open source and comparisons. Expert Systems, 43(3), e70222. https://doi.org/10.1111/exsy.70222
Maheshwari, R. U., & Paulchamy, B. (2024). Securing online integrity: A hybrid approach to deepfake detection and removal using Explainable AI and Adversarial Robustness Training. Automatika, 65(4), 1517–1532. https://doi.org/10.1080/00051144.2024.2400640
Masood, M., Nawaz, M., Malik, K. M., Javed, A., & Irtaza, A. (2023). Deepfakes generation and detection: State-of-the-art, open challenges, countermeasures, and way forward. Applied Intelligence, 53(4), 3974–4026. https://doi.org/10.1007/s10489-021-02977-2
Mirsky, Y., & Lee, W. (2021). The creation and detection of deepfakes. ACM Computing Surveys, 54(1), Article 7, 1–41. https://doi.org/10.1145/3425780
Momin, M. S., Sufian, A., Barman, D., Leo, M., Distante, C., & Damer, N. (2025). Explainable deepfake detection across different modalities: An overview of methods and challenges. Image and Vision Computing, 163, 105738. https://doi.org/10.1016/j.imavis.2025.105738
Moyo, B. V. C., Tuyikeze, T., Matsebula, F., & Obagbuwa, I. C. (2026). An AI-driven conceptual framework for detecting fake news and deepfake content: A systematic review. Frontiers in Artificial Intelligence, 9, 1737790. https://doi.org/10.3389/frai.2026.1737790
Nguyen, T. T., Nguyen, Q. V. H., Nguyen, D. T., Nguyen, D. T., Huynh-The, T., Nahavandi, S., Pham, Q.-V., & Nguyen, C. M. (2022). Deep learning for deepfakes creation and detection: A survey. Computer Vision and Image Understanding, 223, 103525. https://doi.org/10.1016/j.cviu.2022.103525
Nguyen-Le, H.-H., Tran, V.-T., Nguyen, D.-T., & Le-Khac, N.-A. (2024). Passive deepfake detection across multi-modalities: A comprehensive survey. arXiv. https://doi.org/10.48550/arXiv.2411.17911
Passos, L. A., Jodas, D., Costa, K. A. P., Souza Júnior, L. A., Rodrigues, D., Del Ser, J., Camacho, D., & Papa, J. P. (2024). A review of deep learning-based approaches for deepfake content detection. Expert Systems, 41(8), e13570. https://doi.org/10.1111/exsy.13570
Qureshi, S. M., Saeed, A., Almotiri, S. H., Ahmad, F., & Al Ghamdi, M. A. (2024). Deepfake forensics: A survey of digital forensic methods for multimodal deepfake identification on social media. PeerJ Computer Science, 10, e2037. https://doi.org/10.7717/peerj-cs.2037
Rana, M. S., Nobi, M. N., Murali, B., & Sung, A. H. (2022). Deepfake detection: A systematic literature review. IEEE Access, 10, 25494–25513. https://doi.org/10.1109/ACCESS.2022.3154404
Sandotra, N., & Arora, B. (2024). A comprehensive evaluation of feature-based AI techniques for deepfake detection. Neural Computing and Applications, 36(8), 3859–3887. https://doi.org/10.1007/s00521-023-09288-0
Schwalbe, G., & Finzel, B. (2024). A comprehensive taxonomy for explainable artificial intelligence: A systematic survey of surveys on methods and concepts. Data Mining and Knowledge Discovery, 38(5), 3043–3101. https://doi.org/10.1007/s10618-022-00867-8
Shao, R., Wu, T., Wu, J., Nie, L., & Liu, Z. (2024). Detecting and grounding multi-modal media manipulation and beyond. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(8), 5556–5574. https://doi.org/10.1109/TPAMI.2024.3367749
Shoaib, M. R., Wang, Z., Ahvanooey, M. T., & Zhao, J. (2023). Deepfakes, misinformation, and disinformation in the era of frontier AI, generative AI, and large AI models. In 2023 International Conference on Computer and Applications (ICCA) (pp. 1–7). IEEE. https://doi.org/10.1109/ICCA59364.2023.10401723
Tolosana, R., Vera-Rodríguez, R., Fiérrez, J., Morales, A., & Ortega-Garcia, J. (2020). Deepfakes and beyond: A survey of face manipulation and fake detection. Information Fusion, 64, 131–148. https://doi.org/10.1016/j.inffus.2020.06.014
Tsigos, K., Apostolidis, E., Baxevanakis, S., Papadopoulos, S., & Mezaris, V. (2024). Towards quantitative evaluation of explainable AI methods for deepfake detection. In C. L. Stanciu, L. Cuccovillo, B. Ionescu, G. Kordopatis-Zilos, S. Papadopoulos, A. Popescu, & R. Caldelli (Eds.), Proceedings of the 3rd ACM International Workshop on Multimedia AI against Disinformation (MAD '24) (pp. 37–45). Association for Computing Machinery. https://doi.org/10.1145/3643491.3660292
Verdoliva, L. (2020). Media forensics and DeepFakes: An overview. IEEE Journal of Selected Topics in Signal Processing, 14(5), 910–932. https://doi.org/10.1109/JSTSP.2020.3002101
Wang, T., Liao, X., Chow, K.-P., Lin, X., & Wang, Y. (2024). Deepfake detection: A comprehensive survey from the reliability perspective. ACM Computing Surveys, 57(3), 1–35. https://doi.org/10.1145/3699710
Yang, X., Li, Y., & Lyu, S. (2022). Exposing deep fakes using inconsistent head poses: A reliability perspective. ACM Computing Surveys.
Zhang, B., Zhou, J. P., Shumailov, I., & Papernot, N. (2020). On attribution of deepfakes. arXiv. https://arxiv.org/abs/2008.09194
Zhang, T., Weng, J., Li, S., Wang, Z., & Zhu, Y. (2020). DeepFake generation and detection: A survey. IEEE Access, 8, 108254–108272.
Zhang, X., Karaman, S., & Chang, S.-F. (2020). Detecting and simulating artifacts in GAN fake images for deepfake attribution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops.
The authors and co-authors warrant that the article is their original work, does not infringe any copyright, and has not been published elsewhere. By submitting the article to GPH-International Journal of Computer Science and Engineering (GPH-IJCSE), the authors agree that the journal has the right to retract or remove the article in case of proven ethical misconduct.