A Framework for Multilingual Phishing Detection: Integrating Hybrid Feature Engineering and Transformer Models — ICACNC 2025 | TechShield Publications
ICACNC 2025 · Conference Article

A Framework for Multilingual Phishing Detection: Integrating Hybrid Feature Engineering and Transformer Models

Authors: Basit Ali, Ali Hassan, Hassam Mehmood, Farhan Hassan

Abstract

Phishing attacks have become a global cybersecurity threat, extending beyond English-speaking regions. Traditional detection methods relying on blacklists, English-only datasets, or manual features struggle with multilingual phishing. This study presents a hybrid detection framework that combines feature-based machine learning with multilingual Transformer models, such as mBERT, leveraging its semantic understanding. The framework evaluates phishing detection across English, German, and Chinese datasets, integrating mBERT’s contextual analysis with lexical, host-based, and content-level phishing indicators. Unlike prior research focused on English datasets, this work emphasizes cross-lingual adaptability and real-time deployment via a web application. Experimental results demonstrate over 95% accuracy, supporting scalability for browser extensions and enterprise solutions.

Multilingual Phishing Detection Hybrid Models Transformer (BERT) NLP Cybersecurity

Cite This Paper

B. Ali, A. Hassan, H. Mehmood, and F. Hassan, “A Framework for Multilingual Phishing Detection: Integrating Hybrid Feature Engineering and Transformer Models,” Proc. Int. Conf. on AI, Cybersecurity, and Next-Gen Computing (ICACNC 2025), The Islamia University of Bahawalpur, Jun. 2025, doi: 10.67535/tsp.000001.043.