A Framework for Multilingual Phishing Detection: Integrating Hybrid Feature Engineering and Transformer Models
Authors: Basit Ali, Ali Hassan, Hassam Mehmood, Farhan Hassan
Abstract
Phishing attacks have become a global cybersecurity threat, extending beyond English-speaking regions. Traditional detection methods relying on blacklists, English-only datasets, or manual features struggle with multilingual phishing. This study presents a hybrid detection framework that combines feature-based machine learning with multilingual Transformer models, such as mBERT, leveraging its semantic understanding. The framework evaluates phishing detection across English, German, and Chinese datasets, integrating mBERT’s contextual analysis with lexical, host-based, and content-level phishing indicators. Unlike prior research focused on English datasets, this work emphasizes cross-lingual adaptability and real-time deployment via a web application. Experimental results demonstrate over 95% accuracy, supporting scalability for browser extensions and enterprise solutions.
