PhishShield: Machine-learning Based Detection of Phishing Websites
Authors: Afifa Yaseen, Sana Younas, Iram Haider, Sana Tariq
Abstract
According to reports, there were over 1.5 million phishing attacks — fraudulent attempts to gain private information by posing as trustworthy companies — in 2024, costing billions of dollars every year. Traditional rule-based detection systems struggle to adapt to sophisticated phishing techniques like polymorphic URLs and zero-day exploits, while current machine learning (ML) models remain susceptible to adversarial attacks. This paper proposes PhishShield, a novel machine learning framework that utilizes URL feature extraction with Random Forest and ongoing learning for adaptability. This approach surpasses conventional methods in accuracy and resilience while addressing deficiencies in adversarial robustness and real-time detection. PhishShield integrates mechanisms to counter adversarial attacks and enables continuous learning to adjust to new threats, overcoming limitations in real-time detection and adversarial durability while achieving high accuracy. PhishShield was developed with a particular focus on protecting vulnerable groups, such as small businesses and senior citizens, who are disproportionately affected by phishing attacks due to a lack of adequate cybersecurity resources.
