Detection of Data Poisoning Attacks On NSL-KDD Using Isolation Forest and Random Forest
Authors: Aqsa Ghaffar, Sidra Ghaffar, Sana Tariq, Muhammad Aoun
Abstract
Data poisoning attacks are a serious threat to machine learning models used in cybersecurity. In this paper, we performed three types of data poisoning attacks — Label Flipping, Targeted Poisoning, and Clean Label Boundary Attack — on the NSL-KDD cybersecurity dataset at two poisoning rates of 20% and 50%. We used the Isolation Forest algorithm to detect and remove poisoned samples and then retrained a Random Forest classifier to recover model accuracy. Our results show that at 20% poisoning, Clean Label Boundary Attack caused the biggest accuracy drop of 5.92%, reducing accuracy from 91.87% to 85.95%. At 50% poisoning, Label Flipping caused the most severe degradation with accuracy dropping by 19.00% from 91.33% to 72.33%. Isolation Forest provided meaningful but incomplete recovery across all attacks, performing best against Targeted Poisoning at 50% where accuracy was restored to 91.36%. These findings show that cybersecurity datasets are vulnerable to data poisoning attacks and stronger defense methods are needed to fully protect intrusion detection systems.
