Detecting Data Poisoning Attacks on Security Datasets using Machine Learning
Authors: Hassan Raza Khan, Muhammad Awais Khan, Aoun Muhammad, Umar Fayyaz, Sehrish Raza
Abstract
Machine Learning (ML) has been established as an important area within the context of present-day cybersecurity solutions, particularly in the context of intrusions and attack detection systems. The mechanism associated with the algorithms has been proved to be highly dependent on the quality associated with the data on which it operates. The impact associated with the poisoning carried out on the data by the attacker has been established to be highly negative within the context of the mechanism associated with these algorithms. This paper has addressed the problem related to the impact on Machine Learning-based intrusions detectors. A publicly available dataset within the context of the concerned area has been acquired from Kaggle. Simulated poisoning attacks have been done on the acquired dataset to represent the behavior of the attacker. The purpose has been achieved by the use of the Isolation Forest Algorithm in combination with the analysis of label consistency. The analysis has disclosed the fact that the poisoning of the data has adversely affected the performance, while the elimination of anomalous points has improved the performance. This has established the concept that integrity is the most important factor in the realm of ML-enabled cybersecurity systems.
