XAI for cybersecurity: State of the Art, Challenges, Open Issues and Future Directions
1. INTRODUCTION
Cybersecurity involves protecting data, programs, networks, and systems from potential cyber attacks. The cost of a data breach is high, including direct impact, investigation, response, loss of revenue, downtime, and reputational brand damage. Cybersecurity strategies include layered protection, IAM, and comprehensive data security platforms. The implementation of these strategies aims to prevent cases of financial extortion from users or reputed organizations, which prevent normal business operations. Cybersecurity frameworks and related best practices enable efficient reduction in cyber vulnerabilities by protecting confidential information without compromising on users’ privacy and customer experience. Cybersecurity risk management helps to understand the various characteristics of security threats and the relevant internal interactions at the individual and organizational level. The ALARP is a concept that can help organizations manage cybersecurity risks.
1.1 How AI help can in Cybersecurity
Machine learning (ML) algorithms can be trained to make decisions similar to human behavior and how they are used to detect security threats and breaches. Automated ML-based security tools have been developed to perform tasks such as clustering, classification, and regression. Cluster analysis groups data based on similarities, classification predicts the class of data, and regression establishes relationships between dependent and independent predictor variables. AI and ML have also been used for pro-active vulnerability management, such as analyzing user interactions to detect anomalous behavior. AI techniques have played a significant role in antivirus detection, with heuristic techniques, data mining, agent techniques, and artificial neural networks being used. Biometric-based authentication systems using physical and behavioral identification and recognition systems, which are difficult for hackers to compromise, have also been developed with the help of AI.
1.2 Limitations of AI that necessitates XAI in Cybersecurity
The use of AI in cybersecurity poses many challenges and introduces counter-indications and secondary risks that can be exploited by malicious actors. These risks include evasion attacks, false positives, false negatives, and vulnerability to interception of communication, failure of services, accidents, disasters, legal issues, attacks, power outages, and physical damages. AI-based systems require heavy computational power, data, and expertise to build and maintain, making them expensive to deploy. However, these systems can be compromised by hackers who mutate malware to become AI immune, enabling them to foil security algorithms and manipulate data without detection. The lack of justifiability and interpretability of decisions made by AI-based systems is also a challenge. To address this, there is a need for statistical AI algorithms and XAI to provide interpretability to the results generated by the AI-based statistical models, enabling researchers and experts to understand causal reasoning and primary data evidence. In healthcare, explainable AI (XAI) allows doctors and healthcare providers to understand how a particular decision is made. In the manufacturing industry, AI-based natural language processing (NLP) helps technicians make better decisions by analyzing unstructured data related to equipment and maintenance standards along with structured data such as work orders and sensor readings.
AI threat landscape
1.3 How XAI can help
The use of AI models has become increasingly popular in complex domains, enabling solutions in mission-critical settings. However, issues such as overfitting and the black-box nature of traditional AI systems have prevented the ability to provide justifiable solutions and make accurate decisions, which can lead to financial losses, physical harm, or even death. Therefore, the implementation of explainable AI (XAI) is necessary to provide transparency and eliminate biases in datasets, ensuring impartial decision-making and meaningful variable inference. Decision Trees (DT) can support explanations on the prediction of results and help achieve insight on security policies in intrusion detection systems (IDS).
Several studies have explored the potential of XAI as the primary control against AI risk in cybersecurity. For example, a study proposed a novel black-box attack that implemented XAI to compromise the privacy and security of relevant classifiers. Another study proposed a methodology for explaining inaccurate classifications generated by data-oriented IDSs, while a third study proposed a deep learning-based intrusion detection framework that achieved transparency at each level of the ML model. XAI has made significant progress in various industry verticals, and since 2018, the number of researchers who have adopted XAI in their research has kept increasing.
XAI methods used in these studies include SHAP, BRCG, and the Protodash method, which enable complete understanding of the model's behavior and identification of differences and similarities between training dataset samples. The implementation of XAI in cybersecurity can help manage over-flooding issues in alarm systems of SIEM and IDS, using techniques such as zero-shot learning.
2. BACKGROUND AND MOTIVATION
2.1 Motivation to Integrate XAI for Cybersecurity
As our society has evolved, being more interconnected, cyberattacks have become even more sophisticated.Therefore, as data breaches grow more widespread, it is critical to have a thorough grasp of modern cyberattacks. Overview of some of the most prominent cyberattacks:
Malware is malicious software that can infect a computer system when an individual clicks on a bad link or opens an email attachment mistakenly. Once malware gains access to a computer system, it can restrict access to key network resources, collect and transmit sensitive information, and potentially disable a large number of system components. The following are some of the popular types of malware: Ransomware, Spyware, Botnet, Fileless Malware
Phishing is a type of internet fraud where fraudulent emails that seem to come from reputable companies are used to steal sensitive information or infect a victim's computer with malicious software. There are various types of phishing attacks, including spear phishing, clone phishing, and whaling. These attacks have become increasingly popular in recent years. There are a variety of phishing attacks: Spear phishing, Whaling, Smishing and Vishing, Pharming.
Other types of cyberattacks : Man-in-the-middle attack, Denial-of-service attack, SQL injection, Zero-day exploit, DNS Tunneling, Evasion attacks, Poisoning attack.,
2.2 XAI Frameworks
The aim of XAI is to make AI models more transparent and understandable to humans, in order to build trust in their decision-making capabilities. XAI frameworks are tools that generate reports on the model's operation and attempt to explain how it works. Some of the most popular XAI frameworks are: SHAP, LIME, ELI5, Skater, DALEX, ALE.
2.3 Motivation to Integrate XAI for Cybersecurity
The use of AI and ML in various industries has increased, but the lack of transparency in these systems can lead to difficulty in defending their outcomes, understanding their logic during security breaches, and making decisions in a black-box environment. XAI, or explainable AI, provides a framework that generates reports and attempts to explain the system's operation, making it beneficial in the field of cybersecurity. XAI can improve threat detection, response capability, and risk discovery and prioritization, and help with incident response coordination and malware threat detection. XAI appears to be a promising solution in circumstances requiring explainability, interpretability and accountability.
3. APPLICATIONS OF XAI IN CYBERSECURITY FOR DIFFERENT INDUSTRY VERTICALS
3.1 XAI for Cybersecurity in Smart Healthcare
The healthcare industry's sensitive data makes it an attractive target for attackers, and AI solutions can provide improved accuracy, low latency, and processing of diverse data. However, AI models are often black-box and do not provide explainable outcomes, which is where XAI can help. XAI can make AI outcomes interpretable and enhance learning from complex sensory data in IoMT, while addressing security and privacy concerns in a white-box way. The involvement of healthcare workers in XAI medical solutions can produce highly accurate predictions.
3.2 XAI for Cybersecurity in Smart Banking
The popularity of smart banking systems is increasing, but so is the risk of cyber attacks, making it important to minimize fraud and scams. AI-based algorithms can provide higher efficiency and security for banking data, but there are still security issues that need to be addressed. XAI can help by providing transparency and interpretability to end-users and regulatory bodies, which can lead to more fruitful decisions and services. XAI can also efficiently address data breaches, trust issues, and fraud detection through its ability for detailed reasoning and complex pattern identification.
4. RESEARCH PROJECTS
Research projects of XAI across various types of cyber attacks and relevant applications:
4.1 FeatureCloud
The FeatureCloud project is funded by the European Union's Horizon 2020 program and aims to implement XAI in medical services. It is an interdisciplinary project that covers machine learning, cybersecurity, federated learning, blockchain, XAI, and medicine. The project uses feature spaces for medical data composed by FeatureCloud, applies federated learning approaches for preserving the privacy of medical data, and offers XAI to interpret the results for medical practitioners and other end-users. The explanation strategy can be chosen based on the user's preference.
4.2. SPATIAL (Security and Privacy Accountable Technology Innovations, Algorithms and machine Learning)
SPATIAL is an EU-funded project aimed at addressing issues faced by blackbox AI and data management in cybersecurity. The project focuses on transparency, explainability, resilience, privacy, societal impact, education, skills building, and transparent AI applications. It seeks to develop appropriate methods and techniques to ensure AI explainability, build a system that can recover from failures in uncontrolled environments, define organizational policies for trustworthy AI solutions, provide awareness about current and future AI, and design XAI for cyber systems that are transparent and understandable between users and service providers. The project is funded with 5 million EUR under the Grant agreement number 101021808 for EU Horizon 2020.
4.3 Blue Hexagon-Hexnet
Blue Hexagon has developed a project called Hexnet, which incorporates XAI into their AI-based cybersecurity framework to enhance the SOC's understanding of malware threats. Hexnet extracts metadata from network objects to classify samples as malware or non-malware, but the conventional method of classification is not enough to fully understand the threat. Hexnet maps the indicators of compromise (IOC) of a malicious sample to the MITRE ATT&CK framework to explain decisions made by the neural net. Hexnet is resistant to evasion techniques and provides real-time detection and analysis for organizations with large SOC teams.
4.4. LEMNA
LEMNA is a project supported by NSF grants which aims to improve the accuracy and robustness of machine learning models for security systems. The project includes methods to mitigate misclassifications caused by contaminated training data, understand human errors in the data, and improve the robustness of the models against adversarial attacks. However, LEMNA is not suitable for classifier models for cyber systems built over anonymized feature sets, as the original feature set needs to be mapped to the anonymized feature set to make LEMNA's output explainable.
Research projects on XAI and their relevance to cybersecurity
5. LESSONS LEARNT, CHALLENGES, OPEN ISSUES AND FUTURE DIRECTIONS
This section discusses the challenges, open issues, and future directions of XAI methodologies in smart cities applications. The five main challenges and their possible alternatives are:
Security attacks on XAI integrated cybersecurity frameworks: XAI methodologies need to protect existing AI-enabled cybersecurity systems from various security attacks, including poisoning attacks, adversarial attacks, and counterfactual gaming attacks. A detailed experimental study of the impacts of these attacks can be done on a variety of cybersecurity structures and data types.
Managing balance between security and usability of XAI integrated cybersecurity systems: Cybersecurity experts need to maintain the trade-off between security and usability aspects of the newly introduced XAI-enabled cybersecurity systems. To maintain the balance between security and usability, there is an immediate need to introduce tools to analyze and debug the model and its data internals.
Legal Compliance: To ensure the safety and security of newly introduced cybersecurity frameworks against attackers, terrorists, and wrongdoers, appropriate rules and regulations need to be formulated to ensure the constructive use of XAI-enabled cybersecurity systems for the safety of the citizens.
Performance of XAI Algorithms: XAI models face several challenges in balancing interpretability and performance aspects, especially when employed for decision-making in IoT systems. To address the issues related to optimization of XAI algorithms, the algorithms can be optimized by considering multiple constraints such as computational resources, model accuracy, and computational power. In addition, accurate decision-making can be targeted by considering data diversity, data rate, and data quality.
Privacy and Data Protection: XAI models need to consider privacy issues explicitly during system life cycle, and data privacy should be protocol-driven and indicative of the users and circumstances under which the data can be accessed. To address the limitations of XAI with post-hoc XAI techniques, model fusion can become a powerful technique. Meanwhile, while designing XAI models, it is essential to ensure that the model is not reversible.
6. CONCLUSION
XAI is a framework which helps to enable understanding and interpretations of predictions with the help of ML models. It provides understandability, transparency and justifiability to the results generated by the traditional AI systems. Cybersecurity is one such area where the application of AI helps to analyse dataset and track wide spectrum of security threats and malicious activities. he presented work provides a state-of-the art review of the application of XAI in cybersecurity. At the outset, a brief introduction to cybersecurity is presented highlighting the different types of attacks and its implications. The traditional risk management systems and relevant standards are discussed in details. Although these preventive measures have given promising results but they have associated challenges in terms of their lags in proactive risk management. This set the stage for the implementation of ML algorithms which help to detect attacks and anomalies with enhanced accuracy resembling human behaviour. The ability of AI-based applications in effectively managing the vulnerability of the network by identification of various threats, is established. However, there still exist associated challenges wherein attacker strategically foil he AI algorithms, making it immune, acting as a vector for various forms of attacks. Additionally, these AI-based systems require sophisticated expertise to develop and maintain, contrarily consuming computational power, data and resources. The results generated are black-boxed wherein understanding and visualization is constrained to input and output having inability to provide traceability between both the parameters. This drawback lays the foundation of XAI which provides opaqueness and explainability to the results leading to accurate and justifiable decision making.


Comentarii
Trimiteți un comentariu