Skip to content

Latest commit

 

History

81 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

logo

DOI DOI

Before delving into the primary datasets, it's essential to grasp the significance of cybersecurity and why these datasets play a critical role in safeguarding our digital realm. In our interconnected world, cybersecurity threats pose substantial risks to individuals, enterprises, and governments. With the surge in cybercrimes, ranging from data breaches to cyberattacks, having access to trustworthy and current cybersecurity datasets is paramount. These datasets empower us to detect and thwart potential threats effectively.

Main datasets used in cybersecurity

The cybersecurity field is vast, encompassing a wide range of topics and challenges. The community provides a selection of datasets designed for specific research and analysis within the realm of cybersecurity. Here are some key datasets used in cybersecurity, along with brief descriptions and links (as for my study):

1. DARPA Intrusion Detection Data:

The DARPA Intrusion Detection Evaluation datasets were collected as part of the 1998 and 1999 DARPA intrusion detection evaluations. These datasets contain a variety of network traffic data for evaluating intrusion detection systems.

2. KDD Cup 1999:

The KDD Cup 1999 dataset is one of the earliest datasets used for intrusion detection research. It contains data from the DARPA Intrusion Detection Evaluation experiment.

3. NSL-KDD:

The NSL-KDD dataset is a modified version of the well-known KDD Cup 1999 dataset, addressing issues such as redundancy and balance. The new dataset is reduced to the unique values and balanced representation of the different types of the described attacks.

4. CTU-13:

The CTU-13 dataset is particularly notable for its comprehensive representation of botnet traffic, allowing researchers to analyze and develop detection methods for botnet-related activities. CTU-13 features traffic data related to a total of 7 distinct botnets, featuring a mix of both real-world botnet and user-simulated traffic in a university-like environment.

5. ISCXIDS2012:

The ISCXIDS2012 dataset consists of network traffic data, including both normal network traffic and various types of simulated and real-world cyberattacks.

6. CIC-IDS2017:

The Canadian Institute for Cybersecurity Intrusion Detection Systems (CICIDS2017) dataset contains network traffic data specific to machine learning for intrusion detection system (IDS) research, describing various attack scenarios, such as DoS, DDoS, and port scanning.

7. CSE-CIC-IDS2018:

The colaboorative project between the Communications Security Establishment (CSE) and the Canadian Institute for Cybersecurity (CIC) resulted in a comprehensive dataset, describing various attack scenarios, such as DoS, DDoS, and port scanning, that can be used for machine learning intrusion detection system (IDS) research.

aws s3 sync s3://cse-cic-ids2018/dir/ ./localdir

8. CIDDS-001 & CIDDS-002:

The colaboorative project between the Communications Security Establishment (CSE) and the Canadian Institute for Cybersecurity (CIC) resulted in a comprehensive dataset, describing various attack scenarios, such as DoS, DDoS, and port scanning, that can be used for machine learning intrusion detection system (IDS) research.

9. Kyoto2006+:

The Kyoto (Kyoto 2006+) dataset, also known as the Kyoto University Honeypot Dataset, is a collection of network traffic data specifically designed for cybersecurity research, with a focus on honeypot-based intrusion detection and analysis.

Other datasets to be considered

10. Hornet:

The Hornet datasets consist of a collection of data sets created to explore the potential influence of geographic factors on the occurrence of network attacks. This data was gathered during April and May 2021 from eight identically configured honeypot servers strategically positioned in various regions spanning North America, Europe, and Asia.

11. UNSW-NB15:

The UNSW-NB15 dataset is another network intrusion detection dataset containing diverse network traffic data. It includes normal traffic and a wide range of attacks, making it suitable for evaluating intrusion detection systems.

12. MAWILab:

The MAWILab dataset consists of real-world traffic data collected from the MAWI (Monitoring and Analysis of Internet Wide-Area Network Traffic) project. It is used for anomaly detection and network traffic analysis.

13. Microsoft Malware Classification Challenge (BIG 2015):

This dataset is used for malware classification tasks. It contains a large collection of files, each labeled as benign or malicious, making it suitable for machine learning-based malware detection.

14. AWID (Aegean WiFi Intrusion Dataset):

The AWID project aims to offer robust tools, methodologies, and datasets to help researchers create advanced security solutions for present and future wireless networks. The AWID2 dataset includes a large packet set (F) and a smaller one (R), focusing on WEP-based infrastructure with over 150 different attributes. AWID3 targets WPA2 Enterprise, 802.11w, and Wi-Fi 5, featuring multi-layer and contemporary attacks like Krack and Kr00k.

15. The H23Q dataset:

The H23Q dataset is a extensive, labeled 802.3 corpus features traces of ten different attacks targeting HTTP/2, HTTP/3, and QUIC services, including modern attacks specific to HTTP/3. The dataset, which is 30 GB in size, is accessible in both pcap and CSV formats.

16. Malware Traffic Analysis Knowledge Dataset 2019 (MTA-KDD-19):

The MTA-KDD'19 is a curated dataset designed for training and evaluating machine learning algorithms in malware traffic analysis. It was developed from extensive online network traffic databases, emphasizing relevant features while minimizing size and noise through cleaning and preprocessing. This dataset is versatile, not tailored to any particular application, and can be automatically updated to remain current.

17. The UGR'16 dataset:

This dataset is constructed using real traffic and contemporary attacks, sourced from various NetFlow v9 collectors positioned strategically within the network of a Spanish ISP. It consists of two distinct sets of data, each pre-divided on a weekly basis.

18. The Contagio Mobile Mini-Dump dataset:

The Contagio Mobile Mini-Dump is an extensive collection of mobile malware samples, essential for cybersecurity research and the development of mobile security solutions. It provides a diverse range of real-world malware examples, aiding in understanding and combating mobile threats.

19. VAST Challenge datasets:

The MC2 - Computer Networking Operations, part of the VAST Challenge 2011, presents a scenario involving detailed analysis of intricate computer network operations. Participants are tasked with sifting through extensive data to identify unusual patterns, potential cyber threats, and anomalies in network behavior. This challenge simulates a real-world situation the dataset may be used for developing purposes.

20. NIST Software Assurance Reference Dataset (SARD)s:

The Software Assurance Reference Dataset (SARD) is a growing collection of test programs with documented weaknesses. Test cases vary from small synthetic programs to large applications. The programs are in C, C++, Java, PHP, and C#, and cover over 150 classes of weaknesses.

21. Draper VDISC Dataset - Vulnerability Detection in Source Codes:

The Draper dataset consists of the source code of 1.27 million functions mined from open source software, labelled by static analysis for potential vulnerabilities. This dataset can be used for developing vulnerability prediction models.

22. DEFCON Challenge Datasets:

DEFCON regularly releases cybersecurity-focused challenge datasets, engaging participants with diverse machine learning security problems. Their DEFCON30 event, hosted by the AI Village, exemplified this with a unique competition on Kaggle.

23. CDX 2009:

The CDX 2009 dataset, provided by the Cyber Research Center at West Point, captures network traffic and system logs from the 2009 Cyber Defense Exercise. It's a valuable resource for cybersecurity research, offering real-world data for studying network attacks, defense strategies, and system vulnerabilities. This dataset is especially useful for developing and testing intrusion detection systems and other security tools.

24. The Twente Labeled Data Set For Flow-based Intrusion Detection:

The dataset from the University of Twente focuses on flow-based intrusion detection in high-speed networks (1-10 Gbps). It was built on a honeypot connected directly to the Internet, ensuring data relevance. This dataset is ideal for tuning, training, and evaluating intrusion detection systems.

25. ASNM-CDX-2009 Dataset:

The ASNM-CDX-2009 dataset, part of the Advanced Security Network Metrics & CDX 2009 initiative, features ASNM features extracted from tcpdump captures of both malicious and legitimate TCP communications. It focuses on network services susceptible to buffer overflow attacks, providing a rich dataset for network traffic analysis.

26. ADFA IDS Datasets:

The ADFA IDS Datasets, available through UNSW Research, are tailored for system call-based Host Intrusion Detection Systems (HIDS) evaluation. These datasets encompass both Linux and Windows environments, providing a comprehensive resource for developing and assessing HIDS capabilities.

27. LBNL/ICSI Enterprise Tracing Project Dataset:

This dataset consists of full-header network traffic from a medium-sized location, excluding payload. The dataset underwent extensive anonymization to eliminate any details that could reveal individual IP identities. Nevertheless, the LBNL consist of 100 hours of activity specifying the traces of packet header for identifying malicious traffic.

28. University of Massachusetts Datasets:

The UMass Network Datasets, available on the UMass Trace Repository website, are a comprehensive collection of network-related data, designed for a wide array of research applications. These datasets include detailed network traffic traces from a variety of sources and settings, offering valuable insights for studies in network security, performance analysis, and protocol development.

image

Online resources for dataset creation

VirusTotal - Malware samples Phishtank - Phishing samples Kaggle - Challenges and public datasets Scispace - Find research papers relevant to your study

Download recommendation

To download the files mentioned above, you may access the provided URL directly, or just call the below command.

wget --mirror -np -nH --cut-dirs=1 -P /path/to/save/directory <URL>
# Example
wget --mirror -np -nH --cut-dirs=1 -P Kyoto2006+ -N -r -A '*' https://www.takakura.com/Kyoto_data/new_data201704/
# `-r``: Recursively download files.
# `-A '*'`: Download files with any extension.
# `-N`: Download only  files that are newer.

Please note that the availability and specifics of these datasets may change over time, and it's important to review the dataset documentation and terms of use before using them for research or analysis.

Intrusion Detection System (IDS) Public Datasets Benchmarking

In cybersecurity, the design, development, and implementation of effective Intrusion Detection Systems (IDS) are important for safeguarding IT&C infrastructures from unauthorized access, data breaches, and various forms of malicious activities. The selection of an appropriate ML/DL algorithm plays a essential role in ensuring the security and integrity of protected systems.

But before we can dive in the development of a new-edge algorithm, we shoud have the appropriate data, that needs to be studied and analysed in order to undestant the reality and challenges of our ML problem. In accordance with this paradigm, we chosed to study the early created datasets designed for IDS systems in order to derive leasons learn for feature dataset development.

This experiment aims to comprehensively evaluate the performance of different ML and DL algorithms on a variety of datasets, encompassing a wide range of network traffic scenarios. The datasets used for this analysis include well-known benchmark datasets such as KDD, NSL-KDD, CTU-13, ISCXIDS2012, CIC-IDS2017, CSE-CIC-IDS2018, CIDDS-001/CIDDS-002, and Kyoto 2006+. Each dataset represents a distinct set of challenges and characteristics, making this evaluation both diverse and insightful.

The experiment is divided into three main phases:

  1. Data Acquisition and Preprocessing:
  • In this phase, we acquire the selected datasets from reputable sources, ensuring the integrity and accuracy of the data.
  • Data preprocessing tasks include handling missing values, selecting the most relevant features using feature selection techniques, normalizing the data, and, if necessary, performing feature engineering to enhance the dataset's suitability for machine learning.
  1. Algorithm Evaluation:
  • We evaluate the performance of a range of ML/DL algorithms on each dataset. The chosen algorithms include baseline methods like ZeroRule and OneRule, traditional machine learning approaches like Naive Bayes and Random Forest, as well as some of the most used anomaly detection deep learning algorithms.
  • Cross-validation is applied to ensure the robustness of our results. Performance metrics such as precision, variance, and Mean Absolute Error (MAE) are calculated for each algorithm and dataset.
  1. Results and Insights:
  • The results of this evaluation provide valuable insights into the strengths and weaknesses of different IDS algorithms under various conditions.
  • We analyze the performance of algorithms on both the original datasets and balanced datasets to address the challenge of class imbalance in intrusion detection.
  • Observations and additional details regarding the algorithms' performance are documented, providing a comprehensive overview of their behavior.

By conducting this experiment, we aim to contribute to the understanding of cyber domain dataset generation. The findings will assist in making informed decisions when developing a cybersecurity AI application, by deriving necesary steps and procedures in selecting the appropriate learning data.

The following Jupyter notebooks will provide a detailed walkthrough of the experiments, including code snippets, visualizations, and discussions of the results:

  1. KDD99-BM
  2. NSL-KDD-BM
  3. CTU-13-BM
  4. ISCXIDS2012-BM
  5. CIC-IDS2017-BM
  6. CSE-CIC-IDS2018-BM
  7. CIDDS-001-BM
  8. CIDDS-002-BM
  9. Kyoto2006+-BM

The results are also saved under the pickle files mentioned below:

  1. KDD99-BM
  2. NSL-KDD-BM
  3. CTU-13-BM
  4. ISCXIDS2012-BM
  5. CIC-IDS2017-BM
  6. CSE-CIC-IDS2018-BM
  7. CIDDS-001-BM
  8. CIDDS-002-BM
  9. Kyoto2006+-BM

Spread the Word

We’d love your help to grow CY0P5_ML_Datasets!
If you find this project helpful, please share it with your friends, colleagues, or communities—on social media, newsletters, or blogs. Every mention helps us reach more people who might benefit from it.

You can also support the project directly and Support CY0P5_ML_Datasets

—it means a lot and helps keep everything running!

Thank you for being part of our journey and for helping CY0P5_ML_Datasets thrive 🚀

About

Datasets for cybersecurity

Resources

Stars

18 stars

Watchers

2 watching

Forks

Releases

Packages

Used by

Contributors

Languages