Skip to content

Latest commit

 

History

202 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Public Data Documentation

A collection of public healthcare datasets

Type of datasets include

  1. Tabular
  2. Images
  3. Free Text
  4. Audio

Table of Contents

  1. Cancer Datasets

  2. Diabetes Dataset

  3. Monkeypox

  4. Parkinson Disease

  5. Brain Related Datasets

  6. Eye Related Datasets

  7. Gastrointestinal Related Datasets

  8. Heart Related Datasets

  9. Kidney Related Datasets

  10. Lung Related Datasets

  11. Medical Transcripts

  12. Others


Cancer Datasets

Breast Cancer Data set

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Building a predictive machine learning model for breast cancer

  • Size: 19 KB

  • Example of data + Special Instructions(if any):

    Download breast-cancer.data stored in Data Folder from Source

    Example: CSV version

    Screenshot 2022-07-14 at 4 33 48 PM

    Columns present (from left to right):

    • Class: no-recurrence-events, recurrence-events
    • age: 10-19, 20-29, 30-39, 40-49, 50-59, 60-69, 70-79, 80-89, 90-99.
    • menopause: lt40, ge40, premeno.
    • tumor-size: 0-4, 5-9, 10-14, 15-19, 20-24, 25-29, 30-34, 35-39, 40-44, 45-49, 50-54, 55-59.
    • inv-nodes: 0-2, 3-5, 6-8, 9-11, 12-14, 15-17, 18-20, 21-23, 24-26, 27-29, 30-32, 33-35, 36-39.
    • node-caps: yes, no.
    • deg-malig: 1, 2, 3.
    • breast: left, right.
    • breast-quad: left-up, left-low, right-up, right-low, central.
    • irradiat: yes, no.

  • Note: Dataset from Source is of .data file. If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad).

    To download as .csv file

    Right click on breast-cancer.data as shown :

    Screenshot 2022-07-14 at 4 44 53 PM

    Proceed to add the .csv extention in file name and save to desired location

    Screenshot 2022-07-14 at 4 46 13 PM

  • Citations

    - This breast cancer domain was obtained from the University Medical Centre, Institute of Oncology, Ljubljana, Yugoslavia. Thanks go to M. Zwitter and M. Soklic for providing the data. Please include this citation if you plan to use this database.
    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Breast Cancer Wisconsin (Original) Data Set

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Building a predictive machine learning model for breast cancer

  • Size: 20 KB

  • Example of data + Special Instructions(if any):

    Download breast-cancer-wiscoinsin.data stored in Data Folder from Source

    Example: CSV version

    Screenshot 2022-07-14 at 6 11 27 PM

    Columns present (from left to right):

    • Sample code number: id number
    • Clump Thickness: 1 - 10
    • Uniformity of Cell Size: 1 - 10
    • Uniformity of Cell Shape: 1 - 10
    • Marginal Adhesion: 1 - 10
    • Single Epithelial Cell Size: 1 - 10
    • Bare Nuclei: 1 - 10
    • Bland Chromatin: 1 - 10
    • Normal Nucleoli: 1 - 10
    • Mitoses: 1 - 10
    • Class: (2 for benign, 4 for malignant)

  • Note: Dataset from Source is of .data file. If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad).

    To download as .csv file

    Right click on breast-cancer-wiscoinsin.data as shown :

    Screenshot 2022-07-14 at 5 58 13 PM

    Proceed to add the .csv extention in file name and save to desired location

    Screenshot 2022-07-14 at 5 58 32 PM

  • Citations

    - This breast cancer databases was obtained from the University of Wisconsin Hospitals, Madison from Dr. William H. Wolberg. If you publish results when using this database, then please include this information in your acknowledgements. Also, please cite one or more of:
    
      1. O. L. Mangasarian and W. H. Wolberg: "Cancer diagnosis via linear programming", SIAM News, Volume 23, Number 5, September 1990, pp 1 & 18.
          
      2. William H. Wolberg and O.L. Mangasarian: "Multisurface method of pattern separation for medical diagnosis applied to breast cytology", Proceedings of the National Academy of Sciences, U.S.A., Volume 87, December 1990, pp 9193-9196.
          
      3. O. L. Mangasarian, R. Setiono, and W.H. Wolberg: "Pattern recognition via linear programming: Theory and application to medical diagnosis", in: "Large-scale numerical optimization", Thomas F. Coleman and Yuying Li, editors, SIAM Publications, Philadelphia 1990, pp 22-30.
          
      4. K. P. Bennett & O. L. Mangasarian: "Robust linear programming discrimination of two linearly inseparable sets", Optimization Methods and Software 1, 1992, 23-34 (Gordon & Breach Science Publishers).
     - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Breast Cancer Wisconsin (Prognostic) Data Set

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Building a predictive machine learning model for breast cancer

  • Size: 44 KB

  • Example of data + Special Instructions(if any):

    Download wpbc.data stored in Data Folder from Source

    Example (No headers)

    119513 N 31 18.02 27.6 117.5 1013 0.09489 0.1036 0.1086 0.07055 0.1865 0.06333 0.6249 1.89 3.972 71.55 0.004433 0.01421 0.03233 0.009854 0.01694 0.003495 21.63 37.08 139.7 1436 0.1195 0.1926 0.314 0.117 0.2677 0.08113 5 5
    8423 N 61 17.99 10.38 122.8 1001 0.1184 0.2776 0.3001 0.1471 0.2419 0.07871 1.095 0.9053 8.589 153.4 0.006399 0.04904 0.05373 0.01587 0.03003 0.006193 25.38 17.33 184.6 2019 0.1622 0.6656 0.7119 0.2654 0.4601 0.1189 3 2

    Columns present (from left to right):

    • ID number

    • Outcome (R = recur, N = nonrecur)

    • Time (recurrence time if field 2 = R, disease-free time if field 2 = N)

    • Ten real-valued features are computed for each cell nucleus (Column 4 to Column 33):

      • radius (mean of distances from center to points on the perimeter)
      • texture (standard deviation of gray-scale values)
      • perimeter
      • area
      • smoothness (local variation in radius lengths)
      • compactness (perimeter^2 / area - 1.0)
      • concavity (severity of concave portions of the contour)
      • concave points (number of concave portions of the contour)
      • symmetry
      • fractal dimension ("coastline approximation" - 1)

      In other words, column 4 to column 13 represent the above 10 values for the first cell nucleus and so on.

  • Note: Dataset from Source is of .data file. If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad).

    To download as .csv file

    Right click on wpbc.data as shown :

    Screenshot 2022-07-14 at 9 07 11 PM

    Proceed to add the .csv extention in file name and save to desired location

    Screenshot 2022-07-14 at 9 08 03 PM

  • Citations

    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Breast Cancer Wisconsin (Diagnostic) Data Set

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Building a predictive machine learning model for breast cancer

  • Size: 124 KB

  • Example of data + Special Instructions(if any):

    Download wdbc.data stored in Data Folder from Source

    Example(No headers)

    842302 M 17.99 10.38 122.8 1001 0.1184 0.2776 0.3001 0.1471 0.2419 0.07871 1.095 0.9053 8.589 153.4 0.006399 0.04904 0.05373 0.01587 0.03003 0.006193 25.38 17.33 184.6 2019 0.1622 0.6656 0.7119 0.2654 0.4601 0.1189
    842517 M 20.57 17.77 132.9 1326 0.08474 0.07864 0.0869 0.07017 0.1812 0.05667 0.5435 0.7339 3.398 74.08 0.005225 0.01308 0.0186 0.0134 0.01389 0.003532 24.99 23.41 158.8 1956 0.1238 0.1866 0.2416 0.186 0.275 0.08902

    Columns present (from left to right):

    • ID number

    • Diagnosis (M = malignant, B = benign)

    • Ten real-valued features are computed for each cell nucleus (Column 3 to Column 32):

      • radius (mean of distances from center to points on the perimeter)
      • texture (standard deviation of gray-scale values)
      • perimeter
      • area
      • smoothness (local variation in radius lengths)
      • compactness (perimeter^2 / area - 1.0)
      • concavity (severity of concave portions of the contour)
      • concave points (number of concave portions of the contour)
      • symmetry
      • fractal dimension ("coastline approximation" - 1)

      In other words, column 3 to column 12 represent the above 10 values for the first cell nucleus and so on.

  • Note: Dataset from Source is of .data file. If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad).

    To download as .csv file

    Right click on wdbc.data as shown : Screenshot 2022-07-14 at 11 08 42 PM

    Proceed to add the .csv extention in file name and save to desired location

    Screenshot 2022-07-14 at 11 09 00 PM

  • Citations

    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Breast Cancer Dataset (SEER)

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Building a predictive machine learning model for breast cancer

  • Size: 520KB

  • Example of data + Special Instructions(if any):

    Download SEER Breast Cancer Dataset .csv from Source.

    Example

    Screenshot 2022-07-27 at 4 46 46 PM

  • Note: NA

  • Citations

    - JING TENG. (2019). SEER Breast Cancer Data. IEEE Dataport. https://dx.doi.org/10.21227/a9qy-ph35
    

Lymphography Data Set

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Building a predictive machine learning model for malignant tumor

  • Size: 6 KB

  • Example of data + Special Instructions(if any):

    Download lymphography.data stored in Data Folder from Source

    Example: CSV version

    Screenshot 2022-07-18 at 3 13 38 PM

    Columns present (from left to right):

    • class: normal find, metastases, malign lymph, fibrosis
    • lymphatics: normal, arched, deformed, displaced
    • block of affere: no, yes
    • bl. of lymph. c: no, yes
    • bl. of lymph. s: no, yes
    • by pass: no, yes
    • extravasates: no, yes
    • regeneration of: no, yes
    • early uptake in: no, yes
    • lym.nodes dimin: 0-3
    • lym.nodes enlar: 1-4
    • changes in lym.: bean, oval, round
    • defect in node: no, lacunar, lac. marginal, lac. central
    • changes in node: no, lacunar, lac. margin, lac. central
    • changes in stru: no, grainy, drop-like, coarse, diluted, reticular, stripped, faint,
    • special forms: no, chalices, vesicles
    • dislocation of: no, yes
    • exclusion of no: no, yes
    • no. of nodes in: 0-9, 10-19, 20-29, 30-39, 40-49, 50-59, 60-69, >=70

    All attribute values in the database have been entered as numeric values corresponding to their index in the list of attribute values for that attribute domain as given above.

    For example, a '3' in first column (refer to first row of image above) will correspond to class = malign lymph

  • Note: Dataset from Source is of .data file. If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad).

    To download as .csv file

    Right click on lymphography.data as shown :

    Screenshot 2022-07-18 at 3 10 34 PM

    Proceed to add the .csv extention in file name and save to desired location

    Screenshot 2022-07-18 at 3 10 55 PM

  • Citations

    - This lymphography domain was obtained from the University Medical Centre, Institute of Oncology, Ljubljana, Yugoslavia. Thanks go to M. Zwitter and M. Soklic for providing the data. Please include this citation if you plan to use this database.
    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Colorectal Cancer Dataset

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Gland segmentation problem with images of Haematoxylin and Eosin (H&E) stained slides

  • Size: 259.4MB

  • Example of data + Special Instructions(if any):

    Download by clicking 'Warwick-QU-dataset' on Source

    Example

    train_1.bmp (converted to jpg here):

    Each image will have its corresponding annotated image (filename_anno.bmp) and its grade in the Grade.csv file as shown:

    train_1_anno.bmp (converted to jpg here) :

    Grade: Screenshot 2022-07-18 at 9 12 13 PM

  • Note:

    Comment regarding the annotated images can be found in the Q&A section in the same source link.

    View screenshot

    Screenshot 2022-07-18 at 9 16 42 PM

  • Citations

    - K. Sirinukunwattana, J. P. W. Pluim, H. Chen, X Qi, P. Heng, Y. Guo, L. Wang, B. J. Matuszewski, E. Bruni, U. Sanchez, A. Böhm, O. Ronneberger, B. Ben Cheikh, D. Racoceanu, P. Kainz, M. Pfeiffer, M. Urschler, D. R. J. Snead, N. M. Rajpoot, "Gland Segmentation in Colon Histology Images: The GlaS Challenge Contest" (http://arxiv.org/abs/1603.00275)
    - K. Sirinukunwattana, D.R.J. Snead, N.M. Rajpoot, "A Stochastic Polygons Model for Glandular Structures in Colon Histology Images," in IEEE Transactions on Medical Imaging, 2015 doi: 10.1109/TMI.2015.2433900 
    

Cervical cancer (Risk Factors) Data Set

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Building a predictive machine learning model for cervical cancer

  • Size: 102 KB

  • Example of data + Special Instructions(if any):

    Download risk_factors_cervical_cancer.csv from Data Folder from Source

    Example

    Age Number of sexual partners First sexual intercourse Num of pregnancies Smokes Smokes (years) Smokes (packs/year) Hormonal Contraceptives Hormonal Contraceptives (years) IUD IUD (years) STDs STDs (number) STDs:condylomatosis STDs:cervical condylomatosis STDs:vaginal condylomatosis STDs:vulvo-perineal condylomatosis STDs:syphilis STDs:pelvic inflammatory disease STDs:genital herpes STDs:molluscum contagiosum STDs:AIDS STDs:HIV STDs:Hepatitis B STDs:HPV STDs: Number of diagnosis STDs: Time since first diagnosis STDs: Time since last diagnosis Dx:Cancer Dx:CIN Dx:HPV Dx Hinselmann Schiller Citology Biopsy
    18 4 15 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 ? ? 0 0 0 0 0 0 0 0
    15 1 14 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 ? ? 0 0 0 0 0 0 0 0

  • Note: NA

  • Citations

    - Kelwin Fernandes, Jaime S. Cardoso, and Jessica Fernandes. 'Transfer Learning with Partial Observability Applied to Cervical Cancer Screening.' Iberian Conference on Pattern Recognition and Image Analysis. Springer International Publishing, 2017.
    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

CT Medical Images

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Examining the trends in CT image data associated with using contrast and patient age/Build tools for automatically classifying these images when they have been misclassified

  • Size of files

    • 52.8MB for dicom_dir folder
    • 209.7MB for tiff_images folder
  • Example of data + Special Instructions(if any):

    Download relevant folders from Source.

    Example

    ID_0000_AGE_0060_CONTRAST_1_CT.dcm (converted to jpg here):

    ID_0000_AGE_0060_CONTRAST_1_CT.tif (converted to jpg here):

    Screenshot 2022-07-19 at 1 32 27 PM

  • Note: The dataset from Source is only a subset from the cancer imaging archive. For more, click here

  • Citations

    - Albertina, B., Watson, M., Holback, C., Jarosz, R., Kirk, S., Lee, Y., … Lemmerman, J. (2016). Radiology Data from The Cancer Genome Atlas Lung Adenocarcinoma [TCGA-LUAD] collection. The Cancer Imaging Archive. http://doi.org/10.7937/K9/TCIA.2016.JGNIHEP5
    - Clark K, Vendt B, Smith K, Freymann J, Kirby J, Koppel P, Moore S, Phillips S, Maffitt D, Pringle M, Tarbox L, Prior F. The Cancer Imaging Archive (TCIA): Maintaining and Operating a Public Information Repository, Journal of Digital Imaging, Volume 26, Number 6, December, 2013, pp 1045-1057. (paper) 
    

Leukemia Classification

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Creating a model to classify normal from abnormal cell images

  • Size: 10.42GB

  • Example of data + Special Instructions(if any):

    Download from Source

    Example

    Acute lymphoblastic leukemia (ALL) (converted to jpg here):

    Normal (Images in hem folder) (converted to jpg here):

    Labels for images in validation_data can be found in the csv file that is stored in the same folder.

    For example:

    1.bmp (converted to jpg here):

    The label for this image can be found in its corresponding row in the csv file. The image file name will be the value of new_names in the csv file as shown below (second column):

    Screenshot 2022-07-18 at 5 28 01 PM

  • Note: The labels for the testing data was not given. Click here for more information

  • Citations

    - Gupta, A., & Gupta, R. (2019). ALL Challenge dataset of ISBI 2019 [Data set]. The Cancer Imaging Archive.(https://doi.org/10.7937/tcia.2019.dc64i46r)
    - Publication Citation (if required): 
      - Anubha Gupta, Rahul Duggal, Ritu Gupta, Lalit Kumar, Nisarg Thakkar, and Devprakash Satpathy, “GCTI-SN: Geometry-Inspired Chemical and Tissue Invariant Stain Normalization of Microscopic Medical Images,”, under review. Ritu Gupta, Pramit Mallick, Rahul Duggal, Anubha Gupta, and Ojaswa Sharma, "Stain Color Normalization and Segmentation of Plasma Cells in Microscopic Images as a Prelude to Development of Computer Assisted Automated Disease Diagnostic Tool in Multiple Myeloma," 16th International Myeloma Workshop (IMW), India, March 2017.
      - Rahul Duggal, Anubha Gupta, Ritu Gupta, Manya Wadhwa, and Chirag Ahuja, “Overlapping Cell Nuclei Segmentation in Microscopic Images UsingDeep Belief Networks,” Indian Conference on Computer Vision, Graphics and Image Processing (ICVGIP), India, December 2016. Rahul Duggal, Anubha Gupta, and Ritu Gupta, “Segmentation of overlapping/touching white blood cell nuclei using artificial neural networks,” CME Series on Hemato-Oncopathology, All India Institute of Medical Sciences (AIIMS), New Delhi, India, July 2016.
      - Rahul Duggal, Anubha Gupta, Ritu Gupta, and Pramit Mallick, "SD-Layer: Stain Deconvolutional Layer for CNNs in Medical Microscopic Imaging," In: Descoteaux M., Maier-Hein L., Franz A., Jannin P., Collins D., Duchesne S. (eds) Medical Image Computing and Computer-Assisted Intervention − MICCAI 2017, MICCAI 2017. Lecture Notes in Computer Science, Part III, LNCS 10435, pp. 435–443. Springer, Cham. DOI: https://doi.org/10.1007/978-3-319-66179-7_50 .
    

Lung and Colon Cancer Histopathological Images

  • Source: Click here to proceed to Kaggle site

    • The data set can also be downloaded from here
    • Relevant github
  • Data Type: Images

  • Possible uses: Creating a model to identify lung/colon cancer images from normal lung/colon images

  • Size of files

    • 959.8MB for colon_image_sets
    • 929.7MB for lung_image_sets
  • Example of data + Special Instructions(if any):

    For colon images download colon_image_sets from Source. For lung images, download lung_image_sets from Source.

    Example

    Colon adenocarcinoma (from colon_aca folder in colon_image_sets):

    Colon benign tissue (from colon_n folder in colon_image_sets):

    Lung benign tissue(from lung_n folder in lung_image_sets):

    Lung adenocarcinoma(from lung_aca folder in lung_image_sets):

    Lung squamous cell carcinoma(from lung_scc folder in lung_image_sets):

  • Note: NA

  • Citations

    - Borkowski AA, Bui MM, Thomas LB, Wilson CP, DeLand LA, Mastorides SM. Lung and Colon Cancer Histopathological Image Dataset (LC25000). arXiv:1912.12142v1 [eess.IV], 2019 (https://arxiv.org/abs/1912.12142v1)
    

Skin Cancer MNIST: HAM10000

  • Source: Click here to proceed to site

    • The original challenge can be found here and the dataset can also be downloaded here after accepting their terms and conditions.
  • Data Type: Images + Tabular

  • Possible uses: Create a model to predict diagnosis for pigmented lesions

  • Size of files

    • 1.37GB for HAM10000_images_part_1
    • 1.4GB for HAM10000_images_part_2
    • 563KB for HAM10000_metadata.csv
  • Example of data + Special Instructions(if any):

    Download HAM10000_images_part_1.zip, HAM10000_images_part_2 and HAM10000_metadata.csv from Source

    Example: ISIC_0024306.jpg

    Each image has its corresponding row in HAM10000_metadata.csv where its file name is its image_id as shown:

    Screenshot 2022-07-15 at 10 29 29 AM

  • Note: NA

  • Citations

    - Noel Codella, Veronica Rotemberg, Philipp Tschandl, M. Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, Harald Kittler, Allan Halpern: “Skin Lesion Analysis Toward Melanoma Detection 2018: A Challenge Hosted by the International Skin Imaging Collaboration (ISIC)”, 2018; https://arxiv.org/abs/1902.03368
    - Tschandl, P., Rosendahl, C. & Kittler, H. The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Sci. Data 5, 180161 doi:10.1038/sdata.2018.161 (2018).
    

Thoracic Surgery Data Data Set

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Classification problem related to the post-operative life expectancy in the lung cancer patients

  • Size: 24 KB

  • Example of data + Special Instructions(if any):

    Download ThoraricSurgery.arff from Data Folder from Source

    Example

    Screenshot 2022-07-15 at 2 42 52 PM

    Columns present (from left to right):

    • DGN
    • PRE4
    • PRE5
    • PRE6
    • PRE7
    • PRE8
    • PRE9
    • PRE10
    • PRE11
    • PRE14
    • PRE17
    • PRE19
    • PRE25
    • PRE30
    • PRE32
    • AGE
    • Risk1Yr

  • Note: Dataset from Source is of .arff file.If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad)

    To convert to csv file

    Remove the highlighted text as shown and save:

    Screenshot 2022-07-15 at 2 45 46 PM

    Rename the new file by changing ".arff" to ".csv" extension.

    Updated file:

    Screenshot 2022-07-15 at 2 53 11 PM

  • Citations

    - Zięba, M., Tomczak, J. M., Lubicz, M., & Świątek, J. (2013). Boosted SVM for extracting rules from imbalanced data in application to prediction of the post-operative life expectancy in the lung cancer patients. Applied Soft Computing (https://www.sciencedirect.com/science/article/abs/pii/S0166432811002221?via%3Dihub)
    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Diabetes Dataset

Association of body mass index and age with incident diabetes in Chinese adults

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Building a model to study association of BMI and Age with Diabetes

  • Size: 28.3 KB

  • Example of data + Special Instructions(if any):

    Download RC Health Care Data-20180820.xlsx from Source.

    Example

    id Age (y) Gender(1, male; 2, female) site height(cm) weight(kg) BMI(kg/m2) SBP(mmHg) DBP(mmHg) FPG (mmol/L) Cholesterol(mmol/L) Triglyceride(mmol/L) HDL-c(mmol/L) LDL(mmol/L) ALT(U/L) AST(U/L) BUN(mmol/L) CCR(umol/L) FPG of final visit(mmol/L) Diabetes diagnosed during followup(1,Yes) censor of diabetes at followup(1, Yes; 0, No) year of followup smoking status(1,current smoker;2, ever smoker;3,never smoker) drinking status(1,current drinker;2, ever drinker;3,never drinker) family histroy of diabetes(1,Yes;0,No)
    1 43 2 16 166.4 53.5 19.3 96 57 4.99 5.13 0.78 10 3.08 50.3 4.97 0 2.15195 3 3 1
    2 34 1 2 169 57 20 124 69 3.51 4.61 1.75 1.09 3.13 29.1 6.13 83.7 5.5 0 3.96988 0
    3 32 2 2 157 51 20.7 98 68 4.25 4.73 0.47 6.9 19.5 4.45 42.8 4.9 0 3.93977 0

  • Note: NA

  • Citations

    - Chen, Ying, Zhang, Xiao-Ping, Yuan, Jie, Cai, Bo, Wang, Xiao-Li, Wu, Xiao-Li, Zhang, Yue-Hua, Zhang, Xiao-Yi, Yin, Tong, Zhu, Xiao-Hui, Gu, Yun-Juan, Cui, Shi-Wei, Lu, Zhi-Qiang, & Li, Xiao-Ying. (2018). Data from: Association of body mass index and age with incident diabetes in Chinese adults: a population-based cohort study [Data set]. https://doi.org/10.5061/dryad.ft8750v
    

Clinical characteristics, types and complications of diabetics with young age at the onset

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: To determine the distribution, clinical features and complications of the different types of diabetes in young age

  • Size: 22 KB

  • Example of data + Special Instructions(if any):

    Download 9. Master Chart.xlsx from Source.

    Example

    Screenshot 2022-07-19 at 4 45 06 PM

  • Note: NA

  • Citations

    - SAHU, PRANABANANDA; Das, Sidhartha (2019), “Data for: Clinical characteristics, types and complications of diabetics with young age at the onset ( 14 to 25 years).”, Mendeley Data, V1, doi: 10.17632/jf429jpgwt.1
    

Diabetes 130-US hospitals for years 1999-2008 Data Set

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Identifying significant features in predicting readmission rates of diabetic patients

  • Size of files

    • 19.2MB for diabetic_data.csv
    • 3KB for IDs_mapping.csv
  • Example of data + Special Instructions(if any):

    Download dataset_diabetes.zip from Data Folder from Source

    Example

    encounter_id patient_nbr race gender age weight admission_type_id discharge_disposition_id admission_source_id time_in_hospital payer_code medical_specialty num_lab_procedures num_procedures num_medications number_outpatient number_emergency number_inpatient diag_1 diag_2 diag_3 number_diagnoses max_glu_serum A1Cresult metformin repaglinide nateglinide chlorpropamide glimepiride acetohexamide glipizide glyburide tolbutamide pioglitazone rosiglitazone acarbose miglitol troglitazone tolazamide examide citoglipton insulin glyburide-metformin glipizide-metformin glimepiride-pioglitazone metformin-rosiglitazone metformin-pioglitazone change diabetesMed readmitted
    2278392 8222157 Caucasian Female [0-10) ? 6 25 1 1 ? Pediatrics-Endocrinology 41 0 1 0 0 0 250.83 ? ? 1 None None No No No No No No No No No No No No No No No No No No No No No No No No No NO
    149190 55629189 Caucasian Female [10-20) ? 1 1 7 3 ? ? 59 0 18 0 0 0 276 250.01 255 9 None None No No No No No No No No No No No No No No No No No Up No No No No No Ch Yes >30
  • Note: NA

  • Citations

    - Beata Strack, Jonathan P. DeShazo, Chris Gennings, Juan L. Olmo, Sebastian Ventura, Krzysztof J. Cios, and John N. Clore, “Impact of HbA1c Measurement on Hospital Readmission Rates: Analysis of 70,000 Clinical Database Patient Records,” BioMed Research International, vol. 2014, Article ID 781670, 11 pages, 2014. (https://www.hindawi.com/journals/bmri/2014/781670/)
    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Diabetic Retinopathy Debrecen Data Set Data Set

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Build a model to predict whether an image contains signs of diabetic retinopathy or not

  • Size: 117 KB

  • Example of data + Special Instructions(if any):

    Download messidor_features.arff from Data Folder from Source

    Example

    Screenshot 2022-07-15 at 4 05 20 PM

    Columns present (from left to right):

    • The binary result of quality assessment. 0 = bad quality 1 = sufficient quality.
    • The binary result of pre-screening, where 1 indicates severe retinal abnormality and 0 its lack.
    • The results of MA detection. Each feature value stand for the number of MAs found at the confidence levels alpha = 0.5, . . . , 1, respectively. (Columns 2-7)
    • contain the same information as 2-7) for exudates. However, as exudates are represented by a set of points rather than the number of pixels constructing the lesions, these features are normalized by dividing the number of lesions with the diameter of the ROI to compensate different image sizes (Columns 8-15)
    • The euclidean distance of the center of the macula and the center of the optic disc to provide important information regarding the patient's condition. This feature is also normalized with the diameter of the ROI.
    • The diameter of the optic disc.
    • The binary result of the AM/FM-based classification.
    • Class label. 1 = contains signs of DR (Accumulative label for the Messidor classes 1, 2, 3), 0 = no signs of DR.

  • Note: Dataset from Source is of .arff file.If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad)

    To convert to csv file

    Remove the highlighted text as shown and save:

    Screenshot 2022-07-15 at 4 05 56 PM

    Rename the new file by changing ".arff" to ".csv" extension.

    Updated file:

    Screenshot 2022-07-15 at 4 06 37 PM

  • Citations

    - Balint Antal, Andras Hajdu: An ensemble-based system for automatic screening of diabetic retinopathy, Knowledge-Based Systems 60 (April 2014), 20-27.
        - The dataset is based on features extracted from the Messidor image dataset. However, link provided is no longer available: http://messidor.crihan.fr/index-en.php 
    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Early Classification of Diabetes

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Create a classification model to predict diabetes/ Explore the most common features associated with diabetic risk

  • Size: 21 KB

  • Example of data + Special Instructions(if any):

    Download diabetes_data.csv from Source.

    Example

    age gender polyuria polydipsia sudden_weight_loss weakness polyphagia genital_thrush visual_blurring itching irritability delayed_healing partial_paresis muscle_stiffness alopecia obesity class
    40 Male 0 1 0 1 0 0 0 1 0 1 0 1 1 1 1
    58 Male 0 0 0 1 0 0 1 0 0 0 1 0 1 0 1
    41 Male 1 0 0 1 1 0 0 1 0 1 0 1 1 0 1
  • Note: Viewing file on Microsoft Excel might be an issue as the delimeter used is a semicolon.

  • Citations

    - Islam M.M.F., Ferdousi R., Rahman S., Bushra H.Y. (2020) Likelihood Prediction of Diabetes at Early Stage Using Data Mining Techniques. In: Gupta M., Konar D., Bhattacharyya S., Biswas S. (eds) Computer Vision and Machine Intelligence in Medical Image Analysis. Advances in Intelligent Systems and Computing, vol 992. Springer, Singapore. https://doi.org/10.1007/978-981-13-8798-2_12
    

Monkeypox

Monkeypox Skin Lesion Dataset

  • Source: Click here to proceed to site

  • Data Type: Image + Tabular

  • Possible uses: Build a model to classify monkeypox images from non monkeypox images

  • Size of files

    • 1.5MB for Original Images folder
    • 5KB for Monkeypox_Dataset_metadata.csv
    • 25.9MB for Augmented_Images.zip
    • 20.9MB for Fold1
  • Example of data + Special Instructions(if any):

    Download relevant files from Source

    • Users may choose to use the folds/augmented images directly (Refer to source for more information)
    Example

    Monkeypox:

    Others:

  • Note: There is a csv file that indicates the labels of the images. This may not be needed as the images are already sorted into their respective folders in main folder

    CSV file

    Screenshot 2022-07-18 at 11 46 15 AM

  • Citations

    - If this dataset helped your research, please cite: Ali, S. N., Ahmed, M. T., Paul, J., Jahan, T., Sani, S. M. Sakeef, Noor, N., & Hasan, T. (2022). Monkeypox Skin Lesion Detection Using Deep Learning Models: A Preliminary Feasibility Study(https://arxiv.org/abs/2207.03342). arXiv preprint arXiv:2207.03342.
    

Clinical features of 21 human monkeypox cases seen at NDUTH 2017

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Study clinical characteristics of confirmed cases of monkeypox

  • Size: 11KB

  • Example of data + Special Instructions(if any):

    Download relevant files from Source

    Example

    Screenshot 2022-07-27 at 2 02 10 PM

  • Note: Small dataset with only 21 records (All positive: 18 laboratory-confirmed and three probable cases). Two of the 21 cases had laboratory evidence of concomitant chicken pox (by PCR/serology)

  • Citations

    - Ogoina, Dimie; Izibewule, James Hendris; Ogunleye, Adesola; Ederiane, Ebi; Anebonam, Uchenna; Neni, Aworabhi; et al. (2019): Clinical features of 21 human monkeypox cases seen at NDUTH.. PLOS ONE. Dataset. https://doi.org/10.1371/journal.pone.0214229.s001 
    

Parkinson Disease

Parkinson Speech Dataset with Multiple Types of Sound Recordings Data Set

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Build a model to identify parkinson disease from audio details

  • Size of files

    • 194KB for train_data.txt
    • 47KB for test_data.txt
  • Example of data + Special Instructions(if any):

    Download Parkinson_Multiple_Sound_Recording.rar from Data Folder from Source

    Example

    Screenshot 2022-07-15 at 3 28 31 PM

    Columns present (from left to right):

    • Subject id
    • features(column 2 to 27)
      • features 1-5: Jitter (local),Jitter (local, absolute),Jitter (rap),Jitter (ppq5),Jitter (ddp)
      • features 6-11: Shimmer (local),Shimmer (local, dB),Shimmer (apq3),Shimmer (apq5), Shimmer (apq11),Shimmer (dda)
      • features 12-14: AC,NTH,HTN
      • features 15-19: Median pitch,Mean pitch,Standard deviation,Minimum pitch,Maximum pitch
      • features 20-23: Number of pulses,Number of periods,Mean period,Standard deviation of period
      • features 24-26: Fraction of locally unvoiced frames,Number of voice breaks,Degree of voice breaks
    • UPDRS
    • class information (not present in test_data.txt)

  • Note: Click here for more information on how to open .rar files

  • Citations

    - Erdogdu Sakar, B., Isenkul, M., Sakar, C.O., Sertbas, A., Gurgen, F., Delil, S., Apaydin, H., Kursun, O., 'Collection and Analysis of a Parkinson Speech Dataset with Multiple Types of Sound Recordings', IEEE Journal of Biomedical and Health Informatics, vol. 17(4), pp. 828-834, 2013.
    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Parkinson Tappy Keystroke Data

  • Source: Click here to proceed to site

  • Data Type: Text

  • Possible uses: Build a model for the detection of early Parkinson's Disease using multiple characteristics of finger movement while typing

  • Size of files

    • 43KB for Archived users
    • 549.7MB for Tappy Data
  • Example of data + Special Instructions(if any):

    Download files from Source.

    Example

    Each user will have their corresponding text file in Archived users with the following information:

    Screenshot 2022-07-20 at 5 10 10 PM

    In Tappy Data folder, a participant may have more than one file. The filename comprises the 10 character code (matching the user details file) and the YYMM of the data:

    Screenshot 2022-07-20 at 5 12 33 PM

  • Note: NA

  • Citations

    - Goldberger, A., Amaral, L., Glass, L., Hausdorff, J., Ivanov, P. C., Mark, R., ... & Stanley, H. E. (2000). PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation [Online]. 101 (23), pp. e215–e220.
    - Adams WR (2017) High-accuracy detection of early Parkinson's Disease using multiple characteristics of finger movement while typing. PLOS ONE 12(11): e0188226. https://doi.org/10.1371/journal.pone.0188226
    

Safety and Preliminary Efficacy of Intranasal Insulin for Cognitive Impairment in Parkinson Disease and Multiple System Atrophy

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Study the effects of intranasal insulin (INI) on cognition and motor performance in Parkinson Disease

  • Size of files

    • 3KB for PD-Table1.csv
    • 818 bytes for PD-Table2.csv
  • Example of data + Special Instructions(if any):

    Download files from Source.

    Example

    PD-Table1.csv:

    Study No Pharmacy excluded IversusP Pre-Post Age Gender Etnicity Smoking Height (in) Weight BMI Disease Disease Duration SBP DBP HR Walk-Duration Walk-NoSteps AverStride (in) MOCA Glucose SBP DBP HR BMI 2 Walk-Duration 2 Walk-NoSteps 2 AverStrid (in) 2 Gait Speed 2 HnH PGI-I BDI MOCA 2 F#words A#words S#words FAS Pre-Post Glucose 2 MOCA 3 Walk-Duration 3 Walk-NoSteps 3 AverStrid (in) 3 Gait Speed 3 HnH 2 PGI-I 2 BDI 2 F#words 2 A#words 2 S#words 2 FAS 2
    1 n/a excluded/screen failure Post
    2 1-PD 0 Pre 79 1 w 0 69 176 26.02 PD 3 154 88 66 5.81 9 16 26 84 148 82 62 26.02 4.76 7 21 1.05042017 2.5 3 17 27 16 7 8 31 Post 73 24 5.45 9 16 0.91743119 2.5 3 13 6 4 10 20
    3 2-PD 1 Pre 64 1 w 0 76.5 184 22.1 PD 9 129 83 70 3.72 7 21 30 89 129 83 70 22.1 3.9 7 21 1.28205128 2.5 3 3 30 15 16 15 46 Post 83 29 4.08 7 21 1.2254902 2.5 3 3 15 17 16 48
    4 3-PD 0 Pre 53 2 w 0 67 155 24.3 PD 5 113 69 70 3.84 8 18 28 98 113 69 70 24.3 4.4 8 18 1.13636364 2.5 3 15 29 11 11 12 34 Post 98 30 4.55 8 18 1.0989011 2.5 3 13 11 8 9 28

    PD-Table2.csv:

    Study_No Group Pharmacy Disease Disease_Dur Age Gender Ethnicity updrs1_base updrs1_tre1 updrs2_base updrs2_tre1 updrs3_base updrs3_tre1 brady_base brady_tre1
    1 excluded -99
    2 0 1-PD PD 3 79 1 w 15 9 14 11 32 24 13.5 11
    3 1 2-PD PD 9 64 1 w 9 9 15 15 23 23 5 5
    4 0 3-PD PD 5 53 2 w 15 15 13 13 26 23 7 7

  • Note: Small dataset with only 16 rows

  • Citations

    - Novak, V., & Novak, P. (2019). Safety and Preliminary Efficacy of Intranasal Insulin for Cognitive Impairment in Parkinson Disease and Multiple System Atrophy (version 1.0). PhysioNet. https://doi.org/10.13026/bn1x-my19.
       
    - Goldberger, A., Amaral, L., Glass, L., Hausdorff, J., Ivanov, P. C., Mark, R., ... & Stanley, H. E. (2000). PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals. Circulation [Online]. 101 (23), pp. e215–e220.
    

Brain Related Datasets

Brain tumor data

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Build a model to to identify the types of brain tumor

  • Size: 733.1MB

  • Example of data + Special Instructions(if any):

    Download relevant files from Source

    Example

    meningioma(1):

    glioma(2):

    pituitary tumor(3):

  • Note: NA

  • Citations

    - Cheng, Jun (2017): brain tumor dataset. figshare. Dataset. https://doi.org/10.6084/m9.figshare.1512427.v5
    

MRI and Alzheimers

  • Source: Click here to proceed to site

  • Data Type: Tabular (According to discussion, images was said to be found here, where a request for access has to be made to OASIS)

  • Possible uses: Build a model to predict Alzheimer based on MRI data

  • Size of files

    • 22KB for oasis_coss-sectional.csv
    • 28KB for oasis_longitudinal.csv
  • Example of data + Special Instructions(if any):

    Download relevant files from Source

    Example

    oasis_coss-sectional.csv: Screenshot 2022-07-19 at 2 18 24 PM

    oasis_longitudinal.csv: Screenshot 2022-07-19 at 2 17 45 PM

  • Note: NA

  • Citations

    - When publishing findings that benefit from OASIS data, please include the following grant numbers in the acknowledgements section and in the associated Pubmed Central submission: P50 AG05681, P01 AG03991, R01 AG021910, P20 MH071616, U24 RR0213
    

Stroke Prediction

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Build a model to predict stroke from given features

  • Size: 2.6 MB

  • Example of data + Special Instructions(if any):

    Download dataset.rar from Source

    Example

    Screenshot 2022-07-19 at 4 00 39 PM

  • Note: Click here for more information on how to access .rar files

  • Citations

    - Liu, Tianyu; Fan, Wenhui; Wu, Cheng (2019), “Data for: A hybrid machine learning approach to cerebral stroke prediction based on imbalanced medical-datasets”, Mendeley Data, V1, doi: 10.17632/x8ygrw87jw.1
    

Eye Related Datasets

Bajwa Hospital (Multi Eye Disease Dataset)

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Build a model that can provide a highly accurate diagnosis from an eye image

  • Size: 3.59GB

  • Example of data + Special Instructions(if any):

    Download file from Source

    Example

    Normal:

    Cataract:

    Glaucoma:

    Retina Disease:

  • Note: Click here for more information on how to access .rar files

  • Citations

    - Kaur, Palwinder (2022), “Bajwa Hospital (Multi Eye Disease Dataset)”, Mendeley Data, V2, doi: 10.17632/rgwpd4m785.2
    

Retinal OCT Images (optical coherence tomography)

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Building a model to identify normal OCT images from those with diseases

  • Size: 7.19GB

  • Example of data + Special Instructions(if any):

    Download from Source

    Example

    NORMAL:

    CNV:

    DME:

    DRUSEN:

  • Note: In the same folder, a folder named chest_xray can be found. Information regarding this image dataset can be found at the Chest X-ray Images (Pneumonia) section.

  • Citations

    - Kermany, Daniel; Zhang, Kang; Goldbaum, Michael (2018), “Large Dataset of Labeled Optical Coherence Tomography (OCT) and Chest X-Ray Images”, Mendeley Data, V3, doi: 10.17632/rscbjbr9sj.3
    

Gastrointestinal Related Datasets

Bowel Sounds

  • Source: Click here to proceed to site

  • Data Type: Audio

  • Possible uses: Build an algorithm such as deep neural network to detect bowel sounds

  • Size: 283.8MB

  • Example of data + Special Instructions(if any):

    Download from Source

    Example:

    0_a.mp4

    Each audio file has its corresponding csv file:

    Screenshot 2022-07-20 at 4 23 48 PM

  • Note: NA

  • Citations

    - Jakub Ficek and Kacper Radzikowski and Jan Nowak and Osamu Yoshie and Jarosław Walkowiak and Robert Nowak, "Analysis of gastrointestinal acoustic activity using deep neural networks", MDPI Sensors, 2021, doi:10.3390/s21227602, http://dx.doi.org/10.3390/s21227602
    

Kvasir Dataset

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Build a model to classify gastrointenstinal disease

  • Size of files:

    • 2.52GB for kvasir-dataset-v2
    • 40.2MB for kvasir-dataset-v2-features
  • Example of data + Special Instructions(if any):

    Download relevant folders from Source

    Example

    dyed-lifted-polyps:

    dyed-resection-margins:

    esophagitis:

    normal-cecum:

    normal-pylorus:

    normal-z-line:

    polyps:

    ulcerative-colitis:

  • Note: NA

  • Citations

    - Pogorelov, K., Randel, K., Griwodz, C., Eskeland, S., Lange, T., Johansen, D., Spampinato, C., Dang-Nguyen, D.T., Lux, M., Schmidt, P., Riegler, M., & Halvorsen, P. (2017). KVASIR: A Multi-Class Image Dataset for Computer Aided Gastrointestinal Disease Detection. In Proceedings of the 8th ACM on Multimedia Systems Conference (pp. 164–169). ACM.
    

Heart Related Datasets

An Extensive Dataset for the Heart Disease Classification System

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Build a model to predict heart disease/Identify important characteristics/factors that contribute to a heart attack

  • Size: 51 KB

  • Example of data + Special Instructions(if any):

    Download Medicaldataset.arff from Source

    Example

    Screenshot 2022-07-19 at 4 17 35 PM

    Columns present (from left to right):

    • age
    • gender
    • impulse
    • pressurehight
    • pressurelow
    • glucose
    • kcm
    • troponin
    • class

  • Note: Dataset from Source is of .arff file.If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad)

    To convert to csv file

    Remove the highlighted text as shown and save:

    Screenshot 2022-07-19 at 4 18 48 PM

    Rename the new file by changing ".arff" to ".csv" extension.

    Updated file:

    Screenshot 2022-07-19 at 4 19 37 PM

  • Citations

    - Maghdid, Sozan; Rashid, Tarik A. (2022), “An Extensive Dataset for the Heart Disease Classification System ”, Mendeley Data, V2, doi: 10.17632/65gxgy2nmg.2
    

Heart Failure Prediction

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Create a model for predicting mortality caused by Heart Failure

  • Size: 12KB

  • Example of data + Special Instructions(if any):

    Download heart_failure_clinical_records_dataset.csv from Source

    Example

    Screenshot 2022-07-18 at 12 24 40 PM

  • Note: NA

  • Citations

    - Davide Chicco, Giuseppe Jurman: Machine learning can predict survival of patients with heart failure from serum creatinine and ejection fraction alone. BMC Medical Informatics and Decision Making 20, 16 (2020). https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-020-1023-5
    

Kidney Related Datasets

A Brazilian dataset for screening the risk of the Chronic Kidney Disease

  • Source: Click here to proceed to site

  • Data Type: Text/Tabular

  • Possible uses: Build a model to predict level of risk of Chronic Kidney Disease

  • Size of files:

    • 2 KB for Original.arff
    • 2 KB for Augmented.arff
  • Example of data + Special Instructions(if any):

    Download relevant file from Source

    • In Augmented.arff, authors have augmented the dataset to decrease the impact of imbalanced data.
    Example

    Screenshot 2022-07-19 at 5 27 05 PM

  • Note: Dataset from Source is of .arff file.If you are unable to view the file after downloading directly, right click on the file and open with any text editor application(eg Notepad)

    To convert to csv file

    Remove the highlighted text as shown and save:

    Screenshot 2022-07-19 at 5 28 38 PM

    Rename the new file by changing ".arff" to ".csv" extension.

    Updated file:

    Screenshot 2022-07-19 at 5 29 09 PM

  • Citations

    - Sobrinho, Alvaro; Dias da Silva, Leandro ; Perkusich, Angelo; Queiroz, Andressa; Eliete Pinheiro, Maria (2021), “A Brazilian dataset for screening the risk of the Chronic Kidney Disease”, Mendeley Data, V3, doi: 10.17632/2gkg7vvcrm.3
    

Risk Factor prediction of Chronic Kidney Disease Data Set

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Create a model for predicting Chronic Kidney Disease

  • Size: 34KB

  • Example of data + Special Instructions(if any):

    Download ckd-dataset-v2.csv from Data Folder from Source

    Example

    bp (Diastolic) bp limit sg al class rbc su pc pcc ba bgr bu sod sc pot hemo pcv rbcc wbcc htn dm cad appet pe ane grf stage affected age
    discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete discrete
    class meta
    0 0 1.019 - 1.021 1 - 1 ckd 0 < 0 0 0 0 < 112 < 48.1 138 - 143 < 3.65 < 7.31 11.3 - 12.6 33.5 - 37.4 4.46 - 5.05 7360 - 9740 0 0 0 0 0 0 ≥ 227.944 s1 1 < 12
    0 0 1.009 - 1.011 < 0 ckd 0 < 0 0 0 0 112 - 154 < 48.1 133 - 138 < 3.65 < 7.31 11.3 - 12.6 33.5 - 37.4 4.46 - 5.05 12120 - 14500 0 0 0 0 0 0 ≥ 227.944 s1 1 < 12
    0 0 1.009 - 1.011 ≥ 4 ckd 1 < 0 1 0 1 < 112 48.1 - 86.2 133 - 138 < 3.65 < 7.31 8.7 - 10 29.6 - 33.5 4.46 - 5.05 14500 - 16880 0 0 0 1 0 0 127.281 - 152.446 s1 1 < 12
    1 1 1.009 - 1.011 3 - 3 ckd 0 < 0 0 0 0 112 - 154 < 48.1 133 - 138 < 3.65 < 7.31 13.9 - 15.2 41.3 - 45.2 4.46 - 5.05 7360 - 9740 0 0 0 0 0 0 127.281 - 152.446 s1 1 < 12

  • Note: According to site, pre-processing has to be done for any machine learning algorithm to be applied. Some special characters might be observed in Microsoft Excel. Using other applications such as Numbers(Macbook), these special characters correspond to mathematical symbols such as >= or <=

  • Citations

    - M. A. Islam, S. Akter, M. S. Hossen, S. A. Keya, S. A. Tisha and S. Hossain, 'Risk Factor Prediction of Chronic Kidney Disease based on Machine Learning Algorithms,' 2020 3rd International Conference on Intelligent Sustainable Systems (ICISS), Thoothukudi, India, 2020, pp. 952-957, doi: 10.1109/ICISS49785.2020.9315878.(https://ieeexplore.ieee.org/document/9315878)
    - Dua, D. and Graff, C. (2019). UCI Machine Learning Repository [http://archive.ics.uci.edu/ml]. Irvine, CA: University of California, School of Information and Computer Science
    

Lung Related Datasets

A dataset of lung sounds recorded from the chest wall using an electronic stethoscope

  • Source: Click here to proceed to site

  • Data Type: Audio

  • Possible uses: Create a model to predict disease from lung sounds

  • Size of files:

    • 46.8MB for Audio Files
    • 13KB for Data annotation.xlsx
  • Example of data + Special Instructions(if any):

    Download by Audio Files.zip and Data annotation.xlsx from Source

    • "Stethoscope Files.zip" contains the original records imported from the stethoscope
    Example

    BP1_Asthma,I E W,P L L,70,M (converted to mp4 here):

    BP1_Asthma.I.E.W.P.L.L.70.M.mp4

    Filename correspond to values in its corresponding row in Data annotation.xlsx.

    Screenshot 2022-07-20 at 2 51 39 PM Screenshot 2022-07-20 at 2 58 32 PM

  • Note: If .wav files are not supported, try using online converter tools to convert to a audio file type that is supported by your device.

  • Citations

    - Fraiwan, Mohammad; Fraiwan, Luay; Khassawneh, Basheer; Ibnian, Ali (2021), “A dataset of lung sounds recorded from the chest wall using an electronic stethoscope”, Mendeley Data, V3, doi: 10.17632/jwyy9np4gv.3
    

NIH Chest X-ray Dataset

  • Source: Click here to proceed to site

  • Data Type: Images + Tabular

  • Possible uses: Create a model to predict diseases from chest x-rays

  • Size of files

    • 2 to 4 GB per image zip folder depending on number of images in the folder
    • 7.9 MB for Data_Entry_2017.csv
    • 92 KB for BBox_List_2017.csv
  • Example of data + Special Instructions(if any):

    Download image folders images_001 to images_012 from Source. Class labels are to be obtained from Data_Entry_2017.csv. Extra information regarding the patient and image can be found in the file as well. Coordinates of the disease region bounding boxes can be obtained from BBox_List_2017.csv.

    Example: 00000032_037.png

    Each image has its corresponding row in Data_Entry_2017.csv where its file name (with extension) is its Image Index as shown:

    Screenshot 2022-07-13 at 3 55 31 PM In this example, the image is classified to 3 different diseases, namely Cadiomegaly, Edema and Infiltration. There are a total of 15 class labels as follows:

    • Atelectasis
    • Consolidation
    • Infiltration
    • Pneumothorax
    • Edema
    • Emphysema
    • Fibrosis
    • Effusion
    • Pneumonia
    • Pleural_thickening
    • Cardiomegaly
    • Nodule Mass
    • Hernia

    Data in BBox_List_2017.csv is very limited. An image may be classified to one of the diseases but do not have a corresponding record in BBox_List_2017.csv file. Referencing to the same example above shows that there is only one disease region bounding box identified in BBox_List_2017.csv file.

    Screenshot 2022-07-13 at 4 19 17 PM

  • Note: Some image labels may not be accurate. A more accurate version is said to be available here. You would be required to submit the following google form to obtain the updated data.

    View form

    Screenshot 2022-07-13 at 2 24 23 PM

  • Citations

    - Wang X, Peng Y, Lu L, Lu Z, Bagheri M, Summers RM. ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases. IEEE CVPR 2017 (https://openaccess.thecvf.com/content_cvpr_2017/papers/Wang_ChestX-ray8_Hospital-Scale_Chest_CVPR_2017_paper.pdf)
    - NIH News release: https://www.nih.gov/news-events/news-releases/nih-clinical-center-provides-one-largest-publicly-available-chest-x-ray-datasets-scientific-community
    - Original source files and documents: https://nihcc.app.box.com/v/ChestXray-NIHCC/folder/36938765345
    

Chest X-ray Images (Pneumonia)

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Create a model to identify classify Pneumonia from chest x-ray images

  • Size: 1.27GB

  • Example of data + Special Instructions(if any):

    Download file from Source. Images can be found in the chest_xray folder. They have been split into train, val and test folders. Within each folder, there are two folders NORMAL and PNEUMONIA.

    Example

    NORMAL:

    PNEUMONIA:

  • Note: NA

  • Citations

    -  Kermany, Daniel; Zhang, Kang; Goldbaum, Michael (2018), “Large Dataset of Labeled Optical Coherence Tomography (OCT) and Chest X-Ray Images”, Mendeley Data, V3, doi: 10.17632/rscbjbr9sj.3
    

Covid19 Coswara Dataset

  • Source: Click here to proceed to site

  • Data Type: Audio + Tabular

  • Possible uses: Build an algorithm to distinguish cough between non-covid19 patients and covid19 patients

  • Size: -

  • Example of data + Special Instructions(if any):

    Download relevant files from Source

    Example

    Audio sample to be included soon.

    Summary of metadata found in combined_data.csv in source:

    id a cold record_date covid_status ctScan dT ep fV fever g l_c l_l l_s others_resp rU smoker testType test_date test_status um vacc cough ftg mp st diabetes ht bd cld diarrhoea ctDate ctScore asthma loss_of_smell others_preexist ihd pneumonia
    eK8ikIYnLQPWGetLBHzkJVCGfpq2 27 TRUE 18/1/22 positive_mild n web y 2 TRUE male India Shimoga Karnataka TRUE n n rtpcr 17/1/22 p y y
    AeP4E7hKFtOmcWye2MghbvDfGlo2 41 29/1/22 positive_mild n web y 2 TRUE male India Bangalore rural Karnataka n n rtpcr 27/1/22 p y y TRUE TRUE

  • Note: Audio files are in a zip file (example 20220224.tar.gz.aa). A extract_data.py is provided in source to extract the audio files which are of .wav format.

  • Citations

    - Coswara - A Database of Breathing, Cough, and Voice Sounds for COVID-19 Diagnosis (https://arxiv.org/abs/2005.10548)
    

QaTa COV19 Dataset

  • Source: Click here to proceed to site

  • Data Type: Images

  • Possible uses: Build a model to classify covid19 patients using chest x-ray images

  • Size of files

    • 215.8MB for QaTa-COV19-v2 folder
    • 4.45GB for Control_Group
  • Example of data + Special Instructions(if any):

    Download the relevant files from Source

    Example

    QaTa-COV19-v2

    covid_1:

    Each image for covid cases has its corresponding image, named mask_FILENAME.png in the Ground-truths folder

    mask_covid_1:

    Control_Group

    In Control_Group_I, all are normal.

    normal_1:

    In Control_Group_II, there is a variety of cases. More details can be found in their readme.txt file in the downloaded folder.

          Control_Group: The chest X-rays that are in the control group can be found under Control_Group/ folder. 
          There are two control groups in this folder. Control Group-I  consists  of  only  normal  (healthy) chest X-rays with 12,544 images. 
          On the other hand, Control Group-II consists of 116,365 chest X-rays from normal and 14 different thoracic disease images. 
          In Control Group II, the CHESTXRAY-14 folder includes train and test sets of ChestXray-14 dataset. 
          In addition to this, bacterial and viral pneumonia from pediatric patients can be found in this folder. 
          However, for the pediatric patient data, there are no train and test sets predefined.
    

  • Note: If you would like to use other versions in the source, please cite accordingly as stated in the source.

  • Citations

    - A. Degerli, S. Kiranyaz, M. E. H. Chowdhury, and M. Gabbouj, "OSegNet: Operational Segmentation Network for COVID-19 Detection using Chest X-ray Images," arXiv:2202.10185, 2022.
    

Respiratory Sound Database

  • Source: Click here to proceed to site

  • Data Type: Audio + Text/Tabular

  • Possible uses: Build an algorithm to give an accurate diagnosis based on respiratory sounds

  • Size of files

    • 1KB for diagnosis file
    • 3KB for demographic data file
    • 2.18GB for database folder
  • Example of data + Special Instructions(if any):

    Download the relevant files from Source:

    To obtain class labels/diagnosis of each subject: Screenshot 2022-07-16 at 11 51 25 PM To obtain demographic information regarding subjects: Screenshot 2022-07-16 at 11 57 06 PM To obtain recordings: Screenshot 2022-07-16 at 11 59 00 PM

    Example

    Diagnosis file

    Screenshot 2022-07-17 at 12 03 45 AM

    Demographic data

    Screenshot 2022-07-17 at 12 04 46 AM

    Columns present (from left to right):

    • Participant ID
    • Age
    • Sex
    • Adult BMI (kg/m2)
    • Child Weight (kg)
    • Child Height (cm)

    Recording (101_1b1_Al_sv_Meditron.wav)

    File was converted to mp4 here

    101_1b1_Al_sc_Meditron.mp4

    File name is divided into 5 elements. Refer to image below for more details: Screenshot 2022-07-17 at 12 13 01 AM

    A corresponding text file (in this case 101_1b1_Al_sv_Meditron.txt) can also be found in the database folder: Screenshot 2022-07-17 at 12 14 14 AM

    Columns present (from left to right):

    • Beginning of respiratory cycle(s)
    • End of respiratory cycle(s)
    • Presence/absence of crackles (presence=1, absence=0)
    • Presence/absence of wheezes (presence=1, absence=0)

  • Note: If .wav files are not supported, try using online converter tools to convert to a audio file type that is supported by your device.

  • Citations

    - Rocha BM et al. (2019) "An open access database for the evaluation of respiratory sound classification algorithms" Physiological Measurement 40 035001
    

Medical Transcripts

Adapting Phrase-based Machine Translation to Normalise Medical Terms in Social Media Messages

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: Develop a tool to translate laymen's terms to a particular medical concept (Text normalization)

  • Size: 11MB

  • Example of data + Special Instructions(if any):

    Download EMNLP_gold_standard.txt from Source.

    Example

    Screenshot 2022-07-26 at 11 58 05 PM

  • Note: Some twitter phrases in dataset may contain inappropriate language

  • Citations

    - Limsopatham, N., & Collier, N. (2015). Adapting Phrase-based Machine Translation to Normalise Medical Terms in Social Media Messages [Data set]. Conference on Empirical Methods in Natural Language Processing (EMNLP 2015), Lisboa, Portugal. Zenodo. https://doi.org/10.5281/zenodo.27354
    

ADE Corpus V2

  • Source: Click here to proceed to site

    • Entire zip file is also available here
  • Data Type: Free Text

  • Possible uses: Build a model to extract information regarding Adverse Drug Effects from medical text

  • Size of files:

    • 2.3MB for ADE-NEG.txt
    • 1.4MB for DRUG-AE.rel
    • 60KB for DRUG-DOSE.rel
  • Example of data + Special Instructions(if any):

    Download relevant files from Source

    Example:

    ADE-NEG.txt

    ADE-NEG.txt provides all sentences in the ADE corpus that DO NOT contain any drug-related adverse effects.

    ADE-NEG.txt: Screenshot 2022-07-22 at 10 20 24 PM

    DRUG-AE.rel

    DRUG-AE.rel provides relations between drugs and adverse effects.

    DRUG-AE.rel: Screenshot 2022-07-22 at 10 21 12 PM

    The format of DRUG-AE.rel is as follows with pipe delimiters:

    • Column-1: PubMed-ID
    • Column-2: Sentence
    • Column-3: Adverse-Effect
    • Column-4: Begin offset of Adverse-Effect at 'document level'
    • Column-5: End offset of Adverse-Effect at 'document level'
    • Column-6: Drug
    • Column-7: Begin offset of Drug at 'document level'
    • Column-8: End offset of Drug at 'document level'

    DRUG-DOSE.rel

    DRUG-DOSE.rel provides relations between drugs and dosages.

    DRUG-DOSE.txt Screenshot 2022-07-22 at 10 22 11 PM

    The format of DRUG-DOSE.rel is as follows with pipe delimiters:

    • Column-1: PubMed-ID
    • Column-2: Sentence
    • Column-3: Dose
    • Column-4: Begin offset of Dose at 'document level'
    • Column-5: End offset of Dose at 'document level'
    • Column-6: Drug
    • Column-7: Begin offset of Drug at 'document level'
    • Column-8: End offset of Drug at 'document level'

  • Note: NA

  • Citations

    - If you use this corpus for any publication purposes, you are requested to cite the source article:
    
     Gurulingappa et al., Benchmark Corpus to Support Information Extraction for Adverse Drug Effects, JBI, 2012.
     http://www.sciencedirect.com/science/article/pii/S1532046412000615
    
    

CDR BioCreative V Chemical Disease Relation corpus

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: Disease named entity recognition and chemical-induced disease relation extraction

  • Size: 11.2MB

  • Example of data + Special Instructions(if any):

    Download folder from CDR section from Source.

    Example

    Screenshot 2022-07-25 at 5 15 23 PM

  • Note: Medical Subject Headings(MeSH) is used in this dataset

  • Citations

    - Li, J., Sun, Y., Johnson, R. J., Sciaky, D., Wei, C. H., Leaman, R., Davis, A. P., Mattingly, C. J., Wiegers, T. C., & Lu, Z. (2016). BioCreative V CDR task corpus: a resource for chemical disease relation extraction. Database : the journal of biological databases and curation, 2016, baw068. https://doi.org/10.1093/database/baw068
    

i2b2 dataset (Limited preview)

MeDAL Dataset

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: Medical Dataset for Abbreviation Disambiguation for Natural Language Understanding (MeDAL) is a large medical text dataset curated for abbreviation disambiguation, designed for natural language understanding pre-training in the medical domain. It was published at the ClinicalNLP workshop at EMNLP.

  • Size: 21.06GB

    • 15.16GB for full_data.csv
    • 5.9GB for pretrain_subset
      • 3.54GB for train.csv
      • 1.18GB for valid.csv
      • 1.18GB for test.csv
  • Example of data + Special Instructions(if any):

    Download relevant folder from Source.

    Example

    ABSTRACT_ID TEXT LOCATION LABEL
    5017844 different electrocardiographic changes have been described during thrombolytic therapy for AIM to indicate successful reperfusion the occluded coronary artery also can be reopened by percutaneous TCA ptca this study was performed to compare electrocardiographic changes during primary or rescue ptca and thrombolytic therapy the electrocardiographic changes were studied directly at the moment of reperfusion during ptca 25 transluminal coronary angioplasty
    8923525 there are limited data regarding cdx expression in rectal carcinoma the ckck immunoprofile of CRC has been described in studies which have mostly lumped Tc and rectal PT together in this study we investigated the diagnostic utility of immunohistochemical stains for ck ck and cdx in a series of rectal adenocarcinoma fiftyfive specimens of rectal adenocarcinomas were retrieved and immunostained for ck dakom ck novocastra ncllck and cdx novocastra nclcdx thirty cases of pancreatic adenocarcinoma and CC were also studied as a comparison group ck was expressed in and ck in cases of rectal adenocarcinoma the ckck immunophenotype was identified in ckck in and ckck in rectal adenocarcinoma cdx showed moderatestrong positivity in all cases and was not related to RT differentiation benign rectal mucosa was available in cases and showed the following results ckck in ckck in and ckck in cases in pancreatic adenocarcinomas and cholangiocarcinomas were ckck and were ckck cdx was positive in only of these cases all were pancreatic adenocarcinomas in conclusion ck can be expressed in rectal adenocarcinoma and should not be used as the sole basis for excluding a rectal primary cdx is a CS marker for rectal origin of adenocarcinoma it can be helpful in cases with metastatic rectal carcinoma especially those with ckck or ckck immunophenotype in this study cdx expression was not influenced by the grade differentiation of rectal adenocarcinoma 14 colorectal adenocarcinoma

  • Note: The files might take some time to open due to its large size

  • Citations

    - Wen, Z., Lu, X., & Reddy, S. (2020). MeDAL: Medical Abbreviation Disambiguation Dataset for Natural Language Understanding Pretraining. In Proceedings of the 3rd Clinical Natural Language Processing Workshop (pp. 130–135). Association for Computational Linguistics.
    

MedMCQA

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: AI for answering medical questions

  • Size: 151.6MB

  • Example of data + Special Instructions(if any):

    Download folder from Source.

    Example

    Screenshot 2022-07-27 at 12 38 03 AM

  • Note: NA

  • Citations

    - Pal, A., Umapathi, L., & Sankarasubbu, M. (2022). MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering. In Proceedings of the Conference on Health, Inference, and Learning (pp. 248–260). PMLR.
    

MedMentions

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: NLP research on Biomedical text/Build a tool to recognise biomedical concepts in long medical texts

  • Size: 22.5MB

  • Example of data + Special Instructions(if any):

    Download corpus_pubtator.txt.gz from Source.

    Example

    Screenshot 2022-07-25 at 2 53 03 PM

    Format:

    PMID | t | Title text

    PMID | a | Abstract text

    PMID TAB StartIndex TAB EndIndex TAB MentionTextSegment TAB SemanticTypeID TAB EntityID

  • Note: The IDs are linked to UMLS concept

  • Citations

    - Sunil Mohan and Donghui Li. 2019. MedMentions: A Large Biomedical Corpus Annotated with UMLS Concepts. In Proceedings of the 2019 Conference on Automated Knowledge Base Construction (AKBC 2019). Amherst, Massachusetts, USA. May 2019 (https://arxiv.org/abs/1902.09476)
    

MedQA

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: Building models for open domain question answering(OpenQA) tasks in the medical field(This is a free-form multiple-choice OpenQA dataset for solving medical problems)

  • Size: 391.4MB

  • Example of data + Special Instructions(if any):

    Download data from Google Drive folder from Source.

    Example:

    US_qbank.jsonl from Questions US folder

    Screenshot 2022-07-25 at 3 43 35 PM

    Textbook

    There is a list of text files regarding various medical topics in this folder:

    Screenshot 2022-07-25 at 3 42 13 PM

    Example of abstract from Anatomy_Gray.txt:

    Screenshot 2022-07-25 at 3 42 37 PM

  • Note: NA

  • Citations

    - Jin, D., Pan, E., Oufattole, N., Weng, W.H., Fang, H., & Szolovits, P. (2020). What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams. arXiv preprint arXiv:2009.13081.
    

MedQuAD

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: Develop a tool that could answer medical questions automatically by mapping new questions to formally answered questions that are "similar"

  • Size of folders:

    • 3MB for CancerGov
    • 9.6MB for GARD
    • 6.3MB for GHR
    • 1.5MB for MPlus_Health_Topics
    • 2.3MB for NIDDK
    • 1MB for NINDS
    • 1.3MB for SeniorHealth
    • 1.9MB for NHLBI
    • 455KB for CDC
    • 5.4MB for MPlus_ADAM
    • 2.8MB MPlusDrugs
    • 175KB for MPlusHerbsSupplements
  • Example of data + Special Instructions(if any):

    Download relevant files from Source. There are 12 categories available.

    Example:

    CancerGov_QA

    Screenshot 2022-07-25 at 1 55 16 PM

  • Note: NA

  • Citations

    - Asma Ben Abacha, & Dina Demner-Fushman (2019). A Question-Entailment Approach to Question Answering. BMC Bioinform., 20(1), 511:1–511:23. (https://bmcbioinformatics.biomedcentral.com/articles/10.1186/s12859-019-3119-4)
    

NCBI Disease Corpus

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: Disease name recognition and normalization for biomedical research

  • Size of files:

    • 1.3MB for NCBI Disease Corpus(Mention Level)
    • 1MB for NCBI Disease Corpus(Complete - Train set)
    • 178KB for NCBI Disease Corpus(Complete - Development set)
    • 189KB for NCBI Disease Corpus(Complete - Test set)
  • Example of data + Special Instructions(if any):

    Download relevant files from Source

    Example:

    NCBI Disease Corpus (Mention Level)

    Training: Screenshot 2022-07-25 at 4 55 14 PM

    NCBI Disease Corpus(Complete - Train set)

    Screenshot 2022-07-25 at 4 56 27 PM

  • Note: Medical Subject Headings (MeSH) and Online Mendelian Inheritance in Man (OMIM) are used in this dataset

  • Citations

    - Rezarta Islamaj Doğan, Robert Leaman, & Zhiyong Lu (2014). NCBI disease corpus: A resource for disease name recognition and concept normalization. Journal of Biomedical Informatics, 47, 1-10.
    

Patient Physician Medical Interviews

  • Source: Click here to proceed to site

  • Data Type: Audio + Free Text

  • Possible uses: Speech recognition detection for speech-to-text errors/ Training NLP models to extract symptoms, detect diseases/ For educational purposes such as training an avatar to converse with healthcare professional students as a standardized patient during clinical examinations

  • Size: 1.05 GB

  • Example of data + Special Instructions(if any):

    Download from source

    Example

    CAR0001.mp3 in Audio Recordings will correspond to CAR0001.txt in Clean Transcripts

    (File was converted to mp4 here)

    CAR0001.mp4

    Click here to expand full clean transcript

        D: What brought you in today?
    
        P: Sure, I'm I'm just having a lot of chest pain and and so I thought I should get it checked out.
    
        D: OK, before we start, could you remind me of your gender and age? 
    
        P: Sure 39, I'm a male.
    
        D: OK, and so when did this chest pain start?
    
        P: It started last night, but it's becoming sharper.
    
        D: OK, and where is this pain located? 
    
        P: It's located on the left side of my chest.
    
        D: OK, and, so how long has it been going on for then if it started last night?
    
        P: So I guess it would be a couple of hours now, maybe like 8.
    
        D: OK. Has it been constant throughout that time, or uh, or changing? 
    
        P: I would say it's been pretty constant, yeah.
    
        D: OK, and how would you describe the pain? People will use words sometimes like sharp, burning, achy. 
    
        P: I'd say it's pretty sharp, yeah.
    
        D: Sharp OK. Uh, anything that you have done tried since last night that's made the pain better?
    
        P: Um not laying down helps.
    
        D: OK, so do you find laying down makes the pain worse?
    
        P: Yes, definitely.
    
        D: OK, do you find that the pain is radiating anywhere?
    
        P: No.
    
        D: OK, and is there anything else that makes the pain worse besides laying down? 
    
        P: Not that I've noticed, no.
    
        D: OK, so not like taking a deep breath or anything like that?
    
        P: Maybe taking a deep breath. Yeah.
    
        D: OK. And when the pain started, could you tell me uh, could you think of anything that you were doing at the time?
    
        P: I mean, I was moving some furniture around, but, that I've done that before.
    
        D: OK, so you didn't feel like you hurt yourself when you were doing that?
    
        P: No.
    
        D: OK, and in regards to how severe the pain is on a scale of 1 to 10, 10 being the worst pain you've ever felt, how severe would you say the pain is?
    
        P: I'd say it's like a seven or eight. It's pretty bad.
    
        D: OK, and with the pain, do you have any other associated symptoms?
    
        P: I feel a little lightheaded and I'm having some trouble breathing.
    
        D: OK. Have you had any loss of consciousness?
    
        P: No.
    
        D: OK. Uh, have you been experiencing any like racing of the heart? 
    
        P: Um, a little bit, yeah.
    
        D: OK. And have you been sweaty at all?
    
        P: Just from the from having issues breathing.
    
        D: OK, have you been having issues breathing since the pain started? 
    
        P: Yes.
    
        D: OK. Um recently have you had any periods of time where you like have been immobilized or or, you haven't been like able to move around a lot?
    
        P: No no. 
    
        D: OK. And have you been feeling sick at all? Any infectious symptoms? 
    
        P: No. 
    
        D: OK, have you had any nausea or vomiting? 
    
        P: No. 
    
        D: Any fevers or chills?
    
        P: No. 
    
        D: OK, how about any abdominal pain?
    
        P: No.
    
        D: Any urinary problems?
    
        P: No.
    
        D: Or bowel problems?
    
        P: No.
    
        D: OK, have you had a cough?
    
        P: No.
    
        D: OK. You haven't brought up any blood?
    
        P: No. 
    
        D: OK, have you had a wheeze with your difficulty breathing?
    
        P: No, not that I've heard.
    
        D: OK, any changes to the breath sounds at all like any noisy breathing?
    
        P: No. Well, I guess if when I'm really having trouble breathing, yeah.
    
        D: OK. Has anything like this ever happened to you before?
    
        P: No. 
       
        D: No, OK. And have you had any night sweats?
       
        P: No
       
        D: Alright, and then how about any rashes or skin changes?
       
        P: No rashes, but I guess like my neck seems to be a little swollen. 
    
        D: OK, do you have any neck pain? 
    
        P: No. 
    
        D: OK, have you had any like accidents like a car accident or anything where you really jerked your neck? 
    
        P: No.
    
        D: OK. Um any any trauma at all to the chest or or back?
    
        P: No.
    
        D: OK, so just in regards to past medical history, do you have any prior medical conditions? 
    
        P: No.
    
        D: OK, any recent hospitalizations?
    
        P: No. 
    
        D: OK, any prior surgeries?
    
        P: No.
    
        D: OK, do you take any medications regularly? Are they prescribed or over the counter?
    
        P: No. 
    
        D: Alright, how about any allergies to medications? 
    
        P: None.
    
        D: Alright, any immunizations or are they up to date? 
    
        P: They are all up to date.
    
        D: Excellent. Alright, and could you tell me a little bit about your living situation currently? 
    
        P: Sure, I live in an apartment by myself. I, uh, yep, that's about it.
    
        D: OK, and how do you support yourself financially?
    
        P: I'm an accountant.
    
        D: OK, sounds like a pretty stressful job or that it can be. Do you smoke cigarettes? 
    
        P: I do.
    
        D: OK, and how much do you smoke?
    
        P: I smoke about a pack a day.
    
        D: OK, how long have you been smoking for?
    
        P: For the past 10 to 15 years.
    
        D: OK, and do you smoke cannabis?
    
        P: Uh sometimes. 
    
        D: Uh, how much marijuana would you smoke per per week? 
    
        P: Per week, maybe about 5 milligrams. Not that much.
    
        D: OK, and do you use any other recreational drugs like cocaine, crystal, meth, opioids?
    
        P: No.
    
        D: OK. Have you used IV drugs before?
    
        P: No. 
    
        D: OK. And do you drink alcohol?
    
        P: I do.
    
        D: OK. How much alcohol do you drink each week?
    
        P: Uhm about I would say I have like one or two drinks a day, so about 10 drinks a week. 
    
        D: OK, uh, yeah and um alright, and then briefly, could you tell me a little bit about your like diet and exercise?
    
        P: Sure, I try to eat healthy for dinner at least, but most of my lunches are, uh I eat out. And then in terms of exercise, I try to exercise every other day, I run for about half an hour.
    
        D; OK, well that's great that you've been working on the the activity and the diet as well. So has anything like this happened in your family before?
    
        P: No. 
    
        D: OK, has anybody in the family had a heart attack before?
    
        P: Actually, yes, my father had a heart attack when he was 45.
    
        D: OK, and anybody in the family have cholesterol problems?
    
        P: I think my father did.
    
        D: I see OK, and how about anybody in the family have a stroke?
    
        P: No strokes.
    
        D: OK, and then any cancers in the family?
    
        P: No.
    
        D: OK, and is there anything else that you wanted to tell me about today that that I on on history? 
    
        P: No, I don't think so. I think you asked me everything.
    

  • Note: First 3 characters of the file name represent the categories the case belong to.

    • Respiratory cases (designated “RES”)
    • Musculoskeletal cases (designated “MSK”)
    • Cardiac cases (designated “CAR”)
    • Dermatological case (designated “DER”)
    • Gastrointestinal cases (designated “GAS”)
  • Citations

    - Smith, Christopher William; Fareez, Faiha; Parikh, Tishya; Wavell, Christopher; Shahab, Saba; Chevalier, Meghan; et al. (2022): Collection of simulated medical exams. figshare. Dataset. 
    

PubMed 200k RCT dataset

  • Source: Click here to proceed to site

  • Data Type: Free Text

  • Possible uses: Develop an algorithm that can summarise a long medical abstract accurately

  • Size of folders (each folder has a train.txt, test.txt and dev.txt):

    • 39.2MB for PubMed_20k_RCT
    • 38.6MB for PubMed_20k_RCT_numbers_replaced_with_at_sign
    • 367.1MB for PubMed_200k_RCT
    • 361MB for PubMed_200k_RCT_numbers_replaced_with_at_sign
  • Example of data + Special Instructions(if any):

    Download relevant files from Source

    Example

    Screenshot 2022-07-22 at 9 14 59 PM

  • Note: NA

  • Citations

    - Franck Dernoncourt, Ji Young Lee. PubMed 200k RCT: a Dataset for Sequential Sentence Classification in Medical Abstracts. International Joint Conference on Natural Language Processing (IJCNLP). 2017. https://arxiv.org/abs/1710.06071
    

Other Diseases

Chula RBC-12-Dataset

  • Source: Click here to proceed to site

  • Data Type: Images + Text/Tabular

  • Possible uses: Build a model to classify type of red blood cell disease from red blood cell images

  • Size: Depends. One image is about 79KB.

    • 73.5MB for Dataset folder
    • 216KB for Labels folder
  • Example of data + Special Instructions(if any):

    Download images in Dataset folder and corresponding text files in Label folder from Source.

    You could also download the entire folder by downloading the zip folder as shown below

    Screenshot 2022-07-18 at 10 23 20 PM

    Example: 1.jpg in Dataset folder will correspond to 1.txt in Label folder

    Screenshot 2022-07-13 at 6 38 07 PM

    Each row contains 3 values which corresponds to the x coordinate, y coordinate and the type of red blood cell found at the coordinate. There are 13 types as follows:

    Number Type of Red Blood Cell
    0 Normal cell
    1 Macrocyte
    2 Microcyte
    3 Spherocyte
    4 Target cell
    5 Stomatocyte
    6 Ovalocyte
    7 Teardrop
    8 Burr cell
    9 Schistocyte
    10 uncategorised
    11 Hypochromia
    12 Elliptocyte

  • Note: There are other images that correspond to other red blood diseases that can be found in the RBC Diseases folder from Source.

  • Citations

    - Naruenatthanaset, K., Chalidabhongse, T. H., Palasuwan, D., Anantrasirichai, N., &amp; Palasuwan, A. (2021, November 3). Red blood cell segmentation with overlapping cell separation and classification on Imbalanced Dataset. arXiv.org. Retrieved July 14, 2022, from https://arxiv.org/abs/2012.01321   
    

Fetal Health Classification

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Build a model to classify fetal health in order to prevent child and maternal mortality

  • Size: 229KB

  • Example of data + Special Instructions(if any):

    Download from Source

    Example

    baseline value accelerations fetal_movement uterine_contractions light_decelerations severe_decelerations prolongued_decelerations abnormal_short_term_variability mean_value_of_short_term_variability percentage_of_time_with_abnormal_long_term_variability mean_value_of_long_term_variability histogram_width histogram_min histogram_max histogram_number_of_peaks histogram_number_of_zeroes histogram_mode histogram_mean histogram_median histogram_variance histogram_tendency fetal_health
    120 0 0 0 0 0 0 73 0.5 43 2.4 64 62 126 2 0 120 137 121 73 1 2
    132 0.006 0 0.006 0.003 0 0 17 2.1 0 10.4 130 68 198 6 1 141 136 140 12 0 1
    133 0.003 0 0.008 0.003 0 0 16 2.1 0 13.4 130 68 198 5 1 141 135 138 13 0 1

  • Note: NA

  • Citations

    - Ayres de Campos et al. (2000) SisPorto 2.0 A Program for Automated Analysis of Cardiotocograms. J Matern Fetal Med 5:311-318(https://onlinelibrary.wiley.com/doi/10.1002/1520-6661(200009/10)9:5%3C311::AID-MFM12%3E3.0.CO;2-9)
    

Disease-symptom associations generated by an automated method based on information in textual discharge summaries of patients at New York Presbyterian Hospital admitted during 2004

  • Source: Click here to proceed to site

  • Data Type: Tabular

  • Possible uses: Build a model to predict disease from symptoms

  • Size: 84KB

  • Example of data + Special Instructions(if any):

    You can download the data here

    Example

    Screenshot 2022-07-27 at 2 55 51 PM

  • Note: NA

  • Citations

    - Wang X, Chused A, Elhadad N, Friedman C, Markatou M. Automated knowledge acquisition from clinical narrative reports. AMIA Annu Symp Proc. 2008 Nov 6;2008:783-7. PMID: 18999156; PMCID: PMC2656103.
    

About

A collection of public healthcare datasets

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors