On the Feature Selection Methods and Reject Option Classifiers for Robust Cancer Prediction

Abstract

Cancer is the second leading cause of mortality across the globe. Approximately 9.6 million people are estimated to have died due to cancer disease in 2019. Accurate and early prediction of cancer can assist healthcare professionals to devise timely therapeutic innervations to control sufferings and the risk of mortality. Generally, a machine learning (ML) based predictive system in healthcare uses data (genetic profile or clinical parameters) and learning algorithms to predict target values for cancer detection. However, optimization of predictive accuracy is an important endeavor for accurate decision making. Reject Option (RO) classifiers have been used to improve the predictive accuracy of classifiers for cancer like complex problems. In a gene profile all of the features are not important and should be shaved off. ML offers different techniques with their own methodology for feature selection (FS) and the classification results are dependent on the datasets each having its own distribution and features. Therefore, both FS methods and ML algorithms with RO need to be considered for robust classification. The main objective of this study is to optimize three parameters (learning algorithm, FS method and rejection rate) for robust cancer prediction rather than considering two traditional parameters (learning algorithm and rejection rate). The analysis of different FS methods (including t-Test, Las Vegas Filter (LVF), Relief, and Information Gain (IG)) and RO classifiers on different rejection thresholds is performed to investigate the robust predictability of cancer. The three cancer datasets (Colon cancer, Leukemia and Breast cancer) were reduced using different FS methods and each of them were used to analyze the predictability of cancer using different RO classifiers. The results reveal that for each dataset predictive accuracies of RO classifiers were different for different FS methods. The findings based on proposed scheme indicate that, the ML algorithms along with their dependence on suitable FS methods need to be taken into consideration for accurate prediction.

Publication DOI: https://doi.org/10.1109/ACCESS.2019.2944295
Additional Information: © 2019. This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see http://creativecommons.org/licenses/by/4.0/ .
Uncontrolled Keywords: Cancer,classification,feature selection,genetic profile,machine learning,reject option,Computer Science(all),Materials Science(all),Engineering(all)
Publication ISSN: 2169-3536
Last Modified: 27 Mar 2024 08:25
Date Deposited: 02 Dec 2022 16:25
Full Text Link:
Related URLs: https://ieeexpl ... ocument/8851142 (Publisher URL)
PURE Output Type: Article
Published Date: 2019-10-09
Published Online Date: 2019-09-27
Authors: Waseem, Muhammad Hammad
Nadeem, Malik Sajjad Ahmed
Abbas, Assad
Shaheen, Aliya
Aziz, Wajid
Anjum, Adeel
Manzoor, Umar (ORCID Profile 0000-0001-7602-1914)
Balubaid, Muhammad A.
Shim, Seong O.

Download

[img]

Version: Published Version

License: Creative Commons Attribution

| Preview

Export / Share Citation


Statistics

Additional statistics for this record