資料載入中.....
|
請使用永久網址來引用或連結此文件:
https://ir.lib.ncu.edu.tw/handle/987654321/106896
|
| 題名: | On mining incomplete medical datasets: Ordering imputation and classification |
| 作者: | 柯士文;Chen, Chih-Wen;Lin, Wei-Chao;Ke, Shih-Wen;Tsai, Chih-Fong;Hu, Ya-Han |
| 貢獻者: | 管理學院資訊管理學系 |
| 關鍵詞: | Algorithms;Data Accuracy;Data Interpretation, Statistical;Data Mining - methods;Data Mining - standards;Humans;Support Vector Machine |
| 日期: | 2015-01-01 |
| 上傳時間: | 2026-04-23 13:48:07 (UTC+8) |
| 出版者: | IOS Press;London, England: SAGE Publications |
| 摘要: | 摘要: BACKGROUND: To collect medical datasets, it is usually the case that a number of data samples contain some missing values. Performing the data mining task over the incomplete datasets is a difficult problem. In general, missing value imputation can be approached, which aims at providing estimations for missing values by reasoning from the observed data. Consequently, the effectiveness of missing value imputation is heavily dependent on the observed data (or complete data) in the incomplete datasets. OBJECTIVE: In this paper, the research objective is to perform instance selection to filter out some noisy data (or outliers) from a given(complete) dataset to see its effect on the final imputation result. Specifically, four different processes of combining instance selection and missing value imputation are proposed and compared in terms of data classification. METHODS: Experiments are conducted based on 11 medical related datasets containing categorical, numerical, and mixed attribute types of data. In addition, missing values for each dataset are introduced into all attributes (the missing data rates are 10%, 20%, 30%, 40%, and 50%). For instance selection and missing value imputation, the DROP3 and k-nearest neighbor imputation methods are employed. On the other hand, the support vector machine (SVM) classifier is used to assess the final classification accuracy of the four different processes. RESULTS: The experimental results show that the second process by performing instance selection first and imputation second allows the SVM classifiers to outperform the other processes. CONCLUSIONS: For incomplete medical datasets containing some missing values, it is necessary to perform missing value imputation. In this paper, we demonstrate that instance selection can be used to filter out some noisy data or outliers before the imputation process. In other words, the observed data for missing value imputation may contain some noisy information, which can degrade the quality of the imputation result as well as the classification performance. 其他題名: Technol Health Care 出版者: London, England: SAGE Publications 出版日期: 2015-01-01 出處: Technology and health care, 2015-01, Vol.23 (5), p.619-625 資源來源: Academic Search Premier (Ebsco) 版權: IOS Press and the authors. All rights reserved 識別號: ISSN: 0928-7329 識別號: ISSN: 1878-7401 識別號: EISSN: 1878-7401 識別號: DOI: 10.3233/THC-151018 識別號: PMID: 26410122 |
| 顯示於類別: | [資訊管理學系] 期刊論文
|
文件中的檔案:
| 檔案 |
描述 |
大小 | 格式 | 瀏覽次數 |
| index.html | | 0Kb | HTML | 19 | 檢視/開啟 |
|
在NCUIR中所有的資料項目都受到原著作權保護.
|