中大學術數位典藏-NCU Institutional Repository-提供博碩士論文、考古題、期刊論文、研究計畫等下載:Item 987654321/107082
English  |  正體中文  |  简体中文  |  Items with full text/Total items : 94274/94274 (100%)
Visitors : 82857895      Online Users : 1781
RC Version 7.0 © Powered By DSPACE, MIT. Enhanced by NTU Library IR team.
Scope Tips:
  • please add "double quotation mark" for query phrases to get precise results
  • please goto advance search for comprehansive author search
  • Adv. Search
    HomeLoginUploadHelpAboutAdminister Goto mobile version


    Please use this identifier to cite or link to this item: https://ir.lib.ncu.edu.tw/handle/987654321/107082


    Title: The distance function effect on k-nearest neighbor classification for medical datasets
    Authors: 蔡志豐;Hu, Li-Yu;Huang, Min-Wei;Ke, Shih-Wen;Tsai, Chih-Fong
    Contributors: 管理學院資訊管理學系
    Keywords: Case Study;Classification;Computer Science;Humanities and Social Sciences;multidisciplinary;Science;Science (multidisciplinary)
    Date: 2016-12-01
    Issue Date: 2026-04-23 13:55:45 (UTC+8)
    Publisher: Springer Science and Business Media Deutschland GmbH;Cham: Springer Science and Business Media LLC
    Abstract: 摘要: Introduction K-nearest neighbor (k-NN) classification is conventional non-parametric classifier, which has been used as the baseline classifier in many pattern classification problems. It is based on measuring the distances between the test data and each of the training data to decide the final classification output. Case description Since the Euclidean distance function is the most widely used distance metric in k-NN, no study examines the classification performance of k-NN by different distance functions, especially for various medical domain problems. Therefore, the aim of this paper is to investigate whether the distance function can affect the k-NN performance over different medical datasets. Our experiments are based on three different types of medical datasets containing categorical, numerical, and mixed types of data and four different distance functions including Euclidean, cosine, Chi square, and Minkowsky are used during k-NN classification individually. Discussion and evaluation The experimental results show that using the Chi square distance function is the best choice for the three different types of datasets. However, using the cosine and Euclidean (and Minkowsky) distance function perform the worst over the mixed type of datasets. Conclusions In this paper, we demonstrate that the chosen distance function can affect the classification accuracy of the k-NN classifier. For the medical domain datasets including the categorical, numerical, and mixed types of data, K-NN based on the Chi square distance function performs the best.
    其他題名: SpringerPlus
    其他題名: Springerplus
    出版者: Cham: Springer Science and Business Media LLC
    出版日期: 2016-08-09
    出處: SpringerPlus, 2016-08, Vol.5 (1), p.1304-1304, Article 1304
    資源來源: Agricultural & Environmental Science Collection
    版權: The Author(s) 2016
    版權: SpringerPlus is a copyright of Springer, 2016.
    識別號: ISSN: 2193-1801
    識別號: EISSN: 2193-1801
    識別號: DOI: 10.1186/s40064-016-2941-7
    識別號: PMID: 27547678
    Appears in Collections:[Department of Information Management] journal & Dissertation

    Files in This Item:

    File Description SizeFormat
    index.html0KbHTML38View/Open


    All items in NCUIR are protected by copyright, with all rights reserved.

    社群 sharing

    ::: Copyright National Central University. | 國立中央大學圖書館版權所有 | 收藏本站 | 設為首頁 | 最佳瀏覽畫面: 1024*768 | 建站日期:8-24-2009 :::
    DSpace Software Copyright © 2002-2004  MIT &  Hewlett-Packard  /   Enhanced by   NTU Library IR team Copyright ©   - 隱私權政策聲明