中大機構典藏-NCU Institutional Repository-提供博碩士論文、考古題、期刊論文、研究計畫等下載:Item 987654321/8640
English  |  正體中文  |  简体中文  |  Items with full text/Total items : 78818/78818 (100%)
Visitors : 34695625      Online Users : 1121
RC Version 7.0 © Powered By DSPACE, MIT. Enhanced by NTU Library IR team.
Scope Tips:
  • please add "double quotation mark" for query phrases to get precise results
  • please goto advance search for comprehansive author search
  • Adv. Search
    HomeLoginUploadHelpAboutAdminister Goto mobile version


    Please use this identifier to cite or link to this item: http://ir.lib.ncu.edu.tw/handle/987654321/8640


    Title: 中文資料擷取系統之設計與研究;Mining Relevant Syntactic Patterns for Chinese Text Extraction
    Authors: 吳東軒;Dong-shun Wu
    Contributors: 資訊工程研究所
    Keywords: 資訊擷取系統;iformation extraction
    Date: 2002-06-20
    Issue Date: 2009-09-22 11:32:13 (UTC+8)
    Publisher: 國立中央大學圖書館
    Abstract: 在資訊化時代的今日,大量的資料正在慢慢的電子化中,再加上網際網路的蓬勃發展,新的資訊正每天不斷地在網路的脈絡中流動及累積,面對隨著時間不斷地增加的訊息,要從中尋得個人所需的資訊是相當困難的。以電子化的優點,從龐大的資料中,利用電腦快速且精準地找到我們所需的資訊,這正是資訊擷取(Information Extraction)的精神所在。 資訊擷取系統在英文的處理方面,已經發展有一段時間了,但是對於中文的處理方面,仍然有很大的發展空間。由於中文文法中,句型結構相對於英文來說是較為鬆散的,因此中文資訊擷取系統很難利用英文資訊擷取系統中常使用的句型分析來幫助資訊的擷取。在本篇論文中,我們針對中文的純文字資料的擷取問題,提出了一套流程,希望透過這一流程,順利地從中文純文字資料中,擷取出我們所需的資訊。 IE is a research topic related to TREC (Text Retrieval Conference) and MUC (Message Understanding Conference). The target of Information extraction (IE) is to extract specific types of information from text. The IE systems for free text form written in English are different from the systems for Chinese. In this paper we propose a simple method for extracting information from free text from written in Chinese. We use training examples and encode them with the responding targets. Then we find the repeated substrings within the encoded text. These repeated substrings play the role in our IE system for Chinese which is likes the role of the sentence analyzers in some IE systems for free text form in English. In the phrase for extracting information from testing data, we first encode them and then extract the interesting target by the repeated substrings fined previously.
    Appears in Collections:[Graduate Institute of Computer Science and Information Engineering] Electronic Thesis & Dissertation

    Files in This Item:

    File SizeFormat


    All items in NCUIR are protected by copyright, with all rights reserved.

    社群 sharing

    ::: Copyright National Central University. | 國立中央大學圖書館版權所有 | 收藏本站 | 設為首頁 | 最佳瀏覽畫面: 1024*768 | 建站日期:8-24-2009 :::
    DSpace Software Copyright © 2002-2004  MIT &  Hewlett-Packard  /   Enhanced by   NTU Library IR team Copyright ©   - 隱私權政策聲明