頁籤選單縮合
題 名 | Noisy Channel Models for Corrupted Chinese Text Restoration and GB-to-Big5 Conversion |
---|---|
作 者 | Chang,Chao-huang; | 書刊名 | International Journal of Computational Linguistics & Chinese Language Processing |
卷 期 | 3:2 1998.08[民87.08] |
頁 次 | 頁79-91 |
分類號 | 312.23 |
關鍵詞 | |
語 文 | 英文(English) |
英文摘要 | In this article, we propose a noisy channel/information restoration model for error recovery problems in Chinese natural language processing. A language processing system is considered as an information restoration process executed through a noisy channel. By feeding a large-scale standard corpus C into a simulated noisy channel, we can obtain a noisy version of the corpus N. Using N as the input to the language processing system (i.e., the information restoration process), we can obtain the output results C". After that, the automatic evaluation module compares the original corpus C and the output results C", and computes the performance index (i.e., accuracy) automatically. The proposed model has been applied to two common and important problems related to Chinese NLP for the Internet: corrupted Chinese text restoration and GB-to-BIG5 conversion. Sinica Corpora version 1.0 and 2.0 are used in the experiment. The results show that the proposed model is useful and practical. |
本系統中英文摘要資訊取自各篇刊載內容。