篇名 | Noisy Channel Models for Corrupted Chinese Text Restoration and GB-to-Big5 Conversion |
---|---|
卷期 | 3:2 |
作者 | Chang, Chao-huang |
頁次 | 079-091 |
關鍵字 | THCI Core |
出刊日期 | 199808 |
In this article, we propose a noisy channel/information restoration model for error recovery problems in Chinese natural language processing. A language processing system is considered as an information restoration process executed through a noisy channel. By feeding a large-scale standard corpus C into a simulated noisy N. Using N as the input to the
language processing system (i.e., the information restoration process), we can obtain the output results C'. After that, the automatic evaluationmodule compares the original corpus C and the output results C', and computes the performance index (i.e., accuracy) automatically. The proposed model has been applied to two common and important problems related to Chinese NLP for the Internet: corrupted Chinese text restoration and GB-to-BIG5 conversion. Sinica Corpora version 1.0 and 2.0 are used in the experiment. The results show that the proposed model is useful and practical.