TY - JOUR
T1 - A hybrid language model based on a combination of N-grams and stochastic context-free grammars
AU - Linares, Diego
AU - Benedí, José Miguel
AU - Sánchez, Joan Andreu
PY - 2004/6
Y1 - 2004/6
N2 - In this paper, a hybrid language model is defined as a combination of a word-based n-gram, which is used to capture the local relations between words, and a category-based stochastic context-free grammar (SCFG) with a word distribution into categories, which is defined to represent the long-term relations between these categories. The problem of unsupervised learning of a SCFG in General Format and in Chomsky Normal Form by means of estimation algorithms is studied. Moreover, a bracketed version of the classical estimation algorithm based on the Earley algorithm is proposed. This paper also explores the use of SCFGs obtained from a treebank corpus as initial models for the estimation algorithms. Experiments on the UPenn Treebank corpus are reported. These experiments have been carried out in terms of the test set perplexity and the word error rate in a speech recognition experiment.
AB - In this paper, a hybrid language model is defined as a combination of a word-based n-gram, which is used to capture the local relations between words, and a category-based stochastic context-free grammar (SCFG) with a word distribution into categories, which is defined to represent the long-term relations between these categories. The problem of unsupervised learning of a SCFG in General Format and in Chomsky Normal Form by means of estimation algorithms is studied. Moreover, a bracketed version of the classical estimation algorithm based on the Earley algorithm is proposed. This paper also explores the use of SCFGs obtained from a treebank corpus as initial models for the estimation algorithms. Experiments on the UPenn Treebank corpus are reported. These experiments have been carried out in terms of the test set perplexity and the word error rate in a speech recognition experiment.
KW - Language model
KW - Stochastic context-free grammar
UR - http://www.scopus.com/inward/record.url?scp=10044298606&partnerID=8YFLogxK
U2 - 10.1145/1034780.1034783
DO - 10.1145/1034780.1034783
M3 - Article
AN - SCOPUS:10044298606
SN - 1530-0226
VL - 3
SP - 113
EP - 127
JO - ACM Transactions on Asian Language Information Processing
JF - ACM Transactions on Asian Language Information Processing
IS - 2
ER -