Jump to content

Factored language model: Difference between revisions

From Wikipedia, the free encyclopedia
Content deleted Content added
References: Improving a reference
Monkbot (talk | contribs)
m Task 18 (cosmetic): eval 1 template: del empty params (1×); hyphenate params (1×);
 
Line 6: Line 6:


==References==
==References==
*{{cite conference | author=J Bilmes and K Kirchhoff | url=http://ssli.ee.washington.edu/people/bilmes/mypapers/hlt03.pdf | title=Factored Language Models and Generalized Parallel Backoff | booktitle=Human Language Technology Conference | pages= | year=2003 | archive-url=https://web.archive.org/web/20120717075838/http://ssli.ee.washington.edu/people/bilmes/mypapers/hlt03.pdf | archive-date=17 July 2012}}
*{{cite conference | author=J Bilmes and K Kirchhoff | url=http://ssli.ee.washington.edu/people/bilmes/mypapers/hlt03.pdf | title=Factored Language Models and Generalized Parallel Backoff | book-title=Human Language Technology Conference | year=2003 | archive-url=https://web.archive.org/web/20120717075838/http://ssli.ee.washington.edu/people/bilmes/mypapers/hlt03.pdf | archive-date=17 July 2012}}


[[Category:Language modeling]]
[[Category:Language modeling]]

Latest revision as of 02:17, 1 December 2020

The factored language model (FLM) is an extension of a conventional language model introduced by Jeff Bilmes and Katrin Kirchoff in 2003. In an FLM, each word is viewed as a vector of k factors: An FLM provides the probabilistic model where the prediction of a factor is based on parents . For example, if represents a word token and represents a Part of speech tag for English, the expression gives a model for predicting current word token based on a traditional Ngram model as well as the Part of speech tag of the previous word.

A major advantage of factored language models is that they allow users to specify linguistic knowledge such as the relationship between word tokens and Part of speech in English, or morphological information (stems, root, etc.) in Arabic.

Like N-gram models, smoothing techniques are necessary in parameter estimation. In particular, generalized back-off is used in training an FLM.

References

[edit]
  • J Bilmes and K Kirchhoff (2003). "Factored Language Models and Generalized Parallel Backoff" (PDF). Human Language Technology Conference. Archived from the original (PDF) on 17 July 2012.