Professional Documents
Culture Documents
Lec-3 Language Modeling N-Grams
Lec-3 Language Modeling N-Grams
Lec-3 Language Modeling N-Grams
P( w1w2 wn ) P( wi | w1w2 wi 1 )
i
•Simplifying assumption:
Andrei Markov
•or maybe
P( the | its water is so transparent that ) P( the | transparent that )
Markov Assumption
P( w1w2 wn ) P( wi | wi k wi 1 )
i
P( w1w2 wn ) P( wi )
i
Some automatically generated sentences from a unigram model
https://sites.google.com/view/ell881-iitd/home
Estimating N-gram
Probabilities
Estimating bigram probabilities
• The Maximum Likelihood Estimate (MLE)
count(wi-1,wi )
P(wi | w i-1) =
count(w i-1 )
c(wi-1,wi )
P(wi | w i-1 ) =
c(wi-1)
An example
• Result:
Bigram estimates of sentence probabilities
P(<s> I want english food </s>) =
P(I|<s>)
× P(want|I)
× P(english|want)
× P(food|english)
× P(</s>|food)
= .000031
What kinds of knowledge?
• P(english|want) = .0011 world
• P(chinese|want) = .0065
• P(to|want) = .66
grammar
• P(eat | to) = .28
• P(food | to) = 0 grammar (contingent zero)
• P(want | spend) = 0 grammar (structural zero)
• P (i | <s>) = .25 The right to food, and its non variations, is a human
right protecting the right for people to feed themselves in dignity
Practical Issues
For bigrams: