Hi guys,

Some of you probably already know this, but just in case you don't: Google recently released their (parsed, tokenized and counted) data that they use to train their language models ( http://googleresearch.blogspot.com/2006/08/all-our-n-gram-are-belong-to-you.html). Some quick numbers:

Number of tokens:    1,024,908,267,229
Number of sentences:    95,119,665,584
Number of unigrams:         13,588,391
Number of bigrams:         314,843,401
Number of trigrams:        977,069,902
Number of fourgrams:     1,313,818,354
Number of fivegrams:     1,176,470,663

Yes, that's greater than a billion 5 grams! (Wonder how long our poor little count.pl will take to chomp through that kind of data!)

Happy n-gram-ing
Bano
__._,_.___


SPONSORED LINKS
School education New york university school education New york university school of education
Natural language Computational linguistics

Your email settings: Individual Email|Traditional
Change settings via the Web (Yahoo! ID required)
Change settings via email: Switch delivery to Daily Digest | Switch to Fully Featured
Visit Your Group | Yahoo! Groups Terms of Use | Unsubscribe

__,_._,___

Reply via email to