Some of you probably already know this, but just in case you don't: Google recently released their (parsed, tokenized and counted) data that they use to train their language models ( http://googleresearch.blogspot.com/2006/08/all-our-n-gram-are-belong-to-you.html). Some quick numbers:
Number of tokens: 1,024,908,267,229
Number of sentences: 95,119,665,584
Number of unigrams: 13,588,391
Number of bigrams: 314,843,401
Number of trigrams: 977,069,902
Number of fourgrams: 1,313,818,354
Number of fivegrams: 1,176,470,663
Yes, that's greater than a billion 5 grams! (Wonder how long our poor little count.pl will take to chomp through that kind of data!)
Happy n-gram-ing
Bano
__._,_.___
SPONSORED LINKS
| School education | New york university school education | New york university school of education |
| Natural language | Computational linguistics |
Your email settings: Individual Email|Traditional
Change settings via the Web (Yahoo! ID required)
Change settings via email: Switch delivery to Daily Digest | Switch to Fully Featured
Visit Your Group | Yahoo! Groups Terms of Use | Unsubscribe
Change settings via the Web (Yahoo! ID required)
Change settings via email: Switch delivery to Daily Digest | Switch to Fully Featured
Visit Your Group | Yahoo! Groups Terms of Use | Unsubscribe
__,_._,___

