In our meeting on Thursday Feb 24 there was some discussion of the British National Corpus versus the Brown Corpus. Prath located a brief description of BNC (which is the source of word distributional data that the McCarthy paper is based on) which you can find below.
========================================================= http://www.natcorp.ox.ac.uk/what/index.html The British National Corpus is a very large (over 100 million words) corpus of modern English, both spoken and written. The Corpus is designed to represent as wide a range of modern British English as possible. The written part (90%) includes, for example, extracts from regional and national newspapers, specialist periodicals and journals for all ages and interests, academic books and popular fiction, published and unpublished letters and memoranda, school and university essays, among many other kinds of text. The spoken part (10%) includes a large amount of unscripted informal conversation, recordeded by volunteers selected from different age, region and social classes in a demographically balanced way, together with spoken language collected in all kinds of different contexts, ranging from formal business or government meetings to radio shows and phone-ins. ------------------------ Yahoo! Groups Sponsor --------------------~--> Help save the life of a child. Support St. Jude Children's Research Hospital's 'Thanks & Giving.' http://us.click.yahoo.com/i8TXDC/5WnJAA/HwKMAA/x3XolB/TM --------------------------------------------------------------------~-> Yahoo! Groups Links <*> To visit your group on the web, go to: http://groups.yahoo.com/group/nlpatumd/ <*> To unsubscribe from this group, send an email to: [EMAIL PROTECTED] <*> Your use of Yahoo! Groups is subject to: http://docs.yahoo.com/info/terms/

