Here is a source of 1.5 million small (32 x 32) images with English labels. This might be an interesting large-ish dataset for clustering and learning experiments.
The total size of the data set is pretty small (only 3+ GB) so it shouldn't be such a big deal to snag. http://people.csail.mit.edu/torralba/tinyimages/ -- ted
