Recently, a Stanford researcher and a Google X team
made headlines with their
study on unsupervised learning. Starting with one of the biggest neural networks ever built (1 billion connections and 16,000 CPU cores), they fed in a still frame from each of 10 million YouTube videos and watched it learn to identify human faces and cats.
But why aren't experiments on a similar scale being done with text data? I can think of three reasons why they should be.