Loading slide

Loading contents...

[░░░░░░░░░░░░░░░░░░][░░░░░░░░░░░░░░░░░░░░░░░░░░░░]0 / 15
<back>next

Module 5 Chapter 3

The ImageNet Turning Point

Fields need a way to settle arguments. Without one, everybody's method is the best method, and progress becomes a matter of who is most persuasive at conferences.

What changed things for machine vision was mundane: a very large pile of photographs, labelled by hand, with the answers to part of it kept secret. Anyone could train on the public part, and everyone was scored on the hidden part, which meant nobody could win by memorising. It sounds like bookkeeping. It was the thing that made the results believable.

Then in 2012 an entry won by a margin that did not look like an ordinary year of progress. It was far enough ahead that the reasonable response was not to be impressed but to be suspicious, and because the benchmark was public, suspicion was cheap to act on. It held up.

Within a couple of years the field had changed direction almost entirely. What is easy to miss is that nothing in the winning entry was new. The ideas had been sitting in the literature for years, some for decades, waiting for enough data and enough speed to arrive at the same time. Being right early is nearly indistinguishable from being wrong.

In this chapter

  • Building a giant picture collectionhow people created the labelled ImageNet dataset
  • Making systems face the same testwhy a shared hidden set made comparison meaningful
  • Reusing one detector across an imagehow the same learned pattern can scan many locations
  • Bringing the pieces togetherhow AlexNet combined several improvements with GPUs
  • What the 2012 result showedwhat the breakthrough proved and what it did not
# citations