Carrot and stick
We said it at the end of 3.3. So far there's been one way to teach an AI: show it the answers.
That works for a cat photo, because you can write "cat" on it. But what about riding a bike? Can you write down "lean 3.7 degrees right now"?
No. And yet people learn to ride bikes. By falling off. That's the bike you asked a grown-up about in 1.2 — nobody wrote the answer down then either.
At first it really does play at random. But after a hundred games or so, winning gets hard.
Nobody wrote it a single rule. Not "the middle is good," not "block when they get two in a row." All it ever experienced was winning and losing.
Rewards instead of answers
Teaching this way is called .
You never say what the right answer is. Instead you reward what works and penalise what doesn't, and let it shift towards whatever earns more.
In the matchbox machine, beads were the reward. Beads added after a win, beads binned after a loss.
That "reward" has a name too: .
Here's where this differs from the checkers program in 1.2. Checkers fixed itself by how wrong it was (the error). The matchboxes fix themselves by whether they won or lost (reward and punishment).
| Every method so far | Reinforcement learning | |
|---|---|---|
| What a person gives | The answers | A reward |
| When you'd use it | You know the answer | Nobody knows the answer |
| What the machine does | Get closer to the answer | Collect more reward |
Keep an eye on that last row. The whole chapter comes out of it.
A neural network instead of beads
Matchboxes worked because there are only 304 board layouts. That count treats boards that turn or flip into the same shape as one board. The machine on the screen doesn't fold them together, so it has more boxes. Driving a car, though, has more situations than anyone could count. And a real road has other cars, people, even the weather.
For that you swap the beads for the neural network from 3.2, and build each new generation out of whichever ones did best.
At first all twenty crash. A few generations later, some last much longer.
Tap one car. All it knows is five distances ahead of it. It can't see the shape of the track and doesn't know where it is. That's enough to work out how to go round.
Reward the wrong thing
Here's the real story of this chapter.
Those three sliders decide what gets rewarded. Push them to the extremes, one at a time.
Set Going fast to 100 and everything else to 0. Then press Learn again from scratch.
The cars crash far sooner. Nothing is measuring how long they last, so flooring it and hitting a wall pays.
Now set only Not crashing to 100.
Almost none of them move. Standing perfectly still is the very best way to not crash.
They did exactly as they were told
Both times, the machine found precisely the way to collect the most reward.
We thought we said "drive well." What we actually handed it was "go fast, get reward." So it went fast, and nothing else.
An AI doesn't learn what you wanted. It learns how to collect the reward you gave it.
It's the twin of 3.3. There, the photos we gave it sent it wrong. Here it's the reward. Both start with something a person handed over, not with the machine.
The road through Part 3. One box, one chapter.
That's Part 3
Three ways of learning: showing the answers, rewarding what works, and all the ways it goes wrong.
But everything so far has been AI that looks — photos, points, tracks.
Part 4 moves to AI that deals in words. The ones you run into most these days.
But words aren't like photos. Photos were numbers — so how do you turn writing into numbers?
Sources for this chapter
- 10 Michie, MENACE (1961) · "Experiments on the Mechanization of Game-Learning", The Computer Journal 6(3) (1963)