# Demis Hassabis, 'Artificial Intelligence and the Future' (The RSA, London, 29 September 2016)

- Speaker: Demis Hassabis
- Event: The RSA (Royal Society of Arts), London: 'Artificial Intelligence and the Future'
- Date: Published 29 September 2016
- Length: Talk 45 min (14:20 to 59:03 in this recording). The recording has 12 minutes of silent pre-roll; the RSA host begins at 12:04 and Hassabis at 14:20. The on-stage conversation and Q&A are not in this recording.
- Video URL: https://www.youtube.com/watch?v=i3lEG6aRGm8 (RSA Replay) ; archive: https://archive.org/details/Artificial_Intelligence_and_the_Future_Demis_Hassabis
- Slides: Yes, a conventional lab deck of roughly 40 slides: DeepMind mission, Atari and AlphaGo results, charts, applications.
- Transcript source: Whisper (faster-whisper, model medium) on the archive.org audio; the archive.org ASR caption file was used as a second check. Timestamps are from the archive.org audio (add nothing to match the RSA Replay YouTube upload only if that upload also carries the pre-roll; otherwise subtract about 12:04).
- Corrections: Corrections applied: the speaker's name where the host says it, and 'Lee Sedol'.
- Format: `[mm:ss]` is the time the line starts in the source recording. Machine transcript: expect small word errors; quotes used in the analysis were checked against the published text where one exists.

---

[12:04] Good evening everyone. I'm Matthew Taylor. I'm Chief Executive of the RSA and it's my enormous
[12:10] pleasure to welcome you here for this evening's special event. Can you make sure your mobile
[12:14] phone is switched to silent? We're filming tonight and live streaming over the web so
[12:19] welcome to everyone who's watching online. The hashtag is RSAAI if you'd like to get involved
[12:27] in the conversation on Twitter. So housekeeping notice is over. It's my enormous pleasure
[12:33] to introduce this evening's distinguished guest speaker,
[12:36] Dr. Demis Hisavis.
[12:37] Dennis is the co-founder and CEO of DeepMind,
[12:41] a neuroscience-inspired AI company
[12:43] acquired by Google in 2014,
[12:45] and he leads the general AI efforts of Google,
[12:50] including the development of AlphaGo,
[12:51] the first program to beat a professional player
[12:54] at the game of Go.
[12:56] Demis is a former child chess and coding prodigy
[12:58] on graduating from Cambridge.
[13:00] He founded the pioneering video games company,
[13:02] Alexia Studios. After a decade leading successful technology startups, he returned to academia
[13:07] to complete a PhD in cognitive neuroscience at UCL, followed by post-docs at MIT and Harvard.
[13:13] His research, connecting memory with imagination, was listed in the top 10 scientific breakthroughs
[13:17] of 2007 by the journal Science. He's also five times world games champion, the recipient
[13:21] of the Royal Society's prestigious Mollard Award, and I'm delighted to say, a fellow
[13:25] of the RSA.
[13:28] Demis has described his DeepMind project
[13:30] as, quote, an Apollo program for the 21st century.
[13:33] And in his talk for us this evening,
[13:35] he'll offer us some of the latest insights
[13:37] in the frontiers of artificial intelligence research.
[13:41] Just as we are here at the RSA,
[13:43] Demis and his team are interested
[13:45] in how we harness the incredible,
[13:46] accelerating power of technology
[13:48] for the benefit of all humanity, not just an elite.
[13:51] And tonight, we're delighted that Demis will focus
[13:53] on AI's potential to help solve
[13:55] some of our greatest global challenges
[13:57] from healthcare to climate change.
[14:01] Demis is going to speak for 20, 25 minutes.
[14:04] We'll have a conversation up here
[14:05] and then we'll bring you in for questions.
[14:07] So without further ado,
[14:09] please join me in welcoming Dr. Demis Asavis.
[14:11] Thanks Matthew for that introduction.
[14:21] It's a great pleasure to be here.
[14:23] I'll always love visiting the Royal Society of Arts
[14:26] and it's great to have these kinds of dialogues
[14:29] between the sciences and the arts.
[14:31] So I'm going to talk about artificial intelligence
[14:34] and what's happening at the cutting edge
[14:35] of artificial intelligence, through the lens of the work
[14:38] that we're doing at DeepMind.
[14:41] So DeepMind was founded in 2010,
[14:44] and we joined forces with Google in 2014.
[14:48] And one way you can think about DeepMind,
[14:50] one way I'd describe it, is as an Apollo program
[14:53] effort for AI.
[14:54] And what that means, what we mean by that,
[14:56] is to bring together the world's best scientists
[14:59] and best engineers and put them in a perfect environment
[15:02] surrounded by all the resources they need
[15:03] to try and make as quick and a rapid progress on the topic
[15:07] of AI as possible.
[15:09] So we're around 350 people now and over 250 research
[15:13] scientists working at DeepMind.
[15:15] So it's probably the biggest collection
[15:17] anywhere in the world of brain power focused on this issue.
[15:23] But we're not just researching AI
[15:26] in terms of pushing the boundaries of AI research.
[15:29] Another thing that we're sort of experimenting on
[15:31] is also a kind of new way of organizing science.
[15:35] And I haven't got time to sort of go into that today.
[15:37] That would be a whole other lecture in itself.
[15:39] But the summary is really trying
[15:42] to create a hybrid organization that
[15:44] combines the best from the sort of top Silicon Valley
[15:47] startup culture, the focus, and the buzz,
[15:49] and the energy they have, with the best
[15:52] from academic institutes when they're operating very well,
[15:56] kind of blue sky thinking, interdisciplinary work,
[15:59] and creative thinking that they encourage,
[16:01] and trying to meld together the best
[16:03] from both of those worlds.
[16:06] So our mission, the way we describe it normally
[16:08] is sort of articulate it as a kind of two-step process.
[16:12] So step one, fundamentally solve intelligence.
[16:15] And then step two, we think that it would naturally
[16:17] follow, if you did step one, that you
[16:19] could use that power of that technology
[16:21] to practically help us as a society solve everything else.
[16:27] Now, that might seem a little bit fanciful and far-fetched,
[16:30] But I hope by the end of this talk,
[16:34] I'll have convinced you that at least it's plausible
[16:36] that this could happen.
[16:39] So in terms of approaches to AI, I
[16:41] think there's at least four dimensions that it's
[16:43] worth thinking about AI in.
[16:46] And then the different types of AI and the different ways
[16:50] different industrial groups and academic institutes
[16:53] approach trying to build AI kind of spans
[16:56] these four dimensions.
[16:58] So the first and probably most important one
[17:00] is the idea of learning systems, systems
[17:03] that learn directly from their experience or from raw data,
[17:06] versus handcrafted heuristic systems.
[17:09] So systems where they've been specifically
[17:12] pre-programmed with a particular solution
[17:14] to a problem.
[17:16] So there's learning versus handcrafted.
[17:19] Second dimension is the idea of generality
[17:22] versus a special casing or specific purpose.
[17:26] So what many systems are is they're handcrafted
[17:30] and they're built for one particular purpose in mind.
[17:33] What we're interested in at DeepMind
[17:34] is the idea of generality.
[17:37] One system that out of the box
[17:39] can do a wide range of tasks.
[17:44] The third dimension that we think a lot about
[17:45] at DeepMind is the idea of groundedness
[17:48] versus logic-based.
[17:49] And what we mean here is that
[17:52] we think that for a true thinking machine
[17:54] to be able to think about and achieve high-end tasks,
[17:59] they need to be grounded in a sensory motor reality.
[18:02] So they have to experience the world around them
[18:04] through their senses and ground the knowledge
[18:07] that they acquire grounded in this sensory motor experience.
[18:12] And opposed to that, a logic-based systems
[18:14] or symbolic systems that are encoded.
[18:18] And the problem with those systems
[18:19] is when they interact with the outside world,
[18:23] with real events, they find it very difficult
[18:26] to map the logic or knowledge they have
[18:28] to these real world messy situations
[18:31] that they find themselves in.
[18:33] And in fact, a lot of the AI research
[18:35] that went on in the 80s and 90s at places like MIT
[18:38] were actually based around these logic-based systems
[18:42] and expert systems.
[18:44] And they only got so far.
[18:45] And then they couldn't encode things
[18:47] like a common sense because they weren't ultimately
[18:51] grounded in perceptual experience.
[18:56] And then the fourth dimension we think a lot about
[18:59] is the idea of active learning versus passive observation.
[19:04] So a lot of AI systems that you'll be using today,
[19:06] that you use every day, like image recognition or voice
[19:08] recognition, they're kind of passive observation systems.
[19:12] They get in some data, and then they try and classify
[19:14] that data in some sense.
[19:17] What we're interested in instead is active agents.
[19:20] So agents that have a goal in mind,
[19:23] that actually have actions they undertake
[19:26] and direct their own learning.
[19:28] So actually decide what they should explore.
[19:34] So at DeepMind, we are committed to the left-hand side
[19:39] on all four of these dimensions, these continuum.
[19:43] And so what we're interested in building
[19:45] is general purpose learning systems.
[19:48] So systems that are at their heart,
[19:51] they learn how to do things and master tasks
[19:54] rather than being given the solution
[19:56] by the team of programmers.
[19:58] And this kind of AI is often dubbed
[20:00] artificial general intelligence to distinguish it from the
[20:04] normal everyday AI that's out there.
[20:08] Now, one of the biggest and obviously most famous sort of
[20:11] watershed moments in AI history was when in the late
[20:14] 90s, IBM's Deep Blue beat Garry Kasparov.
[20:19] And of course, this is a huge technical feat, very
[20:23] impressive technical feat.
[20:25] But actually, Deep Blue was an example of one of these
[20:28] what we would call narrow AI systems that were special
[20:31] cased and pre-programmed for one particular purpose.
[20:34] In this case, playing chess.
[20:37] And the way Deep Blue worked was by a team of very smart
[20:40] programmers working hand in hand with a team of very
[20:43] smart chess abilities in chess to end up playing the game
[20:47] of chess very well.
[20:49] But Deep Blue, although it was incredible at chess, it
[20:52] was not able to do anything else.
[20:54] So not even, for example, play a much simpler game like
[20:58] Noughts and Crosses.
[20:59] It would have to be, none of the knowledge that it had
[21:03] was useful for anything else.
[21:05] It would have to be pre-programmed
[21:06] with a whole new set of rules and heuristics
[21:08] in order to do anything else.
[21:10] So in some sense, this was slightly unsatisfying
[21:13] from an AI point of view.
[21:15] In a way, is that really intelligence?
[21:18] Where does the intelligence of that software reside?
[21:20] It's not really in the machine or the algorithm,
[21:22] it's actually in the mind of the programmers
[21:24] who solved that problem.
[21:27] So we wanted to go sort of beyond this kind of narrow AI
[21:31] and work on this kind of general purpose learning system.
[21:37] So we came up and pioneered with a technique called
[21:39] deep reinforcement learning, which has now become
[21:42] all the rage in AI research.
[21:45] And we've been working on this sort
[21:47] of since the beginning, since the founding of DeepMind.
[21:49] And it's sort of combining two techniques together,
[21:51] deep learning, which is hierarchical neural networks
[21:55] that are used to perceive the world around them.
[21:58] So this is what if you interact with any image recognition
[22:02] through photo search on your phone or voice recognition
[22:05] by speaking into your phone, it will most likely
[22:07] be using deep learning systems at the back end.
[22:11] We combine that with another technique
[22:13] called reinforcement learning,
[22:14] which is about trying to select the right action
[22:17] from the set of available options to you at that moment
[22:20] that will best get you towards its goal.
[22:23] So we combine these two techniques together
[22:25] and scale them up.
[22:28] Now, we used computer games as a training platform
[22:32] for developing and testing our AI algorithms.
[22:35] Partly that was my background from computer games.
[22:37] That's what I used to do prior to DeepMind,
[22:40] was design and program AI for computer games.
[22:44] And I realized that virtual environments, including games,
[22:49] are a much more efficient way
[22:51] to test out the capabilities of your AI systems
[22:53] than, say, using something like real world robotics,
[22:58] which are much slower and messier and more
[22:59] expensive to deal with.
[23:01] So we're very interested in robotics as an application,
[23:03] but we don't usually use it as a development platform.
[23:06] We use these virtual environments.
[23:09] And you can see why they'd be much more efficient,
[23:11] because you can obviously run millions of simulations
[23:14] at once in the cloud and train your algorithms
[23:18] on a much more fast iteration.
[23:20] So it gives us much quicker feedback
[23:22] as to whether these algorithms are working.
[23:25] Another important thing about games
[23:26] is that when you're on a very long-term mission,
[23:29] like we are, a multi-decade mission,
[23:31] then it's even more important to know in the short term
[23:34] that you're heading in the right direction.
[23:36] And because games lend themselves very nicely
[23:38] to measuring progress in the terms of scores,
[23:41] or it's often very easy to have a proxy to a score
[23:44] if there isn't a direct score,
[23:46] you can know very quickly if the changes you've
[23:48] made to the AI algorithms are actually
[23:51] heading you in the right direction.
[23:55] So we started with probably the most
[23:56] iconic of the game consoles that really started off the
[24:00] whole gaming boom, which was the Atari 2600
[24:04] computer from the 80s.
[24:07] And there's lots of very good emulators now for these.
[24:10] And we took one of those open source
[24:11] emulators and sped it up.
[24:14] And then we used it as our proving ground for our
[24:16] initial algorithms.
[24:18] So we took 50 classic 8-bit games.
[24:21] For those of you old enough in the audience who might
[24:23] remember some of these, some of the most iconic games
[24:26] ever, like Space Invaders, these kinds of games.
[24:28] And before, I just want to show you a video of how these
[24:31] agents perform.
[24:32] But before I do that, I just want to explain what it is
[24:34] you're going to see.
[24:35] So the agents here, all they get as their input are the
[24:38] raw pixels on the screen.
[24:40] So it's around 30,000 numbers per frame, because
[24:43] the screen is 250 by 150 pixels in size.
[24:47] And the goal here that we set the agent is to
[24:50] maximize the score.
[24:51] Everything else is learned completely from scratch, from
[24:53] first principles.
[24:55] So it doesn't have any idea what it's controlling, what
[24:57] gets it points, what the rules of the game are.
[25:00] It doesn't even have any idea about how video streams work.
[25:03] So the idea that pixels next to each other
[25:05] are correlated in time.
[25:06] It has to find and learn about all this structure for
[25:09] itself from first principles by experimenting and
[25:12] experiencing the game, by playing the game.
[25:15] And then we add additional constraint that we want one
[25:18] system to play all the different games out of the
[25:21] box, so the same system with the same parameter
[25:23] settings can actually master all these 50 different games,
[25:27] some of them very, very different in terms of their
[25:29] objectives and the way they look.
[25:31] So this is the notion, again, of generality coming in.
[25:36] So I'll show you this one.
[25:36] I've only got time today to show you one video.
[25:38] And I'll show you my favorite one, which is from
[25:40] a game called Breakout.
[25:42] Our algorithm's called DQN.
[25:44] And in this game Breakout, the player plays.
[25:48] You control this pink bat at the bottom of the screen.
[25:51] And there's this little pink ball, this pixel ball,
[25:53] that's bouncing off this rainbow-colored wall.
[25:56] And what you've got to try and do is knock out
[25:58] the bricks from the rainbow-colored wall, brick by brick,
[26:00] but not let the ball go past your bat.
[26:02] If the ball goes past your bat, you lose a life.
[26:05] And what I'm gonna show you now as I roll the video
[26:07] is the agent getting better over time
[26:09] after it's played a certain number of games
[26:11] and it will carry on improving
[26:13] as it gets more and more experience with the game.
[26:17] So this is after 100 games.
[26:19] So it's not very good yet at playing the game,
[26:21] but it's started to get the hang of the idea
[26:23] that it should move this bat towards the ball,
[26:26] and that letting the ball go past the bat
[26:28] is probably a bad idea.
[26:30] Now, after 300 games, and like another couple hundred games,
[26:33] now it's almost perfectly mastered the game
[26:36] and can kind of play it as well as any human could play
[26:38] this, and it gets the ball back most of the time,
[26:41] even if it's coming back at very fast angles.
[26:44] So we thought that was pretty cool,
[26:45] but what would happen if we let the agent play
[26:49] for another 200 games?
[26:50] And what it did was this unexpected,
[26:52] It found that the best strategy was to dig a tunnel
[26:55] around the left-hand side of the wall
[26:57] and send the ball right around the back,
[26:59] and with some kind of incredible efficiency and accuracy,
[27:02] sort of superhuman accuracy.
[27:04] And obviously realized that was the most effective strategy
[27:08] with the least risk, and so would aim for that strategy
[27:11] from the beginning of the game.
[27:13] And the funny story about that is that
[27:15] obviously the researchers who created DQN
[27:18] are absolutely amazing researchers and programmers,
[27:21] but they're not so good at playing these Atari games
[27:23] themselves.
[27:24] So they actually learned something about,
[27:25] they didn't know about this strategy,
[27:27] and they learned something from their own system.
[27:29] So that's pretty funny.
[27:31] If you think about the power of general learning systems
[27:34] like this, they can actually master things, complex things,
[27:38] that even the programmers don't necessarily
[27:41] know how to codify.
[27:44] So then we took this a big step further,
[27:47] these kinds of systems.
[27:48] And earlier this year, we created
[27:50] a program called AlphaGo.
[27:52] And AlphaGo is a program to play the ancient game of Go.
[27:57] Now, for those of you who don't know, this is what a Go board
[27:59] looks like.
[28:00] There are two sides, black and white.
[28:03] And players take turns to put these stones on the
[28:07] vertices of a 19 by 19 board.
[28:10] So this is what the board looks like.
[28:11] And the board initially starts empty.
[28:13] And then it starts filling up with stones.
[28:16] And the aim of the game is to surround your opponent's
[28:18] stones, with your ones, or to wall off empty parts of the
[28:23] board and surround empty parts of the board and capture it as
[28:25] territory.
[28:27] And the aim of the game is to, at the end of the game, it
[28:30] looks something like this, where the board has been
[28:31] mostly filled up.
[28:33] And then you count up the amount of territory that white
[28:36] and black have walled off with their stones.
[28:39] And you add that to the number of capture stones that
[28:43] you've taken.
[28:44] And the player with the most points wins the game.
[28:47] So in this case, it was a very close game,
[28:49] but white would win this game by one point after the count up.
[28:54] Now the thing about Go is it has only two rules,
[28:57] incredibly simple.
[28:57] I could teach it to you in sort of five minutes.
[29:00] But it leads to the most profound complexity.
[29:03] So the history of Go is actually a long and storied
[29:05] one.
[29:05] So it's over 3,000 years old.
[29:06] It's invented in China.
[29:08] And it's considered to be much more than just a game.
[29:11] In China, and ancient China especially,
[29:13] it was considered to be something more
[29:15] into poetry or music. So a kind of art form. And in fact, it almost had a spiritual sort
[29:22] of dimension in terms of embodying the mysteries of the universe within this game. And Confucius
[29:29] wrote about it as one of the four arts that had to be mastered for any true scholar.
[29:35] So for 3,000 years, there's been this massive tradition of play and go. And today
[29:39] it's more as popular as ever. It's 40 million active players and 2,000 professionals.
[29:44] And in fact, in Japan and Korea and China, where basically they play Go instead of in
[29:50] the West where we play chess, there are professional Go schools that if you're a kid who's six
[29:57] or seven or eight years old and you're spotted to have talent at your normal school, then
[30:02] you'll be taken out of your normal school and put into one of these Go schools to
[30:05] play Go 12 hours a day, seven days a week, sort of live and breathe Go all through
[30:10] your adolescence.
[30:12] So it's kind of amazing how seriously they take this
[30:15] and how important this game is to the culture
[30:18] of these Asian countries.
[30:21] And one way to just quickly illustrate
[30:23] the complexity of the game that arises out
[30:25] of these very simple, elegant rules
[30:27] is the fact that there are more possible board configurations
[30:31] in Go than there are atoms in the universe.
[30:34] So there's 10 to the power 170 possible positions in Go.
[30:39] And there are only around 10 to the power 80 atoms
[30:43] in the observable universe.
[30:44] That's estimated.
[30:46] So there's no way you can brute force calculate
[30:49] what's going to happen in Go.
[30:50] Even if you took all the compute power in the world
[30:52] and ran it for a million years,
[30:54] that wouldn't be enough compute power
[30:56] to enumerate all the possibilities in Go.
[31:01] So that explains why Go is so much harder for computers
[31:04] to play than something like chess.
[31:06] It's because this brute force search
[31:09] that was used to master chess is not tractable for Go.
[31:14] And this breaks down into two main challenges.
[31:17] Firstly, this search space of possibilities is really huge.
[31:20] In Go, in an average position, there are 200 possible moves.
[31:24] Whereas in chess, in an average position,
[31:26] there are about 20.
[31:27] So Go has an order of magnitude larger branching factor
[31:31] that's called in computer chess.
[31:33] So Go is an order of magnitude more complex
[31:36] from that point of view.
[31:38] But the second challenge, which is even harder,
[31:40] is that it was thought until AlphaGo came along
[31:42] that it was impossible to write
[31:44] what's called an evaluation function
[31:46] to determine who is winning in a particular position,
[31:49] whether black or white is winning and by how much.
[31:52] Now in chess, it's relatively easy
[31:54] to write an evaluation function
[31:55] because chess has a concept of materiality.
[31:58] So a queen in chess is worth nine points
[32:01] and a rook is five and knight is three and so on.
[32:04] So if you add up the pieces on both sides of the board
[32:08] And you count how many points there are.
[32:10] At a first approximation, that will indicate
[32:12] who is winning the game.
[32:14] And so that's already a good heuristic for an evaluation
[32:17] function.
[32:17] And there are many others that you can codify that build
[32:20] up and taken together can give you a very accurate picture
[32:22] of who's winning.
[32:24] In Go, all the pieces are the same.
[32:26] They're just stones.
[32:28] So there's no concept of materiality.
[32:30] So there's no shortcut to figuring out who is winning.
[32:33] The other problem with Go is that it's a
[32:35] constructive game.
[32:37] The board starts empty, and it fills up.
[32:39] So if you're trying to evaluate a complex mid-game position,
[32:42] you also have to project into the future
[32:45] about what it might happen.
[32:46] Whereas in chess, chess is a sort of destructive game,
[32:49] where in the sense that the board starts off
[32:52] full with all the pieces.
[32:54] And as the game goes on, it simplifies.
[32:56] Pieces come off the board.
[32:58] So at any moment, you can evaluate the position right
[33:00] now, and that has the full information.
[33:02] You don't have to sort of project into the future
[33:05] as to how the board might fill up.
[33:08] So to get around on this complexity,
[33:10] the way humans get around it is that they use their intuition
[33:14] and their instinct.
[33:15] So in fact, Go is a much more intuitive game than chess.
[33:18] Chess is much more about calculation
[33:20] and enumerating the possibilities.
[33:22] Go is much more about feel and instinct and intuition.
[33:26] And in fact, if you ask a top Go player why
[33:29] it is they made a particular move in a complex position,
[33:31] often they'll tell you it just felt right.
[33:33] Whereas a great chess player will never say that.
[33:35] They'll tell you, you know, I'm planning A because I thought B was going to happen and
[33:39] then I would do C. Now that plan may not pan out.
[33:42] It may not be a good plan, but they'll usually have an explicit plan in mind.
[33:46] Whereas with Go players, it's much more intuitive and implicit.
[33:51] So these things, so even that idea of intuition, trying to codify that, as you
[33:56] all know, intuition is not normally associated with computers, whereas calculation is.
[34:02] And that's one of the reasons why it makes Go so much harder.
[34:06] So to get around these two challenges, we turned to neural networks, and we trained
[34:11] two neural networks, one each to get over each of these two challenges.
[34:16] So the first network we trained was called a policy network, and what this did was we
[34:21] downloaded 100,000 games from an internet Go server from strong amateurs playing Go
[34:26] on the internet.
[34:28] And we tried to train a neural network to predict in any position what a human player
[34:33] would do next.
[34:34] move would they make next?
[34:37] And what that means is that instead of having to look at all the 200 possibilities every
[34:41] time you're in a position of legal moves, you can actually just narrow it down to
[34:46] the top three or top four most probable moves and concentrate on looking at those.
[34:53] The second neural network is what allowed us to have this fabled evaluation function.
[34:59] We call it the value network.
[35:01] And the way we created this is we took that first network and we had it playing
[35:04] against itself 30 million times to create 30 million training games.
[35:10] And with each game, we know who wins, we know all the intermediate positions.
[35:15] And from that, the system, AlphaGo, learns over time how to predict the end result from
[35:21] any of the earlier positions.
[35:25] And over time, after each of these games it's playing, it gets incrementally slightly
[35:29] better each time.
[35:30] It learns from its mistakes and its prediction errors that it makes.
[35:33] And each time it slightly improves itself, almost
[35:36] imperceptibly with each game.
[35:38] But after a few million games, you end up with a highly
[35:42] accurate evaluation function that's sort of as good as any
[35:46] human can evaluate positions.
[35:49] So armed with these two neural networks, this cuts
[35:51] down that huge search space I was talking about to
[35:54] something much more tractable.
[35:56] So it was time to sort of challenge some of the top
[35:59] players in the world.
[36:01] And in March earlier this year, we challenged a $1 million
[36:06] challenge match against Lee Sedol, who is a South Korean
[36:10] grandmaster and considered to be the best player of the
[36:13] past decade.
[36:14] So he's a kind of complete legend of the game.
[36:17] He's won 18 world titles.
[36:19] And I like describing him as the sort of Roger Federer of
[36:21] Go.
[36:22] So he's sort of been at the top for the last 10
[36:25] years, won the most grand slams, but is still at the
[36:27] top even now.
[36:30] And so AlphaGo, we challenged Lisa Dole, and it was a huge
[36:35] deal in Asia, and especially in Korea and China, where the
[36:39] whole country actually came to sort of a standstill
[36:42] watching this match.
[36:43] Here are some pictures of these two pictures here of
[36:47] the press conferences we had.
[36:48] Absolutely completely filled these huge
[36:50] ballrooms in the hotel.
[36:52] I think it was the biggest ever press conference that
[36:54] Google had ever had.
[36:56] I think it was 1,000 journalists or something.
[36:58] There was also, at one point during the middle of the match,
[37:03] it was on every single national TV station live.
[37:06] So literally, we were in the hotel
[37:07] flicking through all the TV stations,
[37:09] and they were all covering AlphaGo.
[37:12] And it was even on these jumbo screens
[37:15] in the shopping centers and other things.
[37:17] It was quite an amazing experience, actually.
[37:20] And many of the Go world and the AI world
[37:25] had thought that it was at least gonna be
[37:27] another decade before a program like AlphaGo came along that
[37:30] could master the game of Go and beat the world's best players
[37:33] because of these complexities I've just talked about.
[37:36] So it's thought to be at least a decade away.
[37:38] And even on the evening of the match when they interviewed
[37:40] Lee Sedol, he felt confident he would win 5-0.
[37:45] And even losing one game would be unthinkable for him.
[37:50] But in the end, as some of you will know, AlphaGo
[37:53] actually emerged victorious, winning 4-1, and therefore
[37:58] causing a huge splash in both the Go worlds and the AI
[38:02] worlds a decade before its time.
[38:04] And if you want to read about the details of how the
[38:08] algorithm works, we published a front cover article in
[38:11] Nature that details out all the scientific advances we
[38:14] had to do to create AlphaGo.
[38:18] Now, the culture of the impact of the match was huge.
[38:21] 280 million viewers watched the five games overall,
[38:26] 60 million in China just for game one.
[38:28] And so there's more viewers, I think, than the Super Bowl.
[38:32] 35,000 press articles were written about Alpha Go
[38:35] in this match over that week.
[38:37] And the thing I'm actually most pleased about
[38:39] is it really had a big breakthrough
[38:44] in terms of the consciousness in the West about this great game
[38:47] Go, and actually boosting its popularity.
[38:49] And one measure of that, that there were more than 10 times
[38:52] more boards and pieces sold online from the online places
[38:56] you can buy these than normally around that time of year.
[38:59] So I hope this is going to lead to an explosion of people
[39:02] taking up Go, especially in the West.
[39:05] Now, not only did it win this match,
[39:08] the other thing that was pretty amazing was the way
[39:11] that AlphaGo won and the kinds of strategies
[39:13] that it came up with.
[39:15] So I haven't got time to go into many.
[39:17] There were so many amazing moments.
[39:18] But I'm going to try and explain to you the significance of one of the famous moves that
[39:23] it did.
[39:24] And just as another point to the tradition of Go, because the game's so intuitive and
[39:31] so revered, and the best players are like legends of the past, often as well, when
[39:39] a really amazing move is played in a very important game, that move will go down
[39:43] in legend, too, and it will get a name, and that game will get a name, and it will
[39:46] we sort of studied over hundreds of years
[39:49] by many, many students of the game.
[39:50] And we think that some of the games in this AlphaGo match
[39:54] will end up being remembered like that.
[39:56] And this is my favorite move.
[39:57] This is move 37 from game two.
[40:00] And AlphaGo here is black.
[40:02] And Lee Sedol is white.
[40:05] And it's quite early in the game.
[40:07] And AlphaGo on move 37 played this move here
[40:10] on the right-hand side, that black stone
[40:12] with the white triangle there showing you
[40:13] where it's been put down.
[40:15] Now I just want to try and explain to you in one minute
[40:17] why that was so potentially historical.
[40:22] So we'll see if this works.
[40:25] So there's two very important lines in Go
[40:28] that basically the whole game revolves around.
[40:31] So there's the third line, which is here.
[40:34] Now if you play on the third line,
[40:36] what you're telling your opponent is
[40:38] that you're trying to capture territory,
[40:41] wall off empty parts of the board
[40:42] to the side of the board.
[40:44] So this sort of side.
[40:46] If instead of that you play on the fourth line, what you're
[40:49] telling your opponent is that you're giving up territory on
[40:52] the side of the board, but instead what you're going for
[40:55] is influence and power into the center of the board.
[40:58] And the idea of playing on the fourth line is that
[41:00] that influence and power that you pick up later on
[41:04] can be converted into territory
[41:05] elsewhere on the board.
[41:07] And that territory that you get elsewhere on the board
[41:09] will make up for the territory that you're going to
[41:11] lose on that side of the board.
[41:14] And for 3,000 years, the received wisdom
[41:16] has been that the third and fourth lines are a fair trade.
[41:21] So if one player plays on the third line
[41:22] and the other player plays on the fourth line,
[41:24] the influence you get balances out the territory
[41:26] that you lose.
[41:28] And that's been the kind of proverb, if you like,
[41:30] of how to play Go, one of the key things of playing Go.
[41:34] So instead of that, what AlphaGo has done with move 37
[41:36] is it's played on the fifth line.
[41:38] So this is kind of completely unthinkable to do this.
[41:43] It's hard to explain how unthinkable this is,
[41:46] to the extent that no human Go player would even
[41:49] consider this move.
[41:51] Because the fifth line means that you're giving away
[41:54] territory from the fourth line, which is a huge amount
[41:57] more area to the side of the board.
[42:00] And we think the reason why AlphaGo likes this
[42:02] is that perhaps humans have been undervaluing influence
[42:05] in the center of the board for 3,000 years.
[42:07] And maybe it's more useful than all these human experts
[42:11] have thought.
[42:12] And it turned out in game two that this, of course,
[42:15] anyone can come up with a random new move, right?
[42:18] We could all do that too.
[42:19] But the interesting thing about Go as an art form
[42:22] is that it's kind of like objective art
[42:24] in the sense of you can say,
[42:26] oh, well, I think this move is a good move,
[42:28] but ultimately what it comes down to
[42:30] is did you win the game
[42:31] and was it related to that move?
[42:34] And in game two, this is actually what happened
[42:36] is that around 50 moves later,
[42:39] so a long time later,
[42:41] that stone ended up affecting the fighting that happened
[42:44] in the bottom left of the corner,
[42:45] the other side of the board.
[42:47] And these two stones here marked by the white triangle,
[42:49] they ended up spilling into the center of the board
[42:52] and connecting up around 50 moves later with that stone,
[42:55] that move 37.
[42:57] So it was almost like AlphaGo
[42:58] presciently seen this sequence of events
[43:01] and how that stone was gonna be positioned
[43:03] just perfectly to help that fighting
[43:05] in another part of the board.
[43:08] So, I said that that was a pretty astounding move,
[43:11] But you don't have to take my word for it.
[43:13] What I want to just show you is this quite amusing clip from
[43:16] the live commentary of the match.
[43:19] So just to explain who these people are, this is the
[43:21] English commentary stream.
[43:22] Obviously, there was the Chinese, and the Korean, and
[43:24] the Japanese commentary stream.
[43:25] So this is the English commentary stream.
[43:27] And these two guys are both professional Go players.
[43:30] And on the right is Michael Redmond, who is nine dan.
[43:33] Nine dan is the highest level of Go you can get to.
[43:35] And he's the only Westerner to have ever
[43:38] reached a nine dan level.
[43:40] I think he left the US when he was sort of 18
[43:43] and has lived and studied in Japan ever since.
[43:46] So he's the greatest Westerner sort of ever,
[43:48] and we had him commentating on this move.
[43:50] So let's just see his reaction to move 37.
[43:53] And bear in mind, so this is the commentary room,
[43:55] and he's watching what's going on in the game room
[43:57] via that laptop in front of him.
[44:00] The Google team was talking about
[44:04] is this kind of evaluation, a value.
[44:09] That's a very surprising move.
[44:14] I thought it was a mistake.
[44:18] So you can see that. He's not even sure where the piece is.
[44:21] And later on, they go to Don to talk, speculate that maybe our lead programmer who entered the moves into the computer
[44:28] had misclicked and put it in the wrong place.
[44:31] That was how unthinkable that move was.
[44:34] So that was the reaction.
[44:36] the reaction and since then the whole go world has been studying all these five
[44:40] games for these nuances and now we're seeing human players, expert players
[44:45] starting to play on the fifth line and use these kinds of motifs. So Lisa
[44:52] Dole's reaction was quite interesting too. So this is a this is a view of the
[44:56] match room over there on the right and this is the top-down view of the
[44:59] actual board with the move 37 played and you can see there's our programmer
[45:03] Aja Huang on the left who's entering the moves but on the right
[45:06] There's an empty chair where Lisa doll was sitting and he basically got up and left for 15 minutes after this move
[45:12] No one know where he no one knew where he went. We thought he might have left the building
[45:16] So it turned out that maybe he just went to wash his face or something
[45:19] So it's it's it is you know, that's how sort of shocked he was, too
[45:25] But then of course this whole match spurred Lisa doll onto even greater heights and he came back in game four
[45:31] which is the game that he won, and he played his own incredible move on move 78.
[45:36] And I haven't got time to explain why this move was amazing, this white move here with the black triangle in the center of the board,
[45:41] but it was an amazing move that really surprised AlphaGo, and later on when we looked at the logs of what AlphaGo was thinking,
[45:48] it only gave this move a less than one in ten thousand chance of being, of happening.
[45:53] So it hadn't really analyzed this move at all, and it caused some error and it's an evaluation.
[45:58] So this, I think this move also will go down in history as well, most likely.
[46:02] So there's some really incredible go that was played during these matches.
[46:06] And what was great is after, you know, Lissadal was this amazing guy and we had
[46:11] several sort of dinners together during the match and before the match and after the match.
[46:15] And he told me that this AlphaGo match was one of the greatest experiences of his life and had renewed his passion for the game.
[46:21] You know, he's coming towards the end of his career now, you know, he's been playing this game for, I think he's 33,
[46:26] He's been playing for nearly 30 years.
[46:28] And this completely renewed his passion for the game
[46:32] and sort of made him feel like
[46:33] there was this huge green field spaces to explore still.
[46:37] And he was gonna start exploring that.
[46:39] And what's interesting is he's some direct quotes
[46:41] from his interviews afterwards in the Korean press.
[46:44] My thoughts have become more flexible after this game.
[46:46] Have a lot of new ideas and I expect good results.
[46:49] And I decided to more accurately predict the next move
[46:51] instead of depending on my intuition.
[46:53] So that's really interesting.
[46:54] playing against AlphaGo sort of made him think
[46:57] that he shouldn't just automatically go with his intuition
[47:00] that he's trained over 30 years,
[47:01] but should more explicitly think about new ideas.
[47:05] And in fact, he's ended up over the last few months
[47:08] since the match having an amazing win record against
[47:11] and won a couple more titles
[47:12] against the other top human opponents.
[47:16] So I'm just gonna sort of end this section a little bit.
[47:19] We're just talking about intuition and creativity.
[47:21] So I've talked quite a lot about intuition,
[47:23] but what I mean by it,
[47:24] And what I wanna just unpack a little bit,
[47:26] especially for this audience
[47:27] as we're at the Royal Society of Arts here.
[47:29] And I'm not saying these are full,
[47:31] encompass everything about intuition and creativity,
[47:34] but just sort of operationally within the domain of go,
[47:38] I think these are reasonable definitions.
[47:40] So by intuition, what I'm meaning is
[47:42] the implicit knowledge that's acquired through experience,
[47:45] but it's not consciously expressible or accessible.
[47:49] So you can't explain how you know this stuff
[47:51] to someone else,
[47:52] And you can't even explain it to yourself or access it yourself.
[47:56] So you might ask, well, if you can't explain it or access it
[47:59] consciously, how do we know it's even there?
[48:02] Well, of course, we can test the existence and the quality
[48:05] of this knowledge behaviorally by giving probing tests.
[48:09] So in Go, this is easy.
[48:10] You can just give someone a Go board and evaluate the
[48:13] quality of the output that their intuition is going to
[48:15] come up, what the kind of move they come up with is.
[48:19] Secondly, with creativity, one way to operationally define
[48:22] at would be the ability to synthesize the knowledge you currently have, these building
[48:26] blocks you have, to produce a novel or original idea in the service of some goal.
[48:33] And I think, you know, if you've seen with Move 37, I think it's pretty clear that
[48:37] AlphaGo demonstrated both of these abilities to quite a high level, albeit of course
[48:42] with the caveat that this is still a very constrained domain of the board game Go.
[48:50] So one other important aspect of our research at DeepMind is systems neuroscience and
[48:54] getting inspiration not just from mathematics and machine
[48:58] learning, but also from how the brain works.
[49:01] And we really stress the word systems,
[49:03] because what we're interested in is the algorithms
[49:05] and the representations and the architectures
[49:08] the brain uses, not necessarily the low-level biological
[49:12] details, the implementation details.
[49:14] Because obviously, the substrates are different.
[49:17] You've got computers on silicon,
[49:19] and you've got our wetware, which is made out of carbon.
[49:23] So there's no reason why necessarily the implementation
[49:26] should be the same.
[49:28] But certainly the algorithms and the architectures
[49:31] could have some important clues
[49:32] as to the capabilities the brain has
[49:34] that we would like these machines to have.
[49:37] So we talk a lot at DeepMind
[49:38] about systems neuroscience inspired AI.
[49:41] And here's a list of,
[49:43] I haven't got time to go into these,
[49:44] of areas, high level areas
[49:46] that we're actively researching now
[49:48] that are partially inspired by the latest thinking
[49:51] about how the brain accomplishes these things.
[49:54] memory, attention, concepts, planning, navigation,
[49:57] even imagination.
[49:58] And I studied quite a few of these areas,
[50:00] memory and imagination, for my PhD,
[50:02] which was in cognitive neuroscience.
[50:04] And in fact, I homed in on one of the areas of the brain
[50:07] called the hippocampus.
[50:08] The hippocampus is at the center of your brain,
[50:10] and it's well known that if you have damage
[50:12] to the hippocampus, then you become amnesic.
[50:16] But actually, the hippocampus is responsible
[50:19] and heavily involved in all of these other capabilities
[50:22] as well, not just memory.
[50:24] And one of the projects that we have ongoing at DeepMind
[50:27] and one of the ones I'm personally involved in,
[50:29] you can think of loosely as trying
[50:30] to create an artificial hippocampus with all
[50:33] the capabilities that it has to link in
[50:36] with these neural networks that we've already
[50:39] built for things like the Atari games.
[50:44] So I just want to end with by talking a little bit
[50:46] about real world applications.
[50:48] So I've talked a lot about games,
[50:51] and I've explained why we use games,
[50:52] because they're a very efficient domain to develop
[50:55] these algorithms in.
[50:56] But of course, we're not just interested in solving games
[50:59] for its own sake.
[51:00] What we want to do is actually build these general
[51:02] purpose learning algorithms in such a way that they're
[51:06] general enough that they can be transferred into the
[51:09] real world and actually apply to all sorts of really
[51:12] important problems that have high impact for the good of
[51:14] society.
[51:16] We've already announced some collaborations with the NHS
[51:20] looking at things like image recognition
[51:23] to help diagnosis of head and neck cancers.
[51:25] Also, eye retinal scans to look at macular degeneration.
[51:29] We're interested in robotics.
[51:30] And we're also interested in sort of slightly more left field
[51:33] things that you might not expect these sorts of AI
[51:36] algorithms to appear in, including, most recently,
[51:39] which I think is pretty cool,
[51:41] is our work with data centers.
[51:43] And what we did, of course, Google
[51:45] has huge data centers, some of the biggest ones
[51:47] in the world, perhaps the biggest ones in the world.
[51:50] And they consume a lot of power.
[51:52] I think the latest estimate I read somewhere
[51:54] was about 2% of the world's power
[51:56] is currently used in data centers around the world,
[51:59] by data centers, cloud computing.
[52:01] And that's, of course, only going to get more with as
[52:03] more and more of our compute power is going into the cloud.
[52:08] And what we did is we used a similar system to AlphaGo.
[52:12] But instead of playing Go, we applied it
[52:15] to the cooling systems in the data centers
[52:17] to try and increase the energy efficiency
[52:20] of these data centers.
[52:21] And what we managed to do, which was quite surprising
[52:24] to the data center engineers,
[52:25] was we managed to save 40% of the energy
[52:28] that was used by the cooling systems,
[52:31] which ends up meaning that the whole data center's
[52:33] 15% less power usage.
[52:36] And obviously, that's worth tens of millions
[52:37] of dollars a year, but it's also very good
[52:40] for the environment, of course.
[52:42] And what AlphaGo does,
[52:43] or our data center optimization system does,
[52:46] is it controls all the cooling systems, the fans,
[52:49] opening the windows, the cooling water,
[52:51] even where the compute power is being routed to
[52:55] within the data center, the calculations.
[52:57] All of those things are optimized,
[52:59] and its inputs are all the sensory data,
[53:02] the thermometers, the temperature gauges,
[53:05] and the fan speeds and so on.
[53:08] And here's the graph of what happened.
[53:11] This is the power that's being used in the data center,
[53:13] the power usage efficiency,
[53:15] and you can see the sharp spike down
[53:17] when we turn on this AI system, and now that's
[53:20] controlling the data center.
[53:21] And then when we turn it off, it spikes back up to where it
[53:25] was originally the power.
[53:27] So it makes a huge difference.
[53:29] And what we're thinking now is that, well, why don't we
[53:32] optimize something like the grid, the energy grid, at
[53:36] national scale?
[53:37] There's no need to just think about a data center.
[53:40] There must be huge inefficiencies
[53:41] also at grid scale.
[53:43] So if that's true, and we're investigating this now,
[53:46] Maybe we could save 10% of the energy consumption
[53:50] of a country, which obviously has huge implications
[53:53] for climate change and other things.
[53:56] I'm just going to end with our most recent work, which
[53:59] I thought would be fun for this audience as well,
[54:01] because it's sort of creative too.
[54:04] We also, just last couple of weeks, hot off the press,
[54:07] we announced that we now have one of our models
[54:10] called WaveNet is now the best text-to-speech system
[54:14] in the world.
[54:15] So text-to-speech systems are used for speech synthesis.
[54:18] So if you speak to your phone and it tells you back
[54:20] the result, the voice that it uses is speech synthesis.
[54:25] And most of the time, the current state of the art models
[54:27] are called concatative models.
[54:29] And what they are is that, basically,
[54:32] is that they get an actor to speak for 30 hours,
[54:34] lots of dialogue, and then they chop up
[54:37] that dialogue into syllables.
[54:39] And then when you ask it to speak something new,
[54:43] it stitches back those syllables together.
[54:45] And that's why they sound a bit warbly,
[54:48] and also why they sound a little bit robotic.
[54:53] Instead of that, WaveNet, actually our system,
[54:56] learning system, actually learns to model the raw waveform,
[54:59] the raw audio waveform, directly.
[55:01] And it actually generates these raw waveforms,
[55:03] rather than stitching together syllables.
[55:06] And a waveform looks like this, and it sounds like this.
[55:09] That's one small strip for man.
[55:18] Obviously, that's a very famous waveform.
[55:20] And that's what Neil Armstrong's voice,
[55:24] as a waveform looks like.
[55:26] And what our system was able to do
[55:28] is 50% better than these concatenative systems.
[55:32] So perfect would be a human actually reading it out
[55:37] directly, and that's the green bar.
[55:38] And the current state of the art is the red bar.
[55:41] And you've seen that WaveNet, which is the blue bar,
[55:44] reduces the distance between the best models out there
[55:47] today to perfect human level by more than 50%.
[55:51] So I'm just going to play you a couple of samples
[55:53] So you can hear for yourself the difference, hopefully.
[55:56] So there's two samples.
[55:58] One is going to be from the concatenative system, and then from WaveNet.
[56:02] And it's about a strange film, I think, because this is one of the hundred test phrases
[56:07] that Google uses to evaluate how good a system is.
[56:10] So this is the concatenative.
[56:11] The Blue Lagoon is a 1980 American romance and adventure film directed by Randall
[56:16] Kleiser. This is WaveNet. The Blue Lagoon is a 1980 American romance and adventure
[56:21] film directed by Randall Kleiser. So hopefully you agree that the second one
[56:26] is better and at least 86% of people did in our user tests so and more
[56:31] natural sounding. But just a fun thing that I haven't shown before is you
[56:37] can do use this to not only model waveforms of speech but you can also
[56:41] model waveforms of music. And one of the, you know, kind of holy grails, I guess, of AI has always
[56:47] been, can we generate music and composition? And I think we may be on the cusp of that now.
[56:54] And the beginning, this one I'm going to show you is bits of piano music. It was trained on,
[56:58] I think, a bit of Rachmaninoff and that kind of music. And it's just free form, constructing,
[57:07] creating. So we haven't really told it to, it doesn't understand a thing about long-term
[57:10] melody yet or these kinds of things. That's coming. But this is a sort of free-form composition
[57:16] that it's doing. And then the second one. So it's a bit like a drunken pianist at the
[57:41] moment, but it's creating that sound. That's not a synthesized sample. It's creating
[57:47] that sound, that piano sound you can hear, all the nuances of that sound. And it's
[57:51] creating what, to our surprise, we've just started working on this, it's already
[57:55] creating sort of musically interesting sounding compositions.
[58:00] Of course, it's changing style too quickly, but if we can make it sort of stick to one
[58:05] style for longer, we think we might have something there.
[58:09] So I just want to end by coming back to my last slide on this idea of solving intelligence
[58:16] and then using it to solve other problems.
[58:18] We really think about AI as a kind of meta-solution to all these other issues.
[58:22] And if we can sort of fundamentally solve AI and make intelligence abundant in this
[58:27] learning fashion, then we think there are all sorts of areas where we can apply this
[58:31] to.
[58:32] The two I'm most excited about are science and healthcare, and actually having systems
[58:37] that can deal with huge amounts of data, more data than any human experts can keep
[58:41] in mind and make sense of, and actually try and surface interesting insights that
[58:45] then the top human experts can take on and theorise about and make quicker breakthroughs
[58:50] with.
[58:51] AI is one of the most powerful tools we can create to help aid these research
[58:57] programs and the best scientists and clinicians to do their jobs even
[59:02] better. Thanks for listening.
