Skip to main content
發布日期
<Aja Huang: AlphaGo Zero Covered Thousands of Years of Human Go Research in Just Three Days> Featured Article

Aja Huang: AlphaGo Zero Covered Thousands of Years of Human Go Research in Just Three Days 

AlphaGo, DeepMind, Artificial Intelligence, Go, Aja Huang

by  2017/11/10  Li Bo-feng

Source: https://www.inside.com.tw/2017/11/10/aja-alphago-zero

Image source: pixabay

Dr. Aja Huang, Senior Researcher at DeepMind, returned to Taiwan today to deliver a lecture titled "AlphaGo: The Triumph of Deep Learning and Reinforcement Learning" at the inaugural Artificial Intelligence Annual Conference. The event drew considerable attention from Taiwan's industry, government, and academic communities, with crowds filling the lecture hall at Academia Sinica well before nine o'clock. Beyond sharing his research on AI and the game of Go, Huang also discussed the recently published AlphaGo Zero and how it can teach itself Go without any human knowledge, ultimately surpassing its predecessor — the version that defeated human champions.

From a Taiwanese PhD Student to an Employee of DeepMind, Acquired by Google

Aja Huang was the first student admitted to the Graduate Institute of Computer Science and Information Engineering at National Taiwan Normal University, where he studied from master's through doctoral level, marrying during his fifth year of the PhD program. The Go software he developed during his doctorate was called Erica — his wife's name. In its standalone version, it defeated Zen, the strongest AI Go program at the time, which ran on six machines. This achievement caught the attention of DeepMind, and David Silver personally invited Huang to join — making him the 40th employee.

During the interview, David Silver asked Huang what it felt like to have developed Erica. Huang replied: "It was deeply fulfilling — to be able to build an AI on my own." After joining DeepMind, he found that this sense of fulfillment was shared throughout the company, and that DeepMind's dream was to build "artificial general intelligence." In 2014, DeepMind was acquired by Google, and the greatest benefit of joining Google was access to enormous computing resources.

Returning to Go: The Birth of AlphaGo

After becoming a researcher at DeepMind, Huang did not immediately begin developing AlphaGo. It was not until 2014–2015 that the Go AI project was relaunched — and it was not a continuation of Erica, since its limits had already been reached. The team had to start fresh using deep learning technology, while continuously recruiting the world's top talent, including Chris Maddison and Ilya Sutskever from Canada's DNNresearch, which had also been acquired by Google, creating the opportunity for collaboration.

With talent and computing resources in place, the AlphaGo project officially began. Huang shared that the first breakthrough came from applying neural network technology — the team was not initially certain it would work, but when the experimental results came in, the win rate against the original version was 100%, energizing the entire team. The second breakthrough was the value network. Simulations at the time suggested AlphaGo would win 70–80% of competitive matches, already placing it at the top of the world — but DeepMind's ambitions went far beyond that, so the team continued to expand, enabling more research and the resolution of more problems.

Huang also shared that the day-to-day work of developing AlphaGo involved training neural networks, testing, examining win rates, and observing whether changes were effective. Many ideas and problems required constant testing: How many layers should the deep learning architecture have? What structure should be used? Was there anything wrong with the training data? Ultimately, the true measure was always whether AlphaGo's playing strength had improved.

During this observation process, the team also discovered that AlphaGo had an overfitting problem. Once resolved, AlphaGo grew stronger, achieving a 95% win rate against its previous version — which is why the lecture was titled "AlphaGo's Success: The Triumph of Deep Learning and Reinforcement Learning."

Playing Against Humans and Publishing the First Nature Paper

Having confirmed AlphaGo's capabilities, DeepMind decided to test it against a human player. The first opponent was French 2-dan professional Fan Hui. In October 2015, AlphaGo won all five games. An editor from Nature attended the fifth game to verify whether the paper AlphaGo was about to submit was truly as impressive as claimed. Fan Hui became the first professional Go player to be officially defeated by an AI, but after his defeat, he took a positive view of AI's role in the development of Go and went on to provide considerable assistance to the AlphaGo team.

DeepMind, however, is better described as a "research institution" than a "for-profit enterprise." After going to great lengths to develop an AI capable of defeating professional players, why publish a paper disclosing all the details — especially when, after defeating Fan Hui, they had already publicly challenged 9-dan professional Lee Sedol, making disclosure an apparent strategic disadvantage? Huang admitted he did not understand the company's decision at the time, feeling the effort should have gone toward match preparation rather than writing papers.

Because the paper was to be published, Nature required that the news of defeating Fan Hui not be made public before the paper appeared — which is why the general public only learned of it several months later.

Huang also reiterated that after DeepMind joined Google, the computing hardware resources Google provided were enormously helpful — especially when TPUs later replaced GPUs, which made a tremendous difference, as many tasks would have been impossible otherwise. AlphaGo was also among the first programs within Google to make extensive use of TPUs. For further details, Huang noted that the documentary film AlphaGo covers them in depth.

Learning from Defeat Against Lee Sedol to Identify Weaknesses and Further Strengthen the Model

The outcome of the Korea match is well known. But after defeating Lee Sedol, was it time to stop? In fact, during the match itself, AlphaGo displayed a glaring problem in game four — making a mistake that even an amateur player would not commit. Huang, who was responsible for placing the stones, even felt he might have done better himself at that moment, and Lee Sedol looked at the screen in disbelief to confirm Huang hadn't placed the stone in the wrong position by mistake.

Since AlphaGo still had issues, further research was clearly necessary. Resolving all the problems comprehensively took eight months, during which newcomer Karen Simonyan also joined the team. The solution ultimately involved reinforcing the learning capabilities built on deep learning and reinforcement learning techniques.

The first step was expanding the original 13-layer network to 40 layers and switching to ResNet. The second step was combining the Policy Network and Value Network into a Dual Network, training AlphaGo's intuition and judgment simultaneously. The third step was strengthening the training pipelines. Beyond AI learning capabilities, Huang also resolved Go-specific problems such as mirror Go and recurring ko situations. Against the version that had defeated Lee Sedol, the upgraded AlphaGo could give three stones (without komi) and still achieve a win rate exceeding 50%.

Master Goes from Playing Quietly in Tainan to Capturing the World's Attention

After confirming that all identifiable problems had been resolved, the AlphaGo team decided to quietly go online and challenge professional players — this was what later became known as the Master version. Of course, after winning continuously, staying under the radar was no longer possible. The final result was a clean sweep against the top players from China, Japan, Korea, and Taiwan.

At the time, Huang had returned to Taiwan and, from his room in Tainan, opened a new account and invited Go players to compete. Famous players initially declined — though of course it soon became a matter of Huang turning others away — and with each game, more and more spectators tuned in to watch. Throughout the matches, Huang continuously monitored AlphaGo's win rate chart. Aside from Ke Jie, no one had a chance of beating AlphaGo any longer.

AlphaGo Zero Covered Thousands of Years of Human Go Research in Just Three Days

By this point, the AlphaGo team had split into two groups. While Huang was busy using Master to compete against Ke Jie, another group was developing AlphaGo Zero. Huang's task was to strip AlphaGo of all accumulated Go knowledge and triple-check that this had been done — because AlphaGo Zero is an AI that learns entirely through self-play without any prior human knowledge, meaning it could only have knowledge of the rules, not of Go strategy.

The team was not initially certain whether it would succeed, but AlphaGo Zero did indeed go on to defeat Master, once again demonstrating the remarkable power of deep learning and reinforcement learning. In the beginning, AlphaGo Zero played completely randomly and frequently got stuck after a period of learning; it took various adjustments before progress could continue. But with Google's powerful computing resources — 2,000 TPUs — in just three short days, AlphaGo Zero succeeded. And not only did its learning capabilities surpass expectations, but its energy consumption during matches was dramatically lower compared to the computation used during the match against Fan Hui. Many of Zero's moves are now beyond Huang's own comprehension.

After further refinements and improvements, AlphaGo traveled to China for the match against Ke Jie. Huang noted that compared to the Korea match, where the team very much wanted to win every game, the atmosphere in China was more relaxed — because winning or losing was no longer the central focus (the team felt defeat was unlikely). The match had become an exploration of how humans and AI can collaborate, which was reflected in its title: "Exploring the Future of Go Together." Huang remarked that AI will no longer lose to humans at Go, but at this stage, AI's role is to expand human players' thinking and to collaborate with humans in exploring the uncharted frontiers of the game.

Conclusion:

Looking back on everything achieved along the way — two Nature papers published, two human-versus-AI competitions and 60 online games, the opportunity to bring both AI and Go (Huang's two greatest passions) to the attention of the entire world, a feature in Time magazine, and a documentary film — Huang expressed deep satisfaction. The following are the five concluding points Huang summarized in his presentation slides:

1. AlphaGo's success is a triumph of deep learning and reinforcement learning.

2. From start to finish, AlphaGo proved that unity of effort multiplies strength.

3. In AlphaGo's development, TPUs and hardware resources played a critically important role.

4. AlphaGo Zero demonstrates the enormous potential of reinforcement learning.

5. In the foreseeable future, AI will become an important tool for humanity, collaborating alongside humans.

During the Q&A session, an audience member asked whether the emergence of AlphaGo Zero implies that human knowledge has become irrelevant. Huang responded that this is a question worth studying. AlphaGo Zero only answers whether AI can function without human knowledge — whether it needs human knowledge is a question that cannot yet be answered. Having human knowledge does shorten the time AI needs to learn, but without it, could AI develop an entirely different body of knowledge?

There are currently no plans to open-source AlphaGo Zero, but Huang noted that the papers published in Nature are written in considerable detail, and someone has already built an open-source version of AlphaGo Zero based on those papers — so whether DeepMind chooses to open-source it makes little practical difference.

 

Please add some content in Sliding Sidebar block region.

For more information please refer to this tutorial page:
Add content in sliding sidebar