DevOps At Massive Scale

DevOps At Massive Scale

February 16, 2017
By Derek Weeks

3 minute read time

When you have a billion users, people notice. That's where our story about DevOps and Yahoo! starts. For Kishore Jalleda and Gopal Mor, both engineers at Yahoo!, when something goes wrong on a Yahoo! page, people will notice. Correction: many people will notice.

Of course, Yahoo!, like all services on the Internet, constantly improves its products. In fact, they have 100+ iterations and experiments happening at any given time. Some changes bring new innovation to the forefront, and others alter the user experience.

When iterations and experiments are served in front of loyal users who have become comfortable with a specific user experience, they sometimes react with a natural resistance to the change. When change causes, or appears to cause, breaks in service, the backlash can be crippling. Frequent spurts of backlash resulting from these changes became known to some Yahoo! Insiders as "the bad years."

At the recent All Day DevOps conference, Kishore and Gopal shared how Yahoo! turned to DevOps practices to recover from "the bad years." In their presentation, Launching Products at Massive Scale: The DevOps Way, Kishore said, "DevOps is about eliminating technical, process, and cultural barriers between idea and execution using software."

Specific to each of these, Kishore recommends the following:

At Yahoo!, the DevOps practice is built on three functional pillars: deliver products to market quickly; prevent defects from reaching customers; and repair production issues quickly.

Speaking of repairing production issues quickly, Gopal discussed resiliency. With their user base, downtime (or lack thereof) is critical to Yahoo! and the challenges are many:

While it may be counterintuitive, Gopal demonstrates how the combined system is weaker than the weakest subsystem.

To ensure optimum uptime, Gopal tells us we need to:

  1. Analyze the entire range of failure types.
  2. Understand their rate and impact level.
  3. Plan to cover all failure types.
  4. Conduct fire drills: test, test, and test.

Specifically, how does Yahoo! ensure high availability? They maintain four layers of resiliency in the serving stack.

There is a lot here, and you don't have to have a service at the scale of Yahoo! to benefit from the experience. You can dive into Kishore and Gopal's full All Day DevOps conference session (just 30 minutes) to learn more about their learnings from the DevOps front lines. The other 56 presentations from the All Day DevOps Conference are also available online, free-of-charge here.

Written by Derek Weeks
Derek serves as vice president and DevOps advocate at Sonatype and is the co-founder of All Day DevOps, an online community of 65,000 IT professionals.