Normalization of Deviance

Launch of STS-51-L

Source: NASA

The 28th of January 1986 was a cold day in the US South. Icicles had been photographed hanging from railings on Pad 39A at Kennedy Space Center. School had been called off in Metro Atlanta on account of snow.

Icicles on the launch pad. Source: Rogers Commission Report

I was at home playing with Lego bricks when my mother came in and said, “I was watching TV and the space shuttle just exploded!” We went to the den and watched the replay of the launch and accident on TV.

When I was in college taking software engineering courses, we talked a little bit about the accident. I looked up the Rogers Commission Report and read it.

Engineers had pointed out that the O-rings on the Solid Rocket Boosters were less effective at sealing under cold temperatures. An earlier mission, STS-9, had found 0.5″ of water in the gap where the O-ring sat after rain while restacking a booster due to issues with the exhaust nozzle.

On STS-51L, Challenger faced both issues – it was colder than previous launches and 7″ of rain had fallen in the 38 days it sat on the launch pad – more rain than Columbia’s booster experienced on STS-9.

The result was that, on launch, the O-ring failed to seal at a critical point and exhaust leaking from the joint cut the lower Solid Rocket Booster attachment, allowing it to swivel on the upper strut, and then strike the External Tank, igniting the fuel within which led to the midair high-speed breakup of the Space Shuttle and the loss of the 7 crew members aboard.

Hole burnt through Solid Rocket Booster by leak.

Source: Rogers Commission Report

Despite the concerns raised by engineers, management authorized the launch. Previous missions had experienced O-ring erosion and other warning signs, yet those missions had been successful. Over time, what should have been treated as evidence of a problem became accepted as normal.

Sociologist Diane Vaughan later coined the term “normalization of deviance” to describe this phenomenon. When an abnormal condition does not immediately result in failure, organizations can gradually begin to treat it as acceptable. Success becomes evidence that the risk is manageable rather than evidence that disaster was narrowly avoided.

The warning signs were visible – engineers had data from several launches that the O-rings became less effective at sealing as temperatures decreased. There had been no accidents previously. Over time, this absence of failure was accepted as evidence that the risk was acceptable.

This isn’t something limited to astronautics or aerospace. It’s a concept that applies to every engineering discipline – including IT.

Normalization of deviance in IT isn’t a major incident – it’s the little things that get overlooked because they haven’t turned into a big incident:

  • The successful backup job that always reports a warning.

  • The recurring hardware alert that’s ignored because it hasn’t failed yet.

  • The high resource utilization on a server that’s ignored because the services are still up.

  • The repeated failed login alert that admins learn to ignore.

The alerts become familiar and accepted. When they’re accepted, they become normal, they become background noise. Everything appears to be working normally.

But eventually, luck runs out.

The backup job starts failing. The hardware fails. The server goes offline due to high resource utilization. A real security incident is missed because failed logins happen every day.

It’s no longer an annoying alert to track down and resolve – it’s a real incident.

The best time to fix a problem is when it’s still only a warning.

Next
Next

The Importance of Curiosity and Experience, Part II