undefined | Better HN

Thaxll8mo ago· 2 in thread

Best practice does not include plan for when AWS going down. Netflix does not plan for it and they have a very strong eng org.

immibis8mo ago

Did they stop their Chaos Gorilla, which simulates a region outage?

pjmlpOP8mo ago

It was only one region.

spyspy8mo ago· 1 in thread

Eh, the "best practices" that would've prevented this aren't trivial to implement and are definitely far beyond what most engineering teams are capable of, in my experience. It depends on your risk profile. When we had cloud outages at the freemium game company I worked at, we just shrugged and waited for the systems to come back online - nobody dying because they couldn't play a word puzzle. But I've also had management come down and ask what it would take to prevent issues like that from happening again, and then pretend they never asked once it was clear how much engineering effort it would take. I've yet to meet a product manager that would shred their entire roadmap for 6-18 months just to get at an extra 9 of reliability, but I also don't work in industries where that's super important.

pjmlpOP8mo ago

Indeed, yet one would expect AWS to lead by example, including all of those that are only using a single region.

1 more reply

mrbungie8mo ago

> It just goes to show the difference between best practices in cloud computing, and what everyone ends up doing in reality,

Well, inter-region DR/HA is a expensive thing to ensure (whether on salaries, infra or both), specially when you are in AWS.

1 more reply

esafak8mo ago

Does AWS follow its own Well-Architected Framework!?

0 comments

0 comments