Amazon Web Services

In this AWS re:Invent 2023 session, Jim Roskind and Ankit Chadha explore the phenomenon of congestion collapse in large distributed systems and how to prevent it. They discuss real-world examples, including the 2018 Amazon Prime Day incident, to illustrate how systems can reach 100% CPU utilization while providing zero productive work. The speakers delve into strategies for avoiding congestion collapse, such as implementing proper retry mechanisms, throttling upstream traffic, and utilizing AWS services like CloudWatch, WAF, and SQS. The session also covers testing methodologies, including crash testing and chaos engineering principles, to proactively identify and mitigate potential issues. This comprehensive talk provides valuable insights for developers and system architects looking to build more resilient and efficient distributed systems on AWS.

cloud-trends-and-knowledge
skills-and-how-to
resilience
mgmt-govern
networking
Show 5 more

Up Next

VideoThumbnail
29:37

Builders 온라인 시리즈 | Amazon CloudWatch로 모니터링 손쉽게 시작하기

Jun 27, 2025
VideoThumbnail
26:19

Builders 온라인 시리즈 | AWS 파트너와 클라우드 여정 함께하기

Jun 27, 2025
VideoThumbnail
38:14

Builders 온라인 시리즈 | AWS re:Invent recap - 2024년 AWS가 선보이는 혁신적인 클라우드 서비스

Jun 27, 2025
VideoThumbnail
27:19

Builders 온라인 시리즈 | 클릭 몇 번으로 Amazon RDS 손쉽게 구성하기

Jun 27, 2025
VideoThumbnail
35:29

Splunk로 AWS에서 운영 인텔리전스와 인사이트 확보하기 - AWS TechCamp

Jun 26, 2025