Reliability Engineering: Membangun Sistem Cloud yang Tetap Stabil Kendati Terjadi Gangguan

Reliability Engineering: Membangun Sistem Cloud yang Tetap Stabil Kendati Terjadi Gangguan

Banyak organisasi masih mengukur reliability hanya dari uptime. Jika dashboard menunjukkan sistem masih berjalan, maka layanan dianggap baik-baik saja. Padahal, cara mengukur reliability dari perspektif pengguna dapat menghasilkan gambaran yang sangat berbeda.

Mengingat saat ini arsitektur aplikasi makin terdistribusi, menjaga sistem tetap reliable menjadi tantangan yang makin kompleks. Satu layanan dapat bergantung pada database, API, load balancer, jaringan, identity service, hingga layanan pihak ketiga. Akibatnya, ketika salah satu komponen mengalami gangguan maka dapat berdampak pada layanan lainnya.

Inilah alasan reliability engineering tidak hanya berbicara tentang menjaga server tetap hidup. Fokusnya adalah memastikan layanan tetap memenuhi ekspektasi pengguna dan mencegah failure, bahkan ketika latency meningkat, dependency gagal, atau sebagian komponen mengalami gangguan. Karena sistem cloud dan distributed systems tidak mungkin sepenuhnya bebas dari failure.

Table of Contents

Apa itu Reliability Engineering?

Reliability engineering adalah pendekatan engineering untuk memastikan sistem dapat memenuhi tingkat reliability yang dibutuhkan oleh pengguna dan bisnis. Dalam konteks cloud dan distributed system, reliability dan mencakup aspek availability, latency, error rate, throughout, durability, hingga kemampuan sistem untuk pulih dari gangguan.

Reliability engineering membantu tim mengubah konsep yang mengharuskan sistem untuk selalu stabil, menjadi target yang lebih terukur. Dengan begitu, ketimbang memastikan aplikasi harus tersedia, Tim dapat menentukan target yang lebih spesifik. Misalnya, 99,9 persen request pengguna harus berhasil dalam periode tertentu atau 95 persen request harus mendapatkan response dalam waktu kurang dari 300 milidetik.

Dengan memiliki target yang jelas, reliability tidak hanya menjadi persepsi. Tim dapat mengukur, mengevaluasi, dan mengambil tindakan nyata berdasarkan data.

Bagaimana Hubungannya dengan Resilience?

Reliability dan resilience saling berkaitan, tetapi memiliki fokus yang berbeda. Jika reliability berfokus pada kemampuan sistem untuk secara konsisten memberikan layanan sesuai target yang telah ditentukan. Maka resilience berfokus pada kemampuan sistem untuk menghadapi, menoleransi, dan pulih dari gangguan.

Sebuah sistem mungkin memiliki availability tinggi dalam kondisi normal, tetapi belum tentu resilient ketika salah satu dependency gagal. Sebaliknya, sistem yang resilient dirancang agar gangguan pada satu komponen tidak langsung menyebabkan seluruh layanan berhenti. Karena itu, reliability engineering dan resilience perlu berjalan bersama.

Reliability membantu menjawab pertanyaan tentang kemampuan layanan dalam memenuhi target yang dijanjikan kepada pengguna. Sedangkan resilience membantu menjawab pertanyaan tentang kemungkinan yang terjadi jika sistem mengalami kegagalan.

Sistem yang Reliable Bukan Berarti Tidak Pernah Gagal

Salah satu kesalahpahaman terbesar dalam reliability adalah menganggap sistem yang reliable tidak boleh mengalami kegagalan. Padahal faktanya, kegagalan pada sistem cloud tidak selalu dapat dihilangkan.

Server dapat mengalami masalah. Jaringan dapat mengalami latency. Database dapat mencapai batas kapasitas. API dependency dapat gagal. Bahkan kesalahan konfigurasi atau deployment dapat memicu incident.

Karena itu, reliability engineering tidak berusaha menciptakan sistem yang mustahil gagal. Fokusnya adalah membangun sistem yang memiliki kemampuan:

Mengurangi kemungkinan failure

Membatasi dampak failure

Mendeteksi gangguan lebih cepat

Memulihkan layanan dengan lebih efektif

Mencegah kegagalan kecil berkembang menjadi outage yang lebih besar

Baca Juga: Bisnis Ikut Lumpuh Akibat Gangguan API? Kenali Risikonya Sekarang

Mengenal SLI, SLO, dan Error Budget

SLI, SLO, dan error budget merupakan tiga konsep utama yang membantu organisasi mengelola reliability secara lebih objektif.

Secara sederhana, ketiganya dapat dianalogikan sebagai berikut:

SLI mengukur kondisi aktual service.

SLO menentukan target reliability.

Error budget menentukan batas kegagalan yang masih dapat ditoleransi.

Ketiganya membantu tim menjawab tiga pertanyaan penting terkait sejauh mana layanan dapat berjalan, seberapa reliable suatu layanan, dan berapa banyak kegagalan yang masih dapat diterima oleh suatu sistem.

SLI

Service Level Indicator atau SLI adalah indikator yang digunakan untuk mengukur performa atau kondisi aktual sebuah layanan. SLI dapat berupa persentase request yang berhasil, response time, error rate, availability, latency, dan throughput.

Hal terpenting adalah memilih SLI yang mencerminkan pengalaman pengguna. Karena tidak semua metrik teknis harus menjadi SLI.

CPU utilization, misalnya, dapat berguna untuk monitoring infrastruktur. Namun, tingginya CPU tidak selalu berarti pengguna mengalami gangguan. Sebaliknya, peningkatan error rate atau latency dapat secara langsung memengaruhi pengalaman pengguna.

Contohnya:

99,95 persen requests berhasil dalam periode 30 hari.

Atau:

95 persen requests mendapatkan respons dalam waktu kurang dari 500 milidetik.

SLO

Service Level Objective atau SLO adalah target reliability yang ingin dicapai berdasarkan SLI. Jika SLI mengukur kondisi aktual, SLO menentukan target yang ingin dipenuhi.

Contohnya:

99,9% request harus berhasil setiap bulan.

SLO membantu organisasi menghindari target yang terlalu abstrak seperti:

“Sistem harus selalu cepat.” atau:

“Aplikasi tidak boleh down.”

Dengan SLO, ekspektasi reliability menjadi lebih spesifik dan dapat diukur. Namun, target 100 persen tidak selalu menjadi pilihan terbaik.

Target reliability yang terlalu tinggi membutuhkan investasi yang sangat besar dan memperlambat kemampuan tim untuk melakukan perubahan. Karena itu, target reliability perlu disesuaikan dengan kebutuhan bisnis dan dampaknya terhadap pengguna.

Error Budget

Error budget adalah batas kegagalan yang masih dapat ditoleransi berdasarkan target SLO.

Secara sederhana:

Error Budget = 100 persen – SLO

Maka jika sebuah service memiliki SLO 99,9 persen, maka tersedia error budget sebesar 0,1 persen.

Error budget bukan berarti perusahaan membiarkan sistem gagal. Sebaliknya, error budget membantu organisasi menentukan batas risiko yang dapat ditolerir.

Selama layanan masih berada dalam batas error budget, tim dapat terus melakukan deployment, eksperimen, atau perubahan sistem. Namun, ketika error budget telah habis, reliability perlu menjadi prioritas.

Baca Juga: Mengapa API Integration Menjadi Tulang Punggung Transformasi Digital Perusahaan Modern

Bagaimana SLO Membantu Menentukan Prioritas Reliability?

Tanpa target yang jelas, setiap gangguan dapat memicu perdebatan, di antaranya:

Apakah masalah ini cukup serius?

Apakah deployment perlu dihentikan?

Apakah tim harus memperbaiki reliability atau melanjutkan pengembangan fitur?

SLO membantu mengurangi subjektivitas tersebut. Jika performa layanan masih berada dalam target, tim mungkin masih memiliki ruang untuk melakukan perubahan.

Namun, jika layanan secara konsisten gagal memenuhi SLO, reliability perlu menjadi prioritas. Dengan demikian, SLO membantu organisasi menentukan kapan harus fokus pada stabilitas, memperbaiki technical debt, mengurangi risiko deployment, meningkatkan observability, hingga meredesain dependency.

Selain itu, SLO juga membantu tim agar masih memiliki ruang untuk merilis fitur baru, melakukan eksperimen, meningkatkan kecepatan deployment, hingga mengembangkan inovasi.

Dengan kata lain, reliability bukan hanya menjadi tanggung jawab satu tim, tetapi menjadi bagian dari proses pengambilan keputusan bersama.

Bagaimana Error Budget Menyeimbangkan Reliability dan Kecepatan Inovasi?

Salah satu tantangan dalam engineering adalah adanya ketegangan antara dua kebutuhan. Di satu sisi, bisnis ingin bergerak lebih cepat dan merilis fitur baru. Di sisi lain, tim engineering perlu menjaga stabilitas sistem.

Jika setiap perubahan dianggap berisiko, inovasi dapat berjalan terlalu lambat. Namun, jika perubahan dilakukan tanpa kontrol, sistem dapat menjadi semakin tidak stabil. Error budget membantu menciptakan keseimbangan di antara dua tantangan tersebut.

Selama error budget masih tersedia, tim memiliki ruang untuk melakukan perubahan. Ketika error budget habis, organisasi dapat menerapkan kebijakan tertentu, misalnya:

Menunda feature release

Membatasi deployment non-kritis

Memprioritaskan reliability improvement

Melakukan root cause analysis

Memperbaiki monitoring dan testing

Pendekatan ini membuat reliability menjadi tanggung jawab bersama antara development, operations, dan product team. Keputusan tidak lagi hanya berdasarkan opini. Data reliability menjadi dasar untuk menentukan tingkat risiko yang dapat diterima.

Prinsip Membangun Sistem yang Lebih Resilient

Membangun sistem yang resilient membutuhkan lebih dari sekadar menambahkan backup. Arsitektur perlu dirancang dengan asumsi bahwa kegagalan dapat terjadi.

Beberapa prinsip penting meliputi:

Redundancy untuk Mengurangi Ketergantungan Pada Satu Komponen

Single point of failure dapat menyebabkan satu komponen menjadi titik kritis bagi seluruh sistem. Redundancy membantu menyediakan alternatif ketika salah satu komponen gagal.

Contohnya:

Multiple availability zones

Redundant instances

Replicated database

Multiple network paths

Namun, redundancy saja tidak cukup. Komponen cadangan juga perlu diuji untuk memastikan sistem benar-benar dapat melakukan failover ketika diperlukan.

Failure Isolation untuk Membatasi Dampak Gangguan

Gangguan pada satu layanan seharusnya tidak selalu menyebabkan seluruh sistem berhenti. Dengan melakukan isolasi terhadap failure domain, organisasi dapat membatasi dampak kegagalan.

Pendekatan ini dapat dilakukan melalui berbagai cara, mulai dari service isolation, queue, circuit breaker, bulkhead pattern, hingga dependency isolation. Tujuannya adalah mencegah cascading failure.

Graceful Degradation untuk Menjaga Layanan Tetap Berjalan

Ketika salah satu fitur gagal, seluruh aplikasi tidak selalu harus berhenti. Sistem dapat dirancang untuk tetap memberikan fungsi utama meskipun beberapa kapabilitas tidak tersedia.

Misalnya, aplikasi e-commerce mungkin tetap memungkinkan pengguna melihat katalog meskipun sistem rekomendasi sedang mengalami gangguan. Pendekatan ini membantu mengurangi dampak gangguan terhadap pengalaman pengguna.

Observability untuk Memahami Apa yang Terjadi

Tim tidak dapat memperbaiki masalah yang tidak dapat dilihat oleh mata. Untuk itu, observability membantu organisasi memahami kondisi sistem melalui data seperti metrics, logs, traces, dan events. Dengan visibilitas yang lebih baik, tim dapat mempercepat proses detection, diagnosis, dan recovery.

Peran Timeout, Retry, Exponential Backoff, dan Jitter

Dalam distributed systems, banyak kegagalan bersifat sementara atau transient failure. Misalnya, request gagal karena gangguan jaringan sesaat atau dependency sedang mengalami lonjakan beban.

Dalam kondisi seperti ini, retry dapat membantu request berhasil. Namun, retry yang dirancang secara tidak tepat justru dapat memperburuk gangguan.

Timeout

Timeout menentukan berapa lama sebuah service akan menunggu respons. Tanpa timeout yang tepat, request dapat terus menunggu dan menggunakan resource.

Ketika banyak request menunggu secara bersamaan, resource seperti memory, thread, atau connection dapat habis. Karena itu, timeout perlu disesuaikan dengan karakteristik layanan. Karena jika timeout terlalu lama membuat resource tertahan, sementara jika terlalu pendek dapat memicu terlalu banyak retry.

Retry

Retry memungkinkan aplikasi mencoba kembali request yang gagal. Mekanisme ini efektif untuk beberapa jenis transient failure. Namun, retry perlu dibatasi.

Jika sebuah dependency sedang overload, ribuan request yang melakukan retry secara bersamaan sehingga dapat menamnbah beban dan memperlambat proses recovery.

Exponential Backoff

Exponential backoff menambahkan waktu tunggu yang semakin panjang di antara setiap retry. Alih-alih mencoba kembali secara terus-menerus, aplikasi memberikan waktu bagi sistem untuk pulih. Pendekatan ini membantu mengurangi tekanan terhadap dependency yang sedang mengalami masalah.

Jitter

Masalah dapat muncul ketika banyak client melakukan retry pada waktu bersamaan. Jika semua request gagal dan mencoba kembali secara bersamaan, lonjakan traffic baru dapat kembali membebani sistem.

Jitter menambahkan variasi atau random delay pada waktu retry. Dengan demikian, retry tersebar dalam periode waktu yang berbeda dan risiko lonjakan serentak dapat dikurangi.

Intinya, timeout, retry, backoff, dan jitter harus dirancang sebagai satu strategi. Retry bukan selalu solusi. Dalam kondisi yang salah, retry justru dapat memperbesar failure.

Mengapa Testing terhadap Failure Scenario Penting?

Reliability tidak hanya dibangun ketika incident terjadi. Jika organisasi baru mengetahui kelemahan sistem setelah outage, maka pengguna telah menjadi bagian dari proses pengujian.

Karena itu, tim perlu menguji bagaimana sistem merespons berbagai failure scenario. Beberapa contoh skenario yang dapat diuji, antara lain:

Apa yang terjadi jika database tidak tersedia?

Bagaimana aplikasi merespons peningkatan latency?

Apa yang terjadi ketika API dependency gagal?

Apakah sistem dapat melakukan failover?

Apakah retry memperbesar beban?

Apakah alert memberikan informasi yang benar?

Seberapa cepat layanan dapat dipulihkan?

Testing terhadap failure scenario membantu tim menemukan kelemahan sebelum berdampak besar terhadap pengguna. Pendekatan ini dapat mencakup failure testing, load testing, chaos engineering, disaster recovery testing, dan game day exercise.

Tujuannya tak lain untuk membangun keyakinan bahwa ketika failure benar-benar terjadi, sistem dan tim sudah memiliki kesiapan untuk menghadapinya.

Baca Juga: Membangun Platform Thinking dalam Cloud Environment: Dari Cloud Infrastructure ke Platform Engineering

Pelajari Lebih Dalam Mengenai Reliability Engineering Bersama iCCom

Reliability engineering merupakan perjalanan yang terus berkembang di tengan sistem cloud yang semakin kompleks. Untuk itu, penting bagi engineer memahami bagaimana reliability, resilience, observability, dan operasional excellence saling berkaitan satu sama lain.

Saatnya memperluas pemahaman Anda tentang reliability engineering dan membangun sistem cloud yang dapat bertahan, bahkan ketika terjadi gangguan. Mulai perjalanan pemahaman Anda mengenai cloud dengan bergabung bersama komunitas iCCom (Indonesia Cloud Community), ruang untuk para cloud enthusiast belajar, berdiskusi, dan berbagi best practices seputar teknologi cloud, reliability engineering, resilience, serta perkembangan teknologi cloud lainnya.

Keanggotaan iCCOm GRATIS dan terbuka bagi siapa pun, baik pemula maupun profesional yang ingin memperluas wawasan seputar data, cloud, dan AI.

Bergabung bersama iCCOm untuk mengikuti berbagai kegiatan komunitas, mendapatkan teknologi terbaru, dan saling terhubung dengan sesama cloud enthusiast serta praktisi teknologi lainnya di Indonesia.

ENGLISH VERSION

Reliability Engineering: Building Cloud Systems That Remain Stable Despite Disruptions

Many organizations still measure reliability solely based on uptime. If the dashboard shows that the system is still running, the service is considered to be functioning properly. However, measuring reliability from the user’s perspective can provide a very different picture.

As application architectures become increasingly distributed, keeping systems reliable has become a more complex challenge. A single service may depend on a database, API, load balancer, network, identity service, and third-party services. As a result, when one component experiences a disruption, it can affect other services.

This is why reliability engineering is not merely about keeping a server running. Its focus is on ensuring that services continue to meet user expectations and preventing failures, even when latency increases, dependencies fail, or some components experience disruptions. This is because cloud systems and distributed systems cannot be completely free from failures.

What Is Reliability Engineering?

Reliability engineering is an engineering approach to ensuring that systems can meet the level of reliability required by users and businesses. In the context of cloud and distributed systems, reliability encompasses aspects such as availability, latency, error rate, throughput, durability, and the system’s ability to recover from disruptions.

Reliability engineering helps teams transform the expectation that systems must always remain stable into more measurable targets. Instead of simply ensuring that an application is available, teams can define more specific targets. For example, 99.9 percent of user requests must succeed within a certain period, or 95 percent of requests must receive a response in less than 300 milliseconds.
By having clear targets, reliability becomes more than just a perception. Teams can measure, evaluate, and take concrete action based on data.

How Is It Related to Resilience?

Reliability and resilience are closely related, but they have different focuses. Reliability focuses on a system’s ability to consistently deliver services according to predetermined targets, while resilience focuses on the system’s ability to withstand, tolerate, and recover from disruptions.

A system may have high availability under normal conditions but may not necessarily be resilient when one of its dependencies fails. Conversely, a resilient system is designed so that a disruption in one component does not immediately cause the entire service to stop. Therefore, reliability engineering and resilience need to work together.

Reliability helps answer questions about a service’s ability to meet the targets promised to users. Meanwhile, resilience helps answer questions about what may happen when a system experiences a failure.

A Reliable System Does Not Mean It Never Fails

One of the biggest misconceptions about reliability is the assumption that a reliable system must never experience failures. In reality, failures in cloud systems cannot always be eliminated.
Servers can experience problems. Networks can experience latency. Databases can reach their capacity limits. API dependencies can fail. Even configuration errors or deployments can trigger an incident.

Therefore, reliability engineering does not attempt to create a system that can never fail. Instead, its focus is on building systems that can:
• Reduce the likelihood of failures
• Limit the impact of failures
• Detect disruptions more quickly
• Recover services more effectively
• Prevent minor failures from developing into larger outages

Read Also: Can Your Business Be Paralyzed by an API Disruption? Understand the Risks Now

Understanding SLI, SLO, and Error Budgets

SLI, SLO, and error budgets are three key concepts that help organizations manage reliability more objectively.

Simply put, they can be illustrated as follows:
• SLI measures the actual condition of a service.
• SLO defines the reliability target.
• Error budget determines the amount of failure that can still be tolerated.

Together, they help teams answer three important questions: how well a service is performing, how reliable a service is, and how much failure a system can still tolerate.

SLI

Service Level Indicator, or SLI, is an indicator used to measure the actual performance or condition of a service. An SLI can include the percentage of successful requests, response time, error rate, availability, latency, and throughput.

The most important consideration is choosing SLIs that reflect the user experience. Not every technical metric needs to become an SLI.

CPU utilization, for example, can be useful for infrastructure monitoring. However, high CPU utilization does not necessarily mean that users are experiencing disruptions. Conversely, an increase in error rate or latency can directly affect the user experience.

For example:
99.95 percent of requests succeed over a 30-day period.
Or:
95 percent of requests receive a response in less than 500 milliseconds.

SLO

Service Level Objective, or SLO, is the reliability target that an organization aims to achieve based on its SLI. While an SLI measures the actual condition, an SLO defines the target that needs to be met.

For example:
99.9% of requests must succeed each month.
SLOs help organizations avoid overly abstract targets such as:
“The system must always be fast.”
Or:
“The application must never go down.”
With SLOs, reliability expectations become more specific and measurable. However, a 100 percent target is not always the best choice.
An excessively high reliability target requires significant investment and can slow down a team’s ability to make changes. Therefore, reliability targets need to be aligned with business needs and their impact on users.

Error Budget

An error budget is the amount of failure that can still be tolerated based on the SLO target.
Simply put:
Error Budget = 100% – SLO

Therefore, if a service has an SLO of 99.9 percent, it has an available error budget of 0.1 percent.
An error budget does not mean that a company allows its system to fail. Instead, it helps organizations determine the level of risk that can be tolerated.

As long as a service remains within its error budget, teams can continue with deployments, experiments, or system changes. However, when the error budget has been exhausted, reliability needs to become a priority.

Read Also: Why API Integration Is Becoming the Backbone of Digital Transformation for Modern Enterprises

How Do SLOs Help Determine Reliability Priorities?

Without clear targets, every disruption can trigger debates, such as:
• Is this problem serious enough?
• Should the deployment be stopped?
• Should the team improve reliability or continue developing features?
SLOs help reduce this subjectivity. If service performance remains within the target, the team may still have room to make changes.

However, if a service consistently fails to meet its SLO, reliability needs to become a priority. In this way, SLOs help organizations determine when to focus on stability, address technical debt, reduce deployment risks, improve observability, or redesign dependencies.

At the same time, SLOs allow teams to maintain enough room to release new features, conduct experiments, improve deployment speed, and develop innovations.
In other words, reliability is not merely the responsibility of a single team. It becomes part of the shared decision-making process.

How Do Error Budgets Balance Reliability and Innovation Speed?

One of the challenges in engineering is the tension between two needs. On one hand, businesses want to move faster and release new features. On the other hand, engineering teams need to maintain system stability.

If every change is considered risky, innovation may move too slowly. However, if changes are made without proper controls, systems can become increasingly unstable. Error budgets help create a balance between these two challenges.

As long as the error budget remains available, teams have room to make changes. When the error budget is exhausted, organizations can implement specific policies, such as:
• Delaying feature releases
• Limiting non-critical deployments
• Prioritizing reliability improvements
• Conducting root cause analysis
• Improving monitoring and testing

This approach makes reliability a shared responsibility among development, operations, and product teams. Decisions are no longer based solely on opinions. Reliability data becomes the basis for determining an acceptable level of risk.

Principles for Building More Resilient Systems

Building a resilient system requires more than simply adding backups. The architecture needs to be designed with the assumption that failures can occur.
Several key principles include:

Redundancy to Reduce Dependence on a Single Component

A single point of failure can make one component a critical point for the entire system. Redundancy helps provide alternatives when one component fails.

Examples include:
• Multiple availability zones
• Redundant instances
• Replicated databases
• Multiple network paths

However, redundancy alone is not enough. Backup components also need to be tested to ensure that the system can actually perform failover when needed.

Failure Isolation to Limit the Impact of Disruptions

A disruption in one service should not always cause the entire system to stop. By isolating failure domains, organizations can limit the impact of failures.

This approach can be implemented in various ways, including service isolation, queues, circuit breakers, bulkhead patterns, and dependency isolation. The goal is to prevent cascading failures.

Graceful Degradation to Keep Services Running

When one feature fails, the entire application does not necessarily have to stop. Systems can be designed to continue providing core functionality even when some capabilities are unavailable.

For example, an e-commerce application may still allow users to browse its catalog even when the recommendation system is experiencing a disruption. This approach helps reduce the impact of disruptions on the user experience.

Observability to Understand What Is Happening

Teams cannot fix problems that they cannot see. Therefore, observability helps organizations understand system conditions through data such as metrics, logs, traces, and events. With better visibility, teams can accelerate detection, diagnosis, and recovery.

The Role of Timeouts, Retries, Exponential Backoff, and Jitter

In distributed systems, many failures are temporary or transient failures. For example, a request may fail due to a temporary network disruption, or a dependency may be experiencing a sudden spike in load.
Under these conditions, a retry can help a request succeed. However, poorly designed retries can actually make the disruption worse.

Timeout

A timeout determines how long a service will wait for a response. Without an appropriate timeout, a request may continue waiting and consume resources.
When many requests are waiting simultaneously, resources such as memory, threads, or connections can be exhausted. Therefore, timeouts need to be adjusted to the characteristics of the service.

If a timeout is too long, resources may remain occupied for too long. Meanwhile, if it is too short, it may trigger too many retries.

Retry

A retry allows an application to attempt a failed request again. This mechanism is effective for certain types of transient failures. However, retries need to be limited.
If a dependency is overloaded, thousands of requests retrying simultaneously can add more load and slow down the recovery process.

Exponential Backoff

Exponential backoff introduces progressively longer waiting periods between each retry. Rather than continuously retrying, the application gives the system time to recover. This approach helps reduce pressure on a dependency that is experiencing problems.

Jitter

Problems can occur when many clients perform retries at the same time. If all requests fail and retry simultaneously, a new spike in traffic can place additional pressure on the system.
Jitter adds variation or a random delay to the retry timing. As a result, retries are spread across different periods, reducing the risk of simultaneous traffic spikes.
In short, timeouts, retries, backoff, and jitter should be designed as a single strategy. A retry is not always the solution. Under the wrong conditions, a retry can actually amplify a failure.

Why Is Testing Failure Scenarios Important?

Reliability is not built only when an incident occurs. If an organization discovers weaknesses in its system only after an outage, users have already become part of the testing process.
Therefore, teams need to test how systems respond to various failure scenarios. Examples of scenarios that can be tested include:
• What happens if the database becomes unavailable?
• How does the application respond to increased latency?
• What happens when an API dependency fails?
• Can the system perform failover?
• Do retries increase the load?
• Do alerts provide the correct information?
• How quickly can the service be restored?

Testing failure scenarios helps teams identify weaknesses before they have a major impact on users. This approach can include failure testing, load testing, chaos engineering, disaster recovery testing, and game day exercises.

The goal is to build confidence that when a failure actually occurs, both the system and the team are prepared to handle it.

Read Also: Building Platform Thinking in a Cloud Environment: From Cloud Infrastructure to Platform Engineering

Learn More About Reliability Engineering with iCCom

Reliability engineering is an ongoing journey as cloud systems become increasingly complex. Therefore, it is important for engineers to understand how reliability, resilience, observability, and operational excellence are interconnected.

It is time to deepen your understanding of reliability engineering and build cloud systems that can withstand disruptions. Start your cloud learning journey by joining iCCom (Indonesia Cloud Community), a space where cloud enthusiasts can learn, discuss, and share best practices related to cloud technology, reliability engineering, resilience, and other developments in cloud technology.
iCCom membership is FREE and open to everyone, from beginners to professionals who want to expand their knowledge of data, cloud, and AI.

Join iCCom to participate in various community activities, stay updated on the latest technology, and connect with fellow cloud enthusiasts and technology practitioners across Indonesia.

Translated by Malikha Innaya – Cloud Technology Community Specialist Intern CTI Group

Facebook
Twitter
LinkedIn
WhatsApp

Table of Contents

Related Article

Write Your Own Article!