Engineering

How to Start with Chaos Engineering Experiments?

Chaos engineering experiments can be highly exploratory at the beginning. Here are the six acceptance criteria you need to define before you inject a single fault.

Sérgio Martins2 min read
House completely consumed by flames
Photo by Sérgio Martins on Onfido Product and Tech.

Chaos engineering experiments can be highly exploratory at the beginning.

You often have low to no clue of what to look for other than the possible unavailability of the system.

Let me give you a hand, and provide you with the acceptance criteria to get you started:

  • Downtime — Are you permitting downtime during chaos engineering experiments? Perhaps you should. But not for all of the services. Thus, identify the critical and non-critical services and establish their allowed downtime. As a hint, you may want to consider vital services the ones to be directly related to revenue generation.

  • Service degradation — How much time of a degraded service is acceptable? i.e. Mean time to repair (MTTR) to default performance. Separate service degradation into two categories: the latency and the error rate. Be accountable for spikes when initiating the experiment, but ensure that you stipulate a maximum allowed time for the MTTR of your services.

  • Data loss — What is your recovery point objective (RPO)? Are you considering data loss acceptable during chaos engineering experiments? Pay particular attention to requests that are in flight. Thus, if the RPO is zero, ensure that these requests in flight are resolved after service restoration (guaranteeing no loss of data integrity).

  • Incident response — Do you require manual intervention to respond to failures? Or are all of the mechanisms automated? Ensure it is clear when a chaos engineering experiment needs manual intervention for services to recover, even after the maximum allowed time for MTTR. Monitoring should be in place to support you in recognising if these failures are persisting and what is triggering them.

  • Repeatable chaos testing — Is the experiment you’re running repeatable? It most likely should be, especially since these experiments are usually manually intensive. Consider performing chaos experiments “on-demand” with automated playbooks that any engineer can trigger to avoid “siloed engineering” knowledge practices. Furthermore, the experiment results should be published and documented in a form that an audit/compliance team can pick up to use as evidence.

  • Runbooks — How many of these experiments, if they would occur in production, are you ready to provide support for during on-call processes? Document chaos engineering experiments so that any engineer can understand the experiment and act to resolve it in case manual intervention is needed. Track and measure the number of incidents where you didn’t have documentation to back you up during on-calls. Ensure that you document them as soon as possible. You’ll thank me after it saves your product from chaos.


Conduct chaos engineering experiments with proper acceptance criteria.

Share
Written by
Sérgio Martins

Senior Software Engineer at Entrust.

Comments

Hook this up to your favourite commenting platform — Giscus, Disqus, or your own.

Continue reading

Engineering·

Causal Inference at Onfido

How Onfido's data science team built a Structural Causal Model of their document verification pipeline — and used it to simulate the impact of product improvements before shipping a single line of new code.

Anna Borodina · 6 min
Product·

Data-driven Product Launch — Onfido Studio

How a data scientist approaches a B2B product launch from first principles — defining metrics before a line of code ships, building the data model, creating pre-launch dashboards, and learning from the first beta customers.

Anna Borodina · 8 min

A quieter inbox.

One thoughtful letter every other Sunday — new essays, things worth reading, and the occasional photograph.

Free. Unsubscribe in one click.