# Failover routing across providers

A routing layer that decides which external provider handles each request, and stops sending traffic to a provider that is failing.

- **sheet**: 3.3
- **status**: production
- **role**: design, implementation
- **stack**: Java

## Context

> note Implementation details are internal. This note describes the approach and the reasoning, not the code.

Much of the work depends on external providers, and a platform that depends on a single one inherits every one of its bad days. The router is generic: it does not know what a provider does, only how to choose between several that can do the same job.

With more than one provider available, the question becomes which one should take the next request. The answer depends on who is healthy, who has capacity, and what each request costs.

## Requirements

- Keep requests flowing when one provider slows down or starts failing, without anyone flipping a switch.
- Spread load so no single provider becomes the bottleneck.
- Let cost influence the choice when every provider is healthy.
- Make the strategy a choice rather than a rewrite.

## Three strategies

The layer supports three routing strategies. Each is good at something the others are not.

| strategy | good at | weak at |
|---|---|---|
| round-robin | Even distribution with no state to keep | Ignores capacity, cost and health |
| weighted | Encoding capacity or commercial terms as a ratio | Weights go stale if nobody revisits them |
| circuit breaker | Removing a failing provider quickly and bringing it back on its own | Thresholds need tuning, or it trips on noise |

## The circuit breaker

A breaker watches the outcomes of calls to one provider. While things are normal it is *closed* and traffic flows. When failures cross a threshold it *opens* and the provider gets no new requests. After a cool-down it goes *half-open* and lets a probe through: success closes it again, failure reopens it.

```
+--------+   failures >= threshold   +--------+
| closed | ------------------------> |  open  |
+--------+                           +--------+
    ^                                  |    ^
    |                      cool-down   |    | probe
    | probe ok                         v    | fails
    |                              +-----------+
    +------------------------------| half-open |
                                   +-----------+
```

Breaker states for one provider.

The useful property is that recovery is automatic. Nobody has to notice that a provider came back and re-enable it by hand.

## Result

The layer is in production. Requests fail over automatically when a provider degrades, load is distributed across providers, and routing can take cost into account.

## What I would look at next

- Feeding observed latency back into weights, so they cannot drift far from reality.
- Separate thresholds for timeouts and hard errors, which usually mean different things.
- A dashboard of breaker transitions, since a breaker that flaps is a signal in itself.

[← All projects](https://0xshriram.dev/projects.md)

---
Canonical: https://0xshriram.dev/projects/failover-router.html
