
# How to test rate limits (429), retries and outages of LLM APIs

Retry and fallback logic around LLM calls is easy to get wrong and hard to test, because real rate limits don't happen on demand. [llmao](https://github.com/icernigoj/llmao) makes them happen.

## Fail on demand

```ts
import * as llmao from 'llmao/testing';

// Every call fails with a rate limit
llmao.configure({ failures: { rateLimit: 1 } });

// Or some of the time
llmao.configure({ failures: { rateLimit: 0.1, serverError: 0.05, timeout: 0.01 } });

// Or exactly once: the first attempt fails, the retry succeeds
llmao.configure({ script: [{ error: 'rate_limit', once: true }, { text: 'Back online.' }] });

// With a specific retry-after, in seconds (0 by default in test mode)
llmao.configure({ failures: { rateLimit: 1, retryAfter: 20 } });
llmao.configure({ script: [{ error: 'rate_limit', retryAfter: 20, once: true }, { text: 'ok' }] });
```

## The errors your code already handles

| | OpenAI SDK | Anthropic SDK | Vercel AI SDK |
|---|---|---|---|
| Rate limit | `OpenAI.RateLimitError`, `status: 429` | `Anthropic.RateLimitError`, `status: 429` | `APICallError`, `statusCode: 429` |
| Server error | `OpenAI.InternalServerError`, `status: 500` | `Anthropic.InternalServerError` | `APICallError`, `statusCode: 500` |
| Timeout | `OpenAI.APIConnectionTimeoutError` | `Anthropic.APIConnectionTimeoutError` | `APICallError` |

Rate limits come with a `retry-after` header on `error.headers` (a `Headers` object: read it with `error.headers.get('retry-after')`, not `error.headers['retry-after']`). In test mode it is `0` unless you set `retryAfter`, so your backoff doesn't slow the test down. To test that your code honors it, set a value and use fake timers ([recipe for node:test](../test-llm-code-any-runner/#testing-retries-with-fake-timers)).

## Retries

The OpenAI and Anthropic clients retry like the official SDKs (`maxRetries`, default 2, waiting a short `retry-after` when there is one), and the AI SDK runs its own retries. Each attempt is recorded, so you can assert on them:

```ts
llmao.configure({ failures: { rateLimit: 1 } });
await expect(classify('hi')).rejects.toBeInstanceOf(OpenAI.RateLimitError);

expect(llmao.calls).toHaveLength(3); // the first attempt plus two retries
expect(llmao.calls.every((call) => call.failure === 'rate_limit')).toBe(true);
```

If your code wraps the SDK in its own retry loop, this is also how you find out that you are retrying the SDK's retries.

### Testing your own retry loop

The SDK retries first, so a single scripted failure never reaches your code: the SDK's second attempt succeeds. To exercise your loop, make every attempt of the SDK fail, or turn the SDK's retries off in the test:

```ts
// Every request fails: the SDK gives up and your loop sees the error
llmao.configure({ failures: { rateLimit: 1, retryAfter: 20 } });

// Or fail exactly one of your attempts: once per SDK request (maxRetries + 1, so 3 by default)
const rateLimit = { error: 'rate_limit', retryAfter: 20, once: true } as const;
llmao.configure({ script: [rateLimit, { ...rateLimit }, { ...rateLimit }, { text: 'ok' }] });
```

Each `once` rule is consumed by one HTTP request, so copies are needed: the same object would count as one rule. With a real `retry-after`, both the SDK and your loop wait it out: drive the clock with fake timers ([node:test recipe](../test-llm-code-any-runner/#testing-retries-with-fake-timers)). The SDK numbers its attempts in the `x-stainless-retry-count` request header, recorded in `llmao.calls[i].headers`.

## Related

- [Mock OpenAI in Jest](../mock-openai-jest/)
- [Mock the Anthropic SDK](../mock-anthropic-sdk/)
