How to test rate limits (429), retries and outages of LLM APIs
Retry and fallback logic around LLM calls is easy to get wrong and hard to test, because real rate limits don't happen on demand. llmao makes them happen.
Fail on demand
import * as llmao from 'llmao/testing';
// Every call fails with a rate limit
llmao.configure({ failures: { rateLimit: 1 } });
// Or some of the time
llmao.configure({ failures: { rateLimit: 0.1, serverError: 0.05, timeout: 0.01 } });
// Or exactly once: the first attempt fails, the retry succeeds
llmao.configure({ script: [{ error: 'rate_limit', once: true }, { text: 'Back online.' }] });
// With a specific retry-after, in seconds (0 by default in test mode)
llmao.configure({ failures: { rateLimit: 1, retryAfter: 20 } });
llmao.configure({ script: [{ error: 'rate_limit', retryAfter: 20, once: true }, { text: 'ok' }] });The errors your code already handles
| OpenAI SDK | Anthropic SDK | Vercel AI SDK | |
|---|---|---|---|
| Rate limit | OpenAI.RateLimitError, status: 429 |
Anthropic.RateLimitError, status: 429 |
APICallError, statusCode: 429 |
| Server error | OpenAI.InternalServerError, status: 500 |
Anthropic.InternalServerError |
APICallError, statusCode: 500 |
| Timeout | OpenAI.APIConnectionTimeoutError |
Anthropic.APIConnectionTimeoutError |
APICallError |
Rate limits come with a retry-after header on error.headers (a Headers object: read it with error.headers.get('retry-after'), not error.headers['retry-after']). In test mode it is 0 unless you set retryAfter, so your backoff doesn't slow the test down. To test that your code honors it, set a value and use fake timers (recipe for node:test).
Retries
The OpenAI and Anthropic clients retry like the official SDKs (maxRetries, default 2, waiting a short retry-after when there is one), and the AI SDK runs its own retries. Each attempt is recorded, so you can assert on them:
llmao.configure({ failures: { rateLimit: 1 } });
await expect(classify('hi')).rejects.toBeInstanceOf(OpenAI.RateLimitError);
expect(llmao.calls).toHaveLength(3); // the first attempt plus two retries
expect(llmao.calls.every((call) => call.failure === 'rate_limit')).toBe(true);If your code wraps the SDK in its own retry loop, this is also how you find out that you are retrying the SDK's retries.
Testing your own retry loop
The SDK retries first, so a single scripted failure never reaches your code: the SDK's second attempt succeeds. To exercise your loop, make every attempt of the SDK fail, or turn the SDK's retries off in the test:
// Every request fails: the SDK gives up and your loop sees the error
llmao.configure({ failures: { rateLimit: 1, retryAfter: 20 } });
// Or fail exactly one of your attempts: once per SDK request (maxRetries + 1, so 3 by default)
const rateLimit = { error: 'rate_limit', retryAfter: 20, once: true } as const;
llmao.configure({ script: [rateLimit, { ...rateLimit }, { ...rateLimit }, { text: 'ok' }] });Each once rule is consumed by one HTTP request, so copies are needed: the same object would count as one rule. With a real retry-after, both the SDK and your loop wait it out: drive the clock with fake timers (node:test recipe). The SDK numbers its attempts in the x-stainless-retry-count request header, recorded in llmao.calls[i].headers.
