Posts

Showing posts with the label reliability

Circuit Breakers, Hedging, and Provider Fallbacks for Production LLM APIs in September 2026: Keep Latency SLOs When an Upstream Model Stalls

Image
When a production LLM API slows down or returns 5xx errors, your product feels it immediately. Chat UIs hang. Agents stall mid-tool loop. Batch jobs burn their wall-clock budget waiting on a single upstream. Cascades and cost-aware routing help you pick a cheaper model when quality allows. Canary and shadow traffic help you swap models safely. This post is about a different failure mode: the primary provider is healthy most of the time, then suddenly is not. You need timeouts, circuit breakers, hedging, and explicit fallbacks so one slow dependency does not take your SLO with it. Image: mikemacmarketing / photo on flickr via Wikimedia Commons (CC BY 2.0) Image: Midjourney; prompt suggested by Grok via Wikimedia Commons (Public domain) This is a practical playbook for API and platform teams running generative AI in production in September 2026. No fluff: what to measure, what to open and close, and how to wire fallbacks without inventing duplicate answers or double-...