Summary
We are back under max capacity for H100 hardware. Thank you for your patience!
Impact
minor
Timeline
[monitoring] We are seeing high contention on H100 hardware which is resulting in delays on predictions and scale-out.
via statuspage[resolved] We are back under max capacity for H100 hardware. Thank you for your patience!
via statuspageLessons Learned
⚠Replicate has experienced 43 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.
📊Incidents related to api, capacity have occurred 996 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.
💡This incident is categorized as: Capacity Issue, API Issue. Consider implementing preventive measures specific to this failure category.
Similar Incidents
Elevated number of R2 503 errors in Western North America region
Cloudflare · Sep 3, 2026
ChatGPT Work Mode High Error Rates
OpenAI · Sep 3, 2026
Elevated errors for Claude Sonnet 5
Anthropic · Sep 2, 2026
Elevated errors creating new accounts
OpenAI · Sep 2, 2026
Durable Objects increased errors in Western North America
Cloudflare · Sep 2, 2026