Hitting GPU Capacity for H100s creating large queue times for some models
Summary
This issue is resolved
Impact
minor
Timeline
[identified] We are over provisioned on H100s currently which is causing long queue times for some models running on H100s
via statuspage[monitoring] GPU usage is back below capacity. We're continuing to monitor.
via statuspage[monitoring] This issue is now resolved
via statuspage[resolved] This issue is resolved
via statuspageLessons Learned
⚠Replicate has experienced 39 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.
📊Incidents related to capacity have occurred 62 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.
💡This incident is categorized as: Capacity Issue. Consider implementing preventive measures specific to this failure category.
Similar Incidents
Degraded scale-out due to failed setups pulling from huggingface
Replicate · Jul 31, 2026
Latency issues across a number of services
GitHub · Jul 23, 2026
Newly entitled R2 accounts are seeing delays when trying to access the S3-compatible API
Cloudflare · Jul 21, 2026
Service degradation: Increased 5xx Errors
AWS · Jul 16, 2026
H100 GPU shortage resulting in high queue times
Replicate · Jul 15, 2026