Hitting GPU Capacity for H100s creating large queue times for some models

lowReplicateJul 22, 2026 11:55
capacity
Capacity Issue

Summary

This issue is resolved

Impact

minor

Timeline

Jul 22, 2026 11:51

[identified] We are over provisioned on H100s currently which is causing long queue times for some models running on H100s

via statuspage
+5h 20m
Jul 22, 2026 17:11

[monitoring] GPU usage is back below capacity. We're continuing to monitor.

via statuspage
+9h 55m
Jul 23, 2026 03:05

[monitoring] This issue is now resolved

via statuspage
+11h 14m
Jul 23, 2026 14:19

[resolved] This issue is resolved

via statuspage

Lessons Learned

Replicate has experienced 39 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.

📊Incidents related to capacity have occurred 62 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.

💡This incident is categorized as: Capacity Issue. Consider implementing preventive measures specific to this failure category.