Latency issues across a number of services

highGitHubActionsJul 23, 2026 07:53Duration: 1h 45m
apicapacity
Capacity IssueAPI Issue

Summary

On July 23, 2026, between 07:08 and 09:39 UTC, several services experienced delays: 8% of actions workflow runs experienced an average run start delay of 10 minutes, 5% of webhook deliveries exceeded SLO, and code scanning, repos, notifications, issues and pull requests experienced increased latency over the life of the incident. <br /><br />The root cause of the incident was a node of our background job processing system which did not recover after entering scheduled host maintenance. The i

Impact

major

Timeline

Jul 23, 2026 07:53

[investigating] We are investigating reports of degraded availability for Actions, Issues and Webhooks

via statuspage
+31m
Jul 23, 2026 08:25

[investigating] Pull Requests is experiencing degraded performance. We are continuing to investigate.

via statuspage
+9m
Jul 23, 2026 08:34

[investigating] We're currently investigating latency across multiple services. This can show as Actions jobs taking longer to start, Issues search serving stale results, and other listed services being similarly impacted.

via statuspage
+44m
Jul 23, 2026 09:18

[investigating] The degradation affecting Issues has been mitigated. We are monitoring to ensure stability.

via statuspage
+1m
Jul 23, 2026 09:19

[investigating] The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

via statuspage
+3m
Jul 23, 2026 09:22

[investigating] We identified the source of latency affecting multiple services and applied a fix. Issues and Actions are recovering, and remaining affected services are seeing improvement as processing backlogs clear. We are actively monitoring recovery across all services.

via statuspage
+5m
Jul 23, 2026 09:27

[investigating] The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.

via statuspage
+8m
Jul 23, 2026 09:35

[investigating] Webhooks is operating normally.

via statuspage
+4m
Jul 23, 2026 09:39

[monitoring] The degradation has been mitigated. We are monitoring to ensure stability.

via statuspage
+0m
Jul 23, 2026 09:39

[resolved] This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

via statuspage
+0m
Jul 23, 2026 09:39

[resolved] On July 23, 2026, between 07:08 and 09:39 UTC, several services experienced delays: 8% of actions workflow runs experienced an average run start delay of 10 minutes, 5% of webhook deliveries exceeded SLO, and code scanning, repos, notifications, issues and pull requests experienced increased latency over the life of the incident. <br /><br />The root cause of the incident was a node of our background job processing system which did not recover after entering scheduled host maintenance. The incident was mitigated by identifying the problematic shard and restoring its correct state, after which queue backlogs drained and services recovered. <br /><br />To speed mitigation, we have added monitors for nodes in this unhealthy state after maintenance operations. To prevent future recurrence, we are adapting our lifecycle automation to verify host rejoin after a scheduled reboot.

via statuspage

Lessons Learned

GitHub has experienced 127 incidents in the past year. This frequency suggests systemic reliability challenges that may warrant additional monitoring.

📊Incidents related to api, capacity have occurred 865 times across all providers in the past year. This is one of the most common failure categories in cloud infrastructure.

💡This incident is categorized as: Capacity Issue, API Issue. Consider implementing preventive measures specific to this failure category.