Snapshot 15081
Normalized text
Scripts and page chrome removed; this is what change detection compares.
{
"pages": [
{
"end_time": "2026-09-30T23:59:59Z",
"months": [
{
"incidents": [
{
"code": "169pg72gg9jr",
"impact": "critical",
"message": "This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.",
"name": "Copilot Code Review is unable to complete reviews",
"timestamp": "Sep \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e21:16\u003c/var\u003e - \u003cvar data-var='time'\u003e22:08\u003c/var\u003e UTC"
},
{
"code": "1dk955gg3bvz",
"impact": "minor",
"message": "This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.",
"name": "Disruption with billing information updates",
"timestamp": "Sep \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e16:51\u003c/var\u003e - \u003cvar data-var='time'\u003e20:41\u003c/var\u003e UTC"
},
{
"code": "8zc63m64hy36",
"impact": "minor",
"message": "This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.",
"name": "Incident across several services",
"timestamp": "Sep \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e10:11\u003c/var\u003e - Sep \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e04:55\u003c/var\u003e UTC"
},
{
"code": "f6yrxnz5f7bs",
"impact": "minor",
"message": "On September 20, 2026, between 21:46 and 22:24 UTC the Pull Requests service was degraded and pull request merge and test-merge commits were created late, with delays reaching approximately four minutes at peak. Merge commits were delayed rather than lost. Because some Actions workflow runs start only after a pull request's merge commit is created, a subset of workflow runs for pull request events were also delayed. \u003cbr\u003e\u003cbr\u003eThis was due to a routine repository maintenance job for an unusually large repository consuming nearly all of the memory on a single Git storage server, which left that server unable to serve the Git operations used to create merge commits. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by removing the affected server from service at 22:18 UTC, after which the queued merge commits were created within six minutes. \u003cbr\u003e\u003cbr\u003eWe have capped the memory a single repository maintenance job may consume so that one repository cannot exhaust a server, and we have improved monitoring and alerting on storage server health to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Incident with Pull Requests",
"timestamp": "Sep \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e22:13\u003c/var\u003e - \u003cvar data-var='time'\u003e23:22\u003c/var\u003e UTC"
},
{
"code": "nvy2q96mcmyg",
"impact": "minor",
"message": "Between 20:26 and 21:17 UTC on September 17, 2026, GitHub Copilot experienced degradation affecting several GPT models, including GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol, GPT-5.3-Codex, and GPT-6 Astra. Users encountered elevated error rates when using these models.\u003cbr\u003e\u003cbr\u003eThe degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring and coordinated with the provider. Our automated model-warning system activated in-product warnings for affected models during the incident. Service returned to normal after the provider implemented a mitigation.",
"name": "Elevated rate of errors for OpenAI models provided by Copilot",
"timestamp": "Sep \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e20:59\u003c/var\u003e - \u003cvar data-var='time'\u003e21:49\u003c/var\u003e UTC"
},
{
"code": "nlxnbqnkdzdl",
"impact": "major",
"message": "On September 16, 2026, between 04:40 and 11:45 UTC, the Gemini 3.8 Flash model in GitHub Copilot experienced degraded availability. Requests to this model failed at an average rate of 6.4%, and the impact was highest during peak traffic hours. Other Copilot models were not affected. Users could continue to work with a different model or with 'Auto'.\u003cbr\u003e\u003cbr\u003eThe degradation was caused due to a capacity issue with an upstream model provider. Failure rates returned to normal as the daily traffic peak passed. We monitored the model until it was healthy and resolved the incident at 17:48 UTC. We are working to make our systems resilient to cover peak demand for all the Copilot models.",
"name": "Degradation with Gemini 3.8 Flash",
"timestamp": "Sep \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e07:21\u003c/var\u003e - \u003cvar data-var='time'\u003e17:48\u003c/var\u003e UTC"
},
{
"code": "bk7zgdcq7s9t",
"impact": "minor",
"message": "On September 15, 2026 between 15:30 and 20:00 UTC, some Copilot code reviews on pull requests failed to complete. The cause was increased latency in an internal caching service that GitHub Copilot Code Review relies on to coordinate its review jobs. This caused a timeout in lock acquisition, which interrupted the job. We reverted the change to the internal caching service and restored normal operation by 20:00 UTC.\u003cbr\u003e\u003cbr\u003eWe sincerely apologize for the disruption.",
"name": "Disruption with some GitHub services",
"timestamp": "Sep \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e19:11\u003c/var\u003e - \u003cvar data-var='time'\u003e20:00\u003c/var\u003e UTC"
},
{
"code": "qwdwmtqghpk5",
"impact": "minor",
"message": "On September 15, 2026, between 05:45 and 09:50 UTC, the Claude Fable 5.1 model in GitHub Copilot experienced intermittently degraded availability, with an average error rate of 2.8%. During brief recurring 15 minute periods that recurred every ~45 minutes, availability for Claude Fable 5.1 dropped to a maximum of ~40% before recovering completely. Other Copilot models were not affected. Users could continue to work with a different model or with 'Auto'.\u003cbr\u003e\u003cbr\u003eThe cause was an issue with an upstream model provider that intermittently rejected requests while overloaded. GitHub worked with the provider, who acknowledged and then resolved the underlying issue at 9:50 UTC, after which the model returned to constant normal availability. Once recovery was guaranteed, we resolved the incident at 11:17 UTC.\u003cbr\u003e\u003cbr\u003eTo reduce the chance of recurrence and customer impact, GitHub is reviewing per-model availability alerting and automatic in-product fallback so that requests to a degraded model can shift to a healthy alternative more quickly.",
"name": "Disruption with some GitHub services",
"timestamp": "Sep \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e09:47\u003c/var\u003e - \u003cvar data-var='time'\u003e11:17\u003c/var\u003e UTC"
},
{
"code": "kzh5j0mt7qwc",
"impact": "minor",
"message": "On September 14, 2026, between 16:10 and 19:01 UTC, some customers using GitHub Actions larger runners experienced longer-than-normal wait times for jobs to start. During this period, \u003cb\u003e5.7%\u003c/b\u003e of larger-runner jobs were affected. \u003cbr\u003e\u003cbr\u003eA routine expansion of our compute capacity exposed a bug in how our provisioning system handled capacity records when selecting where to create runner virtual machines. This slowed the creation of new runners, leaving insufficient runner capacity to start affected jobs promptly. \u003cbr\u003e\u003cbr\u003eWe restored normal provisioning by correcting the affected capacity records. We have fixed the underlying capacity-selection bug to prevent this failure from recurring. We have also added alerts for VM-record creation failures associated with this capacity issue.",
"name": "Actions Larger Runner Jobs for some customers may be slow to start",
"timestamp": "Sep \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e18:40\u003c/var\u003e - \u003cvar data-var='time'\u003e19:35\u003c/var\u003e UTC"
},
{
"code": "0rn90wk115q9",
"impact": "critical",
"message": "On September 13, 2026, between 08:43 and 10:44 UTC, GitHub experienced degraded availability across approximately 28 services, including Issues, Pull Requests, Actions, Codespaces, Pages, Notifications, Code Scanning, Git LFS, and new account signup. At peak, 8.8% of requests to create GitHub App installation access tokens failed. Token issuance for Actions workflows was also affected, impacting approximately 4% of workflows during the incident time frame. Creating issues through the web interface failed for about 96% of attempts, and signup failures were above 90%. \u003cbr\u003e \u003cbr\u003eThe cause was an internal data-cleanup job that began writing to a shared database cluster at 07:33 UTC. That cluster stores permission data read on nearly every authenticated request. The safeguard that was pacing the background job watched only one health signal — how far the database replicas were lagging — and that signal stayed low the whole time. It did not account for the load building on the primary itself, so the job kept writing while the primary quietly ran toward its limit. \u003cbr\u003e\u003cbr\u003eWhen the primary ran out of available connections, requests that needed it could not complete. First, there was no quick timeout on these database calls, so request handlers waited on the stalled database instead of failing fast, and the shared request-handling capacity degraded into site-wide errors. Second, a retry loop around token creation kept re-sending the writes that were already failing, which held the database saturated rather than letting it recover. \u003cbr\u003e\u003cbr\u003eMonitoring declared the incident at 08:50 UTC, but due to the broad impact and amplification from token creation, it took time to identify the source of the load. First responders mitigated by shedding internal load and pausing the job, and all services recovered by 10:44 UTC. \u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are rate-limiting background jobs against shared, customer-serving databases by default, and adding automatic pausing and paging on primary-server load rather than replication lag alone. We are also surfacing running background work directly alongside database health signals so responders can see and pause it without leaving those dashboards, bounding retries in the token-issuing path, and adding request-level timeouts so one unhealthy database cannot consume shared web server capacity. In addition, we are breaking apart this database cluster to remove the single point of failure. We will be moving various service-specific data, including the authorization data, out of this shared cluster in the next two weeks.",
"name": "Incident with several GitHub Services",
"timestamp": "Sep \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e09:16\u003c/var\u003e - \u003cvar data-var='time'\u003e10:44\u003c/var\u003e UTC"
},
{
"code": "zw1hbx2yyhvr",
"impact": "major",
"message": "On September 4, 2026, between 20:04 and 22:26 UTC, GitHub Copilot code review experienced an increased failure rate. Affected pull request reviews failed to complete or post review comments.\u003cbr\u003e\u003cbr\u003eThe incident was caused by a change to the service’s authentication permissions that prevented it from submitting affected reviews to the GitHub API. We reverted the change and restored normal operation by 22:26 UTC.\u003cbr\u003e\u003cbr\u003eWe apologize for the disruption.",
"name": "Disruption with Copilot Code Review",
"timestamp": "Sep \u003cvar data-var='date'\u003e4\u003c/var\u003e, \u003cvar data-var='time'\u003e20:39\u003c/var\u003e - \u003cvar data-var='time'\u003e22:26\u003c/var\u003e UTC"
},
{
"code": "r05vk75c594j",
"impact": "minor",
"message": "On September 4, 2026, between approximately 21:45 and 22:07 UTC, some users experienced errors and elevated latency for repository operations. The incident was fully resolved at 22:23 UTC.\u003cbr\u003e\u003cbr\u003eThe cause was a capacity change that spread one of our clusters across additional availability zones; our zone-aware traffic routing kept sending requests to the original zone for performance, overloading a small set of servers while the new capacity sat idle. We resolved the incident by reverting the change and letting traffic rebalance.\u003cbr\u003e\u003cbr\u003eWe are improving per-zone capacity guarantees, cross-zone load-shedding, and pre-production testing of multi-zone changes to prevent recurrence.",
"name": "Degradation in repos contents API",
"timestamp": "Sep \u003cvar data-var='date'\u003e4\u003c/var\u003e, \u003cvar data-var='time'\u003e22:02\u003c/var\u003e - \u003cvar data-var='time'\u003e22:23\u003c/var\u003e UTC"
},
{
"code": "ktdr5t0xwnhp",
"impact": "minor",
"message": "Between 13:22 and 17:11 UTC on September 03, 2026, GitHub Copilot experienced degradation affecting several Grok models, including Grok 4.5 and Grok 4.6. Users encountered elevated error rates, but other models were not affected. The degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring, displayed in-product warnings for the affected models, and coordinated with the provider. Service returned to normal after the provider implemented a mitigation.",
"name": "Incident with Grok Copilot AI Model Provider",
"timestamp": "Sep \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e14:17\u003c/var\u003e - \u003cvar data-var='time'\u003e17:11\u003c/var\u003e UTC"
},
{
"code": "jk2dy5h9yp3m",
"impact": "minor",
"message": "On September 1, 2026, between approximately 14:01 and 16:01 UTC, updates in response to pushes were delayed, temporarily showing stale diffs. The median time to refresh a diff after a push rose from the normal level of about 3 seconds to over 2 minutes at the peak, and more than 140,000 customer accounts had at least one delayed refresh during the most affected 75 minutes. Pushing commits and opening pull requests continued to work normally. The incident was caused by a sharp, concentrated surge in push volume that saturated worker pools and job queueing infrastructure. Autoscaling did not increase capacity as intended, so the backlog did not clear on its own. \u003cbr\u003e\u003cbr\u003eThe incident was mitigated by manually scaling the affected worker pools and increasing push-processing capacity. This allowed the system to process the backlog, after which refresh times returned to normal. To reduce the likelihood and impact of similar incidents, we are adding quotas and throttling earlier in the push path so a single concentrated source of load cannot saturate shared capacity, improving worker-pool autoscaling so capacity is added automatically, and improving monitors for background job processing so on-call is paged before customers experience delayed pull request updates.",
"name": "Delays in commit processing",
"timestamp": "Sep \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e15:00\u003c/var\u003e - \u003cvar data-var='time'\u003e16:01\u003c/var\u003e UTC"
}
],
"name": "September",
"year": 2026
},
{
"incidents": [
{
"code": "7fxts6gmq5gr",
"impact": "minor",
"message": "Between 08:37 and 09:41 UTC on August 31, 2026, GitHub Copilot experienced degradation affecting several GPT models, including gpt-5.2, gpt-5.3-codex, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, and the gpt-5.6 family (Luna, Sol, and Terra). Users encountered elevated error rates and interrupted streaming responses. Other models were not affected.\u003cbr\u003e\u003cbr\u003eThe degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring, displayed in-product warnings for the affected models, and coordinated with the provider. Service returned to normal after the provider implemented a mitigation.",
"name": "Elevated rate of errors for OpenAI models provided by Copilot",
"timestamp": "Aug \u003cvar data-var='date'\u003e31\u003c/var\u003e, \u003cvar data-var='time'\u003e09:15\u003c/var\u003e - \u003cvar data-var='time'\u003e09:58\u003c/var\u003e UTC"
},
{
"code": "5bn0vk444m1w",
"impact": "minor",
"message": "On August 26, 2026, between 20:40 UTC and 00:51 UTC on August 27, GitHub Billing experienced degraded performance affecting billing budget pages and GitHub Copilot CLI sessions. Affected customers encountered failed budget page loads or failures when starting or continuing CLI sessions. We confirmed this impact for a small number of customers (\u0026lt;1%). \u003cbr\u003e\u003cbr\u003eThis was caused by a concentrated workload that created processing delays in our data storage layer. Automated retries increased the load and prolonged the degradation. We mitigated the incident by rebalancing traffic within our infrastructure. \u003cbr\u003e\u003cbr\u003eWe are improving workload isolation, retry behavior, and detection of concentrated load to reduce the likelihood of recurrence and shorten our time to detect and mitigate similar incidents.",
"name": "Disruption with GitHub Billing",
"timestamp": "Aug \u003cvar data-var='date'\u003e26\u003c/var\u003e, \u003cvar data-var='time'\u003e23:37\u003c/var\u003e - Aug \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e19:44\u003c/var\u003e UTC"
},
{
"code": "tx9qn4khd664",
"impact": "critical",
"message": "On August 27th, 2026, between approximately 09:20 and 12:14 UTC, the Copilot service experienced a degradation of the Kimi K3 model due to an issue with our upstream provider. Users encountered elevated error rates when using Kimi K3. No other models were impacted. \u003cbr\u003e\u003cbr\u003eThe issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.",
"name": "Incident with Copilot AI Model Providers",
"timestamp": "Aug \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e10:04\u003c/var\u003e - \u003cvar data-var='time'\u003e12:12\u003c/var\u003e UTC"
},
{
"code": "kfspvrz14xr0",
"impact": "minor",
"message": "On August 26, 2026, from 21:55 UTC to 23:58 UTC, 2.6% of workflow runs triggered by pull request events were delayed, with the impact rising as high as 25% at its peak. Some users also experienced delays in pull request merge-commit generation, mergeability information, and merge-button availability. Actions and Pull Requests fully recovered by 23:58 UTC; the incident was resolved at 00:26 UTC after normal operation was confirmed. \u003cbr\u003e\u003cbr\u003eBackground jobs that process pull request updates and generate merge commits were impacted by timeouts reaching a single partition of git data. This resulted in a backlog in pull request merge-commit processing, delaying pull request-triggered GitHub Actions workflows and some mergeability information. \u003cbr\u003e\u003cbr\u003eWe reduced workload, shifted traffic away from affected infrastructure, and restored the affected service component to a healthy state. Together, these actions helped drain the backlog and restore normal operations. \u003cbr\u003e\u003cbr\u003eWe are working to improve resource saturation detection and to eliminate customer impact in this scenario by isolating impact, placing better bounds on retries, and strengthening backpressure to make our systems more resilient under load.",
"name": "Incident with Actions and Pull Requests",
"timestamp": "Aug \u003cvar data-var='date'\u003e26\u003c/var\u003e, \u003cvar data-var='time'\u003e22:56\u003c/var\u003e - Aug \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e00:26\u003c/var\u003e UTC"
},
{
"code": "y1t7p9fzrlj2",
"impact": "critical",
"message": "On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. \u003cbr\u003e\u003cbr\u003eAt 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. \u003cbr\u003e\u003cbr\u003e3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs. \u003cbr\u003e\u003cbr\u003eCustomers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27. \u003cbr\u003e\u003cbr\u003eSome runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued even after service recovery. In the backend, these runs had already failed and will automatically move to canceled state 24 hours after creation. As follow-up, we are fixing the root cause of this queued state and improving our ability to bulk-cancel affected runs. \u003cbr\u003e\u003cbr\u003eSeveral changes to improve the general scalability of this part of Actions were already complete and deploying to production. Rollout of those changes will be complete within the next 24 hours. Further work to improve scale, resiliency, and more graceful degradation of Actions workflows are in flight. We are also taking a repair item to accelerate clearing of stuck queued or waiting jobs in similar future cases.",
"name": "Incident with Actions",
"timestamp": "Aug \u003cvar data-var='date'\u003e26\u003c/var\u003e, \u003cvar data-var='time'\u003e15:11\u003c/var\u003e - \u003cvar data-var='time'\u003e18:01\u003c/var\u003e UTC"
},
{
"code": "hcbtzksccj2f",
"impact": "minor",
"message": "Please refer to the combined summary in this related incident: https://www.githubstatus.com/incidents/y1t7p9fzrlj2",
"name": "Disruption with some GitHub services",
"timestamp": "Aug \u003cvar data-var='date'\u003e26\u003c/var\u003e, \u003cvar data-var='time'\u003e15:09\u003c/var\u003e - \u003cvar data-var='time'\u003e16:07\u003c/var\u003e UTC"
},
{
"code": "lyppgxbq1nyk",
"impact": "minor",
"message": "On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. \u003cbr\u003e \u003cbr\u003eThe incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. \u003cbr\u003e\u003cbr\u003eTo prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.",
"name": "Actions delays in starting runs",
"timestamp": "Aug \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e13:56\u003c/var\u003e - \u003cvar data-var='time'\u003e14:34\u003c/var\u003e UTC"
},
{
"code": "wt3hjqcrczfg",
"impact": "major",
"message": "On August 24th, 2026, between approximately 06:35 and 07:25 UTC, the Copilot service experienced a degradation of the Claude Fable 5 model due to an issue with our upstream provider. Users encountered elevated error rates when using Claude Fable 5, with requests sometimes failing mid-response. No other models were impacted.\u003cbr\u003e\u003cbr\u003eThe issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.",
"name": "Elevated errors on Fable 5 due to upstream provider",
"timestamp": "Aug \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e07:12\u003c/var\u003e - \u003cvar data-var='time'\u003e07:58\u003c/var\u003e UTC"
},
{
"code": "wms44hv62t3p",
"impact": "none",
"message": "On August 21, 2026, between 14:00 and 14:07 UTC, dotcom Git operations over SSH were degraded. Successful Git operations over SSH fell by more than 95% for during the peak impact window, making clone, fetch, or push over SSH effectively unavailable to most users for approximately four minutes. Git operations over HTTPS were not affected. \n\nThe incident was caused by a software defect in our load-balancing infrastructure that was triggered by a configuration change. The defect only occurred when connections passed through multiple layers of load balancers running the new configuration, which meant it was not detected during canary testing. \n\nWe mitigated the incident by rolling back the configuration change. \n\nWe are adding regression coverage for multi-layer load-balancer configurations and improving monitoring and alerting for Git operations over SSH to reduce our time to detection and mitigation of similar issues in the future.",
"name": "Degraded Git Operations over SSH",
"timestamp": "Aug \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e14:00\u003c/var\u003e - \u003cvar data-var='time'\u003e14:00\u003c/var\u003e UTC"
},
{
"code": "bhbcjn4n3jzp",
"impact": "critical",
"message": "Between 13:57 UTC on August 20 and 00:37 UTC on August 21, 2026, some users of the Copilot Cloud Agent experienced delays of up to 60 to 90 minutes in seeing the status and results of their agent tasks. The agent tasks themselves continued to run and complete during this time; only the visibility of their status was delayed.\u003cbr\u003e\u003cbr\u003eThe cause was a regional outage in a third-party cloud database service that Copilot uses to store agent task status. We failed over the affected database to a healthy region, added processing capacity to work through the backlog, and restored normal operation once the underlying service recovered. No task data was lost during the incident.\u003cbr\u003e\u003cbr\u003eTo prevent repetition of similar incidents, we are removing the database configuration that made us vulnerable to this regional outage and improving our database failover procedures.",
"name": "Intermittent failures creating agent tasks",
"timestamp": "Aug \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e14:43\u003c/var\u003e - Aug \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e00:37\u003c/var\u003e UTC"
},
{
"code": "bmpybhnrky3x",
"impact": "minor",
"message": "On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API. \u003cbr\u003e\u003cbr\u003eThe issue was caused by failures in backend requests reading runner and runner group data. The failures were caused by an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents triggered by this operation. \u003cbr\u003e\u003cbr\u003eThe impact was mitigated by completing the enablement of the new certificate in the backend system. We have added additional monitoring to this and other certificates. This service is also in the process of being replaced as part of our availability and scale work, bringing this authentication path and secret management in line with patterns across all GitHub services.",
"name": "Intermittent failures in runner group and runner-related permissions pages",
"timestamp": "Aug \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e07:40\u003c/var\u003e - \u003cvar data-var='time'\u003e11:42\u003c/var\u003e UTC"
},
{
"code": "gx7js8bd0jpz",
"impact": "major",
"message": "On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to run jobs on Actions Larger Runners and were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API. \u003cbr\u003e\u003cbr\u003eThese issues were caused by failures in backend requests resolving essential metadata for starting Larger Runner workflow runs and for reading runner and runner group data. The failures were caused by an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents that had been triggered by this operation. \u003cbr\u003e\u003cbr\u003eWe mitigated the issues by completing the enablement of the new certificate in the backend system. We have added additional monitoring to this and other certificates. The relevant service is also in the process of being replaced as part of our availability and scale work, bringing this authentication path and secret management in line with patterns across all GitHub services.",
"name": "Incident with Actions",
"timestamp": "Aug \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e09:36\u003c/var\u003e - \u003cvar data-var='time'\u003e10:23\u003c/var\u003e UTC"
},
{
"code": "zkxwbgr0cnmx",
"impact": "critical",
"message": "On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02. \u003cbr\u003e\u003cbr\u003eSome of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service. \u003cbr\u003e\u003cbr\u003eThe immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery. \u003cbr\u003e\u003cbr\u003eThe retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed. \u003cbr\u003e\u003cbr\u003eResidual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery. \u003cbr\u003e\u003cbr\u003eComplicating factors that impeded recovery included a number of scraping attacks on codeload endpoints. \u003cbr\u003e\u003cbr\u003eTo prevent recurrence, our follow-up actions include: \u003cbr\u003e\u003cbr\u003e- Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity. \u003cbr\u003e\u003cbr\u003e- Auditing Istio request, concurrency, and scaling limits across affected services. \u003cbr\u003e\u003cbr\u003e- Reviewing retry limits and backoff behavior across gateways and clients. \u003cbr\u003e\u003cbr\u003e- Addressing the VS Code retry behavior that amplified Copilot token traffic. \u003cbr\u003e\u003cbr\u003e- Improving load-balancer capacity monitoring and regional failover safeguards.",
"name": "Incident with GitHub.com",
"timestamp": "Aug \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e13:40\u003c/var\u003e - \u003cvar data-var='time'\u003e21:15\u003c/var\u003e UTC"
},
{
"code": "pf25whpq58hh",
"impact": "minor",
"message": "On August 13, 2026, from 15:31:21 UTC to 18:27:55 UTC, GitHub Enterprise Cloud team synchronization was degraded for enterprises using personal accounts. Organization teams experienced delays of up to 3 to 13 hours (median 8 hours) when syncing with IdP groups, resulting in delayed access grants or removals for enterprise users across 2.8% of teams. \u003cbr\u003e\u003cbr\u003eA temporary change introduced to address a previous issue due to increased usage of this feature remained active after it was intended to be removed, causing synchronization delays during periods of high volume. We removed the temporary change and provisioned additional resources to handle the increased volume.",
"name": "Disruption with GHEC Team Sync",
"timestamp": "Aug \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e16:21\u003c/var\u003e - \u003cvar data-var='time'\u003e18:27\u003c/var\u003e UTC"
},
{
"code": "24t8gsgqx2qb",
"impact": "minor",
"message": "On August 13th, 2026, between approximately 14:06 and 15:47 UTC, the Copilot service experienced a degradation of the Claude Fable 5 model due to an issue with our upstream provider. Users encountered elevated error rates, peaking at 43% and averaging 12%. Users who selected Auto or alternative models were unaffected.\u003cbr\u003e\u003cbr\u003eThe issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.",
"name": "Errors with the Fable 5 Model in Copilot",
"timestamp": "Aug \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e14:43\u003c/var\u003e - \u003cvar data-var='time'\u003e15:47\u003c/var\u003e UTC"
},
{
"code": "k8vbzwqjkxzn",
"impact": "minor",
"message": "Between 14:24 and 14:53 UTC on 13 August 2026, a routine background job to delete an organization overwhelmed a key shared database, causing multiple GitHub services to briefly return elevated errors and slower responses. Most affected was the webhook management API, with smaller impact to Git operations, pull requests, issues, packages, sign-in, and Copilot. Impact cleared on its own at about 14:53 UTC once the job finished; we resolved the incident at 15:36 UTC. \u003cbr\u003e\u003cbr\u003eAffected users may have experienced a brief increase in errors and slower responses, primarily when creating, listing, or updating webhooks, with smaller impacts to pull requests, issues, packages, and Git operations. Failures peaked at about 1% for several minutes around 14:37 UTC. \u003cbr\u003e\u003cbr\u003eTo prevent future incidents, we've already shipped an update that turns on the safer deletion path for organizations, along with caps on deletion holds on databases. Building on these changes, we're auditing all bulk deletion and cleanup jobs that write to shared databases to prevent similar issues in future.",
"name": "Incident with Webhooks",
"timestamp": "Aug \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e14:45\u003c/var\u003e - \u003cvar data-var='time'\u003e15:36\u003c/var\u003e UTC"
},
{
"code": "lsvy8xsf0gxv",
"impact": "major",
"message": "On August 12 and 13, 2026, some anonymous (logged-out) requests to github.com experienced HTTP 5xx errors when loading pages like the sign-in page, and when downloading release assets, due to an unusual traffic pattern that repeatedly overloaded a part of our infrastructure that serves these types of requests. There were three windows of impact: (1) August 12 from 16:34 to 18:34 UTC, with an average error rate of 16.16% that peaked at 28.6%; (2) August 12 from 19:00 to 22:56 UTC, with an average error rate of 16.55% that peaked at 24.18%; and (3) August 13 from 06:19 to 08:05 UTC, with an average error rate of 2.01% that peaked at 7.49%.\u003cbr\u003eRequests from signed-in users were unaffected.\u003cbr\u003e\u003cbr\u003eWe mitigated the incidents by applying traffic controls at our network edge that limited any requests matching the pattern identified previously, thereby preventing overload on our systems.\u003cbr\u003e\u003cbr\u003eSince these incidents occurred, we have tightened our monitoring systems to alert server-side errors that affect logged-out traffic. We are also working to further strengthen our edge protections and reduce the time to detect and mitigate similar incidents.",
"name": "Disruption with Login and Release Asset downloads",
"timestamp": "Aug \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e21:39\u003c/var\u003e - \u003cvar data-var='time'\u003e22:56\u003c/var\u003e UTC"
},
{
"code": "76t89hbfb09h",
"impact": "minor",
"message": "Between 16:03 and 16:29 UTC on August 12, some users encountered errors when viewing pull requests, issues, and search results. During this period, about 1.9% of Pull Request requests and 0.9% of Issues requests failed. During a database migration, two indexes were removed while application settings still referenced them, causing affected requests to fail. We detected the issue after the migration reached one database shard and before it progressed to the remaining shards. We restored service by disabling both settings. We are improving safeguards around database migrations and application configuration to prevent similar mismatches from causing errors.",
"name": "Incident with Pull Requests and Issues",
"timestamp": "Aug \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e16:16\u003c/var\u003e - \u003cvar data-var='time'\u003e16:41\u003c/var\u003e UTC"
},
{
"code": "vm1w8zq95wkt",
"impact": "minor",
"message": "On August 11, 2026, between 14:00 UTC and 16:00 UTC the GraphQL API service was degraded and customers in saw higher than normal timeouts. On average, the timeout rate was 0.06% and peaked at 0.14% of requests routing to the service. \u003cbr\u003e\u003cbr\u003eThis was due to increased utilization at one of our sites which caused resource contention across our dependencies, leading to an increase in timeouts for GraphQL requests. We mitigated the incident by increasing capacity to alleviate the capacity bottleneck. \u003cbr\u003e\u003cbr\u003eWe are working to improve our monitoring so that we can proactively reduce the impact of high consumption requests in addition to scaling up; Additionally, we will improve our time to detection and mitigation of issues like this one in the future.",
"name": "Incident with GraphQL API Requests",
"timestamp": "Aug \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e14:50\u003c/var\u003e - \u003cvar data-var='time'\u003e20:06\u003c/var\u003e UTC"
},
{
"code": "1t7x65n7kvt1",
"impact": "minor",
"message": "On August 10, 2026, between 19:48 UTC and 20:49 UTC, GitHub Copilot users saw an incomplete list of available models. During this window, the service could return as few as one model instead of the full catalog. Requests that tried to use a model missing from that shortened list failed with a \"model not found\" error. Copilot requests that used an available model were not affected. This did not affect customers on data-residency (Proxima) environments.\u003cbr\u003e\u003cbr\u003eThe issue was caused by a change to how model data was published, which our systems could not read back correctly and fell back to a limited default list.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by 20:49 UTC and deployed a fix to prevent immediate recurrence by 21:50 UTC. We are adding validation and retry safeguards so that model data is verified before it is served.\u003cbr\u003e\u003cbr\u003eWe apologize for the disruption.",
"name": "Disruption with Copilot for access to some models",
"timestamp": "Aug \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e20:27\u003c/var\u003e - \u003cvar data-var='time'\u003e21:50\u003c/var\u003e UTC"
},
{
"code": "s19bth7wzkf7",
"impact": "minor",
"message": "On August 10, 2026, between 17:16 and 18:21 UTC, users were unable to create new fine-grained personal access tokens (FG PAT) through the GitHub website. When a user submitted the FG PAT creation form, they were returned to the FG PAT list without an error message and no FG PAT was created. Creating classic personal access tokens, as well as editing or deleting existing FG PAT were not affected.\u003cbr\u003e\u003cbr\u003eThe cause was a change to how the website loads certain front-end JavaScript that was enabled for all users at 17:15 UTC; the change interacted with an issue in the token creation form's confirmation step that prevented it from running, so the final submission that actually creates the token never completed. Because the page still loaded and the server returned a normal response, the failure produced no error message. GitHub mitigated the incident by disabling the change at 18:21 UTC, at which point token creation recovered immediately, and the incident was resolved at 18:46 UTC.\u003cbr\u003e\u003cbr\u003eTo reduce the chance of recurrence, GitHub is adding monitoring and alerting for anomalies in the FG PAT creation success rate and is removing the issue in the FG PAT creation form that prevented the confirmation step from running. GitHub is also adding automated detection of the issue so other areas of the GitHub front end do not repeat the problem.",
"name": "Disruption with creation of fine grained personal access tokens",
"timestamp": "Aug \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e18:02\u003c/var\u003e - \u003cvar data-var='time'\u003e18:46\u003c/var\u003e UTC"
},
{
"code": "qcvjkzcs7j74",
"impact": "critical",
"message": "On August 6, 2026, between 15:05 UTC and 00:14 UTC on August 7, GitHub Actions experienced degraded availability. During the incident, workflow runs failed or remained queued for an extended period of time. Customers using both GitHub-hosted and self-hosted runners were affected. At peak, 71% of workflow runs experienced infrastructure failures and 75% of the remaining workflow runs were delayed by more than 5 minutes. \u003cbr\u003e\u003cbr\u003eThe incident was triggered by a routine deployment to an internal Actions service responsible for processing events and generating Actions jobs. The deployment exposed an existing capacity and concurrency weakness. As pods were replaced during the deployment, remaining capacity became saturated, causing services to crash and triggering a cascading impact across multiple clusters and downstream services. \u003cbr\u003e\u003cbr\u003eThese services recovered at 17:00 after expanding capacity, throttling incoming webhook-triggered work to allow the system to recover, and increasing processing capacity for the backlog of affected events. \u003cbr\u003e\u003cbr\u003eAs the incident progressed, a backlog of work accumulated across the systems responsible for assigning jobs to runners. Due to a latent bug in one of the services responsible for job assignment, runners were getting assigned jobs that were no longer valid and then getting stuck retrying those jobs, preventing them from picking up valid work. \u003cbr\u003e\u003cbr\u003eThis second stage of impact was mitigated by deploying changes to prevent runners from repeatedly attempting to acquire invalid jobs. These mitigations allowed the accumulated queues to drain and Actions to recover to normal operation. \u003cbr\u003e\u003cbr\u003eSome Actions Runner Controller (ARC) runners remained stuck after the incident. A mitigation deployed during the incident inadvertently affected these runners, causing some to remain offline until they were manually recovered. We subsequently rolled back the change and are adding automatic recovery in upcoming Runner and ARC releases. \u003cbr\u003e\u003cbr\u003eSome jobs created during the incident were also left stuck unable to be retried or canceled. CLI and UI solutions for customers to address these were shared at https://github.com/orgs/community/discussions/204152#discussioncomment-17946043. \u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are making improvements to deployment and capacity safeguards for the affected services, strengthening monitoring for the conditions that preceded the incident, improving the resiliency and recovery of queued work and runner assignment, and adding automatic recovery for self-hosted runners affected by similar failure conditions. We are also making additional improvements to reduce the risk of cascading failures and accelerate recovery during large-scale Actions disruptions.",
"name": "Incident with Actions",
"timestamp": "Aug \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e15:22\u003c/var\u003e - Aug \u003cvar data-var='date'\u003e7\u003c/var\u003e, \u003cvar data-var='time'\u003e02:04\u003c/var\u003e UTC"
},
{
"code": "1xqvzyv99skw",
"impact": "minor",
"message": "On August 6, 2026, at 07:00 UTC, a configuration change inadvertently reduced the capacity of the service that processes GitHub Pages deployments. As traffic increased over the following hours, latency in the deployment pipeline progressively increased. \u003cbr\u003e\u003cbr\u003eAt 12:09 UTC, latency crossed the alerting threshold and the team began investigating. We reverted the invalid configuration and applied additional mitigations, including reducing status deployment processing to lower the load on our Redis cluster. Latency returned to normal levels at 15:40 UTC. \u003cbr\u003e\u003cbr\u003eCustomer impact occurred from 11:34 to 15:32 UTC. During this period, we failed to process approximately 128,000 deployments. \u003cbr\u003e\u003cbr\u003eWe have updated our alerts to detect elevated processing latency sooner and to notify us immediately when latency causes deployment processing failures. We've confirmed this incident was not fully captured by our availability metrics. In the coming days, we'll update how GitHub Pages availability is measured so incidents like this are accurately reflected going forward.",
"name": "Incident with Pages - Deployment Lag",
"timestamp": "Aug \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e15:03\u003c/var\u003e - \u003cvar data-var='time'\u003e16:22\u003c/var\u003e UTC"
},
{
"code": "9wzhl5jr80jt",
"impact": "minor",
"message": "On August 5, 2026, between 11:02 and 11:54 UTC, the GitHub Copilot cloud agent service was degraded and new cloud agent jobs were delayed from starting. During this period 100% of newly submitted agent jobs were affected. The incident was limited to delay of cloud agent jobs. No jobs were lost and the queued backlog was processed by 13:00 UTC. This was due to an internal rate limit used to protect service availability that was enabled more broadly than intended delaying more traffic than expected. \u003cbr\u003e \u003cbr\u003eThe service recovered when the rate limit window expired. We then tuned the control so it no longer affected unrelated coding agent traffic. \u003cbr\u003e \u003cbr\u003eWe are working to improve the control's scoping and our monitoring and alerting to reduce our time to detection and mitigation of similar issues in the future.",
"name": "Some Copilot Cloud Agent jobs not starting",
"timestamp": "Aug \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e11:38\u003c/var\u003e - \u003cvar data-var='time'\u003e13:00\u003c/var\u003e UTC"
},
{
"code": "7s119p1yxttr",
"impact": "minor",
"message": "On 2026-08-03, between 06:52 and 11:25 UTC, some GitHub Copilot users experienced errors when using chat and agent features. Requests to list the available models failed, and because every chat or agent interaction begins by retrieving the list of models, affected users saw their requests fail. On average about 3% of these model-listing requests failed during the incident (roughly 97% succeeded), but failures were significantly higher during peak-traffic periods, at times approaching 100% for the affected internal lookups. Approximately 4,066 users were affected in a single 60-minute window, concentrated among IDE-based clients. The underlying AI models themselves remained healthy throughout.\u003cbr\u003e\u003cbr\u003eThe incident was caused by an increase in how often clients requested the model list, which pushed an internal user-authorization lookup past a rate limit; the rate-limited responses were surfaced to users as errors. We mitigated the impact by increasing how long Copilot caches that authorization lookup, which reduced load on the internal service, and we have additional capacity and rate-limit changes in progress. To prevent recurrence we are improving monitoring for this class of failure, adjusting cache and rate-limit settings, and coordinating with client teams on request patterns.",
"name": "Incident with Copilot",
"timestamp": "Aug \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e09:53\u003c/var\u003e - \u003cvar data-var='time'\u003e11:25\u003c/var\u003e UTC"
},
{
"code": "sj1tzyrx599x",
"impact": "minor",
"message": "On August 1, 2026, between 17:47 UTC and 18:20 UTC, users of the Fable 5 model in GitHub Copilot experienced increased request failures and latency. The average failure rate across all Copilot requests was 0.007%, while failures for Fable 5 peaked at 5.6%. Other models remained available. This was caused by degradation of an upstream model provider.\u003cbr\u003e\u003cbr\u003eThe affected endpoint recovered, and we monitored the service until error rates and latency returned to normal levels. We are working to add endpoint redundancy to mitigate similar provider issues in the future.",
"name": "Incident with Copilot AI Model Providers",
"timestamp": "Aug \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e18:03\u003c/var\u003e - \u003cvar data-var='time'\u003e18:44\u003c/var\u003e UTC"
},
{
"code": "kk183dslzdzd",
"impact": "minor",
"message": "On August 1st, 2026, the GPT-5.6 Luna model in GitHub Copilot experienced degraded availability in intermittent time intervals between ~08:05 UTC and ~16:30 UTC. Specifically the timeframes observed were 10:00-10:20 UTC, 10:45-11:50 UTC, 13:00-14:25 UTC, and 16:00-16:30 UTC. During this time, requests to GPT-5.6 Luna in Copilot chat and IDE surfaces frequently failed or timed out. This was caused by an issue with an upstream model provider. Other Copilot models were not affected, and users could continue working by selecting another model or 'Auto'. Availability for GPT-5.6 Luna fully recovered once the provider resolved their outage at 16:30 UTC.",
"name": "Degraded availability GPT 5.6 Luna",
"timestamp": "Aug \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e11:16\u003c/var\u003e - \u003cvar data-var='time'\u003e12:30\u003c/var\u003e UTC"
}
],
"name": "August",
"year": 2026
},
{
"incidents": [
{
"code": "9tpgqq1h4bs7",
"impact": "minor",
"message": "On July 30, 2026, the Claude Fable 5 model in GitHub Copilot experienced degraded availability for approximately 73 minutes, from 08:33 to 09:46 UTC. During this time, requests to Claude Fable 5 in Copilot chat and IDE surfaces frequently failed or timed out. This was caused by an issue with an upstream model provider. Other Copilot models were not affected, and users could continue working by selecting another model or 'Auto'. Availability for Claude Fable 5 fully recovered once the provider resolved their outage at 09:46 UTC, and we confirmed resolution at 10:12 UTC.",
"name": "Copilot model Claude Fable 5 experiencing elevated errors",
"timestamp": "Jul \u003cvar data-var='date'\u003e30\u003c/var\u003e, \u003cvar data-var='time'\u003e09:07\u003c/var\u003e - \u003cvar data-var='time'\u003e10:12\u003c/var\u003e UTC"
},
{
"code": "dsrfymph7my0",
"impact": "minor",
"message": "On July 29, 2026, between 19:45 UTC and 21:51 UTC, users of the Fable 5 model in GitHub Copilot experienced increased request failures and latency. The average failure rate across all Copilot requests was 0.006%, while failures for Fable 5 peaked at 21%. Other models remained available. This was caused by degradation of an upstream model provider.\u003cbr\u003e\u003cbr\u003eThe affected endpoint recovered, and we monitored the service until error rates and latency returned to normal levels. We are working to add endpoint redundancy to mitigate similar provider issues in the future.",
"name": "Incident with Copilot AI Model Providers",
"timestamp": "Jul \u003cvar data-var='date'\u003e29\u003c/var\u003e, \u003cvar data-var='time'\u003e20:07\u003c/var\u003e - \u003cvar data-var='time'\u003e21:51\u003c/var\u003e UTC"
},
{
"code": "75g5xmzptjqb",
"impact": "major",
"message": "On July 29, 2026, from 14:51 UTC to 15:28 UTC, GitHub Actions experienced elevated REST API request timeouts and errors, failures registering runners, and delayed workflow run starts for customers whose traffic was served by a single infrastructure site. This was caused by an under-provisioned internal Actions service in that site: under increased load its instances ran out of memory and became unresponsive, and because Actions API requests wait synchronously on that service, requests routed through the affected site stalled and timed out. During the incident, approximately 2% of workflows were delayed. Requests served by other sites remained unaffected. Both standard and larger hosted runners routed through the affected site could see delayed job starts. \u003cbr\u003e\u003cbr\u003eThe issue was mitigated by scaling out the runner-administration service in the affected site and increasing the replica count, which restored API availability and returned workflow run starts to normal. We are working to add horizontal autoscaling, memory-saturation alerting, and scaling-forecast monitoring for this service, along with responder playbooks, to reduce the likelihood of similar issues in the future.",
"name": "Incident with Actions",
"timestamp": "Jul \u003cvar data-var='date'\u003e29\u003c/var\u003e, \u003cvar data-var='time'\u003e15:26\u003c/var\u003e - \u003cvar data-var='time'\u003e16:00\u003c/var\u003e UTC"
},
{
"code": "fmmsrcg5x638",
"impact": "minor",
"message": "On July 26, 2026 at 21:34 UTC we began seeing intermittent errors on the GitHub GraphQL API. A subset of GraphQL API requests returned HTTP 502 errors in short bursts. During the impact window an average of 0.09% of GraphQL API requests in the affected region failed, with a peak of 0.50% of requests failing during the worst two-minute period at 03:02 UTC on July 27. Requests that failed generally succeeded when retried, and no data was lost or altered. Other GitHub services were not affected.\u003cbr\u003e\u003cbr\u003eThe errors were traced to a single group of servers handling a share of GraphQL API traffic. Application processes on that group intermittently closed connections before completing responses. Impact ended at 03:52 UTC on July 27 when those processes were replaced, and we resolved the incident at 04:09 UTC on July 27 after confirming error rates had returned to normal.\u003cbr\u003e\u003cbr\u003eWe are still investigating why those processes closed connections, and that work is being carried out by the team that owns the underlying compute platform. In the meantime we are adding detection and automated mitigation for when a single group of servers behaves differently from its peers.",
"name": "Incident with GraphQL API Requests",
"timestamp": "Jul \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e03:53\u003c/var\u003e - \u003cvar data-var='time'\u003e04:09\u003c/var\u003e UTC"
},
{
"code": "pz7g535gbs6p",
"impact": "critical",
"message": "Please refer to the combined summary in this related incident: https://www.githubstatus.com/incidents/s65j9gslmfm8",
"name": "Actions run failures and delays",
"timestamp": "Jul \u003cvar data-var='date'\u003e25\u003c/var\u003e, \u003cvar data-var='time'\u003e12:31\u003c/var\u003e - \u003cvar data-var='time'\u003e13:13\u003c/var\u003e UTC"
},
{
"code": "vv9vvksmj4s9",
"impact": "major",
"message": "On July 25, 2026, between 09:07 and 10:04 UTC, the GPT-5.2, GPT-5.3-Codex, GPT-5.4, GPT-5.4 Mini, GPT-5.6 Sol, GPT-5.6 Terra and GPT-5.6 Luna models experienced degraded availability in GitHub Copilot products and IDE surfaces. Requests to these models had an average failure rate of 5.6%. Other Copilot models remained available as alternatives.\u003cbr\u003e\u003cbr\u003eThe degradation was caused by an issue with an upstream model provider. Success rates returned to normal after the upstream issue was mitigated, and we continued monitoring before resolving the incident. We are working on improving the automated failover for the affected models to prevent similar incidents in the future.",
"name": "Several GPT models degraded",
"timestamp": "Jul \u003cvar data-var='date'\u003e25\u003c/var\u003e, \u003cvar data-var='time'\u003e09:42\u003c/var\u003e - \u003cvar data-var='time'\u003e10:11\u003c/var\u003e UTC"
},
{
"code": "s65j9gslmfm8",
"impact": "minor",
"message": "On July 25, 2026, GitHub Actions experienced two related periods of degradation that caused some workflow runs to be delayed by more than 5 minutes or end with infrastructure failures. \u003cbr\u003e\u003cbr\u003eFirst period (08:45 – 09:13 UTC): During planned maintenance on a critical-path Redis cluster for Actions, one participating region was left in a degraded state. Separately, an independent capacity operation temporarily removed another region from the cluster and redirected its traffic to the degraded region. This created cross-region inconsistencies in job-assignment state, causing workflow runs to be delayed, exhaust retries, or fail outright. At peak, about 7% of runs were delayed by more than 5 minutes, and 25% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 09:13 UTC by returning traffic to its normal distribution. \u003cbr\u003e\u003cbr\u003eSecond period (12:08 – 12:48 UTC): As part of mitigating the first incident, traffic was returned to the regional instance that was still undergoing its capacity increase. Multiple Redis nodes in the scaling region experienced failures, increasing traffic to healthy nodes and causing connection limits to be reached on many nodes. At peak, 30% of runs were delayed by more than 5 minutes, and 60% of runs failed with an infrastructure error during the course of the incident. We mitigated the incident at 12:48 UTC by redirecting workflow traffic away from the scaling region. \u003cbr\u003e\u003cbr\u003eWe are adding stronger regional health and capacity checks before maintenance and requiring a stable observation period before restoring traffic. We are also improving automated connection resiliency, and partnering with our platform dependency to automatically detect and remediate unhealthy cluster members and shard imbalance. More generally, we already had work underway to improve the resiliency and scale of this piece of Actions infrastructure.",
"name": "Incident with Actions",
"timestamp": "Jul \u003cvar data-var='date'\u003e25\u003c/var\u003e, \u003cvar data-var='time'\u003e08:59\u003c/var\u003e - \u003cvar data-var='time'\u003e09:25\u003c/var\u003e UTC"
},
{
"code": "jxd617hfwfq8",
"impact": "critical",
"message": "Between July 24, 19:17 UTC and July 24, 20:02 UTC, users were unable to create pull requests due to a database schema change. In total, 113,930 pull request creation attempts were impacted across 50,904 users, with an average error rate of 1.75% and a maximum error rate of 2.25% for all requests to Pull Requests service. Existing pull requests and other GitHub functionality were not affected. The issue was resolved by reverting the change to the affected database, upon which pull request creation immediately resumed.\u003cbr\u003e\u003cbr\u003eThe root cause was related to a backfill workflow into the Vitess keyspace hosting Pull Request data. The backfill Vitess command encountered errors and increased VReplication lag, and the workflow was canceled at 19:17 UTC. The cancellation executed a misunderstood Vitess codepath that dropped the backing table to the target keyspace, leaving a non-existent reference that resulted in errors creating Pull Requests. The mitigation was executing a command to drop the vschema reference to the dropped table, allowing Pull Request creation to resume.\u003cbr\u003e\u003cbr\u003eWe are adding stronger pre-flight validation to our tooling to prevent similar issues and expanding lower-environment support to provide better test coverage end-to-end before promoting them to production. We're also fixing our backfill migration tooling to protect from this specific codepath.",
"name": "Incident with Pull Requests",
"timestamp": "Jul \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e19:37\u003c/var\u003e - \u003cvar data-var='time'\u003e20:23\u003c/var\u003e UTC"
},
{
"code": "yjysg0xrl67m",
"impact": "major",
"message": "On July 24th at 16:04 UTC, a loss of connectivity occurred in network paths in one of our three physical data center availability zones (AZs). This resulted in packet loss due to the remaining active paths becoming saturated. Our data centers use a leaf-spine switch fabric in each compute cage, and an aggregation layer interconnecting the spines from each cage within each AZ. The loss of connectivity affected links between one cage’s spine switches and the aggregation layer within that specific AZ. \u003cbr\u003e\u003cbr\u003eWorkloads depending on compute resources in this cage became degraded due to packet loss, and exhibited intermittent errors: \u003cbr\u003e\u003cbr\u003e- Actions saw 10% of jobs fail during the impact window, and 5% of jobs succeeded but with delayed starts. \u003cbr\u003e- 27% of GitHub issues interactions saw slow requests or timeouts. \u003cbr\u003e- 4% of GitHub Copilot requests experienced errors, though most automatically retry. \u003cbr\u003e- 4% of git push operations saw impacts during the affected window. \u003cbr\u003e- Authentication requests saw increased latency during the affected window, but error rates, while elevated, were \u0026lt; 1% in all cases. \u003cbr\u003e\u003cbr\u003eWe were able to mitigate the outage by re-routing affected connections to available fiber paths that were allocated for future capacity upgrades. Sufficient network capacity to eliminate packet loss was restored at 17:07, with most services showing full recovery by 17:16. All paths were restored and services healthy at 17:36. \u003cbr\u003e\u003cbr\u003eThis incident affected 25% of available network interconnect capacity. Older cages utilize a 100Gbps network interface standard. To remove risk of reoccurrence, a planned upgrade to 400Gbps interfaces is being accelerated as much as possible, ensuring increased bandwidth available at all layers of the switch fabric for resiliency to path or device loss.",
"name": "Disruption with some GitHub services",
"timestamp": "Jul \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e16:17\u003c/var\u003e - \u003cvar data-var='time'\u003e17:36\u003c/var\u003e UTC"
},
{
"code": "594m87r8sw13",
"impact": "none",
"message": "Between July 23, 2026 at 18:45 UTC and July 24, 2026 at 11:19 UTC, an abuse mitigation update caused some legitimate customers whose traffic was routed through our Central Europe and South America edge locations to be incorrectly blocked from GitHub.com. We estimate that approximately 0.25% of GitHub.com requests were affected during this period.\n\nThis was caused by an abuse mitigation configuration that incorrectly classified legitimate traffic. We mitigated the incident by reverting the update. We are adding validation and safeguards to prevent similar incorrect blocking in the future.",
"name": "Incident With Blocked GitHub.com Traffic",
"timestamp": "Jul \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e11:00\u003c/var\u003e - \u003cvar data-var='time'\u003e11:00\u003c/var\u003e UTC"
},
{
"code": "zq3c1jst2vkq",
"impact": "major",
"message": "On July 23, 2026, between 07:08 and 09:39 UTC, several services experienced delays: 8% of actions workflow runs experienced an average run start delay of 10 minutes, 5% of webhook deliveries exceeded SLO, and code scanning, repos, notifications, issues and pull requests experienced increased latency over the life of the incident. \u003cbr\u003e\u003cbr\u003eThe root cause of the incident was a node of our background job processing system which did not recover after entering scheduled host maintenance. The incident was mitigated by identifying the problematic shard and restoring its correct state, after which queue backlogs drained and services recovered. \u003cbr\u003e\u003cbr\u003eTo speed mitigation, we have added monitors for nodes in this unhealthy state after maintenance operations. To prevent future recurrence, we are adapting our lifecycle automation to verify host rejoin after a scheduled reboot.",
"name": "Latency issues across a number of services",
"timestamp": "Jul \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e07:53\u003c/var\u003e - \u003cvar data-var='time'\u003e09:39\u003c/var\u003e UTC"
},
{
"code": "20frdtvv3yg6",
"impact": "minor",
"message": "On July 22, 2026, between 19:36 UTC and 22:04 UTC, GitHub Actions experienced delayed and failed job starts on GitHub-hosted runners. The incident was caused by an unhealthy state in a backend data service responsible for provisioning hosted runners, preventing runner acquisition for a subset of workloads. During most of the incident, approximately 15% of workflow runs on hosted runners were delayed by more than 5 minutes, while roughly 1% failed to start.\u003cbr\u003e\u003cbr\u003eAt 21:49 UTC, we restored the health of the backend data replication system, allowing provisioning to recover and the accumulated workflow backlog to drain. Service performance then returned to expected levels. We are improving provisioning-service resiliency, workload distribution, and capacity balancing to reduce the likelihood and impact of similar incidents.",
"name": "Disruption with actions hosted runners",
"timestamp": "Jul \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e20:43\u003c/var\u003e - \u003cvar data-var='time'\u003e22:09\u003c/var\u003e UTC"
},
{
"code": "g40zcbvchny4",
"impact": "critical",
"message": "On July 21, 2026, between 07:41 UTC and 11:57 UTC, the SSH Authentication service was degraded and some SSH connections failed to authenticate. On average, 12.2% of SSH authentication requests failed, peaking at 15.7%. Both user RSA keys and deploy keys were impacted. This was due to a change in how our SSH service handled one public-key authentication method that caused the affected authentication attempts to be rejected as invalid. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by reverting the change, after which SSH authentication returned to normal. \u003cbr\u003e\u003cbr\u003eWe are working to expand our automated test coverage for our SSH public-key authentication flows to catch more edge cases and to improve observability and alerting on SSH authentication failures, to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Some SSH connections using deploy keys are failing",
"timestamp": "Jul \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e10:31\u003c/var\u003e - \u003cvar data-var='time'\u003e11:57\u003c/var\u003e UTC"
},
{
"code": "fd7j2mw8xw94",
"impact": "minor",
"message": "Between 06:39 and 18:11 UTC on July 20, 2026, the Copilot service experienced a degradation of the GPT 5.3 model due to an issue with our upstream provider. The upstream model provider returned intermittent errors for GPT 5.3 Codex requests, which caused some responses to fail. Auto mode requests that had selected GPT 5.3 Codex were also impacted. On average about 2% of GPT 5.3 Codex requests failed during this window. Copilot automatically routed eligible traffic away from the impacted provider to reduce customer impact. No other models were impacted.\u003cbr\u003e\u003cbr\u003eWe worked with the upstream provider throughout the incident and confirmed sustained recovery before resolving.",
"name": "Disruption with GPT 5.3 Codex",
"timestamp": "Jul \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e16:03\u003c/var\u003e - \u003cvar data-var='time'\u003e18:37\u003c/var\u003e UTC"
},
{
"code": "8vfyvq16hzh9",
"impact": "critical",
"message": "Between July 19, 2026, at 23:05 UTC and July 20, 2026, at 03:55 UTC, Actions self-hosted and larger runners were unable to connect to GitHub. During this period, Actions jobs were delayed or failed when trying to acquire a runner. Jobs using standard and Mac hosted runners were not affected. Reconnection traffic from affected runners also increased load on GitHub APIs, resulting in 3-4 seconds of additional average request latency and elevated 5xx error rates. \u003cbr\u003e\u003cbr\u003eThe incident was caused by a certificate lifecycle management failure in a subset of internal services, resulting in an SSL certificate expiration that disrupted runner connectivity. We restored service by rotating the affected certificate. Recovery began at 02:45 UTC. By 03:55 UTC, queued workflow backlog had been processed and workflow delay rates returned to normal.\u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are strengthening certificate renewal automation, adding fallback expiry monitoring and alerting, and improving circuit-breaker protections during runner API disruptions to reduce the risk of cascading impact to other APIs.",
"name": "Incident with GitHub Actions",
"timestamp": "Jul \u003cvar data-var='date'\u003e19\u003c/var\u003e, \u003cvar data-var='time'\u003e23:34\u003c/var\u003e - Jul \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e04:44\u003c/var\u003e UTC"
},
{
"code": "ph5nns5y4gxj",
"impact": "major",
"message": "Between July 19, 2026, at 23:05 UTC and July 20, 2026, at 03:55 UTC, Actions self-hosted and larger runners were unable to connect to GitHub. During this period, Actions jobs were delayed or failed when trying to acquire a runner. Jobs using standard and Mac hosted runners were not affected. Reconnection traffic from affected runners also increased load on GitHub APIs, resulting in 3-4 seconds of additional average request latency and elevated 5xx error rates.\u003cbr\u003e\u003cbr\u003eThe incident was caused by a certificate lifecycle management failure in a subset of internal services, resulting in an SSL certificate expiration that disrupted runner connectivity. We restored service by rotating the affected certificate. Recovery began at 02:45 UTC. By 03:55 UTC, queued workflow backlog had been processed and workflow delay rates returned to normal.\u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are strengthening certificate renewal automation, adding fallback expiry monitoring and alerting, and improving circuit-breaker protections during runner API disruptions to reduce the risk of cascading impact to other APIs.",
"name": "Disruption with some GitHub services",
"timestamp": "Jul \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e00:25\u003c/var\u003e - \u003cvar data-var='time'\u003e01:46\u003c/var\u003e UTC"
},
{
"code": "gxycch3076xk",
"impact": "major",
"message": "From 22:21 UTC - 23:50 UTC on July 16, 2026, the REST API experienced significant degradation. During this period, about 39% of REST API requests failed with HTTP 500 level responses, with the errors peaking at 44.3%. \u003cbr\u003e\u003cbr\u003eWe identified the issue as an infrastructure change that wrongly marked the majority of API backends in a single region as unhealthy. As a result, requests routed to those backends failed before reaching the application layer. \u003cbr\u003e\u003cbr\u003eTo prevent this from happening again, we're improving our systems to catch this kind of invalid configuration before it reaches production. We'll also audit the related systems to make them more resilient to future changes, and we're increasing our monitoring sensitivity so we're alerted to problems like this sooner.",
"name": "Degraded REST API Availability",
"timestamp": "Jul \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e22:51\u003c/var\u003e - Jul \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e00:14\u003c/var\u003e UTC"
},
{
"code": "0f6hndclsk21",
"impact": "minor",
"message": "On July 16, 2026, GitHub Copilot users experienced elevated errors when using Claude Fable 5 from 17:33 UTC until mitigation at 22:04 UTC. The average error rate was 1.4%, with a maximum error rate of 30.85%. The issue was caused by degradation at an upstream model provider; other Copilot models were not significantly affected, and users could avoid the impact by selecting another model or Auto. Service recovered after the provider mitigated the degradation.",
"name": "Claude Fable 5 experiencing degraded performance",
"timestamp": "Jul \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e21:05\u003c/var\u003e - \u003cvar data-var='time'\u003e22:04\u003c/var\u003e UTC"
},
{
"code": "d35pngcy6sb9",
"impact": "critical",
"message": "On July 16, 2026, between 08:50 UTC and 09:50 UTC, the GitHub MCP Server’s web_search tool experienced elevated failures. The average error rate was 42% and peaked at 82% of requests to the tool. Other GitHub MCP Server tools were unaffected. This was caused by degradation at a downstream web search provider.\u003cbr\u003e\u003cbr\u003eThe incident was mitigated when the downstream provider recovered, after which we confirmed that the tool’s success rate had returned to normal.\u003cbr\u003e\u003cbr\u003eWe are improving the tool’s resilience and failure handling to reduce the customer impact and duration of similar incidents.",
"name": "Disruption with some GitHub services",
"timestamp": "Jul \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e09:13\u003c/var\u003e - \u003cvar data-var='time'\u003e12:20\u003c/var\u003e UTC"
},
{
"code": "ydpk76bj34z8",
"impact": "minor",
"message": "On July 14, 2026, between 15:17 and 15:37 UTC, a rollout to GitHub's internal webhook delivery pipeline caused a subset of webhook delivery records to not be written to our webhook deliveries store after being processed and delivered successfully. Affected deliveries would be missing from the webhook delivery UI and API and won’t be available for redelivery.\n\nThe root cause was an uncoordinated rollout: a change to how delivery records are handed off between pipeline components was deployed before the upstream components producing those records were updated to match. While the rollout was in progress, affected records were silently skipped rather than persisted, with no automatic retry. The impact ended as soon as the rollout was completed.\n\nAbout 2.4M delivery records were skipped (approximately 4% of the 20-minute impact window, 0.04% of a typical 24-hour period). Importantly, 95% of these skipped deliveries reached customer endpoints successfully, only the record of the delivery is missing. Of the ~5% that failed to reach customer endpoints, only ~1.4% (5,463) map to webhooks that retried their deliveries in the past 28 days.\n\nTo prevent recurrence, we are improving our automated detection of unsafe schema changes and tightening rollout coordination for changes that span multiple components in the pipeline.",
"name": "Incident with Webhooks",
"timestamp": "Jul \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e17:38\u003c/var\u003e - \u003cvar data-var='time'\u003e18:01\u003c/var\u003e UTC"
},
{
"code": "dfpfsngcwywf",
"impact": "minor",
"message": "On July 14, 2026, the GitHub Codespaces service was degraded during two periods — between 06:00 UTC and 09:56 UTC, and again between 10:54 UTC and 12:53 UTC — and some users experienced intermittent failures or delays when creating new codespaces. Impact was concentrated in a subset of geographic regions. During the first period, the error rate averaged 0.5% and peaked at 4.6% of codespace creation requests. The second period was more pronounced, peaking at approximately 30% of codespace creation requests in the most-affected region before recovery. Both periods were caused by an unexpected surge in codespace creation from an abusive actor that drained the available compute capacity in the affected regions faster than it could be replenished. \u003cbr\u003e\u003cbr\u003eWe mitigated the impact by identifying and stopping the sources of the excess creation volume, reducing the resources that could be consumed in the affected regions, and rebalancing traffic across regions to restore capacity. Codespace creation success rates returned to normal after each period. \u003cbr\u003e\u003cbr\u003eWe are working to add automated, low-latency controls to throttle abnormal codespace creation and to strengthen our detection and safeguards, so we can reduce our time to detection and mitigation of issues like this in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Jul \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e08:21\u003c/var\u003e - \u003cvar data-var='time'\u003e09:56\u003c/var\u003e UTC"
},
{
"code": "q27ttsnp0x4g",
"impact": "major",
"message": "On July 13, 2026, between 13:11 and 13:53 UTC, some customers experienced failures starting and running GitHub Actions workflows, which also affected Copilot cloud agent sessions and GitHub Pages builds since they depend on Actions. During the peak of the incident, 30% of Actions jobs failed to start and 2% were delayed more than 5 minutes. \u003cbr\u003e\u003cbr\u003eThe incident was triggered by a configuration change in an internal autoscaling component that contained outdated capacity threshold values. This caused a critical Actions service to scale below its required baseline, reducing capacity for workflow processing. We identified the regression, rolled back the change, and restored service capacity. New workflow executions recovered by 13:39 UTC. Full recovery was reached by 13:53 UTC after the queued backlog was drained. \u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we have added deployment guardrails to validate that autoscaling inputs are current and to detect drift between planned and live scaling state before autoscaling changes are applied.",
"name": "Actions runs are experiencing failures to start",
"timestamp": "Jul \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e13:32\u003c/var\u003e - \u003cvar data-var='time'\u003e13:53\u003c/var\u003e UTC"
},
{
"code": "cstx3v63mklm",
"impact": "critical",
"message": "On July 9, 2026, between 03:29 UTC and 13:39 UTC, GitHub Actions experienced delayed and failed job starts on GitHub-hosted runners. The incident was caused by an unhealthy state in a backend data service responsible for provisioning hosted runners, preventing runner acquisition for a subset of workloads. During most of the incident, approximately 8% of workflow runs on hosted runners were delayed by more than 5 minutes, while roughly 2% failed to start.\u003cbr\u003e\u003cbr\u003eAt 13:39 UTC, we restored the health of the backend data replication system, allowing provisioning to recover and the accumulated workflow backlog to drain. Service performance then returned to expected levels. We are improving provisioning-service resiliency, workload distribution, and capacity balancing to reduce the likelihood and impact of similar incidents.",
"name": "Delays starting Actions runs",
"timestamp": "Jul \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e04:34\u003c/var\u003e - \u003cvar data-var='time'\u003e13:52\u003c/var\u003e UTC"
},
{
"code": "fz9sdc2q008p",
"impact": "major",
"message": "On July 7, 2026, between 14:01 UTC and 16:17 UTC the Actions and Codespaces REST APIs were degraded and returned intermittent 500-class errors for a percentage of requests. Error rates peaked at approximately 8% of Actions runner API requests and 13% of Codespaces API requests, though retries were frequently successful. In-progress Actions runs and Codespaces were not impacted and continued successfully. This was due to a recent change that did not deliver the expected performance and, under certain conditions, caused downstream errors.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by rolling back the change, after which the affected services recovered.\u003cbr\u003e\u003cbr\u003eWe are working to improve the resilience of our services to these conditions and to strengthen our monitoring to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Actions and Codespaces APIs experiencing partial failures",
"timestamp": "Jul \u003cvar data-var='date'\u003e7\u003c/var\u003e, \u003cvar data-var='time'\u003e14:14\u003c/var\u003e - \u003cvar data-var='time'\u003e16:17\u003c/var\u003e UTC"
},
{
"code": "5bnwwg9tzd4q",
"impact": "minor",
"message": "On July 2nd, 2026, between approximately 15:00 and 18:30 UTC, the GitHub Pages service experienced degraded deployment performance due to a surge in demand that exceeded available processing capacity. During this period, users publishing to GitHub Pages may have seen their deployments queued or taking substantially longer than usual to go live. No other GitHub services were impacted.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by scaling up Pages deployment workers and provisioning additional storage capacity to clear the backlog.\u003cbr\u003e\u003cbr\u003eGitHub is reviewing capacity planning and autoscaling measures to reduce the likelihood of similar delays in the future.",
"name": "Incident with Pages",
"timestamp": "Jul \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e16:54\u003c/var\u003e - \u003cvar data-var='time'\u003e18:25\u003c/var\u003e UTC"
},
{
"code": "rl7f90w0n0gq",
"impact": "minor",
"message": "On July 1, 2026, between approximately 00:00 UTC and 13:04 UTC, some GitHub Copilot customers whose budget was exhausted before the monthly reset remained incorrectly blocked from paid Copilot usage after the new billing month began, even though their budgets had reset. Some budget changes also took longer than usual to apply. Only customers with an exhausted budget were affected, which limited the impact.\u003cbr\u003e\u003cbr\u003eThis was caused by a caching issue at the monthly reset: for some users, a pre-reset \"budget exhausted\" status was re-saved and served even though their budget had reset, so they stayed blocked. We had built a safeguard ahead of the reset to prevent this, but it did not take effect because an internal configuration service did not load its settings correctly. We resolved the incident by deploying a change that discards the outdated status and recomputes access from current budget data independently of that configuration, and by working through the backlog of budget updates.\u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are ensuring pre-reset status cannot survive the monthly budget reset, adding alerting for this failure mode, and increasing capacity to absorb the monthly surge of budget updates.",
"name": "Delays in copilot budget limits resets for some users",
"timestamp": "Jul \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e10:51\u003c/var\u003e - \u003cvar data-var='time'\u003e13:26\u003c/var\u003e UTC"
}
],
"name": "July",
"year": 2026
}
],
"start_time": "2026-07-01T00:00:00Z",
"time_zone": "UTC"
},
{
"end_time": "2026-06-30T23:59:59Z",
"months": [
{
"incidents": [
{
"code": "mc4yx9f8x1hk",
"impact": "minor",
"message": "Between 15:19 UTC and 15:49 UTC on June 30, 2026, users were unable to complete the signup flow for GitHub.com/signup. Approximately 62% of new user signups failed for about 30 minutes during this window.\u003cbr\u003e\u003cbr\u003eThis was caused by a configuration change to the signup flow that unintentionally blocked users from completing signup.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by reverting the change, which restored successful signups. To reduce the likelihood and impact of similar issues, we are adopting staged, incremental rollouts for changes on the signup path, improving our ability to test these changes before they reach production, and adding checks to verify signup health before and during any change that affects this flow.",
"name": "Disruption with some GitHub services - Signup Flow",
"timestamp": "Jun \u003cvar data-var='date'\u003e30\u003c/var\u003e, \u003cvar data-var='time'\u003e15:38\u003c/var\u003e - \u003cvar data-var='time'\u003e15:49\u003c/var\u003e UTC"
},
{
"code": "0rkjjs2ssp7z",
"impact": "minor",
"message": "From June 26, 2026 at 23:40 UTC through June 28, 2026 at 20:55 UTC, Copilot Cloud Agent was degraded. The agent could fail when reporting progress, replying to pull request comments, or opening pull requests. For affected built-in tool calls, the average error rate was approximately 8%, with hourly error rates peaking around 26%.\u003cbr\u003e\u003cbr\u003eThis was due to a regression introduced during a Copilot Cloud Agent runtime deployment that caused several built-in agent tools to become unavailable. In many cases, the affected tool calls failed silently so agent jobs appeared to succeed. This monitoring gap meant it took longer than expected to identify the failure. We mitigated the incident by reverting the runtime deployment to the previously stable version.\u003cbr\u003e\u003cbr\u003eWe've added monitoring and alerting for this class of tool-availability error to reduce time-to-detection. We're also adding regression tests for these built-in agent tools, and improving the shipping safety for future runtime rollouts to avoid similar issues.",
"name": "Disruption with some GitHub services",
"timestamp": "Jun \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e17:50\u003c/var\u003e - \u003cvar data-var='time'\u003e20:55\u003c/var\u003e UTC"
},
{
"code": "v0b3bpsyvqtk",
"impact": "minor",
"message": "This incident was used to notify for a maintenance event. There is no specific root cause analysis. Work progressed as planned without any issues to report.",
"name": "Disruption with some GitHub services",
"timestamp": "Jun \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e14:02\u003c/var\u003e - \u003cvar data-var='time'\u003e20:33\u003c/var\u003e UTC"
},
{
"code": "9ndxtnrwjf37",
"impact": "major",
"message": "On June 25, 2026, between 17:33 UTC and 17:55 UTC, our background job service experienced degradation which increased delays to pull requests, repository pushes, Actions workflows, and Webhooks, with delays peaking at 7m. The issue was caused by underlying hypervisor issues and an incoming traffic spike, causing service timeouts which led to a connection storm and continual rebalances. \u003cbr\u003e\u003cbr\u003eThe issue was mitigated by replacing the problem node at 17:49, after which all services saw recovery by 18:07.",
"name": "Degradation with Webhooks, Pull Requests and Actions",
"timestamp": "Jun \u003cvar data-var='date'\u003e25\u003c/var\u003e, \u003cvar data-var='time'\u003e17:50\u003c/var\u003e - \u003cvar data-var='time'\u003e18:27\u003c/var\u003e UTC"
},
{
"code": "9n670kvk0vw9",
"impact": "minor",
"message": "On June 23, 2026, between 22:45 and 23:29 UTC, GitHub Copilot Completions and Next Edit Suggestions were degraded for users in all regions. During this window, affected users may have seen failed or missing code completions and Next Edit Suggestions. On average about 25% of Completions and Next Edit Suggestions requests failed during the impact window, peaking at roughly 27%. The cause was a configuration change that prevented the Copilot service from obtaining the authentication tokens it needs to reach its model backends; this both failed requests directly and caused the service to temporarily remove backends from rotation. GitHub engineers detected the elevated error rate within minutes, declared an incident, and mitigated the issue at 23:22 UTC by redeploying the service with a known-good configuration, which restored normal operation. As a follow-up, the team disabled the affected authentication path to prevent a future deployment from re-introducing the problem, and is making the change rollout safer. We apologize for the disruption and are taking steps to reduce the likelihood of similar incidents.",
"name": "We are seeing elevated errors with Next Edit Suggestions and Completions",
"timestamp": "Jun \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e23:04\u003c/var\u003e - \u003cvar data-var='time'\u003e23:29\u003c/var\u003e UTC"
},
{
"code": "5t81zk0vrk3z",
"impact": "minor",
"message": "On June 17, 2026, between 16:57 UTC and 19:14 UTC, Copilot code completions were degraded and users were unable to receive Next Edit Suggestions. Standard ghost text code completions were not affected. This was due to a configuration change that caused the service's routing layer to incorrectly discard all Next Edit Suggestion model endpoints as invalid.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by deploying a corrected configuration change at 18:55 UTC, with full recovery observed at 19:14 UTC.\u003cbr\u003e\u003cbr\u003eWe are working to improve the resilience of our routing layer to limit impact due to a subset of invalid configurations, and to improve our alerting to detect sudden traffic changes that are not captured by standard error rate monitors.",
"name": "Disruption with Copilot next edit suggestions",
"timestamp": "Jun \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e17:57\u003c/var\u003e - \u003cvar data-var='time'\u003e19:28\u003c/var\u003e UTC"
},
{
"code": "kn7gv3tlfc54",
"impact": "none",
"message": "On June 17, 2026, between 11:35 UTC and 19:20 UTC, the Webhooks service was degraded and delivered webhook payloads with missing installation information. On average, 11.3% of webhook deliveries were impacted. Customers relying on the installation field for authentication or routing were unable to process affected webhooks. A smaller subset of deliveries for the security_advisory event (0.04%) were delivered successfully but were not recorded for redelivery. This was due to a defect in a new delivery code path that failed to include installation data in webhook payloads.\n\nWe mitigated the incident by disabling the feature flag controlling the new code path.\n\nWe are working to improve our automated validation of webhook payloads, and introduce automated alerting for webhook payload regressions to reduce our time to detection and mitigation of issues like this one in the future.\n\nThe following events were affected: branch_protection_configuration, code_scanning_alert, commit_comment, custom_property, custom_property_values, dependabot_alert, deploy_key, deployment_protection_rule, deployment_review, dismissal_request_code_scanning, dismissal_request_secret_scanning, installation_target, member, membership, merge_queue_entry, org_block, organization, projects_v2, projects_v2_item, pull_request_review_thread, repository_ruleset, secret_scanning_alert, secret_scanning_alert_location, secret_scanning_scan, security_and_analysis, star, sub_issues, team, team_add, workflow_job.",
"name": "Incident With Webhooks",
"timestamp": "Jun \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e19:00\u003c/var\u003e - \u003cvar data-var='time'\u003e19:00\u003c/var\u003e UTC"
},
{
"code": "rfmjwng33vjf",
"impact": "critical",
"message": "On June 17, 2026, between approximately 03:35 UTC and 04:44 UTC, GitHub Copilot was degraded and most of its frontier chat models were temporarily unavailable across all regions. During this window, affected models either disappeared from the model picker in the web, editor, and CLI experiences, or returned a \"model not available\" error when selected. Customers could continue using GitHub Copilot by selecting one of the models that remained available. The incident occurred during off-peak hours, which limited the number of customers affected.\u003cbr\u003e\u003cbr\u003eThis was due to a configuration change that our production system deemed invalid. We mitigated the incident by reverting the configuration change, after which the affected models returned automatically as the service reloaded the previous configuration.\u003cbr\u003e\u003cbr\u003eWe are working to roll out configuration changes gradually with stronger validations, alerts on sudden drops in the number of available models, and automatically rolls back configuration changes that produce these alerts.",
"name": "Incident with Copilot Availability",
"timestamp": "Jun \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e03:50\u003c/var\u003e - \u003cvar data-var='time'\u003e04:44\u003c/var\u003e UTC"
},
{
"code": "1s444p9sf9wg",
"impact": "major",
"message": "On June 16, 2026, between 17:20 UTC and 18:15 UTC, the Opus 4.8 model experienced degraded availability in GitHub Copilot. During this window, some requests to Opus 4.8 failed or errored. Other Copilot models were not affected and remained available as alternatives. This was caused by an issue with an upstream model provider. The upstream provider resolved the issue, and we monitored Opus 4.8 until success rates returned to normal. The incident is fully resolved.",
"name": "Disruption with some GitHub services",
"timestamp": "Jun \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e17:45\u003c/var\u003e - \u003cvar data-var='time'\u003e18:15\u003c/var\u003e UTC"
},
{
"code": "d9b4dsg8d0r6",
"impact": "minor",
"message": "Between 17:38 UTC and 18:22 UTC on June 15, 2026, approximately 83% of requests to the analytics endpoint serving the /chronicle feature failed. The cause was an internal feature-flag service that encountered a transient error and failed to recover, causing feature flag checks to fail. The analytics endpoint was gated behind one of these flags, resulting in requests being rejected. We restored service health by removing the feature flag gating the analytics endpoint and deploying that change. To avoid recurrence of similar incidents, we have changed the feature-flag client so that errors that are not known to be permanent are retried, and we are improving alerting and startup behavior so this class of failure is detected and recovered from faster.",
"name": "Multiple services have elevated errors and endpoint failures when checking feature flags",
"timestamp": "Jun \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e18:32\u003c/var\u003e - \u003cvar data-var='time'\u003e19:10\u003c/var\u003e UTC"
},
{
"code": "67w72l5v25x3",
"impact": "minor",
"message": "On June 15, 2026, between 15:27 UTC and 16:23 UTC, GitHub webhook deliveries were delayed. During this window, webhook events were delivered later than normal, with average end-to-end delivery latency peaking at approximately 8.8 minutes. No webhook deliveries were lost — delayed events were queued and delivered once processing recovered.\u003cbr\u003e\u003cbr\u003eThis was caused by a temporary throughput constraint in an internal event-processing system that moves webhook events through GitHub's delivery pipeline. The rate at which events were processed for delivery dropped below the incoming volume, creating a backlog. We restarted the affected pipeline service, after which throughput recovered and the backlog fully drained by approximately 16:29 UTC. Webhook delivery latency returned to normal, the incident was mitigated at 16:39 UTC, and fully resolved at 17:37 UTC.\u003cbr\u003e\u003cbr\u003eTo reduce the likelihood and impact of similar incidents, we are working on improving the accuracy of the utilization metrics used to scale our delivery worker pools, reviewing connection and capacity headroom in the delivery pipeline.",
"name": "Increased latency with webhooks",
"timestamp": "Jun \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e15:37\u003c/var\u003e - \u003cvar data-var='time'\u003e17:37\u003c/var\u003e UTC"
},
{
"code": "4yz3c18qmdxn",
"impact": "minor",
"message": "On June 11, 2026, between 19:28 UTC and 21:06 UTC, GitHub webhook deliveries were delayed. Average delivery latency peaked at approximately 3.4 minutes, with some deliveries delayed by as much as 62 minutes at the 99th percentile. No events were lost — delayed events were queued and delivered once processing caught up.\u003cbr\u003e\u003cbr\u003eThis was due to a change in how webhook traffic was distributed across regions: to relieve load on one region, a portion of processing was shifted to another, where higher latency prevented our delivery workers from keeping pace with incoming volume, creating a backlog. We mitigated the incident by rebalancing webhook traffic distribution; as load returned to normal levels, processing caught up and the delivery backlog fully drained.\u003cbr\u003e\u003cbr\u003eWe are working on improving the accuracy of the utilization metrics used to scale our delivery worker pools, and reassess how we distribute webhook traffic across regions, to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Incident with Webhooks",
"timestamp": "Jun \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e19:42\u003c/var\u003e - \u003cvar data-var='time'\u003e22:19\u003c/var\u003e UTC"
},
{
"code": "fcj3088jg1wx",
"impact": "critical",
"message": "Between 15:05 UTC and 16:25 UTC, GitHub API services experienced degraded availability due to sporadic authentication failures affecting approximately 9% of requests. Customers experienced intermittent \"logged out\" behavior as erroneous 401 responses triggered repeated authentication flows in app integrations. Affected requests also experienced approximately 800ms of additional latency. \u003cbr\u003e\u003cbr\u003eA memcached proxy service rollout to our internal API infrastructure caused our authentication service to pick up an incorrect memcached host configuration, leading to intermittent authentication lookup failures. We mitigated the incident by deploying a configuration change to memcached to use the correct host. \u003cbr\u003e\u003cbr\u003eTo prevent similar issues in the future, we plan to migrate our authentication system to the new memcached infrastructure to improve resilience and strengthen overall reliability posture.",
"name": "Authentication issues related to API requests",
"timestamp": "Jun \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e15:20\u003c/var\u003e - \u003cvar data-var='time'\u003e16:39\u003c/var\u003e UTC"
},
{
"code": "jpd6l1jq0r54",
"impact": "none",
"message": "On June 8, 2026, between 14:49 and 14:54 UTC, a subset of requests to GitHub.com, the REST API, GraphQL API, and Webhooks UI/API experienced elevated error rates due to a transient infrastructure capacity issue that self-resolved within approximately 5 minutes.\n\nUsers experienced HTTP 500 errors and timeouts when accessing GitHub.com, the REST API, GraphQL API, and Webhooks UI/API for approximately 5 minutes, with the REST API taking up to 12 minutes to fully recover.",
"name": "Degraded availability for GitHub.com, GraphQL API, and Webhooks UI/API",
"timestamp": "Jun \u003cvar data-var='date'\u003e8\u003c/var\u003e, \u003cvar data-var='time'\u003e15:00\u003c/var\u003e - \u003cvar data-var='time'\u003e15:00\u003c/var\u003e UTC"
},
{
"code": "71hv2q6tk693",
"impact": "minor",
"message": "On June 8, 2026, between 08:40 UTC and 09:30 UTC, the Claude Opus 4.7 model experienced degraded availability with error rates peaking at 8.4% and averaging 1.9%. This was due to an upstream provider issue that caused temporary unavailability and rate limiting on secondary failover systems. Users selecting Auto or alternative models were unaffected. We are improving provider failover mechanisms and monitoring to prevent similar issues.",
"name": "Disruption with Claude Opus 4.7",
"timestamp": "Jun \u003cvar data-var='date'\u003e8\u003c/var\u003e, \u003cvar data-var='time'\u003e09:05\u003c/var\u003e - \u003cvar data-var='time'\u003e10:03\u003c/var\u003e UTC"
},
{
"code": "m7n7sm0sr1pz",
"impact": "critical",
"message": "On June 8, 2026, between approximately 06:30 UTC and 08:36 UTC, signed-out users experienced sustained elevated HTTP 504 errors when accessing Pull Requests, Issues, releases, patch diffs, and other related GitHub.com pages. During the incident, approximately 17% of unauthenticated requests to the affected GitHub.com endpoints returned gateway timeout errors, peaking at roughly 34% of requests at around 06:50 UTC. Some GitHub Actions workflows were also affected when they depended on release downloads or related GitHub.com endpoints. The impact lasted approximately two hours and was isolated to unauthenticated traffic; signed-in users were not affected. \u003cbr\u003e\u003cbr\u003eThe issue was caused by a significant increase in abusive traffic to specific GitHub.com endpoints. This degraded our ability to respond to unauthenticated requests, causing requests to queue beyond timeout thresholds and return gateway timeout errors. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by identifying the anomalous traffic pattern and applying targeted blocks at the load balancer and application layers. Error rates returned to normal and affected services were fully restored by 08:36 UTC. \u003cbr\u003e\u003cbr\u003eTo reduce the likelihood and impact of similar incidents in the future, we are improving automated detection and blocking for these traffic patterns, improving our emergency traffic-blocking deployment path, and evaluating routing changes for endpoints used by both signed-out users and automated workflows.",
"name": "Pull Requests and Issues unavailable for signed-out users",
"timestamp": "Jun \u003cvar data-var='date'\u003e8\u003c/var\u003e, \u003cvar data-var='time'\u003e07:11\u003c/var\u003e - \u003cvar data-var='time'\u003e08:36\u003c/var\u003e UTC"
},
{
"code": "4843jm0lsls6",
"impact": "minor",
"message": "This incident was used to notify for a maintenance event. There is no specific root cause analysis. Maintenance did run longer than expected (we were complete at 18:48 UTC) but the work proceeded as planned.",
"name": "EU Network Maintenance",
"timestamp": "Jun \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e15:31\u003c/var\u003e - \u003cvar data-var='time'\u003e18:49\u003c/var\u003e UTC"
},
{
"code": "b0plzff6yl6f",
"impact": "minor",
"message": "On June 6, 2026 between 16:18 UTC and 17:01 UTC, users experienced elevated error rates when performing Git operations (cloning, fetching, downloading archives) and accessing package registries. The issue affected users whose traffic was routed through our European infrastructure.\u003cbr\u003e\u003cbr\u003eDuring this time, on average 0.95% of Codeload requests and 9.2% of Package Registry requests failed with server errors. At peak, the Codeload error rate reached 1.76% and Package Registry errors reached 27%.\u003cbr\u003e\u003cbr\u003eThe root cause was a planned network circuit migration that disrupted connectivity at one of our points of presence. Our process for shifting traffic away from the site did not operate as expected, resulting in a small amount of production traffic to continue being serviced at the effected site during the maintenance window. The issue was mitigated by rolling back the network change, restoring normal connectivity. Services fully recovered by 17:01 UTC.\u003cbr\u003e\u003cbr\u003eTo reduce the likelihood of similar incidents in the future, we are reviewing our site drain process to make it more verbose and add visibility so any unexpected behavior is caught earlier.",
"name": "Disruption with some GitHub services in the EU region",
"timestamp": "Jun \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e16:53\u003c/var\u003e - \u003cvar data-var='time'\u003e17:07\u003c/var\u003e UTC"
},
{
"code": "2nmfnbknhlnv",
"impact": "minor",
"message": "On June 5, 2026, between 15:35 UTC and 16:45 UTC, 0.11% of authenticated REST API requests incorrectly returned “not found” responses. Impact was concentrated among - and significantly higher for - users authenticating with user-to-server tokens to access organization-owned repositories.\u003cbr\u003e\u003cbr\u003eSome users of our GitHub for Slack and GitHub for Microsoft Teams integrations saw their channel subscriptions removed as those systems interpreted the transient \"not found\" response as durable loss of access. Roughly 12% of organizations with active channel subscriptions were impacted, with ~2% of all channel subscriptions being removed.\u003cbr\u003e\u003cbr\u003eThese issues were triggered by a change to an internal authorization component that did not correctly resolve access for user-to-server tokens against organization-owned repositories. We mitigated the incident by disabling the accompanying feature flag at 16:45 UTC, after which API responses returned to normal. We then restored all impacted Slack and Microsoft Teams channel subscriptions, with restoration completed at 22:21 UTC.\u003cbr\u003e\u003cbr\u003eWe are working to add retry and grace-period logic in the chat integrations so transient errors no longer trigger subscription deletions. In parallel, we are improving observability and gating of authorization changes so downstream impact is detected during scoped, gradual rollouts.",
"name": "Auth issue resulting in API impacts, including some Slack and Teams channel subscriptions",
"timestamp": "Jun \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e17:20\u003c/var\u003e - \u003cvar data-var='time'\u003e22:21\u003c/var\u003e UTC"
},
{
"code": "5h5lmbffp07c",
"impact": "critical",
"message": "On June 4, 2026, from 17:30 UTC to 18:55 UTC, Copilot Code Review experienced elevated failures for review requests on GitHub.com. Affected users saw “Copilot ran into an error” on pull requests when requesting a code review.\u003cbr\u003e\u003cbr\u003eDuring the incident window, an average of 81.6% of Copilot Code Review requests failed, with a peak failure rate of 93.9%. Approximately 36,800 code review requests failed. GitHub Enterprise Cloud with data residency was not impacted.\u003cbr\u003e\u003cbr\u003eThe issue was caused by a newly released dependency used by the Copilot Code Review processing workflow. The release introduced an incompatibility with the runtime environment. Because the workflow automatically consumed the latest release, the incompatible version was picked up without sufficient compatibility validation and caused review processing to fail.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by removing the problematic dependency version and redeploying the affected processing service. New code reviews began recovering at 18:44 UTC, and the failure rate returned to baseline by 18:55 UTC. Remaining timed-out work drained by 19:59 UTC.\u003cbr\u003e\u003cbr\u003eTo reduce the risk of recurrence, we are pinning the dependency version instead of automatically consuming the latest release, adding compatibility checks for future releases, improving fast-failure behavior when the review processor cannot start, adding shorter timeout controls for review workflows, and improving monitoring for review completion failures.",
"name": "Copilot Code Review Failing",
"timestamp": "Jun \u003cvar data-var='date'\u003e4\u003c/var\u003e, \u003cvar data-var='time'\u003e18:02\u003c/var\u003e - \u003cvar data-var='time'\u003e19:59\u003c/var\u003e UTC"
},
{
"code": "tf25qs8wdw5h",
"impact": "minor",
"message": "Between June 1, 2026, 23:00 UTC and June 4, 2026 04:11 UTC, customers experienced delays in Dependabot scheduled version updates. \u003cbr\u003e\u003cbr\u003ePull request creation for version updates was delayed, with delays increasing over time and reaching up to two days. Approximately 1.5 million repositories with active Dependabot version update configurations were affected. Dependabot security updates were not affected. The primary cause was changes to an internal platform service that routes requests for Dependabot and other services. \u003cbr\u003e \u003cbr\u003eWe mitigated the incident by deploying a fix that enables batch enqueuing of update jobs, which significantly increased processing throughput. Once the backlog was drained, Dependabot returned to normal processing times. \u003cbr\u003e \u003cbr\u003eTo reduce the risk of recurrence, we are working on tuning batch size and concurrency limits for Dependabot update job processing. We are also adding monitoring for job processing lag to enable earlier detection and faster mitigation of similar issues.",
"name": "Disruption with some GitHub services",
"timestamp": "Jun \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e19:42\u003c/var\u003e - Jun \u003cvar data-var='date'\u003e4\u003c/var\u003e, \u003cvar data-var='time'\u003e04:11\u003c/var\u003e UTC"
},
{
"code": "j240y90h4g0r",
"impact": "minor",
"message": "On June 2, 2026, between 21:54 UTC and June 3, 2026 06:45 UTC, the Spark service was degraded and users were unable to store or retrieve data for their Spark apps in one of our hosting regions. Users could still make changes to their app configuration during this time. The error rate peaked at 25% of affected requests to the service. Impact was limited to users whose requests were served through a single affected region; 43 users experienced errors during this window.\u003cbr\u003e\u003cbr\u003eThe root cause was a configuration that referenced a service component by a fixed address rather than a dynamic service endpoint. When the component was replaced, requests could no longer reach the fixed address and began to fail. We resolved the incident by updating the configuration to use a our standard service endpoints that are resilient to component replacement. Recovery time was extended because replacing the component required overrides to a temporary deployment safeguard.\u003cbr\u003e\u003cbr\u003eWe are working to add validation that prevents fixed infrastructure addresses from being used in application configuration outside of test environments and to improve our monitoring to reduce our time to detect.",
"name": "Disruption with some GitHub services",
"timestamp": "Jun \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e03:13\u003c/var\u003e - \u003cvar data-var='time'\u003e06:46\u003c/var\u003e UTC"
},
{
"code": "5wdgxw2rvbt3",
"impact": "minor",
"message": "Starting from 13:00 UTC June 1, 2026, to 00:17 UTC June 2, 2026, multiple services experienced delayed job processing due to increased latency in our background job queue service. The root cause was insufficient queue processing capacity to handle a large week-over-week increase in total job traffic.\u003cbr\u003e\u003cbr\u003eUsers saw up to 90 minutes of delay in billing usage updates, 30 minutes of delay for webhook notifications to show, and 15 minutes of delay to see email notifications. Mitigation involved scaling up our background job service capacity to handle the spike in job traffic.\u003cbr\u003e\u003cbr\u003eWe have added queue capacity monitoring to our background job queue service to stay ahead of weekly growth patterns and to reduce time to detect in the future.",
"name": "Delays with Code Scanning and Billing",
"timestamp": "Jun \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e15:17\u003c/var\u003e - Jun \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e00:17\u003c/var\u003e UTC"
}
],
"name": "June",
"year": 2026
},
{
"incidents": [
{
"code": "rhqcgg8lg6mm",
"impact": "critical",
"message": "On May 28th, 2026, between approximately 18:27 and 20:41 UTC, the GitHub Copilot service was degraded due to an issue with the Responses API of an upstream provider affecting the GPT-5.2, GPT-5.3-Codex, GPT-5.4, and GPT-5.5 models. Requests routed to these models via the Responses API returned elevated error rates, which also affected Copilot coding agent and Copilot code review. No other models were impacted. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by shifting traffic away from the affected models while the upstream provider deployed a fix. \u003cbr\u003e\u003cbr\u003eGitHub is working to improve automated failover for the affected models and strengthen monitoring to prevent similar incidents in the future.",
"name": "Disruption with OpenAI Models",
"timestamp": "May \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e19:01\u003c/var\u003e - \u003cvar data-var='time'\u003e20:41\u003c/var\u003e UTC"
},
{
"code": "11n51rf8tf27",
"impact": "none",
"message": "On May 28, 2026, between 19:07 UTC and 19:16 UTC, multiple GitHub services experienced elevated error rates. This was due to a change that was partially deployed to an authentication service, causing errors for dependent services including the web experience, REST API, Git operations, and GitHub Actions. At peak impact, 10% of GitHub Actions runs failed to queue or encountered errors while downloading actions. We mitigated the incident by rolling back the change. \n\nWe are expanding test coverage and improving our deployment validation process to prevent recurrence of this issue in the future.",
"name": "Elevated error rates across multiple services",
"timestamp": "May \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e19:07\u003c/var\u003e - \u003cvar data-var='time'\u003e19:07\u003c/var\u003e UTC"
},
{
"code": "f3kmjlk25vhd",
"impact": "minor",
"message": "On May 28, 2026, between 00:54 UTC and 01:19 UTC, some users experienced errors when interacting with the Webhooks API, including webhook delivery history and configuration endpoints. On average, the error rate was 0.28% and peaked at 0.45%. This was due to a bug that caused a single Kubernetes pod to enter a CrashLoopBackOff after receiving a 500 with an empty response body from Cosmos DB.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by restarting the service. To prevent future incidents, we are pushing a change to handle this response scenario from Cosmos DB appropriately.",
"name": "Webhook APIs and UI Degraded",
"timestamp": "May \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e01:13\u003c/var\u003e - \u003cvar data-var='time'\u003e01:32\u003c/var\u003e UTC"
},
{
"code": "xy1tt3hs572m",
"impact": "minor",
"message": "On May 27, 2026, between 12:07 UTC and 13:16 UTC, users experienced degraded performance for Git operations, Pull Requests, Issues, GraphQL API, and related services on github.com. During this time, operations that depended on Git file servers experienced elevated error rates (3.5% of pushes via HTTPS and 0.2% of pushes via SSH failed; no fetches/clones failed). An internal analytics component generated unexpectedly high load, which caused CPU saturation on the underlying infrastructure. This led to cascading slowdowns and errors across services that depend on Git operations. The issue was mitigated by stopping the offending component. Services began recovering shortly after mitigation and were fully restored by 13:16 UTC. We are taking steps to add resource limits and kill switches for internal analytics components to prevent similar issues in the future.",
"name": "Incident with Pull Requests, Issues, Git Operations and API Requests",
"timestamp": "May \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e12:10\u003c/var\u003e - \u003cvar data-var='time'\u003e13:16\u003c/var\u003e UTC"
},
{
"code": "xflkh26pm7vv",
"impact": "minor",
"message": "On May 26, 2026, between 15:10 UTC and 16:35 UTC the Copilot service was degraded and many models were no longer available for use. On average, the error rate was ~5% and peaked at 11% of requests to the service. This was due to a change that introduced a configuration mismatch in HMAC signing credentials which caused the list of available models to be truncated. This was mitigated by rolling back the change. This rollback was complete by 15:34 UTC though users continued to see impact until cache TTLs expired. \u003cbr\u003e\u003cbr\u003eWe are working to improve our monitoring and error handling to reduce time to detection and better experience for issues like this in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "May \u003cvar data-var='date'\u003e26\u003c/var\u003e, \u003cvar data-var='time'\u003e15:44\u003c/var\u003e - \u003cvar data-var='time'\u003e16:35\u003c/var\u003e UTC"
},
{
"code": "gnftqj9htp0g",
"impact": "critical",
"message": "On May 26, 2026, between 10:40 UTC and 12:56 UTC, GitHub Actions jobs were degraded. From 10:40 to 12:16 UTC, all newly queued Actions runs failed to start. From 12:16 to 12:56 UTC, Actions runs that required downloading actions for their workflows continued to fail. GitHub Pages, Copilot Code Review, Copilot coding agent, Octoshift, and GitHub Enterprise Importer were also impacted due to their dependency on Actions. \u003cbr\u003e\u003cbr\u003eThis was caused by our automated account review system incorrectly suspending the service account used by GitHub Actions to authenticate workflow runs and download actions. \u003cbr\u003e\u003cbr\u003eWe mitigated by restoring the account at 12:16 UTC, marking it exempt from further automated review at 12:20 UTC, and redeploying a related service at 12:48 UTC to flush cached account state. Full recovery was confirmed at 12:56 UTC. \u003cbr\u003e\u003cbr\u003eDuring this incident, a small number of Issues, PRs, Comments, and Discussions were marked as hidden when the service account was disabled. No data was lost. All content hidden because of this incident has been restored and full search index restoration is in progress. \u003cbr\u003e\u003cbr\u003eTo prevent a recurrence, we have added an allowlist of all service accounts that cannot be suspended by automated systems, and ensuring these protections are enforced consistently across all account management tooling. We are also improving diagnostic tooling for accounts and reducing cache propagation delays to shorten time to mitigate similar incidents in the future.",
"name": "Incident with Actions and Pages",
"timestamp": "May \u003cvar data-var='date'\u003e26\u003c/var\u003e, \u003cvar data-var='time'\u003e10:57\u003c/var\u003e - \u003cvar data-var='time'\u003e13:18\u003c/var\u003e UTC"
},
{
"code": "tm84vy58jq0h",
"impact": "major",
"message": "On May 25, 2026, between 09:02 UTC and 09:11 UTC, Git push operations over HTTPS and SSH experienced elevated failures. During this window, an average of 31% and a peak of 43% of push requests failed.\n\nThe incident was caused by a recently enabled code path that issued an unexpectedly expensive database query against a primary database. The resulting load exhausted the database's connection pool, which caused the push failures above. The acute impact resolved automatically as in-flight work completed. We mitigated the incident by disabling the feature flag controlling the new code path. To prevent recurrence, we have updated the affected background workflows to route reads to replica databases instead of the primary, removing the specific code pattern that caused this incident; broader follow-up work is underway to apply the same safeguard to similar workflows across GitHub.",
"name": "Elevated rate of Git push errors",
"timestamp": "May \u003cvar data-var='date'\u003e25\u003c/var\u003e, \u003cvar data-var='time'\u003e09:02\u003c/var\u003e - \u003cvar data-var='time'\u003e09:02\u003c/var\u003e UTC"
},
{
"code": "k5z4d1v1tqmt",
"impact": "minor",
"message": "On May 23, 2026 between 06:00 UTC and 19:12 UTC, GitHub experienced intermittent errors authenticating \u003ca href=\"https://docs.github.com/en/apps/creating-github-apps/authenticating-with-a-github-app/authenticating-as-a-github-app-installation\"\u003eGitHub app installation tokens\u003c/a\u003e. \n\nDuring this time, between 1-5% of app installation token authentication requests failed, with an average of 2.3% and the error rate peaking at approximately 5.4% around 14:00 UTC. Users may have experienced authentication failures when using GitHub Apps, including failures in Git operations and API calls using app installation tokens.\n\nThe issue was caused by an issue in a caching proxy component and was remediated by rolling back that component to a previous version. We are taking steps to improve monitoring for cache miss anomalies to ensure that token authentication remains functional during infrastructure changes and reviewing our protocol for testing and when we upgrade third-party dependencies.",
"name": "Intermittent errors with app installation token authentication",
"timestamp": "May \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e16:00\u003c/var\u003e - \u003cvar data-var='time'\u003e19:32\u003c/var\u003e UTC"
},
{
"code": "g6ffrm0rfvz9",
"impact": "minor",
"message": "On May 20, 2026, between 16:00 UTC and 17:45 UTC, GitHub Actions customers experienced run start delays exceeding 5 minutes. Approximately 4.5% of all runs were delayed during the impact window, with scale set jobs disproportionately affected. 30% of scale set jobs were delayed and 4% failed to start entirely. \u003cbr\u003e\u003cbr\u003eThe incident was caused by a misconfigured health check on an internal service that assigns jobs to runners. A brief latency spike in an upstream dependency triggered health check failures across several pods, removing them from service and concentrating load on the remaining capacity. The added load drove memory pressure that escalated into a cascading failure in one regional cluster, leaving it unable to self-recover. \u003cbr\u003e\u003cbr\u003eResponders mitigated the incident by scaling capacity in the healthy regional clusters and draining traffic away from the impaired one, after which run start latency recovered. To prevent recurrence, we are strengthening our health check configuration to avoid cascading failure scenarios and evaluating automated mitigations to rebalance traffic when a region is degraded.",
"name": "Incident with Actions",
"timestamp": "May \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e16:58\u003c/var\u003e - \u003cvar data-var='time'\u003e20:14\u003c/var\u003e UTC"
},
{
"code": "ykb44v068g8v",
"impact": "none",
"message": "On May 19, 2026, between 05:30 UTC and 14:50 UTC, some Copilot users experienced failures when using code completions, chat sessions, and cloud agent sessions. At peak impact, approximately 13% of Copilot API requests failed, and approximately 24% of remote sessions failed to initialize. A partial mitigation at 08:16 UTC reduced the Copilot API error rate to approximately 0.3%, but intermittent failures persisted until a full fix was deployed at 14:15 UTC and recovery was verified by 14:50 UTC.\n\nThe incident was caused by rate limits being exceeded on a shared infrastructure component. A recently enabled feature increased call volume to this component, and the combined load exceeded capacity limits as traffic increased during business hours.\n\nWe mitigated the incident by deploying a caching layer to reduce load on shared infrastructure. To prevent recurrence, we are separating rate limit scopes between services, adding monitoring for internal dependency rate limiting, and reducing redundant calls.",
"name": "Incident with Copilot",
"timestamp": "May \u003cvar data-var='date'\u003e19\u003c/var\u003e, \u003cvar data-var='time'\u003e05:30\u003c/var\u003e - \u003cvar data-var='time'\u003e05:30\u003c/var\u003e UTC"
},
{
"code": "ctf7nxpq5jzn",
"impact": "critical",
"message": "On May 15, 2026, from approximately 07:43 UTC to 08:48 UTC, GitHub Actions experienced a degradation that caused workflow runs to fail or experience delayed starts for a subset of customers. The incident was triggered by a planned failover of supporting infrastructure used by GitHub Actions. During that operation, an automated service discovery update did not propagate correctly, which caused traffic to be routed incorrectly and increased request timeouts in a core dependency for workflow orchestration. \u003cbr\u003e\u003cbr\u003eAt peak impact, 42% of Actions runs failed. Downstream services that depend on Actions workflow execution were also impacted, including GitHub Pages and Copilot cloud services. At 08:12 UTC, responders manually corrected the service discovery routing issue. Timeout and failure rates recovered shortly after, and we continued monitoring until full stabilization was confirmed across all affected services. The incident was marked resolved at 08:48 UTC. \u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are implementing failover guardrails that validate service discovery state before completing failover operations, strengthening pre-flight and post-flight verification checks, and improving dependency resilience to reduce timeout cascades during infrastructure events.",
"name": "Actions is experiencing degraded availability",
"timestamp": "May \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e08:13\u003c/var\u003e - \u003cvar data-var='time'\u003e08:48\u003c/var\u003e UTC"
},
{
"code": "t5mch8tlq6vl",
"impact": "none",
"message": "Beginning at 02:49 UTC on May 15 2026 and lasting until 03:04 UTC, GitHub.com was unavailable for a subset of customers. This impact has been mitigated and normal service resumed.\n\nThe issue was rooted in a sudden spike in traffic, with intermittent impact. We've identified the source of the traffic and prevented further disruption.",
"name": "[Retroactive] Incident with GitHub.com",
"timestamp": "May \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e02:30\u003c/var\u003e - \u003cvar data-var='time'\u003e02:30\u003c/var\u003e UTC"
},
{
"code": "8wxkrmhm6744",
"impact": "minor",
"message": "On May 13, 2026, between 14:31 and 16:03 UTC, the Code Scanning service experienced processing delays and 12% of check runs took over 15 minutes to complete. The delays were caused by replication lag due to an internal database migration, resulting in insufficient worker capacity for our high rate of job enqueues. \u003cbr\u003e\u003cbr\u003eWe mitigated the impact by scaling our processing workers by 34%. Code Scanning results returned to normal processing times after the mitigation was applied. \u003cbr\u003e\u003cbr\u003eThe capacity increases are permanent, and we are looking into more ways to decrease the load on our workers to help prevent this in the future.",
"name": "Incident with CodeQL",
"timestamp": "May \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e14:41\u003c/var\u003e - \u003cvar data-var='time'\u003e16:03\u003c/var\u003e UTC"
},
{
"code": "z3jhyg3l0dvx",
"impact": "minor",
"message": "On May 12, 2026, between 13:41 and 17:43 UTC, some services experienced delays in processing. For the Code Scanning service, 53% of check runs took over 15 minutes to complete. Additionally, notifications took an average of 22 minutes to be delivered and Slack integration webhooks took an average of 20 minutes to be delivered. The delays were caused by replication lag due to an internal database migration, resulting in insufficient worker capacity for our high rate of job enqueues. \u003cbr\u003e\u003cbr\u003eWe mitigated the impact by scaling our processing workers to handle the increased load. All services returned to normal processing times after the mitigation was applied. \u003cbr\u003e\u003cbr\u003eWe are working to create dedicated worker pools for some of our high usage shared queues to help prevent this in the future.",
"name": "Incident with CodeQL, Webhooks, Notifications, and Slack Integration",
"timestamp": "May \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e14:38\u003c/var\u003e - \u003cvar data-var='time'\u003e17:43\u003c/var\u003e UTC"
},
{
"code": "81b05nqkhylj",
"impact": "minor",
"message": "On May 11th, 2026, between 14:00 UTC and 14:33 UTC, HTTP-based Git read operations were degraded. On average, the error rate was 2.8% and peaked at 7.5% of requests to the service. This was due to resource exhaustion in a networking gateway between GitHub.com’s frontend service for Git operations and a dependency service that performs authentication and authorization. Following the initial spike, the frontend service became stuck in a degraded state in one of our data centers, increasing time to mitigation. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by scaling the networking gateway and re-deploying the frontend service. \u003cbr\u003e\u003cbr\u003eTo reduce our time to detection and mitigation in the future, we are adding auto-scaling to the networking gateway, and resolving a bug which caused the frontend service to remain degraded.",
"name": "Incident with high errors on Git Operations",
"timestamp": "May \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e14:25\u003c/var\u003e - \u003cvar data-var='time'\u003e14:33\u003c/var\u003e UTC"
},
{
"code": "qp0lxr014sw8",
"impact": "critical",
"message": "On May 7, 2026, between 04:12 UTC and 06:13 UTC, Copilot Cloud Agent and Copilot Code Review Agent sessions for pull requests were delayed or failed to start.\u003cbr\u003e\u003cbr\u003eThe issue was caused by follow-up recovery work from a separate Pull Requests incident (https://www.githubstatus.com/incidents/f5pb5d5mr9yh). As part of that recovery, we ran a large database migration, which caused replication delays on several replica hosts.\u003cbr\u003e\u003cbr\u003eAlthough those replicas were not serving user traffic, our safeguards correctly treated the elevated replication lag as a signal to slow down writes to the affected database cluster. As a result, some pull request background processing was temporarily delayed. That processing is responsible for sending the internal events that Copilot agents use to begin work, so affected agents did not start until the database replicas caught up.\u003cbr\u003e\u003cbr\u003eThe system recovered once replication lag returned to normal and pull request processing resumed. We are reviewing how this safeguard interacts with recovery migrations so we can reduce the chance of similar secondary impact during future incident recovery work.",
"name": "CCR and CCA failing to start for PR comments",
"timestamp": "May \u003cvar data-var='date'\u003e7\u003c/var\u003e, \u003cvar data-var='time'\u003e05:02\u003c/var\u003e - \u003cvar data-var='time'\u003e06:56\u003c/var\u003e UTC"
},
{
"code": "f5pb5d5mr9yh",
"impact": "critical",
"message": "On May 6, 2026 between 15:12 and 19:02 UTC creation of new pull request review threads on GitHub.com failed. This included new line comments and file comments on pull requests. Existing PRs and previously created comments were unaffected. \u003cbr\u003e\u003cbr\u003eThis incident was caused by a 32-bit integer key reaching its maximum value in a Vitess lookup table used during PR thread creation. The primary table had been migrated to a 64-bit integer key but the Vitesse lookup table remained 32-bit. Once the values in the primary table passed the available 32-bit ID space in the lookup table, attempts to create new review threads began failing, resulting in near 100% failure rate for new thread creation requests. We mitigated the issue by updating the impacted lookup table definitions across all shards to use 64-bit integer column types, increasing the available ID range and restoring normal operation. Service was fully restored once the schema changes competed globally. \u003cbr\u003e\u003cbr\u003eTo help prevent similar incidents, we are expanding existing monitoring of database columns to include Vitess lookup tables to notify in advance of any tables that is approaching a column size limit. This work is intended to provide earlier detection of columns approaching size limits before customer impact occurs.",
"name": "Incident with Pull Requests",
"timestamp": "May \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e15:25\u003c/var\u003e - \u003cvar data-var='time'\u003e19:04\u003c/var\u003e UTC"
},
{
"code": "n8htfgj6wv3t",
"impact": "critical",
"message": "On May 6, 2026 between 11:02 UTC and 11:13 UTC, users were unable to start or view Copilot Cloud Agent or remote sessions. During this time, requests to the session API returned errors, preventing users from creating new sessions or viewing existing ones. The issue was caused by a configuration change to the service's network routing that inadvertently removed the ingress path for the service. The team reverted the change at 11:13 UTC which restored service. The incident remained open until 11:59 UTC while the team verified full recovery. We are taking steps to improve our deployment validation process to prevent similar configuration changes from impacting production traffic in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "May \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e11:21\u003c/var\u003e - \u003cvar data-var='time'\u003e11:59\u003c/var\u003e UTC"
},
{
"code": "1qzncbrxcsy4",
"impact": "major",
"message": "On May 6, 2026, from approximately 06:45 UTC to 09:15 UTC, GitHub Actions Standard Ubuntu hosted runners were degraded. 17.1% of jobs requesting a standard runner failed.\u003cbr\u003e\u003cbr\u003eThis was caused by an unexpected data shape in the allocation configuration data for standard runners. That data was introduced as part of post-incident remediation work for an incident the previous day and caused new allocations to be blocked as load ramped up for the day. Removing that data at 08:51 allowed allocations to proceed and hosted runner pools to scale up and recover.\u003cbr\u003e\u003cbr\u003eWe are updating the filter logic for this allocation data to be resilient to abnormal data shapes and improving monitoring to alert when allocations are blocked, allowing the team to respond before customer impact starts.",
"name": "Incident with Actions, we are investigating reports of degraded availability",
"timestamp": "May \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e07:19\u003c/var\u003e - \u003cvar data-var='time'\u003e09:44\u003c/var\u003e UTC"
},
{
"code": "8kn8t67gdy36",
"impact": "minor",
"message": "Between approximately 14:00 and 16:10 UTC on May 5, 2026, SSH-based Git operations experienced elevated latency and intermittent failures. On average, the error rate was 0.46% and peaked at 0.6% of SSH write requests. HTTP-based Git operations, including web UI and HTTPS clones, were not affected. \u003cbr\u003e\u003cbr\u003eThe impact was caused by reduced SSH capacity at one of our data center sites. During a period of high traffic, the remaining hosts became overloaded, leading to connection exhaustion and some failures for SSH-based operations. \u003cbr\u003e\u003cbr\u003eAdditional capacity was provisioned to expand SSH capacity and resolve the incident. The expanded capacity was fully online by 18:18 UTC. \u003cbr\u003e\u003cbr\u003eTo reduce the likelihood of similar incidents, we will implement faster scaling solutions for SSH infrastructure and improved alerting for host availability and capacity thresholds.",
"name": "Increased Latency and Failures for SSH Git Operations",
"timestamp": "May \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e16:49\u003c/var\u003e - \u003cvar data-var='time'\u003e18:35\u003c/var\u003e UTC"
},
{
"code": "1j40g94rn22j",
"impact": "critical",
"message": "On May 5, 2026, from approximately 13:22 UTC to 17:05 UTC, GitHub Actions hosted runners in the East US region were degraded. 13.5% of jobs requesting a standard runner failed and ~16% of requested Larger Runners with private networking pinned to East US failed or were delayed by more than 5 minutes. Copilot Code Review requests were also impacted. Approximately 8,500 code review requests timed out during this window. Affected users saw an error comment on their pull requests and were able to retry by re-requesting a review. Most runner requests were picked up by other regions automatically, but a portion of requests still routing to East US were impacted.\u003cbr\u003e\u003cbr\u003eThis was triggered by a scale-up operation for hosted runner VMs in the East US region. This is a regular operation, but the VM create load hit an internal rate limit when VM creates pull images from storage. Existing backoff logic was not triggered because of the response code returned in this case. The rate limiting and VM creation failures were mitigated by reducing load to allow for recovery and allowing queued work to be processed. By 15:34 UTC, queued and failed job assignments were mostly mitigated, with less than 0.5% of runner assignments impacted between 15:34 and full recovery at 17:05.\u003cbr\u003e\u003cbr\u003eWe are improving our system’s throttling behavior when limits occur, improving our controls to more quickly mitigate similar situations in the future, and reviewing all limits end-to-end for similar operations. We also immediately paused all scale and similar operations until these changes are in place and validated.",
"name": "Incident with Actions",
"timestamp": "May \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e13:37\u003c/var\u003e - \u003cvar data-var='time'\u003e17:26\u003c/var\u003e UTC"
},
{
"code": "72q3n8yxthcy",
"impact": "critical",
"message": "On 2026-05-04 at 3:37:17 PM UTC we detected increased latency on issues resulting in timeouts, and elevated 500 errors on webhooks. A scheduled workload drove high utilization on the primary host of a critical datastore, saturating the connection pool. We paused the job to mitigate the problem at 4:40:05 PM UTC and have implemented measures to prevent recurrence.",
"name": "Incident with Issues and Webhooks",
"timestamp": "May \u003cvar data-var='date'\u003e4\u003c/var\u003e, \u003cvar data-var='time'\u003e15:45\u003c/var\u003e - \u003cvar data-var='time'\u003e16:40\u003c/var\u003e UTC"
},
{
"code": "x69zbgdyfzg0",
"impact": "minor",
"message": "On April 28, 2026, at approximately 14:07 UTC, GitHub received reports that pull requests were missing from search results across global and repository /pulls pages. \u003cbr\u003e\u003cbr\u003eThe issue was caused by a manually invoked repair job intended for a single repository, which was executed without the required safety flags. During execution of the repair job, the database query remained correctly scoped to the repo’s PR IDs. However, the Elasticsearch reconciliation logic did not apply the same scope. It interpreted the min and max PR IDs as a continuous range, causing unrelated PR documents across other repos to be marked for deletion. This resulted in the removal of 1,789,756,838 PR documents from the search index, approximately 49% of indexed PR documents. \u003cbr\u003e\u003cbr\u003eCustomer impact was limited to PR search and list discoverability. Primary storage was unaffected, and there was no impact to opening, updating, or merging PRs. \u003cbr\u003e\u003cbr\u003eThe issue was identified ~10 minutes after initial customer reports. Because it affected search index completeness rather than service availability, it was not caught by existing monitoring. \u003cbr\u003e\u003cbr\u003eThe root cause was a flaw in the search document repair framework: it allowed a scoped reconciliation to run without enforcing a matching Elasticsearch query scope. This created a destructive mismatch between the source-of-truth and the index. The issue was compounded by the ability to trigger the job from the production console without safety defaults. Prior testing focused only on safe backfill scenarios and did not cover this reconciliation path. Additionally, there was no automated detection for large-volume deletions in Elasticsearch. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident through three parallel actions: (1) Deployed a MySQL-backed search fallback for the most active repos by traffic to restore PR visibility for highly impacted users (2) Initiated a snapshot restore and reindex process to repopulate missing pull request documents in Elasticsearch (3) Added a degradation notice on PR pages to inform users of incomplete search results while recovery was in progress. The incident was resolved on May 1, 2026 at 4:15 UTC, following completion and validation of the reindex process. \u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are prioritizing improvements to the repair framework and safeguards. These include enforcing scoped query alignment between primary storage and Elasticsearch, preventing destructive operations without explicit opt-in, strengthening guardrails for manual repair jobs, and evaluating restrictions on production console access. \u003cbr\u003e\u003cbr\u003eIn parallel, we are expanding automated test coverage for reconciliation safety invariants and introducing detection for anomalous deletion patterns in Elasticsearch so similar issues can be identified or blocked earlier. \u003cbr\u003e\u003cbr\u003eWe are committed to improving the safety and reliability of our repair systems and ensuring that operational workflows are resilient to both software defects and manual invocation risks.",
"name": "Incomplete pull request results in repositories",
"timestamp": "Apr \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e14:17\u003c/var\u003e - May \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e04:15\u003c/var\u003e UTC"
}
],
"name": "May",
"year": 2026
},
{
"incidents": [
{
"code": "dbypmw7h77l5",
"impact": "minor",
"message": "On April 28, 2026, from approximately 12:41 UTC to 17:09 UTC, GitHub Actions jobs using Standard Ubuntu 22 and Ubuntu 24 hosted runners experienced run start delays. Approximately 8% of hosted runner jobs using Ubuntu 22 and Ubuntu 24 experienced delays greater than 5 minutes or failures. Larger and self-hosted runners were not impacted.\u003cbr\u003e\u003cbr\u003eThis was caused by a performance regression introduced in the VM reimage process. That reimage delay lowered the overall capacity of runners available to pick up new jobs. This was mitigated with a rollback to a known good image version.\u003cbr\u003e\u003cbr\u003eWe are addressing the core issue with reimage performance and improving the granularity of reimage telemetry across our services and our compute provider to more quickly diagnose similar issues in the future. Finally, we are evaluating other rollout changes to automatically detect similar regressions.",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e13:59\u003c/var\u003e - \u003cvar data-var='time'\u003e17:09\u003c/var\u003e UTC"
},
{
"code": "ql942tw29yl6",
"impact": "critical",
"message": "On April 27, 2026 between 16:15 UTC and 22:46 UTC, GitHub search services experienced degraded connectivity due to saturation of the load balancing tier deployed in front of our search infrastructure. This resulted in intermittent failures for services relying on our search data including Issues, Pull Requests, Projects, Repositories, Actions, Package Registry and Dependabot Alerts. The impact was varied by search target, with services seeing up to 65% of searches timing out or returning an error between 16:15 UTC and 18:00 UTC. \u003cbr\u003e\u003cbr\u003eWe detected the drop in search results through our ongoing monitoring and declared an incident at 16:21 UTC when we determined the issues would not self-heal. We tracked the incident as mitigated as of 21:33 UTC and monitored the systems until 22:46 UTC when we declared the incident resolved. Our existing monitoring did not classify the increased scraping as a risk and this dimension of the incident was only discovered while working to mitigate. \u003cbr\u003e\u003cbr\u003eThe saturation was caused by a large influx of anonymous distributed scraping traffic that was crafted to avoid our public API rate limits. This scraping traffic made up 30% of the day’s total search traffic, but it was concentrated within a four-hour period. The traffic originated from over 600,000 Unique IP addresses, with matching actor information across the board. \u003cbr\u003e\u003cbr\u003eTo mitigate, we immediately focused on relieving pressure from the load balancers while simultaneously working on scaling the load balancing tier, blocking the anomalous traffic and applying tuning to the balancers to fully resolve the incident. \u003cbr\u003e\u003cbr\u003eLooking ahead, we’ve not only scaled the load balancer tier, but applied optimizations to improve our connection handling and re-use to reduce the possibility that a saturation event like this can re-occur. We’ve also added new monitors and controls within the platform to allow us to restrict anonymous traffic to mitigate the impact to our registered users. \u003cbr\u003e\u003cbr\u003e",
"name": "GitHub search is degraded",
"timestamp": "Apr \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e16:31\u003c/var\u003e - \u003cvar data-var='time'\u003e22:46\u003c/var\u003e UTC"
},
{
"code": "vq183jvj6vrw",
"impact": "minor",
"message": "On April 22, 2026 from 18:49 to 19:32 UTC , the Copilot Cloud Agent service began failing during session execution for users running the Agent HQ Codex agent. Codex agent sessions failed to start for all entry points (issue assignment, @copilot comment mentions). 0.5% of total Copilot Cloud Agent jobs were impacted (~2,000 failed jobs). Copilot and other agent sessions were unaffected.\u003cbr\u003e\u003cbr\u003eThis was caused by a model resolution mismatch in Codex agent sessions, resulting in an incompatible model being used at runtime. A mitigation was deployed to select a stable default model for Codex agent sessions.\u003cbr\u003e\u003cbr\u003eWe are working to harden the underlying model-resolution path so it correctly scopes to the requesting agent's supported models to prevent similar failure mode in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e16:48\u003c/var\u003e - \u003cvar data-var='time'\u003e19:02\u003c/var\u003e UTC"
},
{
"code": "3dqg0nhwnxs2",
"impact": "minor",
"message": "On April 24, 2026, from approximately 11:39 UTC to April 25, 2026 at 00:15 UTC, GitHub Actions experienced delays and timeouts for Larger Hosted Runner jobs using VNet injection in the East US region without a failover region configured. Standard and Self-hosted runners were not impacted. This was caused by backend failures in our compute provider’s provisioning, scaling, and update operations for VMs in the East US region and mitigated by a rollback across all affected Availability Zones. More detail is available at https://azure.status.microsoft/en-us/status/history/?trackingId=5GP8-W0G.\u003cbr\u003e\u003cbr\u003eWe are working to improve the reliability of our annotations for jobs impacted by regional issues and are adding system log notifications as an additional customer communication channel alongside annotations.\u003cbr\u003e\u003cbr\u003eVNet Failover is also now in public preview, allowing customers to evacuate Larger Hosted Runners using VNet injection in cases like this.",
"name": "Delays with Actions Jobs for Larger Runners using VNet Injection in the East US region",
"timestamp": "Apr \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e19:02\u003c/var\u003e - Apr \u003cvar data-var='date'\u003e25\u003c/var\u003e, \u003cvar data-var='time'\u003e00:36\u003c/var\u003e UTC"
},
{
"code": "zsg1lk7w13cf",
"impact": "minor",
"message": "On April 23, 2026, between 16:05 UTC and 20:43 UTC, the Pull Requests service experienced a regression affecting merge queue operations. PRs merged via merge queue using the squash merge method produced incorrect merge commits when the merge group contained more than one PR. In affected cases, changes from previously merged PRs and prior commits were inadvertently reverted by subsequent merges.\u003cbr\u003e\u003cbr\u003eDuring the impact window 2,092 pull requests were affected. The issue did not affect pull requests merged outside of merge queue, nor merge queue groups using the merge or rebase methods.\u003cbr\u003eIt took approximately 3 hours and 33 minutes to identify the issue. The change completed deployment at approximately 16:05 UTC, and we became aware at 19:38 UTC following an increase in customer support inquiries. Because the issue affected merge commit correctness rather than availability, it was not detected by existing automated monitoring and was identified through customer reports.\u003cbr\u003e\u003cbr\u003eThe regression was introduced by a new code path that adjusted merge base computation for merge queue ref updates. This code path was intended to be gated behind a feature flag for an unreleased feature, but the gating was incomplete.\u003cbr\u003e\u003cbr\u003eAs a result, the new behavior was inadvertently applied to squash merge groups, producing an incorrect three-way merge. This caused subsequent squash merges to revert changes from earlier pull requests and, in some cases, changes between their starting points.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by reverting the code change and force-deploying the fix across all environments. After resolution, we identified affected repositories and sent targeted remediation instructions to repository administrators with step-by-step recovery guidance.\u003cbr\u003e\u003cbr\u003eThe regression was not identified during internal validation. Existing test coverage primarily exercised single-PR merge queue groups, which did not exhibit the faulty base-reference calculation. Because automated checks did not validate merge correctness for multi-PR squash groups, the defect surfaced only in production.\u003cbr\u003e\u003cbr\u003eTo prevent recurrence, GitHub is expanding test coverage for merge correctness validation. We are broadening automated coverage for merge queue operations, including regression checks that validate resulting Git contents across supported configurations, so issues affecting merge correctness are caught before reaching production.\u003cbr\u003e\u003cbr\u003eWe are committed to ensuring the correctness and reliability of merge queue operations. These actions will reduce the risk of similar regressions and improve confidence in future changes to the Pull Requests service.",
"name": "Incident with Pull Requests",
"timestamp": "Apr \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e19:50\u003c/var\u003e - \u003cvar data-var='time'\u003e21:43\u003c/var\u003e UTC"
},
{
"code": "40f3r1jvjwl2",
"impact": "minor",
"message": "Between 18:45 and 19:42 UTC on April 23, users were unable to start new agent tasks using either Claude or Codex agent on github.com. This was caused by a code change to how Copilot mission control routes task creation requests. Ongoing agent tasks and other Copilot agent features were not affected. We mitigated the impact by reverting the breaking change. We are adding extra monitoring and integration test coverage for the task creation path to prevent future recurrence.",
"name": "Disruption with users unable to start Claude and Codex agent task from the web",
"timestamp": "Apr \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e19:28\u003c/var\u003e - \u003cvar data-var='time'\u003e19:42\u003c/var\u003e UTC"
},
{
"code": "myrbk7jvvs6p",
"impact": "critical",
"message": "On April 23, 2026, between 16:03 UTC and 17:27 UTC, multiple GitHub services experienced elevated error rates and degraded performance due to DNS resolution failures originating from our DNS infrastructure in our VA3 datacenter. Approximately 5–7% of overall traffic was affected during the impact window: \u003cbr\u003e\u003cbr\u003e- Webhooks: ~0.35% of API requests returned 5xx (peak ~0.39%). ~0.88% of requests exceeded 3s latency; at peak, \u0026gt;3s responses represented ~10% of Webhooks API traffic. \u003cbr\u003e\u003cbr\u003e- Copilot Metrics: ~9% of Copilot Insights dashboard requests returned 5xx. \u003cbr\u003e\u003cbr\u003e- Copilot cloud agents: ~10% of cloud agent sessions were affected and failing. \u003cbr\u003e\u003cbr\u003e- Octoshift: 0.88% of active repo migrations failed and 79% saw elevated durations (avg. 5.2 min) during this period. \u003cbr\u003e\u003cbr\u003e- Git Operations: averaged 1.25% errors over the duration of the incident, with a peak of 2.07% errors. \u003cbr\u003e\u003cbr\u003e- Actions: Workflow run status updates experienced delays of up to ~8s over the duration of the incident window. \u003cbr\u003e\u003cbr\u003eOur DNS infrastructure in VA3 entered a degraded state and began intermittently returning NXDOMAIN responses and timing out on lookups for both internal service discovery and external endpoints. This caused a cascading impact across the dependent services listed above. \u003cbr\u003e\u003cbr\u003eWe identified a specific load pattern under which our DNS resolvers began failing. The evidence points to a recently introduced traffic-balancing mechanism, rolled out progressively to support our growth, as the root cause. We have since reverted this change. \u003cbr\u003e\u003cbr\u003eWe are immediately prioritizing investments in a more controlled rollout and validation process, including a dedicated environment to safely shadow production DNS traffic and detect these failure modes before they can affect production.",
"name": "Incident with multiple GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e16:12\u003c/var\u003e - \u003cvar data-var='time'\u003e17:30\u003c/var\u003e UTC"
},
{
"code": "3f30ycplyr7l",
"impact": "minor",
"message": "On April 23, 2026 between 14:30 UTC and 15:18 UTC multiple services were degraded on github.com. During this time approximately 1.5% of all web requests resulted in a 5xx status and unicorn pages for github.com users. We also saw elevated error rates across Actions workflow runs, Copilot, Codespaces and Packages, leading to degraded experiences during this timeframe. Codespaces impact peaked at 45% failures for create requests and 65% failures for resume requests. Packages impact was mainly Maven related with 50% failure rates in downloads and 70% failure rates in uploads. Actions experienced a peak of 8% of failed jobs and up to 85% of jobs impacted by run start delays of more than 5 minutes.\u003cbr\u003e\u003cbr\u003eThis was due to a configuration change to an internal billing service that led to a cache being overwhelmed and causing requests to time out. These timeouts cascaded across multiple services and eventually caused requests to queue up and exhaust web request workers.\u003cbr\u003e\u003cbr\u003eThis configuration change was reverted at 14:42 UTC and following this, all services began to see recovery immediately.\u003cbr\u003e\u003cbr\u003eTo prevent this situation in the future, we are taking steps to ensure that failures and timeouts in the billing service don’t cascade to other services causing impact. This includes implementing more aggressive timeouts on callers of these billing services, adding circuit breaker configurations for cache timeouts and using more resilient cache options. We have also decreased max request timeouts within the billing service that caused impact and added more capacity to our cache to prevent traffic spikes from having the same impact.",
"name": "Investigating errors on GitHub",
"timestamp": "Apr \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e14:40\u003c/var\u003e - \u003cvar data-var='time'\u003e15:18\u003c/var\u003e UTC"
},
{
"code": "c53mty24vwg8",
"impact": "minor",
"message": "On April 22, 2026, between 09:00 UTC and 22:05 UTC, the Copilot coding agent and pull request comment event processing were degraded. During this period, approximately 0.5% of total pull request and issue comments mentioned @copilot (~23,000 invocations), explicitly requested work from the Copilot coding agent but were not acted upon.\u003cbr\u003e\u003cbr\u003eCreating, viewing, and replying to pull request comments was unaffected, and other Copilot\u003cbr\u003efunctionality continued to operate normally. The impact was limited to @copilot mentions on pull request comments not triggering Copilot coding agent runs, and to some downstream systems not receiving new pull request comment events during the impact window.\u003cbr\u003e\u003cbr\u003eThe cause was a serialization error that prevented pull request comment events from being published to downstream consumers, including the Copilot coding agent. This was related to the same class of issue as incident #4295 on April 20, affecting a another event type.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by deploying a fix that restored event publishing, after which the Copilot coding agent and other downstream consumers resumed processing pull request comment events normally.\u003cbr\u003e\u003cbr\u003eWe are working to complete our audit of related event schemas, migrate remaining consumers to use\u003cbr\u003ethe updated identifier fields, and improve monitoring to detect drops in publishing on critical event topics, to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e19:53\u003c/var\u003e - \u003cvar data-var='time'\u003e22:43\u003c/var\u003e UTC"
},
{
"code": "8790jgb3ky6t",
"impact": "critical",
"message": "On April 22, 2026, between 15:16 UTC and 19:18 UTC, users experienced errors when interacting with Copilot Chat on github.com and Copilot Cloud Agent. During this time, users were unable to use Copilot Chat or Copilot Cloud Agent. Copilot Memory (in preview) was not available to Copilot agent sessions during this time. The issue was caused by an infrastructure configuration change that resulted in connectivity issues with our databases. The team identified the cause and restored connectivity to the database. Copilot Chat and Cloud Agent for github.com were restored by 18:16 UTC. Remaining regional deployments were restored incrementally, with full resolution at 19:18 UTC. We have taken steps to prevent similar infrastructure changes from causing these kinds of database operations in the future.",
"name": "Disruption with Copilot chat and Copilot Coding Agent",
"timestamp": "Apr \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e15:35\u003c/var\u003e - \u003cvar data-var='time'\u003e19:18\u003c/var\u003e UTC"
},
{
"code": "b4235qyfzx2z",
"impact": "minor",
"message": "On April 21, 2026, between 13:35 UTC and 01:24 UTC the following day the projects service was degraded. During this time period, projects may have been out of sync and users may have experienced delays in changes to projects and their items. Delays in reflected changes peaked at approximately 45 minutes. The delays were caused by serialization errors that failed events and triggered a flood of resyncs, overloading our event processing layers.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by speeding up processing time for incoming changes and otherwise waiting for all changes to be processed.\u003cbr\u003e\u003cbr\u003eWe are working to increase our capacity for processing updates to projects to reduce our time to mitigation of issues like this one in the future.",
"name": "Disruption with projects service",
"timestamp": "Apr \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e15:03\u003c/var\u003e - Apr \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e01:24\u003c/var\u003e UTC"
},
{
"code": "m7pgzw1wlfq7",
"impact": "major",
"message": "On April 20, 2026 between 10:28 UTC and 15:04 UTC GitHub experienced degraded service for code scanning default setup, code quality, and project boards. Repair of affected project boards additionally lasted until April 21, 05:04 UTC \u003cbr\u003e\u003cbr\u003eDuring this time, code scanning default setup and code quality analyses were not triggered on newly opened pull requests. Additionally, newly created issues were not appearing on project boards. \u003cbr\u003e\u003cbr\u003eThe cause was a serialization error that prevented proper triggering of code scanning, code quality analyses, and project board updates. \u003cbr\u003e\u003cbr\u003eWe mitigated the issue by deploying a fix, restoring event publishing for code scanning and code quality. For project boards, an additional code change was deployed to update event consumers, followed by a reindex of affected project items. \u003cbr\u003e\u003cbr\u003eWe are working to prevent recurrence by strengthening our schema validations and improving monitoring for drops in publishing on critical hydro topics.",
"name": "Partial degradation for code scanning default setup and for code quality",
"timestamp": "Apr \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e13:28\u003c/var\u003e - Apr \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e05:04\u003c/var\u003e UTC"
},
{
"code": "kcf0thb1zp30",
"impact": "minor",
"message": "On April 17, 2026, between 14:46 UTC and 15:12 UTC, users experienced a degraded web experience on GitHub.com. During this time, approximately 1.5% of web requests resulted in errors, with some users encountering slow page loads or failed requests. The issue was caused by capacity saturation of a caching component in one of our data center regions. We mitigated the issue by redirecting traffic to an unaffected region and rolling back a recent deployment. The incident was fully resolved at 15:18 UTC. We are taking steps to provide appropriate capacity for this caching path to prevent recurrence.",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e14:56\u003c/var\u003e - \u003cvar data-var='time'\u003e15:18\u003c/var\u003e UTC"
},
{
"code": "bjyfg0yz1lwj",
"impact": "major",
"message": "On April 16, 2026 between 09:30 UTC and 17:15 UTC, users experienced failures when attempting to connect to GitHub Codespaces via the VS Code editor. During this time, approximately 40% of codespace start operations failed. Users connecting via SSH were not impacted. \u003cbr\u003e\u003cbr\u003eThe issue was caused by a failure in an upstream download service that prevented the VS Code Server from being retrieved during codespace startup. The impact was mitigated by implementing a workaround to use an alternative download path when the primary endpoint is degraded. \u003cbr\u003e\u003cbr\u003eWe are working with the upstream dependency to address the root cause of the download service failure, and we are improving our fallback mechanisms to reduce the impact of similar upstream failures in the future.",
"name": "Incident with Codespaces",
"timestamp": "Apr \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e15:06\u003c/var\u003e - \u003cvar data-var='time'\u003e18:28\u003c/var\u003e UTC"
},
{
"code": "gldshbt8tvdg",
"impact": "minor",
"message": "On April 14, between 00:58 UTC and 06:08 UTC, GitHub Enterprise Cloud customers experienced 500 errors when attempting to access Copilot Insights pages which was caused by an authentication failure in our metrics pipeline. We fully mitigated the issue and validated the fix in production. Approximately 709 users were impacted. The total impact duration was approximately 5 hours and 10 minutes. \u003cbr\u003e\u003cbr\u003eOur investigation determined the incident was caused by a change in a tenant credential which caused authentication errors to retrieve the required data needed on our Copilot Insights pages. \u003cbr\u003e\u003cbr\u003eWe understand this disruption impacted customers' ability to access the Copilot Insights page. To prevent similar issues and reduce resolution time in the future, we are investing in improved diagnostics tooling to quickly identify the root cause of failures, enhanced monitoring, and alerting to detect issues at a more granular level. \u003cbr\u003e\u003cbr\u003eGitHub is a critical infrastructure for your work, your teams, and your businesses. We are focused on these remediations and continued reliability improvements for Copilot Insights and related metrics experiences.",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e01:57\u003c/var\u003e - \u003cvar data-var='time'\u003e06:08\u003c/var\u003e UTC"
},
{
"code": "96r9hyz795sv",
"impact": "major",
"message": "On Sunday April 13th, 2026, between 18:53 UTC and 20:30 UTC, the GitHub Pages service experienced elevated error rates. On average, the error rate was 10.58% and peaked at 12.77% of requests to the service, resulting in approximately 17.5 million failed requests returning HTTP 500 errors. This was due to an automated DNS management tool (octodns) erroneously deleting a DNS record for a Pages backend storage host after its upstream data source intermittently failed to return the record, causing the tool to treat it as stale and remove it.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by re-creating the deleted DNS record. To prevent future incidents, we are implementing availability-zone-tolerant routing in the Pages frontend so that an unresolvable backend host triggers failover to healthy hosts rather than returning errors, adding safeguards to prevent automated deletion of DNS records owned by other systems, and improving logging and alerting for DNS resolution failures in the Pages serving path.",
"name": "Incident with Pages",
"timestamp": "Apr \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e19:56\u003c/var\u003e - \u003cvar data-var='time'\u003e20:35\u003c/var\u003e UTC"
},
{
"code": "1ghdqdbyyjmt",
"impact": "minor",
"message": "On April 13, 2026, between 14:41 UTC and 17:29 UTC, the Copilot service experienced degraded performance. All Copilot users were impacted by increased latency, and approximately 20% experienced request failures when interacting with Copilot Cloud Agent (CCA). On average, request latency increased to approximately 950ms. The GitHub User Dashboard also displayed intermittent errors loading Copilot quota information. CCA and the User Dashboard were impacted for approximately 2 hours and 56 minutes. \u003cbr\u003e\u003cbr\u003eThis was due to an infrastructure change that reduced the available compute capacity for a backend service responsible for Copilot rate limiting and quota management. The reduced capacity caused resource exhaustion under normal traffic load, leading to cascading failures in downstream request processing. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by increasing compute resources allocated to the affected service and scaling out the number of service instances to distribute load more effectively. \u003cbr\u003e\u003cbr\u003eWe are working to improve proactive capacity monitoring to detect resource degradation before it impacts users, reviewing retry and timeout configurations across dependent services to reduce amplification during degraded states, and evaluating connection management strategies to improve resilience under constrained resources.",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e16:41\u003c/var\u003e - \u003cvar data-var='time'\u003e17:40\u003c/var\u003e UTC"
},
{
"code": "8hn4vclxz1lq",
"impact": "minor",
"message": "On April 9, 2026, between 22:59 UTC and April 10, 2026, 13:24 UTC, the Copilot Mission Control service was degraded and did not display Claude and Codex Cloud Agent sessions in the agents tab dashboard. Customers were unable to see, list, or manage their third party agent sessions during this period. The underlying agent sessions continued to function normally. This was a visibility and management issue only, and no HTTP errors were generated. The API returned successful responses with incomplete results, with an average error rate of 0% and a maximum error rate of 0%. This was due to a code change that introduced a filter which inadvertently excluded third party agent sessions.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by reverting the problematic code change and deploying the fix to production.\u003cbr\u003e\u003cbr\u003eWe are working to add automated monitoring for dashboard content visibility and improve integration test coverage for third party agent session listing to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Problems with third-party Claude and Codex Agent sessions not being listed in the agents tab dashboard",
"timestamp": "Apr \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e13:07\u003c/var\u003e - \u003cvar data-var='time'\u003e13:28\u003c/var\u003e UTC"
},
{
"code": "2rqwxl8y7m0j",
"impact": "major",
"message": "On April 9, 2026, between 16:05 UTC and 20:36 UTC, the Copilot cloud agent service was degraded, causing new agent sessions to be delayed or fail to start. Users who attempted to start Copilot cloud agent sessions during this period experienced jobs getting stuck in the queue, with wait times peaking at 54 minutes compared to the normal 15–40 seconds. On average, approximately 84% of requests to start agent sessions failed, peaking at 97.5% during the worst period.\u003cbr\u003e\u003cbr\u003eThis was due to an internal service exceeding API rate limits, compounded by a caching bug that persisted the rate-limited state beyond the actual rate limit window, causing recurring outage waves rather than a single recovery.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by deploying a configuration change to bypass the affected cache and shifting API traffic to an alternative authentication path that reduced rate limit exposure. We have since added automated monitoring and alerting for this failure mode, deployed per-endpoint rate limit controls, and added caching for high-traffic API calls to reduce overall load. We are also working on longer-term improvements to rate limit isolation and traffic management to prevent similar issues in the future.\u003cbr\u003e\u003cbr\u003eThis incident shared the same underlying root causes with an incident declared in the time frame https://www.githubstatus.com/incidents/zn1t56bfxdzg",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e16:20\u003c/var\u003e - \u003cvar data-var='time'\u003e20:36\u003c/var\u003e UTC"
},
{
"code": "zn1t56bfxdzg",
"impact": "major",
"message": "On April 9, 2026, between 09:05 UTC and 19:05 UTC, the Copilot coding agent service was degraded and users experienced significant delays starting new agent sessions. Approximately 84% of new agent session requests were delayed across four separate outage waves, with queue wait times peaking at 54 minutes compared to a normal baseline of 15–40 seconds. On average, the error rate was 83.9% and peaked at 97.5% of requests to the service. Approximately 22,700 workflow creations were delayed or failed during the incident.\u003cbr\u003e\u003cbr\u003eThis was due to a bug in our rate limiting logic that incorrectly applied a rate limit globally across all users, rather than scoping it to the individual installation that triggered the limit. A contributing factor was a surge in API traffic from a client update that increased requests to an internal endpoint by 3–4x, which accelerated rate limit exhaustion.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by disabling the faulty rate limit caching mechanism via feature flag and updating our service to use per-installation credentials for API calls, ensuring rate limits are correctly scoped to individual installations.\u003cbr\u003e\u003cbr\u003eWe have since added automated monitoring and alerting to detect this failure mode proactively, deployed fixes to reduce unnecessary API traffic through caching improvements, and are continuing work to further isolate rate limit scoping across client types to prevent similar issues in the future.\u003cbr\u003e\u003cbr\u003eThis incident shared the same underlying root causes with an incident declared in the time frame https://www.githubstatus.com/incidents/2rqwxl8y7m0j",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e09:50\u003c/var\u003e - \u003cvar data-var='time'\u003e10:15\u003c/var\u003e UTC"
},
{
"code": "bnvg7qml7krl",
"impact": "minor",
"message": "On April 9, 2026, between 03:22 UTC and 04:49 UTC, GitHub Notifications experienced degraded availability. During this time, approximately 45% of requests to the notifications service returned errors, with a peak error rate of approximately 54%, preventing affected users from successfully viewing or interacting with their notifications service. The issue was identified and resolved, restoring the service to full availability.\u003cbr\u003e\u003cbr\u003eWe are working to improve our metrics to reduce time to detection and mitigation for similar issues in the future.",
"name": "Disruption with GitHub notifications",
"timestamp": "Apr \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e04:42\u003c/var\u003e - \u003cvar data-var='time'\u003e04:57\u003c/var\u003e UTC"
},
{
"code": "d96l71t3h63k",
"impact": "minor",
"message": "Between 15:20 and 20:18 UTC on Thursday April 2, Copilot Cloud Agent entered a period of reduced performance. Due to an internal feature being developed for Copilot Code Review, the Copilot Cloud Agent infrastructure started to receive an increased number of jobs. This load eventually caused us to hit an internal rate limit, causing all work to suspend for an hour. During this hour, some new jobs would time out, while others would resume once rate limiting ended. Roughly 40% of jobs in this period were affected.\u003cbr\u003e\u003cbr\u003eOnce the cause of this rate limiting was identified, we were able to disable the new CCR feature via a feature flag. Once the jobs that were already in the queue were able to clear, we didn't see additional instances of rate limiting afterwards.",
"name": "Disruption with some GitHub services",
"timestamp": "Apr \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e17:49\u003c/var\u003e - \u003cvar data-var='time'\u003e21:48\u003c/var\u003e UTC"
},
{
"code": "j3sgbdw2lw3c",
"impact": "minor",
"message": "Between 15:20 and 20:18 UTC on Thursday April 2, Copilot Cloud Agent entered a period of reduced performance. Due to an internal feature being developed for Copilot Code Review, the Copilot Cloud Agent infrastructure started to receive an increased number of jobs. This load eventually caused us to hit an internal rate limit, causing all work to suspend for an hour. During this hour, some new jobs would time out, while others would resume once rate limiting ended. Roughly 40% of jobs in this period were affected.\u003cbr\u003e\u003cbr\u003eOnce the cause of this rate limiting was identified, we were able to disable the new CCR feature via a feature flag. Once the jobs that were already in the queue were able to clear, we didn't see additional instances of rate limiting afterwards.\u003cbr\u003e\u003cbr\u003eThis was the same incident declared in https://www.githubstatus.com/incidents/d96l71t3h63k",
"name": "Copilot Coding Agent failing to start some jobs",
"timestamp": "Apr \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e16:18\u003c/var\u003e - \u003cvar data-var='time'\u003e16:30\u003c/var\u003e UTC"
},
{
"code": "xs2p9z3b73r3",
"impact": "major",
"message": "On April 1st, 2026 between 14:40 and 17:00 UTC the GitHub code search service had an outage which resulted in users being unable to perform searches.\u003cbr\u003e\u003cbr\u003eThe issue was initially caused by an upgrade to the code search Kafka cluster ZooKeeper instances which caused a loss of quorum. This resulted in application-level data inconsistencies which required the index to be reset to a point in time before the loss of quorum occurred. Meanwhile, an accidental deploy resulted in query services losing their shard-to-host mappings, which are typically propagated by Kafka.\u003cbr\u003e\u003cbr\u003eWe remediated the problem by performing rolling restarts in the Kafka cluster, allowing quorum to be reestablished. From there we were able to reset our index to a point in time before the inconsistencies occurred.\u003cbr\u003e\u003cbr\u003eThe team is working on ways to improve our time to respond and mitigate issues relating to Kafka in the future.",
"name": "Disruption with GitHub's code search",
"timestamp": "Apr \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e15:02\u003c/var\u003e - \u003cvar data-var='time'\u003e23:45\u003c/var\u003e UTC"
},
{
"code": "00h92ct3ggvv",
"impact": "major",
"message": "On April 1, 2026, between 15:34 UTC and 16:02 UTC, our audit log service lost connectivity to its backing data store due to a failed credential rotation. During this 28-minute window, audit log history was unavailable via both the API and web UI. This resulted in 5xx errors for 4,297 API actors and 127 github.com users. Additionally, events created during this window were delayed by up to 29 minutes in github.com and event streaming. No audit log events were lost; all audit log events were ultimately written and streamed successfully. Customers using GitHub Enterprise Cloud with data residency were not impacted by this incident. \u003cbr\u003e\u003cbr\u003eWe were alerted to the infrastructure failure at 15:40 UTC — six minutes after onset — and resolved the issue by recycling the affected environment, restoring full service by 16:02 UTC. We are conducting a thorough review of our credential rotation process to strengthen its resiliency and prevent recurrence. In parallel, we are strengthening our monitoring capabilities to ensure faster detection and earlier visibility into similar issues going forward.",
"name": "GitHub audit logs are unavailable",
"timestamp": "Apr \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e16:06\u003c/var\u003e - \u003cvar data-var='time'\u003e16:10\u003c/var\u003e UTC"
},
{
"code": "tstkmx7pcs5d",
"impact": "minor",
"message": "On April 1, 2026, between 07:29 and 12:41 UTC, some customers experienced elevated 5xx errors and increased latency when using GitHub Copilot features that rely on `/agents/sessions` endpoints (including creating or viewing agent sessions). The issue was caused by resource exhaustion in one of the Copilot backend services handling these requests, in turn, causing timeouts and failed requests. We mitigated the incident by increasing the service’s available compute resources and tuning its runtime concurrency settings. Service health returned to normal and the incident was fully resolved by 12:41 UTC.",
"name": "Incident with Copilot",
"timestamp": "Apr \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e09:58\u003c/var\u003e - \u003cvar data-var='time'\u003e12:41\u003c/var\u003e UTC"
}
],
"name": "April",
"year": 2026
}
],
"start_time": "2026-04-01T00:00:00Z",
"time_zone": "UTC"
},
{
"end_time": "2026-03-31T23:59:59Z",
"months": [
{
"incidents": [
{
"code": "ml7wplmxbt5l",
"impact": "minor",
"message": "On Monday March 31st, 2026, between 13:53 UTC and 21:23 UTC the Pull Requests service experienced elevated latency and failures. On average, the error rate was 0.15% and peaked at 0.28% of requests to the service. This was due to a change in garbage collection (GC) settings for a Go-based internal service that provides access to Git repository data. The changes caused more frequent GC activity and elevated CPU consumption on a subset of storage nodes, increasing latency and failure rates for some internal API operations.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by reverting the GC changes. To prevent future incidents and improve time to detection and mitigation, we are instrumenting additional metrics and alerting for GC-related behavior, improving our visibility into other signals that could cause degraded impact of this type, and updating our best practices and standards for garbage collection in Go-based services.\u003cbr\u003e",
"name": "Incident with Pull Requests: High percentage of 500s",
"timestamp": "Mar \u003cvar data-var='date'\u003e31\u003c/var\u003e, \u003cvar data-var='time'\u003e15:05\u003c/var\u003e - \u003cvar data-var='time'\u003e21:23\u003c/var\u003e UTC"
},
{
"code": "2yb2tvwny50l",
"impact": "minor",
"message": "On March 31, 2026, between 06:15 UTC and 15:30 UTC, the GitHub billing usage reports feature was degraded due to reduced server capacity. Customers requesting billing usage reports and loading the top usage by organization and repository on the billing overview and usage pages were impacted. The average error rate for usage report requests was 15%, peaking at 98% over an eight-minute window. For the billing pages, an average of 56% of requests failed to load the top usage cards. The root cause was an increase in billing usage report requests with large datasets, which exhausted the capacity of the nodes responsible for reporting data. There was no impact on billing charges. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by adjusting our auto-scaling thresholds to better meet our capacity needs. We are working to improve our metrics to reduce time to detection and mitigation for similar issues in the future.",
"name": "Issues with metered billing report generation",
"timestamp": "Mar \u003cvar data-var='date'\u003e31\u003c/var\u003e, \u003cvar data-var='time'\u003e13:47\u003c/var\u003e - \u003cvar data-var='time'\u003e15:10\u003c/var\u003e UTC"
},
{
"code": "c3ctbhdcvcc8",
"impact": "minor",
"message": "On March 30, 2026, between 10:11 UTC and 13:25 UTC, GitHub Actions experienced degraded performance. During this time, approximately 2.65% of workflow jobs triggered by pull request events experienced start delays exceeding 5 minutes. The issue was caused by replication lag on an internal database cluster used by Actions, which triggered write throttling in our database protection layer and slowed job queue processing. \u003cbr\u003e\u003cbr\u003eThe replication lag originated from planned maintenance to scale the internal database. Newly added database hosts triggered guardrails in the throttling layer, restricting write throughput. The incident was mitigated by excluding the new hosts from replication delay calculations. \u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we have updated our maintenance procedures to ensure new hosts are excluded from throttling assessments during scaling operations. Additionally, we are investing in automation to streamline this type of maintenance activity.",
"name": "Elevated delays in Actions workflow runs and Pull Request status updates",
"timestamp": "Mar \u003cvar data-var='date'\u003e30\u003c/var\u003e, \u003cvar data-var='time'\u003e13:02\u003c/var\u003e - \u003cvar data-var='time'\u003e13:25\u003c/var\u003e UTC"
},
{
"code": "9vyj2jwsk8by",
"impact": "none",
"message": "On March 27, 2026, from 02:30 to 04:56 UTC, a misconfiguration in our rate limiting system caused users on Copilot Free, Student, Pro, and Pro+ plans to experience unexpected rate limit errors. The configuration that was incorrectly applied was intended solely for internal staff testing of rate-limiting experiences. Copilot Business and Copilot Enterprise accounts were not affected.\n\nDuring this period, affected users received error messages instructing them to retry after a certain time. Approximately 32% of active Free users, 35% of active Student users, 46% of active Pro users, and 66% of active Pro+ users were affected.\n\nAfter identifying the root cause, we reverted the change and restored the expected rate limits. We are reviewing our deployment and validation processes to help ensure configurations used for internal testing cannot be inadvertently applied to production environments.",
"name": "Incident with Copilot",
"timestamp": "Mar \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e05:00\u003c/var\u003e - \u003cvar data-var='time'\u003e05:00\u003c/var\u003e UTC"
},
{
"code": "kp06czybl7dw",
"impact": "minor",
"message": "This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.",
"name": "Disruption with some GitHub services",
"timestamp": "Mar \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e20:18\u003c/var\u003e - \u003cvar data-var='time'\u003e20:56\u003c/var\u003e UTC"
},
{
"code": "z7gsp4wd05c5",
"impact": "major",
"message": "On March 24, 2026, between 15:57 UTC and 19:51 UTC, the Microsoft Teams Integration and Teams Copilot Integration services were degraded and unable to deliver GitHub event notifications to Microsoft Teams. On average, the error rate was 37.4% and peaked at 90.1% of requests to the service -- approximately 19% of all integration installs failed to receive GitHub-to-Teams notifications in this time period.\u003cbr\u003e\u003cbr\u003eThis was due to an outage at one of our upstream dependencies, which caused HTTP 500 errors and connection resets for our Teams integration.\u003cbr\u003e\u003cbr\u003eWe coordinated with the relevant service teams, and the issue was resolved at 19:51 UTC when the upstream incident was mitigated.\u003cbr\u003e\u003cbr\u003eWe are working to update observability and runbooks to reduce time to mitigation for issues like this in the future.",
"name": "Teams Github Notifications App is down",
"timestamp": "Mar \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e16:59\u003c/var\u003e - \u003cvar data-var='time'\u003e19:51\u003c/var\u003e UTC"
},
{
"code": "7vwbds3snh28",
"impact": "minor",
"message": "On March 22, 2026, between 09:05 UTC and 10:02 UTC, users may have experienced intermittent errors and increased latency when performing Git http read operations. On average, the error rate was 3.84% and peaked at 15.55% of requests to the service. The issue was caused by elevated latency in an internal authentication service within one of our regional clusters. We mitigated the issue by redirecting traffic away from the affected cluster at 09:39 UTC, after which error rates returned to normal. The incident was fully resolved at 10:02 UTC. \u003cbr\u003e\u003cbr\u003eWe are working to scale the authentication service and reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Mar \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e09:08\u003c/var\u003e - \u003cvar data-var='time'\u003e10:02\u003c/var\u003e UTC"
},
{
"code": "9jyjvxkz735j",
"impact": "major",
"message": "On March 19, 2026, between 01:05 UTC and 02:52 UTC, and again on March 20, 2026, between 00:42 UTC and 01:58 UTC, the Copilot Coding Agent service was degraded and users were unable to start new Copilot Agent sessions or view existing ones. During the first incident, the average error rate was ~53% and\u003cbr\u003epeaked at ~93% of requests to the service. During the second incident, the average error rate was ~99%% and peaked at ~100%% of requests with significant retry amplification. Both incidents were caused by the same underlying system authentication issue that prevented the service from connecting to its\u003cbr\u003ebacking datastore.\u003cbr\u003e\u003cbr\u003eWe mitigated each incident by rotating the affected credentials, which restored connectivity and returned error rates to normal. The mitigation time was 01:24. The second occurrence was due to an incomplete remediation of the first.\u003cbr\u003e\u003cbr\u003eWe are implementing automated monitoring for credential lifecycle events and improving operational processes to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Disruption with Copilot Coding Agent Sessions",
"timestamp": "Mar \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e00:58\u003c/var\u003e - \u003cvar data-var='time'\u003e01:58\u003c/var\u003e UTC"
},
{
"code": "p08cncy05m4k",
"impact": "minor",
"message": "On March 19, 2026 between 16:10 UTC and 00:05 UTC (March 20), Git operations (clone, fetch, push) from the US west coast experienced elevated latency and degraded throughput. Users reported clone speeds dropping from typical speeds to under 1 MiB/s in extreme cases. The root cause was network transport link saturation at our Seattle edge site, where a fiber cut affecting our backbone transport resulted in saturation and packet loss. We had a planned scale-up in progress for the site that was accelerated to resolve the backbone capacity pressure. We also brought online additional edge capacity in a cloud region and redirected some users there. Current scale with the upgraded network capacity is sufficient to prevent reoccurrence, as we upgraded from 800Gbps to 3.2Tbps total capacity on this path. We will continue to monitor network health and respond to any further issues.",
"name": "Git operations for users in the west coast are experiencing an increase in latency",
"timestamp": "Mar \u003cvar data-var='date'\u003e19\u003c/var\u003e, \u003cvar data-var='time'\u003e16:25\u003c/var\u003e - Mar \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e00:05\u003c/var\u003e UTC"
},
{
"code": "3gflh59mjhmf",
"impact": "major",
"message": "On March 19, 2026, between 01:05 UTC and 02:52 UTC, and again on March 20, 2026, between 00:42 UTC and 01:58 UTC, the Copilot Coding Agent service was degraded and users were unable to start new Copilot Agent sessions or view existing ones. During the first incident, the average error rate was ~53% and\u003cbr\u003e peaked at ~93% of requests to the service. During the second incident, the average error rate was ~99%% and peaked at ~100%% of requests with significant retry amplification. Both incidents were caused by the same underlying system authentication issue that prevented the service from connecting to its\u003cbr\u003e backing datastore.\u003cbr\u003e \u003cbr\u003e We mitigated each incident by rotating the affected credentials, which restored connectivity and returned error rates to normal. The mitigation time was 01:24. The second occurrence was due to an incomplete remediation of the first.\u003cbr\u003e \u003cbr\u003e We are implementing automated monitoring for credential lifecycle events and improving operational processes to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Issues with Copilot Coding Agent",
"timestamp": "Mar \u003cvar data-var='date'\u003e19\u003c/var\u003e, \u003cvar data-var='time'\u003e13:44\u003c/var\u003e - \u003cvar data-var='time'\u003e14:32\u003c/var\u003e UTC"
},
{
"code": "djmkyscrj9jh",
"impact": "minor",
"message": "On March 19, 2026, between 01:05 UTC and 02:52 UTC, and again on March 20, 2026, between 00:42 UTC and 01:58 UTC, the Copilot Coding Agent service was degraded and users were unable to start new Copilot Agent sessions or view existing ones. During the first incident, the average error rate was ~53% and\u003cbr\u003e peaked at ~93% of requests to the service. During the second incident, the average error rate was ~99%% and peaked at ~100%% of requests with significant retry amplification. Both incidents were caused by the same underlying system authentication issue that prevented the service from connecting to its\u003cbr\u003e backing datastore.\u003cbr\u003e \u003cbr\u003e We mitigated each incident by rotating the affected credentials, which restored connectivity and returned error rates to normal. The mitigation time was 01:24. The second occurrence was due to an incomplete remediation of the first.\u003cbr\u003e \u003cbr\u003e We are implementing automated monitoring for credential lifecycle events and improving operational processes to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Disruption with Copilot Coding Agent sessions",
"timestamp": "Mar \u003cvar data-var='date'\u003e19\u003c/var\u003e, \u003cvar data-var='time'\u003e02:05\u003c/var\u003e - \u003cvar data-var='time'\u003e02:52\u003c/var\u003e UTC"
},
{
"code": "49xnkj77r7vl",
"impact": "minor",
"message": "On March 19, 2026 between 16:10 UTC and 00:05 UTC (March 20), Git operations (clone, fetch, push) from the US west coast experienced elevated latency and degraded throughput. Users reported clone speeds dropping from typical speeds to under 1 MiB/s in extreme cases. The root cause was network transport link saturation at our Seattle edge site, where a fiber cut affecting our backbone transport resulted in saturation and packet loss. We had a planned scale-up in progress for the site that was accelerated to resolve the backbone capacity pressure. We also brought online additional edge capacity in a cloud region and redirected some users there. Current scale with the upgraded network capacity is sufficient to prevent reoccurrence, as we upgraded from 800Gbps to 3.2Tbps total capacity on this path. We will continue to monitor network health and respond to any further issues.\u003cbr\u003e\u003cbr\u003eThis was the same incident declared in https://www.githubstatus.com/incidents/xs6xtcv196g7",
"name": "Disruption with some GitHub services",
"timestamp": "Mar \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e22:36\u003c/var\u003e - Mar \u003cvar data-var='date'\u003e19\u003c/var\u003e, \u003cvar data-var='time'\u003e01:44\u003c/var\u003e UTC"
},
{
"code": "x1g78jx4sgfk",
"impact": "minor",
"message": "On March 18, 2026, between 18:18 UTC and 19:46 UTC all webhook deliveries experienced elevated latency. During this time, average delivery latency increased from a baseline of approximately 5 seconds to a peak of approximately 160 seconds. This was due to resource constraints in the webhook delivery pipeline, which caused queue backlog growth and increased delivery latency. We mitigated the incident by shifting traffic and adding capacity, after which webhook delivery latency returned to normal. We are working to improve capacity management and detection in the webhook delivery pipeline to help prevent similar issues in the future.",
"name": "Webhook delivery is delayed",
"timestamp": "Mar \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e18:51\u003c/var\u003e - \u003cvar data-var='time'\u003e19:46\u003c/var\u003e UTC"
},
{
"code": "2lh36qcrd7l6",
"impact": "minor",
"message": "On 16 March 2026, between 14:16 UTC and 15:18 UTC, Codespaces users encountered a download failure error message when starting newly created or resumed codespaces. At peak, 96% of the created or resumed codespaces were impacted. Active codespaces with a running VSCode environment were not affected. \u003cbr\u003e\u003cbr\u003eThe error was a result of an API deployment issue with our VS Code remote experience dependency and was resolved by rolling back that deployment. We are working with our partners to reduce our incident engagement time, improve early detection before they impact our customers, and ensure safe rollout of similar changes in the future.",
"name": "Errors starting and connecting to Codespaces",
"timestamp": "Mar \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e15:01\u003c/var\u003e - \u003cvar data-var='time'\u003e15:28\u003c/var\u003e UTC"
},
{
"code": "xsxcyn25nfmq",
"impact": "minor",
"message": "On March 13, 2026, between 13:35 UTC and 16:02 UTC, a configuration change to an internal authorization service reduced its processing capacity below what was needed during peak traffic. This caused intermittent timeouts when other GitHub services checked user permissions, resulting in four to five waves of errors over roughly two hours and forty minutes. In total, 0.4% of users were denied access to actions they were authorized to perform. \u003cbr\u003e\u003cbr\u003eThe root cause was a resource right-sizing change deployed to the authorization service the previous day. It reduced CPU allocation below what was required at peak, causing the service's network gateway to throttle under load. Because the change was deployed after peak traffic on March 12, the reduced capacity wasn't surfaced until the next day's peak. \u003cbr\u003e\u003cbr\u003eThe incident was mitigated by manually scaling up the authorization service and reverting the configuration change. \u003cbr\u003e\u003cbr\u003e \u003cbr\u003eTo prevent recurrence, we are adding further resource utilization monitors across our entire stack to detect throttling and improving error handling so transient infrastructure timeouts are distinguished from authorization failures, enabling quicker detection of the root issue.",
"name": "Degraded performance for various services",
"timestamp": "Mar \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e15:12\u003c/var\u003e - \u003cvar data-var='time'\u003e16:15\u003c/var\u003e UTC"
},
{
"code": "0s7ppykbyvqz",
"impact": "minor",
"message": "On March 12, 2026, between 01:00 UTC and 18:53 UTC, users saw failures downloading extensions within created or resumed codespaces. Users would see an error when attempting to use an extension within VS Code. Active codespaces with extensions already downloaded were not impacted.\u003cbr\u003e\u003cbr\u003eThe extensions download failures were the result of a change introduced in our extension dependency and was resolved by updating the configuration of how those changes affect requests from Codespaces. We are enhancing observability and alerting of critical issues within regular codespace operations to better detect and mitigate similar issues in the future.",
"name": "Degraded Codespaces experience",
"timestamp": "Mar \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e13:06\u003c/var\u003e - \u003cvar data-var='time'\u003e18:53\u003c/var\u003e UTC"
},
{
"code": "02z04m335tvv",
"impact": "minor",
"message": "On March 12, 2026 between 02:30 and 06:02 UTC some GitHub Apps were unable to mint server to server tokens, resulting in 401 Unauthorized errors. During the outage window, ~1.3% of requests resulted in 401 errors incorrectly. This manifested in GitHub Actions jobs failing to download tarballs, as well as failing to mint fine-grained tokens. During this period, approximately 5% of Actions jobs were impacted \u003cbr\u003e\u003cbr\u003eThe root cause was a failure with the authentication service’s token cache layer, a newly created secondary cache layer backed by Redis – caused by Kubernetes control plane instability, leading to an inability to read certain tokens which resulted in 401 errors. The mitigation was to fallback reads to the primary cache layer backed by mysql. As permanent mitigations, we have made changes to how we deploy redis to not rely on the Kubernetes control plane and maintain service availability during similar failure modes. We also improved alerting to reduce overall impact time from similar failures. \u003cbr\u003e",
"name": "Actions failures to download (401 Unauthorized)",
"timestamp": "Mar \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e04:46\u003c/var\u003e - \u003cvar data-var='time'\u003e06:02\u003c/var\u003e UTC"
},
{
"code": "lw6j95nyw3py",
"impact": "minor",
"message": "Between 01:36 and 08:11 UTC on Thursday March 12, GitHub.com experienced elevated error rates across Git operations, web requests, and related services. During a planned infrastructure upgrade, a configuration issue caused newly provisioned Kubernetes nodes to run an incompatible version of etcd, which disrupted cluster consensus across several production clusters. This led to intermittent 5XX errors on git push, git clone, and page loads. Deployments were paused for the duration of the incident.\u003cbr\u003e\u003cbr\u003eOnce the incompatible nodes were identified, they were removed and cluster consensus was restored. A validation deploy confirmed all systems were healthy before normal operations resumed.\u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are adding programmatic enforcement of version compatibility during node replacements, implementing monitoring to detect split-brain conditions earlier, and updating our recovery tooling to reduce restoration time.",
"name": "Disruption with some GitHub services",
"timestamp": "Mar \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e01:54\u003c/var\u003e - \u003cvar data-var='time'\u003e02:45\u003c/var\u003e UTC"
},
{
"code": "78s4l7y4tn2v",
"impact": "minor",
"message": "On March 11, 2026, between 13:00 UTC and 15:23 UTC the Copilot Code Review service was degraded and experienced longer than average review times. On average, Copilot Code Review requests took 4 minutes and peaked at just under 8 minutes. This was due to hitting worker capacity limits and CPU throttling. We mitigated the incident by increasing partitions, and we are improving our resource monitoring to identify potential issues sooner.",
"name": "Degraded experience with Copilot Code Review",
"timestamp": "Mar \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e14:25\u003c/var\u003e - \u003cvar data-var='time'\u003e15:53\u003c/var\u003e UTC"
},
{
"code": "zy1gjfc6x2z9",
"impact": "minor",
"message": "On March 11, 2026, between 14:25 UTC and 14:34 UTC, the REST API platform was degraded, resulting in increased error rates and request timeouts. REST API 5xx error rates peaked at ~5% during the incident window with two distinct spikes: the first impacting REST services broadly, and the second driven by sustained timeouts on a subset of endpoints. \u003cbr\u003e\u003cbr\u003eThe incident was caused by a performance degradation in our data layer, which resulted in increased query latency across dependent services. Most services recovered quickly after the initial spike, but resource contention caused sustained 5xx errors due to how certain endpoints responded to the degraded state. \u003cbr\u003e\u003cbr\u003eA fix addressing the behavior that prolonged impact has already been shipped. We are continuing to work to resolve the primary contributing factor of the degradation and to implement safeguards against issues causing cascading impact in the future.",
"name": "Incident with API Requests",
"timestamp": "Mar \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e14:37\u003c/var\u003e - \u003cvar data-var='time'\u003e15:02\u003c/var\u003e UTC"
},
{
"code": "8jx0t5c136x2",
"impact": "none",
"message": "On March 10, 2026, between 23:00 UTC and 23:40 UTC, the Webhooks service was degraded and ~6% of users experienced intermittent errors when accessing webhook delivery history, retrying webhook deliveries, and listing webhooks via the UI and API. Approximately 0.37% of requests resulted in errors, while at peak 0.5% of requests resulted in errors.\n\nThis was due to unhealthy infrastructure. We mitigated the incident by redeploying affected services, after which service health returned to normal.\n\nWe are working to improve detection of unhealthy infrastructure and strengthen service safeguards to reduce time to detect and mitigate similar issues in the future.",
"name": "Incident With Webhooks",
"timestamp": "Mar \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e23:00\u003c/var\u003e - \u003cvar data-var='time'\u003e23:00\u003c/var\u003e UTC"
},
{
"code": "7k92rcz4f1zq",
"impact": "minor",
"message": "On March 9, 2026, between 15:03 and 20:52 UTC, the Webhooks API experienced was degraded, resulted in higher average latency on requests and in certain cases error responses. Approximately 0.6% of total requests exceeded the normal latency threshold of 3s, while 0.4% of requests resulted in 500 errors. At peak, 2.0% experienced latency greater than 3 seconds and 2.8% of requests returned 500 errors.\u003cbr\u003e\u003cbr\u003eThe issue was caused by a noisy actor that led to resource contention on the Webhooks API service. We mitigated the issue initially by increasing CPU resources for the Webhooks API service, and ultimately applied lower rate limiting thresholds to the noisy actor to prevent further impact to other users.\u003cbr\u003e\u003cbr\u003eWe are working to improve monitoring to more quickly ascertain noisy traffic and will continue to improve our rate-limiting mechanisms to help prevent similar issues in the future.",
"name": "Incident with Webhooks",
"timestamp": "Mar \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e15:50\u003c/var\u003e - \u003cvar data-var='time'\u003e17:03\u003c/var\u003e UTC"
},
{
"code": "tp8m3544w2g8",
"impact": "minor",
"message": "On March 9, 2026, between 01:23 UTC and 03:25 UTC, users attempting to create or resume codespaces in the Australia East region experienced elevated failures, peaking at a 100% failure rate for this region. Codespaces in other regions were not affected.\u003cbr\u003e\u003cbr\u003eThe create and resume failures were caused by degraded network connectivity between our control plane services and the VMs hosting the codespaces. This was resolved by redirecting traffic to an alternate site within the region. While we are addressing the core network infrastructure issue, we have also improved our observability of components in this area to improve detection. This will also enable our existing automated failovers to cover this failure mode. These changes will prevent or significantly reduce the time any similar incident causes user impact.",
"name": "Incident with Codespaces",
"timestamp": "Mar \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e03:04\u003c/var\u003e - \u003cvar data-var='time'\u003e03:51\u003c/var\u003e UTC"
},
{
"code": "tqyg889tl8qz",
"impact": "minor",
"message": "On March 6, 2026, between 16:16 UTC and 23:28 UTC the Webhooks service was degraded and some users experienced intermittent errors when accessing webhook delivery histories, retrying webhook deliveries, and listing webhooks via the UI and API. On average, the error rate was 0.57% and peaked at approximately 2.73% of requests to the service. This was due to unhealthy infrastructure affecting a portion of webhook API traffic.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by redeploying affected services, after which service health returned to normal.\u003cbr\u003e\u003cbr\u003eWe are working to improve detection of unhealthy infrastructure and strengthen service safeguards to reduce time to detection and mitigation of issues like this one in the future.",
"name": "Incident with Webhooks",
"timestamp": "Mar \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e16:58\u003c/var\u003e - \u003cvar data-var='time'\u003e23:28\u003c/var\u003e UTC"
},
{
"code": "g9j4tmfqdd09",
"impact": "major",
"message": "On March 5, between 22:39 and 23:55 UTC, Actions was degraded due to a repeat of an incident a few hours prior. In this case, a Redis cluster topology change made as a follow-up to the earlier incident caused a repeat of the earlier degradation of Actions jobs. Details of both incidents and the follow-ups are shared at https://www.githubstatus.com/incidents/g5gnt5l5hf56.",
"name": "Actions is experiencing degraded availability",
"timestamp": "Mar \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e22:53\u003c/var\u003e - \u003cvar data-var='time'\u003e23:55\u003c/var\u003e UTC"
},
{
"code": "g5gnt5l5hf56",
"impact": "major",
"message": "On Mar 5, 2026, between 16:24 UTC and 19:30 UTC, Actions was degraded. During this time, 95% of workflow runs failed to start within 5 minutes with an average delay of 30 minutes and 10% workflow runs failed with an infrastructure error. This was due to Redis infrastructure updates that were being rolled out to production to improve our resiliency. These changes introduced a set of incorrect configuration change into our Redis load balancer causing internal traffic to be routed to an incorrect host leading to two incidents. \u003cbr\u003e\u003cbr\u003eWe mitigated this incident by correcting the misconfigured load balancer. Actions jobs were running successfully starting at 17:24 UTC. The remaining time until we closed the incident was burning through the queue of jobs. \u003cbr\u003e\u003cbr\u003eWe immediately rolled back the updates that were a contributing factor and have frozen all changes in this area until we have completed follow-up work from this. We are working to improve our automation to ensure incorrect configuration changes are not able to propagate through our infrastructure. We are also working on improved alerting to catch misconfigured load balancers before it becomes an incident. Additionally, we are updating the Redis client configuration in Actions to improve resiliency to brief cache interruptions.",
"name": "Multiple services are affected, service degradation",
"timestamp": "Mar \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e16:35\u003c/var\u003e - \u003cvar data-var='time'\u003e19:30\u003c/var\u003e UTC"
},
{
"code": "nlmpksjb2ll7",
"impact": "minor",
"message": "On March 5, 2026, between 12:53 UTC and 13:35 UTC, the Copilot mission control service was degraded. This resulted in empty responses returned for users' agent session lists across GitHub web surfaces. Impacted users were unable to see their lists of current and previous agent sessions in GitHub web surfaces. This was caused by an incorrect database query that falsely excluded records that have an absent field.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by rolling back the database query change. There were no data alterations nor deletions during the incident.\u003cbr\u003e\u003cbr\u003eTo prevent similar issues in the future, we're improving our monitoring depth to more easily detect degradation before changes are fully rolled out.",
"name": "Disruption with some GitHub services",
"timestamp": "Mar \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e01:13\u003c/var\u003e - \u003cvar data-var='time'\u003e01:30\u003c/var\u003e UTC"
},
{
"code": "t2p9yhsww5lc",
"impact": "minor",
"message": "On March 5th, 2026, between approximately 00:26 and 00:44 UTC, the Copilot service experienced a degradation of the GPT 3.5 Codex model due to an issue with our upstream provider. Users encountered elevated error rates when using GPT 3.5 Codex, impacting approximately 30% of requests. No other models were impacted.\u003cbr\u003e\u003cbr\u003eThe issue was resolved by a mitigation put in place by our provider.",
"name": "Some OpenAI models degraded in Copilot",
"timestamp": "Mar \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e00:47\u003c/var\u003e - \u003cvar data-var='time'\u003e01:13\u003c/var\u003e UTC"
},
{
"code": "nglcmhr2kfb0",
"impact": "minor",
"message": "On March 3, 2026, between 19:44 UTC and 21:05 UTC, some GitHub Copilot users reported that the Claude Opus 4.6 Fast model was no longer available in their IDE model selection. After investigation, we confirmed that this was caused by enterprise administrators adjusting their organization's model policies, which correctly removed the model for users in those organizations. No users outside the affected organizations lost access.\u003cbr\u003e\u003cbr\u003eWe confirmed that the Copilot settings were functioning as designed, and all expected users retained access to the model. The incident was resolved once we verified that the change was intentional and no platform regression had occurred.",
"name": "Claude Opus 4.6 Fast not appearing for some Copilot users",
"timestamp": "Mar \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e20:31\u003c/var\u003e - \u003cvar data-var='time'\u003e21:11\u003c/var\u003e UTC"
},
{
"code": "n07yy1bk6kc4",
"impact": "major",
"message": "On March 3, 2026, between 18:46 UTC and 20:09 UTC, GitHub experienced a period of degraded availability impacting GitHub.com, the GitHub API, GitHub Actions, Git operations, GitHub Copilot, and other dependent services. At the peak of the incident, GitHub.com request failures reached approximately 40%. During the same period, approximately 43% of GitHub API requests failed. Git operations over HTTP had an error rate of approximately 6%, while SSH was not impacted. GitHub Copilot requests had an error rate of approximately 21%. GitHub Actions experienced less than 1% impact. \u003cbr\u003e\u003cbr\u003eThis incident shared the same underlying cause as an incident in early February where we saw a large volume of writes to the user settings caching mechanism. While deploying a change to reduce the burden of these writes, a bug caused every user’s cache to expire, get recalculated, and get rewritten. The increased load caused replication delays that cascaded down to all affected services. We mitigated this issue by immediately rolling back the faulty deployment. \u003cbr\u003e\u003cbr\u003eWe understand these incidents disrupted the workflows of developers. While we have made substantial, long-term investments in how GitHub is built and operated to improve resilience, we acknowledge we have more work to do. Getting there requires deep architectural work that is already underway, as well as urgent, targeted improvements. We are taking the following immediate steps: \u003cbr\u003e\u003cbr\u003e- We have added a killswitch and improved monitoring to the caching mechanism to ensure we are notified before there is user impact and can respond swiftly. \u003cbr\u003e- We are moving the cache mechanism to a dedicated host, ensuring that any future issues will solely affect services that rely on it.",
"name": "Incident with all GitHub services",
"timestamp": "Mar \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e18:59\u003c/var\u003e - \u003cvar data-var='time'\u003e20:09\u003c/var\u003e UTC"
},
{
"code": "0yjbtp1d2cpc",
"impact": "minor",
"message": "Between March 2, 21:42 UTC and March 3, 05:54 UTC project board updates, including adding new issues, PRs, and draft items to boards, were delayed from 30 minutes to over 2 hours, as a large backlog of messages accumulated in the Projects data denormalization pipeline.\u003cbr\u003e\u003cbr\u003eThe incident was caused by an anomalously large event that required longer processing time than expected. Processing this message exceeded the Kafka consumer heartbeat timeout, triggering repeated consumer group rebalances. As a result, the consumer group was unable to make forward progress, creating head-of-line blocking that delayed processing of subsequent project board updates.\u003cbr\u003e\u003cbr\u003eWe mitigated the issue by deploying a targeted fix that safely bypassed the offending message and allowed normal message consumption to resume. Consumer group stability recovered at 04:10 UTC, after which the backlog began draining. All queued messages were fully processed by 05:53 UTC, returning project board updates to normal processing latency.\u003cbr\u003e\u003cbr\u003eWe have identified several follow-up improvements to reduce the likelihood and impact of similar incidents in the future, including improved monitoring and alerting, as well as introducing limits for unusually large project events.",
"name": "Delayed visibility of newly added issues on project boards",
"timestamp": "Mar \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e23:10\u003c/var\u003e - Mar \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e05:54\u003c/var\u003e UTC"
},
{
"code": "vhpmx9vc93m0",
"impact": "minor",
"message": "On March 2nd, 2026, between 7:10 UTC and 22:04 UTC the pull requests service was degraded. Users navigating between tabs on the pull requests dashboard were met with 404 errors or blank pages.\u003cbr\u003e\u003cbr\u003eThis was due to a configuration change deployed on February 27th at 11:03 PM UTC. We mitigated the incident by reverting the change.\u003cbr\u003e\u003cbr\u003eWe’re working to improve monitoring for the page to automatically detect and alert us to routing failures.",
"name": "Incident with Pull Requests /pulls",
"timestamp": "Mar \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e19:11\u003c/var\u003e - \u003cvar data-var='time'\u003e22:04\u003c/var\u003e UTC"
}
],
"name": "March",
"year": 2026
},
{
"incidents": [
{
"code": "kw5jh58jh24k",
"impact": "minor",
"message": "On February 27, 2026, between 22:53 UTC and 23:46 UTC, the Copilot coding agent service experienced elevated errors and degraded functionality for agent sessions. Approximately 87% of attempts to start or interact with agent sessions encountered errors during this period.\u003cbr\u003e\u003cbr\u003eThis was due to an expired authentication credential for an internal service component, which prevented Copilot agent session operations from completing successfully.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by rotating the expired credential and deploying the updated configuration to production. Services began recovering within minutes of the fix being deployed.\u003cbr\u003e\u003cbr\u003eWe are working to improve automated credential rotation coverage across all Copilot service components, add proactive alerting for credentials approaching expiration, and validate configuration consistency to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Incident with Copilot agent sessions",
"timestamp": "Feb \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e23:18\u003c/var\u003e - \u003cvar data-var='time'\u003e23:49\u003c/var\u003e UTC"
},
{
"code": "kv1lzpgzr9yp",
"impact": "minor",
"message": "Starting February 26, 2026 at 22:10 UTC through February 27, 05:50 UTC, the repository browsing UI was degraded and users were unable to load pages for files and directories with non-ASCII characters (including Japanese, Chinese, and other non-Latin scripts). On average, the error rate was 0.014% and peaked at 0.06% of requests to the service. Affected users saw 404 errors when navigating to repository directories and files with non-ASCII names. This was due to a code change that altered how file and directory names were processed, which caused incorrectly formatted data to be stored in an application cache.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by deploying a fix that invalidated the affected cache entries and progressively rolling it out across all production environments.\u003cbr\u003e\u003cbr\u003eWe are working to improve our pre-production testing to cover non-ASCII character handling, establish better cache invalidation mechanisms, and enhance our monitoring to detect this type of failure mode earlier, to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Code view fails to load when content contains some non-ASCII characters",
"timestamp": "Feb \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e03:08\u003c/var\u003e - \u003cvar data-var='time'\u003e06:04\u003c/var\u003e UTC"
},
{
"code": "vd3xqfq36rgm",
"impact": "minor",
"message": "Between February 26, 2026 UTC and February 27, 2026 UTC, customers hitting the webhooks delivery API may have experienced higher latency or failed requests. During the impact window, 0.82% of requests took longer than 3s and 0.004% resulted in a 500 error response.\u003cbr\u003e\u003cbr\u003eOur monitors caught the impact on the individual backing data source, and we were able to attribute the degradation to a noisy neighbor effect due requests to a specific webhook generating excessive load on the API. The incident was mitigated once traffic from the specific hook decreased.\u003cbr\u003e\u003cbr\u003eWe have since added a rate limiter for this webhooks API to prevent similar spikes in usage impacting others and will further refine the rate limits for other webhook API routes to help prevent similar occurrences in the future.",
"name": "High latency on webhook API requests",
"timestamp": "Feb \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e00:01\u003c/var\u003e - \u003cvar data-var='time'\u003e00:04\u003c/var\u003e UTC"
},
{
"code": "f6d6nm7gn304",
"impact": "minor",
"message": "On February 26, 2026, between 09:27 UTC and 10:36 UTC, the GitHub Copilot service was degraded and users experienced errors when using Copilot features including Copilot Chat, Copilot Coding Agent and Copilot Code Review. During this time, 5-15% of affected requests to the service returned errors.\u003cbr\u003e\u003cbr\u003eThe incident was resolved by infrastructure rebalancing.\u003cbr\u003e\u003cbr\u003eWe are improving observability to detect capacity imbalances earlier and enhancing our infrastructure to better handle traffic spikes.",
"name": "Incident with Copilot",
"timestamp": "Feb \u003cvar data-var='date'\u003e26\u003c/var\u003e, \u003cvar data-var='time'\u003e10:22\u003c/var\u003e - \u003cvar data-var='time'\u003e11:06\u003c/var\u003e UTC"
},
{
"code": "l1bt9psnsqpg",
"impact": "minor",
"message": "On February 25, 2026, between 15:05 UTC and 16:34 UTC, the Copilot coding agent service was degraded, resulting in errors for 5% of all requests and impacting users starting or interacting with agent sessions. \u003cbr\u003e\u003cbr\u003eThis was due to an internal service dependency running out of allocated resources (memory and CPU). We mitigated the incident by adjusting the resource allocation for the affected service, which restored normal operations for the coding agent service.\u003cbr\u003e\u003cbr\u003eWe are working to implement proactive monitoring for resource exhaustion across our services, review and update resource allocations, and improve our alerting capabilities to reduce our time to detection and mitigation of similar issues in the future.",
"name": "Incident with Copilot Agent Sessions impacting CCA/CCR",
"timestamp": "Feb \u003cvar data-var='date'\u003e25\u003c/var\u003e, \u003cvar data-var='time'\u003e16:38\u003c/var\u003e - \u003cvar data-var='time'\u003e16:44\u003c/var\u003e UTC"
},
{
"code": "jn8kcmg5ydch",
"impact": "minor",
"message": "Between 2026-02-23 19:10 and 2026-02-24 00:46 UTC, all lexical code search queries in GitHub.com and the code search API were significantly slowed, and during this incident, between 5 and 10% of search queries timed out. This was caused by a single customer who had created a network of hundreds of orchestrated accounts which searched with a uniquely expensive search query. This search query concentrated load on a single hot shard within the search index, slowing down all queries. After we identified the source of the load and stopped the traffic, latency returned to normal.\u003cbr\u003e\u003cbr\u003eTo avoid this situation occurring again in the future, we are making a number of improvements to our systems, including: improved rate limiting that accounts for highly skewed load on hot shards, improved system resilience for when a small number of shards time out, improved tooling to recognize abusive actors, and capabilities that will allow us to shed load on a single shard in emergencies.",
"name": "Code search experiencing degraded performance",
"timestamp": "Feb \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e19:59\u003c/var\u003e - Feb \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e00:46\u003c/var\u003e UTC"
},
{
"code": "47n199g2d88x",
"impact": "minor",
"message": "On February 23, 2026, between 21:01 UTC and 21:30 UTC the Search service experienced degraded performance, resulting in an average of 3.5% of search requests for Issues and Pull Requests being rejected. During this period, updates to Issues and Pull Requests may not have been immediately reflected in search results. \u003cbr\u003e\u003cbr\u003eDuring a routine migration, we observed a spike in internal traffic due to a configuration change in our search index. We were alerted to the increase in traffic as well as the increase in error rates and rolled back to the previous stable index. \u003cbr\u003e\u003cbr\u003eWe are working to enable more controlled traffic shifting when promoting a new index to allow us to detect potential limitations earlier and ensure these operations succeed in a more controlled manner.",
"name": "Incident with Issues and Pull Requests Search",
"timestamp": "Feb \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e21:16\u003c/var\u003e - \u003cvar data-var='time'\u003e21:30\u003c/var\u003e UTC"
},
{
"code": "59rjsk76t3rf",
"impact": "minor",
"message": "On February 23, 2026, between 15:00 UTC and 17:00 UTC, GitHub Actions experienced degraded performance. During the time, 1.8% of Actions workflow runs experienced delayed starts with an average delay of 15 minutes. The issue was caused by a connection rebalancing event in our internal load balancing layer, which temporarily created uneven traffic distribution across sites and led to request throttling. \u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we are tuning connection rebalancing behavior to spread client reconnections more gradually during load balancer reloads. We are also evaluating improvements to site-level traffic affinity to eliminate the uneven distribution at its source. We have overprovisioned critical paths to prevent any impact if a similar event occurs before those workstreams finish. Finally, we are enhancing our monitoring to detect capacity imbalances proactively.",
"name": "Incident with Actions",
"timestamp": "Feb \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e16:17\u003c/var\u003e - \u003cvar data-var='time'\u003e17:03\u003c/var\u003e UTC"
},
{
"code": "h8z3w9ycj0pm",
"impact": "minor",
"message": "On February 23, 2026, between 14:45 UTC and 16:19 UTC, the Copilot service was degraded for Claude Haiku 4.5 model. On average, 6% of the requests to this model failed due to an issue with an upstream provider. During this period, automated model degradation notifications directed affected users to alternative models. No other models were impacted. The upstream provider identified and resolved the issue on their end. \u003cbr\u003eWe are working to improve automatic model failover mechanisms to reduce our time to mitigation of issues like this one in the future.",
"name": "Incident with Copilot",
"timestamp": "Feb \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e14:56\u003c/var\u003e - \u003cvar data-var='time'\u003e16:19\u003c/var\u003e UTC"
},
{
"code": "zhfdlzldgh51",
"impact": "minor",
"message": "On February 20, 2026, between 17:45 UTC and 20:41 UTC, 4.2% of workflows running on GitHub Larger Hosted Runners were delayed by an average of 18 minutes. Standard, Mac, and Self-Hosted Runners were not impacted. \u003cbr\u003e\u003cbr\u003eThe delays were caused by communication failures between backend services for one deployment of larger runners. Those failures prevented expected automated scaling and provisioning of larger hosted runner capacity within that deployment. This was mitigated when the affected infrastructure was recycled, larger runner pools in the affected deployment successfully scaled up, and queued jobs processed. \u003cbr\u003e\u003cbr\u003eWe are working to improve the time to detect and diagnose this class of failures and improve the performance of recovery mechanisms for this degraded network state. In addition, we have architectural changes underway that will enable other deployments to pick up work in similar situations, so there is no customer impact due to deployment-specific infrastructure issues like this.",
"name": "Extended job start delays for larger hosted runners",
"timestamp": "Feb \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e20:00\u003c/var\u003e - \u003cvar data-var='time'\u003e20:41\u003c/var\u003e UTC"
},
{
"code": "8qvx6hqk3nz0",
"impact": "minor",
"message": "On February 20, 2026, between 07:30 UTC and 11:21 UTC, the Copilot service experienced a degradation of the GPT 5.1 Codex model. During this time period, users encountered a 4.5% error rate when using this model. No other models were impacted.\u003cbr\u003eThe issue was resolved by a mitigation put in place by the external model provider. GitHub is working with the external model provider to further improve the resiliency of the service to prevent similar incidents in the future.",
"name": "Incident with Copilot GPT-5.1-Codex",
"timestamp": "Feb \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e10:02\u003c/var\u003e - \u003cvar data-var='time'\u003e11:41\u003c/var\u003e UTC"
},
{
"code": "p1ymhg64hdfq",
"impact": "minor",
"message": "On February 17, 2026, between 17:07 UTC and 19:06 UTC, some customers experienced intermittent authentication failures affecting GitHub Actions, parts of Git operations, and other authentication-dependent requests. On average, the Actions error rate was approximately 0.6% of affected API requests. Git operations ssh read error rate was approximately 0.29%, while ssh write and http operations were not impacted. During the incident, a subset of requests failed due to token verification lookups intermittently failing, leading to 401 errors and degraded reliability for impacted workflows.\u003cbr\u003e\u003cbr\u003eThe issue was caused by elevated replication lag in the token verification database cluster. In the days leading up to the incident, the token store’s write volume grew enough to exceed the cluster’s available capacity. Under peak load, older replica hosts were unable to keep up, replica lag increased, and some token lookups became inconsistent, resulting in intermittent authentication failures.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by adjusting the database replica topology to route reads away from lagging replicas and by adding/bringing additional replica capacity online. Service health improved progressively after the change, with GitHub Actions recovering by ~19:00 UTC and the incident resolved at 19:06 UTC.\u003cbr\u003e\u003cbr\u003eWe are working to prevent recurrence by improving the resilience and scalability of our underlying token verification data stores to better handle continued growth.\u003cbr\u003e\u003cbr\u003eThis was the same incident declared in https://www.githubstatus.com/incidents/xs6xtcv196g7",
"name": "Degraded performance in merge queue",
"timestamp": "Feb \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e18:25\u003c/var\u003e - \u003cvar data-var='time'\u003e19:20\u003c/var\u003e UTC"
},
{
"code": "xs6xtcv196g7",
"impact": "minor",
"message": "On February 17, 2026, between 17:07 UTC and 19:06 UTC, some customers experienced intermittent authentication failures affecting GitHub Actions, parts of Git operations, and other authentication-dependent requests. On average, the Actions error rate was approximately 0.6% of affected API requests. Git operations ssh read error rate was approximately 0.29%, while ssh write and http operations were not impacted. During the incident, a subset of requests failed due to token verification lookups intermittently failing, leading to 401 errors and degraded reliability for impacted workflows.\u003cbr\u003e\u003cbr\u003eThe issue was caused by elevated replication lag in the token verification database cluster. In the days leading up to the incident, the token store’s write volume grew enough to exceed the cluster’s available capacity. Under peak load, older replica hosts were unable to keep up, replica lag increased, and some token lookups became inconsistent, resulting in intermittent authentication failures.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by adjusting the database replica topology to route reads away from lagging replicas and by adding/bringing additional replica capacity online. Service health improved progressively after the change, with GitHub Actions recovering by ~19:00 UTC and the incident resolved at 19:06 UTC.\u003cbr\u003e\u003cbr\u003eWe are working to prevent recurrence by improving the resilience and scalability of our underlying token verification data stores to better handle continued growth.",
"name": "Intermittent authentication failures on GitHub",
"timestamp": "Feb \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e17:46\u003c/var\u003e - \u003cvar data-var='time'\u003e19:06\u003c/var\u003e UTC"
},
{
"code": "lvf95rlzkxrr",
"impact": "minor",
"message": "On February 13, 2026, between 21:46 UTC and 22:58 UTC (72 minutes), the GitHub file upload service was degraded and users uploading from a web browser on GitHub.com were unable to upload files to repositories, create release assets, or upload manifest files. During the incident, successful upload completions dropped by ~85% from baseline levels. This was due to a code change that inadvertently modified browser request behavior and violated CORS (Cross-Origin Resource Sharing) policy requirements, causing upload requests to be blocked before reaching the upload service.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by reverting the code change that introduced the issue.\u003cbr\u003e\u003cbr\u003eWe are working to improve automated testing for browser-side request changes and to add monitoring/automated safeguards for upload flows to reduce our time to detection and mitigation of similar issues in the future.",
"name": "Disruption with some GitHub services regarding file upload",
"timestamp": "Feb \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e22:30\u003c/var\u003e - \u003cvar data-var='time'\u003e22:58\u003c/var\u003e UTC"
},
{
"code": "y7kzv1ps2qrm",
"impact": "minor",
"message": "Between February 11th 21:30 UTC and February 12th 15:40 UTC, users in Western Europe experienced degraded quality for all Next Edit Suggestions requests. Additionally, on February 12th, between 18:40 UTC and 20:30 UTC, users in Australia and South America experienced degraded quality and increased latency of up to 500ms for all Next Edit Suggestions requests. The root cause was a newly introduced regression in an upstream service dependency.\u003cbr\u003e \u003cbr\u003eThe incident was mitigated by failing over Next Edit Suggestions traffic to unaffected regions, which caused the increased latency. Once the regression was identified and rolled back, we restored the impacted capacity. We have improved our quality analysis tooling and are working on more robust quality impact alerting to accelerate detection of these issues in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Feb \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e18:36\u003c/var\u003e - \u003cvar data-var='time'\u003e20:34\u003c/var\u003e UTC"
},
{
"code": "1x0f4x1wwqmb",
"impact": "minor",
"message": "Between February 11th 21:30 UTC and February 12th 15:40 UTC, users in Western Europe experienced degraded quality for all Next Edit Suggestions requests. Additionally, on February 12th, between 18:40 UTC and 20:30 UTC, users in Australia and South America experienced degraded quality and increased latency of up to 500ms for all Next Edit Suggestions requests. The root cause was a newly introduced regression in an upstream service dependency.\u003cbr\u003e\u003cbr\u003eThe incident was mitigated by failing over Next Edit Suggestions traffic to unaffected regions, which caused the increased latency. Once the regression was identified and rolled back, we restored the impacted capacity. We have improved our quality analysis tooling and are working on more robust quality impact alerting to accelerate detection of these issues in the future.",
"name": "Intermittent disruption with Copilot completions and inline suggestions",
"timestamp": "Feb \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e14:06\u003c/var\u003e - \u003cvar data-var='time'\u003e16:50\u003c/var\u003e UTC"
},
{
"code": "5txsjlt9299p",
"impact": "major",
"message": "From Feb 12, 2026 09:16:00 UTC to Feb 12, 2026 11:01 UTC, users attempting to download repository archives (tar.gz/zip) that include Git LFS objects received errors. Standard repository archives without LFS objects were not affected. On average, the archive download error rate was 0.0042% and peaked at 0.0339% of requests to the service. This was caused by deploying a corrupt configuration bundle, resulting in missing data used for network interface connections by the service.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by applying the correct configuration to each site. We have added checks for corruption in this deployment, and will add auto-rollback detection for this service to prevent issues like this in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Feb \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e10:38\u003c/var\u003e - \u003cvar data-var='time'\u003e11:12\u003c/var\u003e UTC"
},
{
"code": "7m8kxmjq0y7m",
"impact": "major",
"message": "On February 12, 2026, between 00:51 UTC and 09:35 UTC, users attempting to create or resume Codespaces experienced elevated failure rates across Europe, Asia and Australia, peaking at a 90% failure rate.\u003cbr\u003e\u003cbr\u003eThe disconnects were triggered by a bad configuration rollout in a core networking dependency, which led to internal resource provisioning failures. We are working to improve our alerting thresholds to catch issues before they impact customers and strengthening rollout safeguards to prevent similar incidents.",
"name": "Incident with Codespaces",
"timestamp": "Feb \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e07:53\u003c/var\u003e - \u003cvar data-var='time'\u003e09:56\u003c/var\u003e UTC"
},
{
"code": "7ylfn6fnwgrq",
"impact": "minor",
"message": "On February 11 between 16:37 UTC and 00:59 UTC the following day, 4.7% of workflows running on GitHub Larger Hosted Runners were delayed by an average of 37 minutes. Standard Hosted and self-hosted runners were not impacted. \u003cbr\u003e\u003cbr\u003eThis incident was caused by capacity degradation in Central US for Larger Hosted Runners. Workloads not pinned to that region were picked up by other regions, but were delayed as those regions became saturated. Workloads configured with private networking in that region were delayed until compute capacity in that region recovered. The issue was mitigated by rebalancing capacity across internal and external workloads and general increases in capacity in affected regions to speed recovery. \u003cbr\u003e\u003cbr\u003eIn addition to working with our compute partners on the core capacity degradation, we are working to ensure other regions are better able to absorb load with less delay to customer workloads. For pinned workflows using private networking, we are shipping support soon for customers to failover if private networking is configured in a paired region.",
"name": "Disruption with some GitHub services",
"timestamp": "Feb \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e18:58\u003c/var\u003e - Feb \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e00:59\u003c/var\u003e UTC"
},
{
"code": "frlwqbqgz113",
"impact": "minor",
"message": "On February 11, 2026, between 13:51 UTC and 17:03 UTC, the GraphQL API experienced degraded performance due to elevated resource utilization. This resulted in incoming client requests waiting longer than normal, timing out in certain cases. During the impact window, approximately 0.65% of GraphQL requests experienced these issues, peaking at 1.06%. \u003cbr\u003e\u003cbr\u003eThe increased load was due to an increase in query patterns that drove higher than expected resource utilization of the GraphQL API. We mitigated the incident by scaling out resource capacity and limiting the capacity available to these query patterns. \u003cbr\u003e\u003cbr\u003eWe're improving our telemetry to identify slow usage growth and changes in GraphQL workloads. We’ve also added capacity safeguards to prevent similar incidents in the future.",
"name": "Incident with API Requests",
"timestamp": "Feb \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e15:26\u003c/var\u003e - \u003cvar data-var='time'\u003e17:15\u003c/var\u003e UTC"
},
{
"code": "531k65vmv284",
"impact": "minor",
"message": "On February 11, 2025, between 14:30 UTC and 15:30 UTC, the Copilot service experienced degraded availability for requests to Claude Haiku 4.5. During this time, on average 10% of requests failed with 23% of sessions impacted. The issue was caused by an upstream problem from multiple external model providers that affected our ability to serve requests. \u003cbr\u003e\u003cbr\u003eThe incident was mitigated once one of the providers resolved the issue and we rerouted capacity fully to that provider. We have improved our telemetry to improve incident observability and implemented an automated retry mechanism for requests to this model to mitigate similar future upstream incidents.",
"name": "Incident with Copilot",
"timestamp": "Feb \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e15:26\u003c/var\u003e - \u003cvar data-var='time'\u003e15:46\u003c/var\u003e UTC"
},
{
"code": "wkgqj4546z1c",
"impact": "minor",
"message": "On February 10th, 2026, between 14:35 UTC and 15:58 UTC web experiences on GitHub.com were degraded including Pull Requests and Authentication, resulting in intermittent 5xx errors and timeouts. The error rate on web traffic peaked at approximately 2%. This was due to increased load on a critical database, which caused significant memory pressure resulting in intermittent errors. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by applying a configuration change to the database to increase available memory on the host. \u003cbr\u003e\u003cbr\u003eWe are working to identify changes in load patterns and are reviewing the configuration of our databases to ensure there is sufficient capacity to meet growth. Additionally, we are improving monitoring and self-healing functionalities for database memory issues to reduce our time to detect and mitigation.",
"name": "Disruption with some GitHub services",
"timestamp": "Feb \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e15:07\u003c/var\u003e - \u003cvar data-var='time'\u003e15:58\u003c/var\u003e UTC"
},
{
"code": "t5qmhtg29933",
"impact": "minor",
"message": "GitHub experienced degraded Copilot policy propagation from enterprise to organizations between February 3 at 21:00 UTC through February 10 at 16:00 UTC. During this period, policy changes could take up to 24 hours to apply. We mitigated the issue on February 10 at 16:00 UTC after rolling back a regression that caused the delays. The propagation queue fully caught up on the delayed items by February 11 at 10:35 UTC, and policy changes now propagate normally.\u003cbr\u003e\u003cbr\u003eDuring this incident, whenever an enterprise updated a Copilot policy (including model policies), there were significant delays before those policy changes reached their child organizations and assigned users. The delay was caused by a large backlog in the background job queue responsible for propagating Copilot policy updates.\u003cbr\u003e\u003cbr\u003eOur investigation determined the incident was caused by a code change shipped on February 3 that increased the number of background jobs enqueued per policy update, in order to accommodate upcoming feature work. When new Copilot models launched on February 5th and 7th, triggering policy updates across many enterprises, the higher job volume overwhelmed the shared background worker queue, resulting in prolonged propagation delays. No policy updates were lost; they were queued and processed once the backlog cleared.\u003cbr\u003e\u003cbr\u003eWe understand these delays disrupted policy management for customers using Copilot at scale and have taken the following immediate steps:\u003cbr\u003e\u003cbr\u003e1. Restored the optimized propagation path and put tests in place to avoid a regression.\u003cbr\u003e2. Ensured upcoming features are compatible with this design. \u003cbr\u003e3. Added alerting on queue depth to detect propagation backlogs immediately.\u003cbr\u003e\u003cbr\u003eGitHub is critical infrastructure for your work, your teams, and your businesses. We are focused on these mitigations and continued improvements so Copilot policy changes propagate reliably and quickly.\u003cbr\u003e",
"name": "Copilot Policy Propagation Delays",
"timestamp": "Feb \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e16:29\u003c/var\u003e - Feb \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e09:57\u003c/var\u003e UTC"
},
{
"code": "lcw3tg2f6zsd",
"impact": "major",
"message": "On February 9, 2026, GitHub experienced two related periods of degraded availability affecting GitHub.com, the GitHub API, GitHub Actions, Git operations, GitHub Copilot, and other services. The first period occurred between 16:12 UTC and 17:39 UTC, and the second between 18:53 UTC and 20:09 UTC. In total, users experienced approximately 2 hours and 43 minutes of degraded service across the two incidents.\n\nDuring both incidents, users encountered errors loading pages on GitHub.com, failures when pushing or pulling code over HTTPS, failures starting or completing GitHub Actions workflow runs, and errors using GitHub Copilot. Additional services including GitHub Issues, pull requests, webhooks, Dependabot, GitHub Pages, and GitHub Codespaces experienced intermittent errors. SSH-based Git operations were not affected during either incident.\n\nOur investigation determined that both incidents shared the same underlying cause: a configuration change to a user settings caching mechanism caused a large volume of cache rewrites to occur simultaneously. During the first incident, asynchronous rewrites overwhelmed a shared infrastructure component responsible for coordinating background work, triggering cascading failures. Increased load caused the service responsible for proxying Git operations over HTTPS to exhaust available connections, preventing it from accepting new requests. We mitigated this incident by disabling async cache rewrites and restarting the affected Git proxy service across multiple datacenters.\n\nAn additional source of updates to the same cache circumvented our initial mitigations and caused the second incident. This generated a high volume of synchronous writes, causing replication delays that cascaded in a similar pattern and again exhausted the Git proxy’s connection capacity, degrading availability across multiple services. We mitigated by disabling the source of the cache rewrites and again restarting Git proxy.\n\nWe know these incidents disrupted the workflows of millions of developers. While we have made substantial, long-term investments in how GitHub is built and operated to improve resilience, GitHub's availability is not yet meeting our expectations. Getting there requires deep architectural work that is already underway, as well as urgent, targeted improvements. We are taking the following immediate steps:\n\n1. We have already optimized the caching mechanism to avoid write amplification and added self-throttling during bulk updates.\n2. We are adding safeguards to ensure the caching mechanism responds more quickly to rollbacks and strengthening how changes to these caching systems are planned, validated, and rolled out with additional checks.\n3. We are fixing the underlying cause of connection exhaustion in our Git HTTPS proxy layer so the proxy can recover from this failure mode automatically without requiring manual restarts.\n\nGitHub is critical infrastructure for your work, your teams, and your businesses. We're focusing on these mitigations and long-term infrastructure work so GitHub is available, at scale, when and where you need it.",
"name": "Incident with Issues, Actions and Git Operations",
"timestamp": "Feb \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e19:01\u003c/var\u003e - \u003cvar data-var='time'\u003e20:09\u003c/var\u003e UTC"
},
{
"code": "54hndjxft5bx",
"impact": "minor",
"message": "On February 9th notifications service started showing degradation around 13:50 UTC, resulting in an increase in notification delivery delays. Our team started investigating. \u003cbr\u003e\u003cbr\u003eAround 14:30 UTC the service started to recover as the team continued investigating the incident. Around 15:20 UTC degradation resurfaced, with increasing delays in notification deliveries and small error rate (below 1%) on UI and API endpoints related to notifications. \u003cbr\u003e\u003cbr\u003eAt 16:30 UTC, we mitigated the incident by reducing contention through throttling workloads and performing a database failover. The median delay for notification deliveries was 80 minutes at this point and queues started emptying. Around 19:30 UTC the backlog of notifications was processed, bringing the service back to normal and declaring the incident closed.\u003cbr\u003e\u003cbr\u003eThe incident was caused by the notifications database showing degradation under intense load. Most notifications-related asynchronous workloads, including notifications deliveries, were stopped to try to reduce the pressure on the database. To ensure system stability, a database failover was executed. Following the failover, we applied a configuration change to improve the performance. The service started recovering after these changes.\u003cbr\u003e\u003cbr\u003eWe are reviewing the configuration of our databases to understand the performance drop and prevent similar issues from happening in the future. We are also investing in monitoring to detect and mitigate this class of incidents faster.",
"name": "Notifications are delayed",
"timestamp": "Feb \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e15:54\u003c/var\u003e - \u003cvar data-var='time'\u003e19:29\u003c/var\u003e UTC"
},
{
"code": "smf24rvl67v9",
"impact": "major",
"message": "On February 9, 2026, GitHub experienced two related periods of degraded availability affecting GitHub.com, the GitHub API, GitHub Actions, Git operations, GitHub Copilot, and other services. The first period occurred between 16:12 UTC and 17:39 UTC, and the second between 18:53 UTC and 20:09 UTC. In total, users experienced approximately 2 hours and 43 minutes of degraded service across the two incidents.\n\nDuring both incidents, users encountered errors loading pages on GitHub.com, failures when pushing or pulling code over HTTPS, failures starting or completing GitHub Actions workflow runs, and errors using GitHub Copilot. Additional services including GitHub Issues, pull requests, webhooks, Dependabot, GitHub Pages, and GitHub Codespaces experienced intermittent errors. SSH-based Git operations were not affected during either incident.\n\nOur investigation determined that both incidents shared the same underlying cause: a configuration change to a user settings caching mechanism caused a large volume of cache rewrites to occur simultaneously. During the first incident, asynchronous rewrites overwhelmed a shared infrastructure component responsible for coordinating background work, triggering cascading failures. Increased load caused the service responsible for proxying Git operations over HTTPS to exhaust available connections, preventing it from accepting new requests. We mitigated this incident by disabling async cache rewrites and restarting the affected Git proxy service across multiple datacenters.\n\nAn additional source of updates to the same cache circumvented our initial mitigations and caused the second incident. This generated a high volume of synchronous writes, causing replication delays that cascaded in a similar pattern and again exhausted the Git proxy’s connection capacity, degrading availability across multiple services. We mitigated by disabling the source of the cache rewrites and again restarting Git proxy.\n\nWe know these incidents disrupted the workflows of millions of developers. While we have made substantial, long-term investments in how GitHub is built and operated to improve resilience, GitHub's availability is not yet meeting our expectations. Getting there requires deep architectural work that is already underway, as well as urgent, targeted improvements. We are taking the following immediate steps:\n\n1. We have already optimized the caching mechanism to avoid write amplification and added self-throttling during bulk updates.\n2. We are adding safeguards to ensure the caching mechanism responds more quickly to rollbacks and strengthening how changes to these caching systems are planned, validated, and rolled out with additional checks.\n3. We are fixing the underlying cause of connection exhaustion in our Git HTTPS proxy layer so the proxy can recover from this failure mode automatically without requiring manual restarts.\n\nGitHub is critical infrastructure for your work, your teams, and your businesses. We're focusing on these mitigations and long-term infrastructure work so GitHub is available, at scale, when and where you need it.",
"name": "Incident with Pull Requests",
"timestamp": "Feb \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e16:19\u003c/var\u003e - \u003cvar data-var='time'\u003e17:40\u003c/var\u003e UTC"
},
{
"code": "tkz0ptx49rl0",
"impact": "minor",
"message": "On February 9th, 2026, between 09:16 UTC and 15:12 UTC GitHub Actions customers experienced run start delays. Approximately 0.6% of runs across 1.8% of repos were affected, with an average delay of 19 minutes for those delayed runs.\u003cbr\u003e\u003cbr\u003eThe incident occurred when increased load exposed a bottleneck in our event publishing system, causing one compute node to fall behind on processing Actions Jobs. We mitigated by rebalancing traffic and increasing timeouts for event processing. We have since isolated performance critical events to a new, dedicated publisher to prevent contention between events and added safeguards to better tolerate processing timeouts.",
"name": "Incident with Actions",
"timestamp": "Feb \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e14:17\u003c/var\u003e - \u003cvar data-var='time'\u003e15:46\u003c/var\u003e UTC"
},
{
"code": "qrlc0jjgw517",
"impact": "minor",
"message": "On February 9, 2026, between ~06:00 UTC and ~12:12 UTC, Copilot Coding Agent and related Copilot API endpoints experienced degraded availability. The primary impact was to agent-based workflows (requests to /agents/swe/*, including custom agent configuration checks), where 154k users saw failed requests and error responses in their editor/agent experience. Impact was concentrated among users and integrations actively using Copilot Coding Agent with VS Code. \u003cbr\u003e\u003cbr\u003eThe degradation was caused by an unexpected surge in traffic to the related API endpoints that exceeded an internal secondary rate limit. That resulted in upstream request denials which were surfaced to users as elevated 500 errors.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by deploying a change that increased the applicable rate limit for this traffic, which allowed requests to complete successfully and returned the service to normal operation.\u003cbr\u003e\u003cbr\u003eAfter the mitigation, we deployed guardrails with applicable caching to avoid a repeat of similar incidents. We also temporarily increased infrastructure capacity to better handle backlog recovery from the rate limiting. We're are improving monitoring around growing agentic API endpoints.",
"name": "Degraded performance for Copilot Coding Agent",
"timestamp": "Feb \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e10:01\u003c/var\u003e - \u003cvar data-var='time'\u003e12:12\u003c/var\u003e UTC"
},
{
"code": "ffz2k716tlhx",
"impact": "minor",
"message": "On February 9, 2026, between 07:05 UTC and 11:26 UTC, GitHub experienced intermittent degradation across Issues, Pull Requests, Webhooks, Actions, and Git operations. Approximately every 30 minutes, users encountered brief periods of elevated errors and timeouts lasting roughly 15 seconds each. During the incident window, approximately 1–2% of requests were impacted across these services, with Git operations experiencing up to 7% error rates during individual spikes. GitHub Actions saw up to 2% of workflow runs delayed by a median of approximately 7 minutes due to backups created during these periods. \u003cbr\u003e\u003cbr\u003eThis was due to multiple resource-intensive workloads running simultaneously, which caused intermittent processing delays on the data storage layer. We mitigated the incident by scaling storage to a larger compute capacity, which resolved the processing delays. \u003cbr\u003e\u003cbr\u003eWe are working to improve detection of resource-intensive queries, identify changes in load patterns, and enhance our monitoring to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Degraded Performance in Webhooks API and UI, Pull Requests",
"timestamp": "Feb \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e08:15\u003c/var\u003e - \u003cvar data-var='time'\u003e11:26\u003c/var\u003e UTC"
},
{
"code": "41mrvyqnmnmb",
"impact": "minor",
"message": "On February 6, 2026, between 17:49 UTC and 18:36 UTC, the GitHub Mobile service was degraded, and some users were unable to create pull request review comments on deleted lines (and in some cases, comments on deleted files). This impacted users on the newer comment-positioning flow available in version 1.244.0 of the mobile apps. Telemetry indicated that the failures increased as the Android rollout progressed. This was due to a defect in the new comment-positioning workflow that could result in the server rejecting comment creation for certain deleted-line positions.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by halting the Android rollout and implementing interim client-side fallback behavior while a platform fix is in progress. The client-side fallback is scheduled to be published early this week. We are working to (1) add clearer client-side error handling (avoid infinite spinners), (2) improve monitoring/alerting for these failures, and (3) adopt stable diff identifiers for diff-based operations to reduce the likelihood of recurrence.",
"name": "Incident with Pull Requests",
"timestamp": "Feb \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e17:49\u003c/var\u003e - \u003cvar data-var='time'\u003e18:36\u003c/var\u003e UTC"
},
{
"code": "y7www8s68myy",
"impact": "minor",
"message": "On February 10, 2026, between 10:28 and 11:54 UTC, Visual Studio Code users experienced a degraded experience on GitHub Copilot when using the Claude Opus 4.6 model. During this time, approximately 50% of users encountered agent turn failures due to the model being unable to serve the volume of incoming requests.\u003cbr\u003e\u003cbr\u003eRate limits set too low for actual demand caused the issue. While the initial deployment showed no concerns, a surge in traffic from Europe on the following day caused VSCode to begin hitting rate limit errors. Additionally, a degradation message intended to notify users of high usage failed to trigger due to a misconfiguration. We mitigated the incident by adjusting rate limits for the model.\u003cbr\u003e\u003cbr\u003eWe improved our rate limiting to prevent future models from experiencing similar issues. We are also improving our capacity planning processes to reduce the risk of similar incidents in the future, and enhancing our detection and mitigation capabilities to reduce impact to customers.",
"name": "Incident with Copilot",
"timestamp": "Feb \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e11:16\u003c/var\u003e - \u003cvar data-var='time'\u003e11:58\u003c/var\u003e UTC"
},
{
"code": "f314nlctbfs5",
"impact": "minor",
"message": "On February 3, 2026, between 14:00 UTC and 17:40 UTC, customers experienced delays in Webhook delivery for push events and delayed GitHub Actions workflow runs. During this window, Webhook deliveries for push events were delayed by up to 40 minutes, with an average delay of 10 minutes. GitHub Actions workflows triggered by push events experienced similar job start delays. Additionally, between 15:25 UTC and 16:05 UTC, all GitHub Actions workflow runs experienced status update delays of up to 11 minutes, with a median delay of 6 minutes.\u003cbr\u003e\u003cbr\u003eThe issue stemmed from connection churn in our eventing service, which caused CPU saturation and delays for reads and writes, with subsequent downstream delivery delays for Actions and Webhooks. We have added observability tooling and metrics to accelerate detection, and are correcting stream processing client configuration to prevent recurrence.",
"name": "Delays in UI updates for Actions Runs",
"timestamp": "Feb \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e16:10\u003c/var\u003e - \u003cvar data-var='time'\u003e19:28\u003c/var\u003e UTC"
},
{
"code": "zr4jt8bj1mg2",
"impact": "minor",
"message": "On February 3, 2026, between 09:35 UTC and 10:15 UTC, GitHub Copilot experienced elevated error rates, with an average of 4% of requests failing.\u003cbr\u003e\u003cbr\u003eThis was caused by a capacity imbalance that led to resource exhaustion on backend services. The incident was resolved by infrastructure rebalancing, and we subsequently deployed additional capacity.\u003cbr\u003e\u003cbr\u003eWe are improving observability to detect capacity imbalances earlier and enhancing our infrastructure to better handle traffic spikes.",
"name": "Incident with Copilot",
"timestamp": "Feb \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e10:16\u003c/var\u003e - \u003cvar data-var='time'\u003e10:56\u003c/var\u003e UTC"
},
{
"code": "xwn6hjps36ty",
"impact": "major",
"message": "On February 2, 2026, between 18:35 UTC and 22:15 UTC, GitHub Actions hosted runners were unavailable, with service degraded until full recovery at 23:10 UTC for standard runners and at February 3, 2026 00:30 UTC for larger runners. During this time, Actions jobs queued and timed out while waiting to acquire a hosted runner. Other GitHub features that leverage this compute infrastructure were similarly impacted, including Copilot Coding Agent, Copilot Code Review, CodeQL, Dependabot, GitHub Enterprise Importer, and Pages. All regions and runner types were impacted. Self-hosted runners on other providers were not impacted. \u003cbr\u003e\u003cbr\u003eThis outage was caused by a backend storage access policy change in our underlying compute provider that blocked access to critical VM metadata, causing all VM create, delete, reimage, and other operations to fail. More information is available at https://azure.status.microsoft/en-us/status/history/?trackingId=FNJ8-VQZ. This was mitigated by rolling back the policy change, which started at 22:15 UTC. As VMs came back online, our runners worked through the backlog of requests that hadn’t timed out. \u003cbr\u003e\u003cbr\u003eWe are working with our compute provider to improve our incident response and engagement time, improve early detection before they impact our customers, and ensure safe rollout should similar changes occur in the future. We recognize this was a significant outage to our users that rely on GitHub’s workloads and apologize for the impact this had.",
"name": "Incident with Actions",
"timestamp": "Feb \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e19:03\u003c/var\u003e - Feb \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e00:56\u003c/var\u003e UTC"
},
{
"code": "t4s6xgbwv60j",
"impact": "major",
"message": "On February 2, 2026, GitHub Codespaces were unavailable between 18:55 and 22:20 UTC and degraded until the service fully recovered at February 3, 2026 00:15 UTC. During this time, Codespaces creation and resume operations failed in all regions. \u003cbr\u003e\u003cbr\u003eThis outage was caused by a backend storage access policy change in our underlying compute provider that blocked access to critical VM metadata, causing all VM create, delete, reimage, and other operations to fail. More information is available at https://azure.status.microsoft/en-us/status/history/?trackingId=FNJ8-VQZ. This was mitigated by rolling back the policy change, which started at 22:15 UTC. As VMs came back online, our runners worked through the backlog of requests that hadn’t timed out. \u003cbr\u003e\u003cbr\u003eWe are working with our compute provider to improve our incident response and engagement time, improve early detection before they impact our customers, and ensure safe rollout should similar changes occur in the future. We recognize this was a significant outage to our users that rely on GitHub’s workloads and apologize for the impact this had.",
"name": "Incident with Codespaces",
"timestamp": "Feb \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e20:17\u003c/var\u003e - Feb \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e00:54\u003c/var\u003e UTC"
},
{
"code": "g9mdtzd4rt72",
"impact": "major",
"message": "From Jan 31, 2026 00:30 UTC to Feb 2, 2026 18:00 UTC Dependabot service was degraded and failed to create 10% of Automated Pull Requests. This was due to a cluster failover that connected to a read-only database.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by pausing Dependabot queues until traffic was properly routed to healthy clusters. We’re working on identifying and rerunning all failed jobs during this time.\u003cbr\u003e\u003cbr\u003eWe’re adding new monitors and alerts to reduce our time to detection and prevent this in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Feb \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e17:41\u003c/var\u003e - \u003cvar data-var='time'\u003e18:46\u003c/var\u003e UTC"
},
{
"code": "vj8q32xs2d32",
"impact": "minor",
"message": "From Feb 2, 2026 17:13 UTC to Feb 2, 2026 17:36 UTC we experienced failures on ~0.02% of Git operations. While deploying an internal service, a misconfiguration caused a small subset of traffic to route to a service that was not ready. During the incident we observed the degradation and statused publicly.\u003cbr\u003e\u003cbr\u003eTo mitigate the issue, traffic was redirected to healthy instances and we resumed normal operation.\u003cbr\u003e\u003cbr\u003eWe are improving our monitoring and deployment processes in this area to avoid future routing issues.",
"name": "Disruption with some GitHub services",
"timestamp": "Feb \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e17:34\u003c/var\u003e - \u003cvar data-var='time'\u003e17:43\u003c/var\u003e UTC"
}
],
"name": "February",
"year": 2026
},
{
"incidents": [
{
"code": "szwg1z6zjsvx",
"impact": "minor",
"message": "Between 2026-01-30 19:06 UTC and 2026-01-30 20:04 UTC, Copilot Coding Agent experienced sessions getting stuck, with a mismatch between the UI-reported session status and the underlying Actions and job execution state. Impacted users could observe Actions finish successfully but the session UI continuing to show in-progress state, or sessions remaining in queued state.\u003cbr\u003e\u003cbr\u003eThe issue was caused by a feature flag that resulted in events being published to a new Kafka topic. Publishing failures led to buffer/queue overflows in the shared event publishing client, preventing other critical events from being emitted. We mitigated the incident by disabling the feature flag and redeploying production pods, which resumed normal event delivery. We are working to improve safeguards and detection around event publishing failures to reduce time to mitigation for similar issues in the future.",
"name": "Degraded Experience - Failing to finalize some CCA Jobs",
"timestamp": "Jan \u003cvar data-var='date'\u003e30\u003c/var\u003e, \u003cvar data-var='time'\u003e20:59\u003c/var\u003e - \u003cvar data-var='time'\u003e21:22\u003c/var\u003e UTC"
},
{
"code": "22zx8fw9qt4t",
"impact": "minor",
"message": "On Jan 28, 2026, between 14:56 UTC and 15:44 UTC, GitHub Actions experienced degraded performance. During this time, workflows experienced an average delay of 49 seconds, and 4.7% of workflow runs failed to start within 5 minutes. The root cause was an atypical load pattern that overwhelmed system capacity and caused resource contention.\u003cbr\u003e\u003cbr\u003eRecovery began once additional resources came online at 15:25 UTC, with full recovery at 15:44 UTC. We are implementing safeguards to prevent this failure mode and enhancing our monitoring to detect and address similar patterns more quickly in the future.",
"name": "Actions Workflows Run Start Delays",
"timestamp": "Jan \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e15:12\u003c/var\u003e - \u003cvar data-var='time'\u003e15:54\u003c/var\u003e UTC"
},
{
"code": "90hj03y5tj3c",
"impact": "minor",
"message": "On Jan 26, 2026, from approximately 14:03 UTC to 23:42 UTC, GitHub Actions experienced job failures on some Windows standard hosted runners. This was caused by a configuration difference in a new Windows runner type that caused the expected D: drive to be missing. About 2.5% of all Windows standard runners jobs were impacted. Re-run of failed workflows had a high chance of succeeding given the limited rollout of the change.\u003cbr\u003e\u003cbr\u003eThe job failures were mitigated by rolling back the affected configuration and removing the provisioned runners that had this configuration. To reduce the chance of recurrence, we are expanding runner telemetry and improving validation of runner configuration changes. We are also evaluating options to accelerate the mitigation time of any similar future events.",
"name": "Regression in windows runners for public repositories",
"timestamp": "Jan \u003cvar data-var='date'\u003e26\u003c/var\u003e, \u003cvar data-var='time'\u003e19:25\u003c/var\u003e - \u003cvar data-var='time'\u003e23:51\u003c/var\u003e UTC"
},
{
"code": "g697qcy5dsks",
"impact": "minor",
"message": "Between January 24, 2026,19:56 UTC and January 25, 2026, 2:50 UTC repository creation and clone were degraded. On average, the error rate was 25% and peaked at 55% of requests for repository creation. This was due to increased latency on the repositories database impacting a read-after-write problem during repo creation. We mitigated the incident by stopping an operation that was generating load on the database to increase throughput. \u003cbr\u003e\u003cbr\u003eWe have identified the repository creation problem and are working to address the issue and improve our observability to reduce our time to detection and mitigation of issues like this one in the future.\u003cbr\u003e",
"name": "Disruption with repo creation",
"timestamp": "Jan \u003cvar data-var='date'\u003e25\u003c/var\u003e, \u003cvar data-var='time'\u003e02:43\u003c/var\u003e - \u003cvar data-var='time'\u003e03:08\u003c/var\u003e UTC"
},
{
"code": "cqb5hcy0gx18",
"impact": "minor",
"message": "On January 22, 2026, our authentication service experienced an issue between 14:00 UTC and 14:50 UTC, resulting in downstream disruptions for users.\u003cbr\u003e\u003cbr\u003eFrom 14:00 UTC to 14:23 UTC, authenticated API requests experienced higher-than-normal error rates, with an average of 16.9% and occasional peaks up to 22.2% resulting in HTTP 401 responses for authenticated API requests. \u003cbr\u003e\u003cbr\u003eFrom 14:00 UTC to 14:50 UTC, git operations over HTTP were impacted, with error rates averaging 3.8% and peaking at 10.8%. As a result, some users may have been unable to run git commands as expected.\u003cbr\u003e\u003cbr\u003eThis was due to the authentication service reaching the maximum allowed number of database connections. We mitigated the incident by increasing the maximum number of database connections in the authentication service.\u003cbr\u003e\u003cbr\u003eWe are adding additional monitoring around database connection pool usage and improving our traffic projection to reduce our time to detection and mitigation of issues like this one in the future.\u003cbr\u003e",
"name": "Disruption with some GitHub services",
"timestamp": "Jan \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e14:12\u003c/var\u003e - \u003cvar data-var='time'\u003e15:22\u003c/var\u003e UTC"
},
{
"code": "6d5kv6l8d8q3",
"impact": "minor",
"message": "On January 21, between 17:50 and 20:53 UTC, around 350 enterprises and organizations experienced slower load times or timeouts when viewing Copilot policy pages. The issue was traced to performance degradation under load due to an issue in upstream database caching capability within our billing infrastructure, which increased query latency to retrieve billing and policy information from approximately 300ms to up to 1.5s.\u003cbr\u003e\u003cbr\u003eTo restore service, we disabled the affected caching feature, which immediately returned performance to normal. We then addressed the issue in the caching capability and re-enabled our use of the database cache and observed continued recovery.\u003cbr\u003e\u003cbr\u003eMoving forward, we’re tightening our procedures for deploying performance optimizations, adding test coverage, and improving cross-service visibility and alerting so we can detect upstream degradations earlier and reduce impact to customers.",
"name": "Policy pages for Copilot are timing out",
"timestamp": "Jan \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e19:31\u003c/var\u003e - \u003cvar data-var='time'\u003e20:53\u003c/var\u003e UTC"
},
{
"code": "qq1gg4klp5vm",
"impact": "minor",
"message": "On Jan 21st, 2025, between 11:15 UTC and 13:00 UTC the Copilot service was degraded for Grok Code Fast 1 model. On average, more than 90% of the requests to this model failed due to an issue with an upstream provider. No other models were impacted.\u003cbr\u003e\u003cbr\u003eThe issue was resolved after the upstream provider fixed the problem that caused the disruption. GitHub will continue to enhance our monitoring and alerting systems to reduce the time it takes to detect and mitigate similar issues in the future.",
"name": "Copilot Chat - Grok Code Fast 1 Outage",
"timestamp": "Jan \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e11:33\u003c/var\u003e - \u003cvar data-var='time'\u003e12:38\u003c/var\u003e UTC"
},
{
"code": "fw294tth5m8y",
"impact": "minor",
"message": "On January 20, 2026, between 19:08 UTC and 20:18 UTC, manually dispatched GitHub Actions workflows saw delayed job starts. GitHub products built on Actions such as Dependabot, Pages builds, and Copilot coding agent experienced similar delays. All jobs successfully completed despite the delays. At peak impact, approximately 23% of workflow runs were affected, with an average delay of 11 minutes.\u003cbr\u003e\u003cbr\u003eThis was caused by a load pattern shift in Actions scheduled jobs that saturated a shared backend resource. We mitigated the incident by temporarily throttling traffic and scaling up resources to account for the change in load pattern. To prevent recurrence, we have scaled resources appropriately and implemented optimizations to prevent this load pattern in the future.",
"name": "Run start delays in Actions",
"timestamp": "Jan \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e19:49\u003c/var\u003e - \u003cvar data-var='time'\u003e20:10\u003c/var\u003e UTC"
},
{
"code": "m4xfyzd8mn2n",
"impact": "minor",
"message": "On January 20, 2026, between 14:39 UTC and 16:03 UTC, actions-runner-controller users experienced a 1% failure rate for API requests managing GitHub Actions runner scale sets. This caused delays in runner creation, resulting in delayed job starts for workflows targeting those runners. The root cause was a service to service circuit breaker that incorrectly tripped for all users when a single user hit rate limits for runner registration. The issue was mitigated by bypassing the circuit breaker, and users saw immediate and full service recovery following the fix.\u003cbr\u003e\u003cbr\u003eWe have updated our circuit breakers to exclude individual customer rate limits from their triggering logic and are continuing work to improve detection and mitigation times.",
"name": "Incident affecting actions-runner-controller",
"timestamp": "Jan \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e16:02\u003c/var\u003e - \u003cvar data-var='time'\u003e16:23\u003c/var\u003e UTC"
},
{
"code": "zltcsbhqmvqq",
"impact": "minor",
"message": "Between 2026-01-16 16:17 and 2026-01-17 02:54 UTC, some Copilot Business users were unable to access and use certain Copilot features and models. This was due to a bug with how we determine if a user has access to a feature, inadvertently marking features and models as inaccessible for users whose enterprise(s) had not configured the policy.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by reverting the problematic deployment. We are improving our internal monitoring and mitigation processes to reduce the risk and extended downtime of similar incidents in the future.\u003cbr\u003e",
"name": "Disruption with some GitHub services",
"timestamp": "Jan \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e23:53\u003c/var\u003e - Jan \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e02:54\u003c/var\u003e UTC"
},
{
"code": "q987xpbqjbpl",
"impact": "major",
"message": "On January 15, 2026, between 16:40 UTC and 18:20 UTC, we observed increased latency and timeouts across Issues, Pull Requests, Notifications, Actions, Repositories, API, Account Login and Alive. An average 1.8% of combined web and API requests saw failure, peaking briefly at 10% early on. The majority of impact was observed for unauthenticated users, but authenticated users were impacted as well.\u003cbr\u003e\u003cbr\u003eThis was caused by an infrastructure update to some of our data stores. Upgrading this infrastructure to a new major version resulted in unexpected resource contention, leading to distributed impact in the form of slow queries and increased timeouts across services that depend on these datasets. We mitigated this by rolling back to the previous stable version.\u003cbr\u003e\u003cbr\u003eWe are working to improve our validation process for these types of upgrades to catch issues that only occur under high load before full release, improve detection time, and reduce mitigation times in the future.",
"name": "Incident with Issues and Pull Requests",
"timestamp": "Jan \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e16:56\u003c/var\u003e - \u003cvar data-var='time'\u003e18:54\u003c/var\u003e UTC"
},
{
"code": "5ccghcfrkv39",
"impact": "minor",
"message": "On January 15th, between 14:18 UTC and 15:26 UTC, customers experienced delays in status updates for workflow runs and checks. Status updates were delayed by up to 20 minutes, with a median delay of 11 minutes.\u003cbr\u003e\u003cbr\u003eThe issue stemmed from an infrastructure upgrade to our database cluster. The new version introduced resource contention under production load, causing slow query times. We mitigated this by rolling back to the previous stable version. We are working to strengthen our upgrade validation process to catch issues that only manifest under high load. We are also adding new monitors to reduce detection time for similar issues in the future.",
"name": "Actions workflow run and job status updates are experiencing delays",
"timestamp": "Jan \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e14:24\u003c/var\u003e - \u003cvar data-var='time'\u003e15:26\u003c/var\u003e UTC"
},
{
"code": "j7yjswzyrktp",
"impact": "minor",
"message": "On January 14, 2026, between 19:34 UTC and 21:36 UTC, the Webhooks service experienced a degradation that delayed delivery of some webhooks. During this window, a subset of webhook deliveries that encountered proxy tunnel errors on their initial delivery attempt were delayed by more than two minutes. The root cause was a recent code change that added additional retry attempts for this specific error condition, which increased delivery times for affected webhooks. Previously, webhook deliveries encountering this error would not have been delivered.\u003cbr\u003e\u003cbr\u003eThe incident was mitigated by rolling back the change, restoring normal webhook delivery. \u003cbr\u003e\u003cbr\u003eAs a corrective action, we will update our monitoring to measure the webhook delivery latency critical path, ensuring that incidents are accurately scoped to this workflow.",
"name": "Incident with Webhooks",
"timestamp": "Jan \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e20:21\u003c/var\u003e - \u003cvar data-var='time'\u003e21:38\u003c/var\u003e UTC"
},
{
"code": "sq20c7lrwjrp",
"impact": "minor",
"message": "From January 14, 2026, at 18:15 UTC until January 15, 2026, at 11:30 UTC, GitHub Copilot users were unable to select the GPT-5 model for chat features in VS Code, JetBrains IDEs, and other IDE integrations. Users running GPT-5 in Auto mode experienced errors. Other models were not impacted.\n\nWe mitigated this incident by deploying a fix that corrected a misconfiguration in available models, rendering the GPT-5 model available again.\n\nWe are improving our testing processes to reduce the risk of similar incidents in the future, and refining our model availability alerting to improve detection time.\n\nWe did not status before we completed the fix, and the incident is currently resolved. We are sorry for the delayed post on githubstatus.com.",
"name": "[Retroactive] Incident with GitHub Copilot (GPT-5 model)",
"timestamp": "Jan \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e18:00\u003c/var\u003e - \u003cvar data-var='time'\u003e18:00\u003c/var\u003e UTC"
},
{
"code": "4htf8bgy1xlz",
"impact": "minor",
"message": "On January 14th, 2026, between approximately 10:20 and 11:25 UTC, the Copilot service experienced a degradation of the Claude Opus 4.5 model due to an issue with our upstream provider. During this time period, users encountered a 4.5% error rate when using Claude Opus 4.5. No other models were impacted.\u003cbr\u003eThe issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.",
"name": "Claude Opus 4.5 model experiencing degraded performance",
"timestamp": "Jan \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e10:56\u003c/var\u003e - \u003cvar data-var='time'\u003e12:23\u003c/var\u003e UTC"
},
{
"code": "j8t05z0nh91f",
"impact": "minor",
"message": "This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.",
"name": "Copilot's GPT-5.1 model has degraded performance",
"timestamp": "Jan \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e09:24\u003c/var\u003e - \u003cvar data-var='time'\u003e10:52\u003c/var\u003e UTC"
},
{
"code": "zwlfnp5z0r4s",
"impact": "minor",
"message": "Between 2026-01-13 22:20 and 2026-01-14 00:18 UTC, GitHub Code Search experienced an increase in latency and request timeouts. This was caused by some network transit links between GitHub and Azure Express Route experiencing a small error rate that contributed to applications requests failing, increasing application latency and timeouts. The incident resulted in less than 1% of requests to fail due to timeouts.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by disabling the links in question. Monitoring each unique network path across providers would have allowed us to mitigate this earlier. We are running root cause analysis with network providers to help us reduce time-to-discover and time-to-mitigate.",
"name": "Disruption with some GitHub services",
"timestamp": "Jan \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e22:21\u003c/var\u003e - Jan \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e00:18\u003c/var\u003e UTC"
},
{
"code": "1lnqb2vk25vn",
"impact": "major",
"message": "On January 13th, 2026, between 09:25 UTC and 10:11 UTC, GitHub Copilot experienced unavailability. During this window, error rates averaged 18% and peaked at 100% of service requests, leading to an outage of chat features across Copilot Chat, VS Code, JetBrains IDEs, and other Copilot-dependent products. \u003cbr\u003e\u003cbr\u003eThis incident was triggered by a configuration error during a model update. We mitigated the incident by rolling back this change. However, a second recovery phase lasted until 10:46 UTC, due to unexpected latency with the GPT 4.1 model. To prevent recurrence, we are investing in new monitors and more robust testing environments to reduce further misconfigurations, and to improve our time to detection and mitigation of future issues.",
"name": "GitHub Copilot failures",
"timestamp": "Jan \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e09:38\u003c/var\u003e - \u003cvar data-var='time'\u003e10:46\u003c/var\u003e UTC"
},
{
"code": "zcxwznlbgzr7",
"impact": "major",
"message": "From January 9 13:11 UTC to January 12 10:17 UTC, new Linux Custom Images generated for Larger Hosted Runners were broken and not able to run jobs. Customers who did not generate new Custom Images during this period were not impacted. This issue was caused by a change to improve reliability of the image creation process. Due to a bug, the change triggered an unrelated protection mechanism which determines if setup has already been attempted on the VM and caused the VM to be marked unhealthy. Only Linux images which were generated while the change was enabled were impacted. The issue was mitigated by rolling back the change.\u003cbr\u003e\u003cbr\u003eWe are improving our testing around Custom Image generation as part of our GA readiness process for the public preview feature.. This includes expanding our canary suite to detect this and similar interactions as part of a controlled rollout in staging prior to any customer impact.",
"name": "Disruption with some GitHub services",
"timestamp": "Jan \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e10:02\u003c/var\u003e - \u003cvar data-var='time'\u003e10:17\u003c/var\u003e UTC"
},
{
"code": "8bcsk6c4prjl",
"impact": "minor",
"message": "From January 5, 2026, 00:00 UTC to January 10, 2026, 02:30 UTC, customers using the AI Controls public preview feature experienced delays in viewing Copilot agent session data. Newly created sessions took progressively longer to appear, initially hours, then eventually exceeding 24 hours. Since the page displays only the most recent 24 hours of activity, once processing delays exceeded this threshold, no recent data was visible. Session data remained available in audit logs throughout the incident.\u003cbr\u003e\u003cbr\u003eInefficient database queries in the data processing pipeline caused significant processing latency, creating a multi-day backlog. As the backlog grew, the delay between when sessions occurred and when they appeared on the page increased, eventually exceeding the 24-hour display window.\u003cbr\u003e\u003cbr\u003eThe issue was resolved on January 10, 2026, 02:30 UTC, after query optimizations and a database index were deployed. We are implementing enhanced monitoring and automated testing to detect inefficient queries before deployment to prevent recurrence.",
"name": "Disruption with some GitHub services",
"timestamp": "Jan \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e17:53\u003c/var\u003e - Jan \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e02:33\u003c/var\u003e UTC"
},
{
"code": "91ryv3glwnvz",
"impact": "minor",
"message": "On January 8th, 2025, between approximately 00:00 and 1:30 UTC, the Copilot service experienced a degradation of the Grok Code Fast 1 model due to an issue with our upstream provider. Users encountered elevated error rates when using Grok Code Fast 1. Approximately 4.5% of requests failed across all users during this time. No other models were impacted.\u003cbr\u003e\u003cbr\u003eThe issue was resolved by a mitigation put in place by our provider.",
"name": "Incident with Copilot",
"timestamp": "Jan \u003cvar data-var='date'\u003e8\u003c/var\u003e, \u003cvar data-var='time'\u003e00:45\u003c/var\u003e - \u003cvar data-var='time'\u003e01:32\u003c/var\u003e UTC"
},
{
"code": "vyxbxqhdt75d",
"impact": "minor",
"message": "On January 7th, 2026, between 17:16 and 19:33 UTC Copilot Pro and Copilot Business users were unable to use certain premium models, including Claude Opus 4.5 and GPT-5.2. This was due to a misconfiguration with Copilot models, inadvertently marking these premium models as inaccessible for users with Copilot Pro and Copilot Business licenses.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by reverting the erroneous config change. We are improving our testing processes to reduce the risk of similar incidents in the future, and refining our model availability alerting to improve detection time.",
"name": "Some models missing in Copilot",
"timestamp": "Jan \u003cvar data-var='date'\u003e7\u003c/var\u003e, \u003cvar data-var='time'\u003e18:32\u003c/var\u003e - \u003cvar data-var='time'\u003e21:07\u003c/var\u003e UTC"
},
{
"code": "266hgcx77zqw",
"impact": "minor",
"message": "On January 6, 2026 between 12:55 UTC and 17:04 UTC, the ability to download Actions artifacts from GitHub’s web interface was degraded. During this time, all attempts to download artifacts from the web interface failed. Artifact downloads via the REST API and GitHub CLI were unaffected.\u003cbr\u003e\u003cbr\u003eThis was due to a client-side change that was deployed to optimize performance when navigating between pages in a repository. We mitigated the incident by reverting the change. \u003cbr\u003e\u003cbr\u003eWe are working to improve testing of related changes and to add monitoring coverage for artifact downloads through the web interface to reduce our time to detection and prevent similar incidents from occurring in the future.",
"name": "Incident with Actions",
"timestamp": "Jan \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e16:41\u003c/var\u003e - \u003cvar data-var='time'\u003e17:06\u003c/var\u003e UTC"
},
{
"code": "d6pkw798c0f0",
"impact": "minor",
"message": "On January 6th, 2026, between approximately 8:41 and 10:07 UTC, the Copilot service experienced a degradation of the GPT-5.1-Codex-Max model due to an issue with our upstream provider. During this time, up to 14.17% of requests to GPT-5.1-Codex-Max failed. No other models were impacted.\u003cbr\u003e\u003cbr\u003eThe issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.",
"name": "Incident with Copilot",
"timestamp": "Jan \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e08:56\u003c/var\u003e - \u003cvar data-var='time'\u003e10:08\u003c/var\u003e UTC"
},
{
"code": "d5t56pdnjxhd",
"impact": "minor",
"message": "On December 31, 2025, between 04:00 UTC and 22:31 UTC, all users visiting https://github.com/features/copilot were unable to load the page and were instead redirected to an error page.\nThe issue was caused by an unexpected content change that resulted in page rendering errors.\nWe mitigated the incident by reverting the change, which restored normal page behavior.\nTo reduce the likelihood and duration of similar issues in the future, we are improving monitoring and alerting for increased error rates on this page and similar pages, and strengthening validation and safeguards around content updates to prevent unexpected changes from causing user-facing errors.",
"name": "Disruption with some GitHub services",
"timestamp": "Jan \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e21:24\u003c/var\u003e - \u003cvar data-var='time'\u003e22:31\u003c/var\u003e UTC"
}
],
"name": "January",
"year": 2026
}
],
"start_time": "2026-01-01T00:00:00Z",
"time_zone": "UTC"
},
{
"end_time": "2025-12-31T23:59:59Z",
"months": [
{
"incidents": [
{
"code": "ccrzb3ms9j2d",
"impact": "minor",
"message": "On December 23, 2025, between 09:15 UTC and 10:32 UTC the Issues and Pull Requests search indexing service was degraded and caused search results to contain stale data up to 3 minutes old for roughly 1.3 million issues and pull requests. This was due to search indexing queues backing up from resource contention caused by a running transition.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by cancelling the running transition.\u003cbr\u003e\u003cbr\u003eWe are working to implement closer monitoring of search infrastructure resource utilization during transitions to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Incident with Issues and Pull Requests",
"timestamp": "Dec \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e09:56\u003c/var\u003e - \u003cvar data-var='time'\u003e10:32\u003c/var\u003e UTC"
},
{
"code": "y2wxzcfbgbn2",
"impact": "major",
"message": "On December 22, 2025, between 22:01 UTC and 22:32 UTC, unauthenticated requests to github.com were degraded, resulting in slow or timed out page loads and API requests. Unauthenticated requests from Actions jobs, such as release downloads, were also impacted. Authenticated traffic was not impacted. This was due to a severe spike in traffic, primarily to search endpoints.\u003cbr\u003e\u003cbr\u003eOur immediate response focused on identifying and mitigating the source of the traffic increase, which along with automated traffic management restored full service for our users.\u003cbr\u003e\u003cbr\u003eWe improved limiters for load to relevant endpoints and are continuing work to more proactively identify these large changes in traffic volume, improve resilience in critical request flows, and improve our time to mitigation.",
"name": "Disruption with some GitHub services",
"timestamp": "Dec \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e22:31\u003c/var\u003e - Dec \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e00:17\u003c/var\u003e UTC"
},
{
"code": "49y7x9g06l4x",
"impact": "major",
"message": "On December 18, 2025, between 16:25 UTC and 19:09 UTC the service underlying Copilot policies was degraded and users, organizations, and enterprises were not able to update any policies related to Copilot. No other GitHub services, including other Copilot services were impacted. This was due to a database migration causing a schema drift.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by synchronizing the schema. We have hardened the service to make sure schema drift does not cause any further incidents, and will investigate improvements in our deployment pipeline to shorten time to mitigation in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Dec \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e17:36\u003c/var\u003e - \u003cvar data-var='time'\u003e19:09\u003c/var\u003e UTC"
},
{
"code": "x696x0g4t85l",
"impact": "major",
"message": "On December 18th, 2025, from 08:15 UTC to 17:11 UTC, some GitHub Actions runners experienced intermittent timeouts for Github API calls, which led to failures during runner setup and workflow execution. This was caused by network packet loss between runners in the West US region and one of GitHub’s edge sites. Approximately 1.5% of jobs on larger and standard hosted runners in the West US region were impacted, 0.28% of all Actions jobs during this period.\u003cbr\u003e\u003cbr\u003eBy 17:11 UTC, all traffic was routed away from the affected edge site, mitigating the timeouts. We are working to improve early detection of cross-cloud connectivity issues and faster mitigation paths to reduce the impact of similar issues in the future.",
"name": "Intermittent networking failures across GitHub-hosted Actions runners",
"timestamp": "Dec \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e16:33\u003c/var\u003e - \u003cvar data-var='time'\u003e17:41\u003c/var\u003e UTC"
},
{
"code": "lctp4bzd37h2",
"impact": "none",
"message": "From 11:50-12:25 UTC, Copilot Coding Agent was unable to process new agent requests. This affected all users creating new jobs during this timeframe, while existing jobs remained unaffected. The cause of this issue was a change to the actions configuration where Copilot Coding Agent runs, which caused the setup of the Actions runner to fail, and the issue was resolved by rolling back this change.\nAs a short term solution, we hope to increase our alerting criteria so that we can be alerted more quickly when an incident occurs, and in the long term we hope to harden our runner configuration to be more resilient against errors.",
"name": "Incident With Copilot",
"timestamp": "Dec \u003cvar data-var='date'\u003e16\u003c/var\u003e, \u003cvar data-var='time'\u003e12:00\u003c/var\u003e - \u003cvar data-var='time'\u003e12:00\u003c/var\u003e UTC"
},
{
"code": "qkr3gcys0lzx",
"impact": "major",
"message": "On December 15, 2025, between 15:15 UTC and 18:22 UTC, Copilot Code Review experienced a service degradation that caused 46.97% of pull request review requests to fail, requiring users to re-request a review. Impacted users saw the error message: “Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.” The remaining requests completed successfully.\u003cbr\u003e\u003cbr\u003eThe degradation was caused by elevated response times in an internal, model-backed dependency, which led to request timeouts and backpressure in the review processing pipeline, resulting in sustained queue growth and failed review completion.\u003cbr\u003e\u003cbr\u003eWe mitigated the issue by temporarily bypassing fix suggestions to reduce latency, increasing worker capacity to drain the backlog, and rolling out a model configuration change that reduced end-to-end latency. Queue depth and request success rates returned to normal and remained stable through peak traffic.\u003cbr\u003e\u003cbr\u003eFollowing the incident, we increased baseline worker capacity, added instrumentation for worker utilization and queue health, and are improving automatic load-shedding, fallback behavior, and alerting to reduce time to detection and mitigation for similar issues.",
"name": "Copilot Code Review is degraded, and not returning responses to users",
"timestamp": "Dec \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e17:43\u003c/var\u003e - \u003cvar data-var='time'\u003e18:22\u003c/var\u003e UTC"
},
{
"code": "jcbsbbfd1x9r",
"impact": "minor",
"message": "On Dec 15th, 2025, between 14:00 UTC and 15:45 UTC the Copilot service was degraded for Grok Code Fast 1 model. On average, 4% of the requests to this model failed due to an issue with our upstream provider. No other models were impacted.\u003cbr\u003e\u003cbr\u003eThe issue was resolved after the upstream provider fixed the problem that caused the disruption. GitHub will continue to enhance our monitoring and alerting systems to reduce the time it takes to detect and mitigate similar issues in the future.",
"name": "Incident with Copilot Grok Code Fast 1",
"timestamp": "Dec \u003cvar data-var='date'\u003e15\u003c/var\u003e, \u003cvar data-var='time'\u003e14:12\u003c/var\u003e - \u003cvar data-var='time'\u003e15:45\u003c/var\u003e UTC"
},
{
"code": "40730vhmg6y8",
"impact": "minor",
"message": "Between 13:25 UTC and 18:35 UTC on Dec 11th, GitHub experienced an increase in scraper activity on public parts of our website. This scraper activity caused a low priority web request pool to increase and eventually exceed total capacity resulting in users experiencing 500 errors. In particular, this affected Login, Logout, and Signup routes, along with less than 1% requests from within Actions jobs. At the peak of the incident, 7.6% of login requests were impacted, which was the most significant impact of this scraping attack.\u003cbr\u003e\u003cbr\u003eOur mitigation strategy identified the scraping activity and blocked it. We also increased the pool of web requests that were impacted to have more capacity, and lastly we upgraded key user login routes to higher priority queues. \u003cbr\u003e\u003cbr\u003eIn future, we’re working to more proactively identify this particular scraper activity and have faster mitigation times.",
"name": "Disruptions in Login and Signup Flows",
"timestamp": "Dec \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e18:40\u003c/var\u003e - \u003cvar data-var='time'\u003e20:05\u003c/var\u003e UTC"
},
{
"code": "xntfc1fz5rfb",
"impact": "minor",
"message": "Between 13:25 UTC and 18:35 UTC on December 11th, GitHub experienced elevated traffic to portions of GitHub.com that exceeded previously provisioned capacity for specific request types. As a result, users encountered intermittent 500 errors. Impact was most pronounced on Login, Logout, and Signup pages, peaking at 7.6% of login requests. Additionally, fewer than 1% of requests originating from GitHub Actions jobs were affected. \u003cbr\u003e\u003cbr\u003eThis incident was driven by the same underlying factors as the previously reported \u003ca href=\"https://www.githubstatus.com/incidents/40730vhmg6y8\"\u003edisruption to Login and Signup flows\u003c/a\u003e\u003cbr\u003e\u003cbr\u003eOur immediate response focused on identifying and mitigating the source of the traffic increase. We increased available capacity for web request handling to relieve pressure on constrained pools. To reduce recurrence risk, we also re-routed critical authentication endpoints to a different traffic pool, ensuring sufficient isolation and headroom for login related traffic.\u003cbr\u003e\u003cbr\u003eIn future, we’re working to more proactively identify these large changes in traffic volume and improve our time to mitigation.\u003cbr\u003e\u003cbr\u003e",
"name": "We are investigating a rise in request failures on several services",
"timestamp": "Dec \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e15:47\u003c/var\u003e - \u003cvar data-var='time'\u003e17:53\u003c/var\u003e UTC"
},
{
"code": "17lxd87h22fj",
"impact": "minor",
"message": "Between December 9th, 2025 21:07 UTC and December 10th, 2025 14:52 UTC, 177 macos-14-large jobs were run on an Ubuntu larger runner VM instead of MacOS runner VMs. The impacted jobs were routed to a larger runner with incorrect metadata. We mitigated this by deleting the runner.\u003cbr\u003e\u003cbr\u003eThe routing configuration is not something controlled externally. A manual override was done previously for internal testing, but left incorrect metadata for a large runner instance. An infrastructure migration caused this misconfigured runner to come online which started the incorrect assignments. We are removing the ability to manually override this configuration entirely, and are adding alerting to identify possible OS mismatches for hosted runner jobs.\u003cbr\u003e\u003cbr\u003eAs a reminder, hosted runner VMs are secure and ephemeral, with every VM reimaged after every single job. All jobs impacted here were originally targeted at a GitHub-owned VM image and were run on a GitHub-owned VM image.",
"name": "Some macOS Actions jobs routing to Ubuntu instead",
"timestamp": "Dec \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e13:34\u003c/var\u003e - \u003cvar data-var='time'\u003e14:52\u003c/var\u003e UTC"
},
{
"code": "znmz4zmsg1rv",
"impact": "minor",
"message": "On December 10, 2025 between 08:50 UTC and 11:00 UTC, some GitHub Actions workflow runs experienced longer-than-normal wait times for jobs starting or completing. All jobs successfully completed despite the delays. At peak impact, approximately 8% of workflow runs were affected.\u003cbr\u003e\u003cbr\u003eDuring this incident, some nodes received a spike in workflow events that led to queuing of event processing. Because runs are pinned to nodes, runs being processed by these nodes saw delays in starting or showing as completed. The team was alerted to this at 8:58 UTC. Impacted nodes were disabled from processing new jobs to allow queues to drain.\u003cbr\u003e\u003cbr\u003eWe have increased overall processing capacity and are implementing safeguards to better balance load across all nodes when spikes occur. This is important to ensure our available capacity can always be fully utilized.",
"name": "Some Actions customers experiencing run start delays",
"timestamp": "Dec \u003cvar data-var='date'\u003e10\u003c/var\u003e, \u003cvar data-var='time'\u003e09:11\u003c/var\u003e - \u003cvar data-var='time'\u003e11:05\u003c/var\u003e UTC"
},
{
"code": "n4fpndtvbmwb",
"impact": "minor",
"message": "On December 8, 2025, between 21:15 and 22:24 UTC, Copilot code completions experienced a significant service degradation. During this period, up to 65% of code completion requests failed.\u003cbr\u003e\u003cbr\u003eThe root cause was an internal feature flag that caused the primary model supporting Copilot code completions to appear unavailable to the backend service. The issue was resolved once the flag was disabled.\u003cbr\u003e\u003cbr\u003eTo prevent recurrence, we expanded test coverage for Copilot code completion models and are strengthening our detection mechanisms to better identify and respond to traffic anomalies.",
"name": "Disruption with some GitHub services",
"timestamp": "Dec \u003cvar data-var='date'\u003e8\u003c/var\u003e, \u003cvar data-var='time'\u003e21:28\u003c/var\u003e - \u003cvar data-var='time'\u003e22:33\u003c/var\u003e UTC"
},
{
"code": "r6fc6tyzr430",
"impact": "major",
"message": "On November 26th, 2025, between approximately 02:24 UTC and December 8th, 2025 at 20:26 UTC, enterprise administrators experienced a disruption when viewing agent session activities in the Enterprise AI Controls page. During this period, users were unable to list agent session activity in the AI Controls view. This did not impact viewing agent session activity in audit logs or directly navigating to individual agent session logs, or otherwise managing AI Agents.\u003cbr\u003e\u003cbr\u003eThe issue was caused by a misconfiguration in a change deployed on November 25th that unintentionally prevented data from being published to an internal Kafka topic responsible for feeding the AI Controls page with agent session activity information.\u003cbr\u003e\u003cbr\u003eThe problem was identified and mitigated on December 8th by correcting the configuration issue. GitHub is improving monitoring for data pipeline dependencies and enhancing pre-deployment validation to catch configuration issues before they reach production.",
"name": "Potential disruption with our Agent Control Plane UI Settings",
"timestamp": "Dec \u003cvar data-var='date'\u003e8\u003c/var\u003e, \u003cvar data-var='time'\u003e19:51\u003c/var\u003e - \u003cvar data-var='time'\u003e21:06\u003c/var\u003e UTC"
},
{
"code": "3hdjjpkvz895",
"impact": "minor",
"message": "On December 5th, 2025, between 12:00 pm UTC and 9:00 pm UTC, our Team Synchronization service experienced a significant degradation, preventing over 209,000 organization teams from syncing their identity provider (IdP) groups. The incident was triggered by a buildup of synchronization requests, resulting in elevated Redis key usage and high CPU consumption on the underlying Redis cluster.\u003cbr\u003e\u003cbr\u003eTo mitigate further impact, we proactively paused all team synchronization requests between 3:00 pm UTC and 8:15 pm UTC, allowing us to stabilize the Redis cluster. Our engineering team also resolved the issue by flushing the affected Redis keys and queues, which promptly stopped runaway growth and restored service health. Additionally, we scaled up our infrastructure resources to improve our ability to process the high volume of synchronization requests. All pending team synchronizations were successfully processed following service restoration.\u003cbr\u003e\u003cbr\u003eWe are working to strengthen the Team Synchronization service by implementing a killswitch, adding throttling to prevent excessive enqueueing of synchronization requests, and improving the scheduler to avoid duplicate job requests. Additionally, we’re investing in better observability to alert when job drops occur. These efforts are focused on preventing similar incidents and improving overall reliability going forward.",
"name": "Team synchronization is experiencing delays for non enterprise managed users",
"timestamp": "Dec \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e18:38\u003c/var\u003e - \u003cvar data-var='time'\u003e22:20\u003c/var\u003e UTC"
},
{
"code": "k95d6yq4tmly",
"impact": "none",
"message": "On December 3, 2025, between 22:21 UTC and 23:44 UTC, the Webhooks service experienced a degradation that delayed writes of webhook delivery records to our database. During this period, many webhook deliveries were not visible in the webhook delivery UI or API for more than an hour after they were sent. As a result, customers were temporarily unable to request redeliveries for those delayed records. The underlying cause was throttling of database writes due to high replication lag.\n\nWe mitigated the incident by temporarily disabling delivery history for a small number of very high‑volume webhook owners to reduce write pressure and stabilize the service. We are contacting the affected customers directly with more details.\n\nWe are improving our webhook delivery storage architecture so it can scale with current and future webhook traffic, reducing the likelihood and impact of similar issues.",
"name": "Webhooks delivery degradation",
"timestamp": "Dec \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e22:30\u003c/var\u003e - \u003cvar data-var='time'\u003e22:30\u003c/var\u003e UTC"
}
],
"name": "December",
"year": 2025
},
{
"incidents": [
{
"code": "d4775b3j5mwm",
"impact": "major",
"message": "On November 28th, 2025, between approximately 05:51 and 08:04 UTC, Copilot experienced an outage affecting the Claude Sonnet 4.5 model. Users attempting to use this model received an HTTP 400 error, resulting in 4.6% of total chat requests during this timeframe failing. Other models were not impacted.\u003cbr\u003e\u003cbr\u003eThe issue was caused by a misconfiguration deployed to an internal service which made Claude Sonnet 4.5 unavailable. The problem was identified and mitigated by reverting the change. GitHub is working to improve cross-service deploy safeguards and monitoring to prevent similar incidents in the future.",
"name": "Incident with Copilot",
"timestamp": "Nov \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e06:59\u003c/var\u003e - \u003cvar data-var='time'\u003e08:23\u003c/var\u003e UTC"
},
{
"code": "v6sx0dv6rv2x",
"impact": "minor",
"message": "On November 24, 2025, between 12:15 and 15:04 UTC, Codespaces users encountered connection issues when attempting to create a codespace after choosing the recently released VS Code Codespaces extension, version 1.18.1. Users were able to downgrade to the 1.18.0 version of the extension during this period to work around this issue. At peak, the error rate was 19% of connection requests. This was caused by mismatching version dependencies for the released VS Code Codespaces extension.\u003cbr\u003e\u003cbr\u003eThe connection issues were mitigated by releasing the VS Code Codespaces extension version 1.18.2 that addressed the issue. Users utilizing version 1.18.1 of the VS Code Codespaces extension are advised to upgrade to version \u0026gt;=1.18.2.\u003cbr\u003e\u003cbr\u003eWe are improving our validation and release process for this extension to ensure functional issues like this are caught before release to customers and to reduce detection and mitigation times for extension issues like this in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Nov \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e13:10\u003c/var\u003e - \u003cvar data-var='time'\u003e15:04\u003c/var\u003e UTC"
},
{
"code": "zzl9nl31lb35",
"impact": "minor",
"message": "Between November 19th, 16:13UTC and November 21st, 12:22UTC, the GitHub Enterprise Importer (GEI) service was in a degraded state, during which time, customers of the service experienced a delay when reclaiming mannequins post-migration.\u003cbr\u003e\u003cbr\u003eWe have taken steps to prevent similar incidents from occurring in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Nov \u003cvar data-var='date'\u003e19\u003c/var\u003e, \u003cvar data-var='time'\u003e16:13\u003c/var\u003e - Nov \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e00:22\u003c/var\u003e UTC"
},
{
"code": "cg3wwz9dw5dg",
"impact": "minor",
"message": "Between November 20, 2025 17:16 UTC to November, 2025 19:08 UTC some users experienced delayed or failed Git Operations for raw file downloads. On average, the error rate was less than 0.2%. This was due to a sustained increase in unauthenticated repository traffic.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by applying regional rate limiting and are taking steps to improve our monitoring and time to mitigation for similar issues in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Nov \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e18:04\u003c/var\u003e - \u003cvar data-var='time'\u003e19:24\u003c/var\u003e UTC"
},
{
"code": "zs5ccnvqv64m",
"impact": "minor",
"message": "On November 19, between 17:36 UTC and 18:04 UTC, GitHub Actions service experienced degraded performance that caused excessive latency in queueing and updating workflow runs and job statuses. Operations related to artifacts, cache, job steps and logs also had significantly increased latency. At peak, 67% of workflow jobs queued during that timeframe were impacted, and the median latency for impacted operations increased by up to 35x.\u003cbr\u003e\u003cbr\u003eThis was caused by a significant change in load pattern on Actions Cache-related operations, leading to a saturated shared resource on the backend. The impact was mitigated by mitigating the new load pattern.\u003cbr\u003e\u003cbr\u003eTo reduce the likelihood of a recurrence, we are improving rate-limiting measures in this area to ensure a more consistent experience for all customers. We are also evaluating changes to reduce the scope of impact.",
"name": "Incident with Actions",
"timestamp": "Nov \u003cvar data-var='date'\u003e19\u003c/var\u003e, \u003cvar data-var='time'\u003e17:48\u003c/var\u003e - \u003cvar data-var='time'\u003e18:07\u003c/var\u003e UTC"
},
{
"code": "5q7nmlxz30sk",
"impact": "major",
"message": "From Nov 18, 2025 20:30 UTC to Nov 18, 2025 21:34 UTC we experienced failures on all Git operations, including both SSH and HTTP Git client interactions, as well as raw file access. These failures also impacted products that rely on Git operations.\u003cbr\u003e\u003cbr\u003eThe root cause was an expired TLS certificate used for internal service-to-service communication. We mitigated the incident by replacing the expired certificate and restarting impacted services. Once those services were restarted we saw a full recovery.\u003cbr\u003e\u003cbr\u003eWe have updated our alerting to cover the expired certificate and are performing an audit of other certificates in this area to ensure they also have the proper alerting and automation before expiration. In parallel, we are accelerating efforts to eliminate our remaining manually managed certificates, ensuring all service-to-service communication is fully automated and aligned with modern security practices.",
"name": "Git operation failures",
"timestamp": "Nov \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e20:39\u003c/var\u003e - \u003cvar data-var='time'\u003e21:59\u003c/var\u003e UTC"
},
{
"code": "ql19qqqmdf99",
"impact": "minor",
"message": "Between November 17, 2025 21:24 UTC and November 18, 2025 00:04 UTC the gists service was degraded and users were unable to create gists via the web UI. 100% of gist creation requests failed with a 404 response. This was due to a change in the web middleware that inadvertently triggered a routing error. We resolved the incident by rolling back the change. We are working on more effective monitoring to reduce the time it takes to detect similar issues and evaluating our testing approach for middleware functionality.",
"name": "Disruption with some GitHub services",
"timestamp": "Nov \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e23:01\u003c/var\u003e - Nov \u003cvar data-var='date'\u003e18\u003c/var\u003e, \u003cvar data-var='time'\u003e00:10\u003c/var\u003e UTC"
},
{
"code": "3bmqzkljz8s3",
"impact": "major",
"message": "From Nov 17, 2025 00:00 UTC to Nov 17, 2025 15:00 UTC Dependabot was hitting a rate limit in GitHub Container Registry (GHCR) and was unable to complete about 57% of jobs.\u003cbr\u003e\u003cbr\u003eTo mitigate the issue we lowered the rate at which Dependabot started jobs and increased the GHCR rate limit.\u003cbr\u003e\u003cbr\u003eWe’re adding new monitors and alerts and looking into more ways to decrease load on GHCR to help prevent this in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Nov \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e16:52\u003c/var\u003e - \u003cvar data-var='time'\u003e19:08\u003c/var\u003e UTC"
},
{
"code": "1jw8ltnr1qrj",
"impact": "minor",
"message": "From Nov 13, 2025 14:50 UTC to Nov 13, 2025 15:01 UTC we experienced failures on all Git Push and SSH operations. An internal service became unhealthy due to a scaling configuration change. We reverted the change and are evaluating our health monitoring and processes to prevent similar incidents.",
"name": "Some users may experience failing git push and pull operations.",
"timestamp": "Nov \u003cvar data-var='date'\u003e13\u003c/var\u003e, \u003cvar data-var='time'\u003e15:00\u003c/var\u003e - \u003cvar data-var='time'\u003e15:13\u003c/var\u003e UTC"
},
{
"code": "br7kz68t38cl",
"impact": "minor",
"message": "On November 12, 2025, between 22:10 UTC and 23:04 UTC, Codespaces used internally at GitHub were impacted. There was no impact to external customers. The scope of impact was not clear in the initial steps of incident response, so it was considered public until confirmed otherwise. One improvement from this will be improved clarity of internal versus public impact for similar failures to better inform our status decisions going forward.",
"name": "Disruption with some GitHub services",
"timestamp": "Nov \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e22:26\u003c/var\u003e - \u003cvar data-var='time'\u003e23:04\u003c/var\u003e UTC"
},
{
"code": "l0brgg302bf5",
"impact": "minor",
"message": "On November 12th, 2025, from 13:10 - 17:40 UTC, notifications service was degraded, showing an increase in web notifications latency and increasing delays in notification deliveries. A change to the notifications settings access path introduced additional load to the settings system, degrading its response times. This impacted both requests to web notifications (with p99 response times as high as 1.5s, while lower percentiles remained stable) and notification deliveries, which reached a peak delay of 24 minutes on average. System capacity was increased around 15:10 UTC and the problematic change was fully reverted soon after that, restoring the latency of web notifications and increasing notification delivery throughput, decreasing the delay in notification deliveries. The notification queue was fully emptied around 17:40 UTC.\u003cbr\u003e\u003cbr\u003eWe are working to adjust capacity in the affected systems and to improve the time needed to address these capacity issues.",
"name": "Delay in notification deliveries",
"timestamp": "Nov \u003cvar data-var='date'\u003e12\u003c/var\u003e, \u003cvar data-var='time'\u003e14:23\u003c/var\u003e - \u003cvar data-var='time'\u003e17:39\u003c/var\u003e UTC"
},
{
"code": "htcm010tcwjq",
"impact": "minor",
"message": "On November 11, 2025, between 16:28 UTC and 20:54 UTC, GitHub Actions larger hosted runners experienced degraded performance, with 0.4% of overall workflow runs and 8.8% of larger hosted runner jobs failing to start within 5 minutes. The majority of impact was mitigated by 18:44, with a small tail of organizations taking longer to recover.\u003cbr\u003e\u003cbr\u003eThe impact was caused by the same database infrastructure issue that caused similar larger hosted runner performance degradation on October 23rd, 2025. In this case, it was triggered by a brief infrastructure event in this incident rather than a database change.\u003cbr\u003e\u003cbr\u003eThrough this incident, we identified and implemented a better solution for both prevention and faster mitigation. In addition to this, a durable solution for the underlying database issue is rolling out soon.",
"name": "Larger hosted runners experiencing delays",
"timestamp": "Nov \u003cvar data-var='date'\u003e11\u003c/var\u003e, \u003cvar data-var='time'\u003e18:02\u003c/var\u003e - \u003cvar data-var='time'\u003e20:54\u003c/var\u003e UTC"
},
{
"code": "gnzclztblsh3",
"impact": "minor",
"message": "Between November 5, 2025 23:27 UTC and November 6, 2025 00:06 UTC, ghost text requests experienced errors from upstream model providers. This was a continuation of the service disruption for which we statused Copilot earlier that day, although more limited in scope.\u003cbr\u003e\u003cbr\u003eDuring the service disruption, users were again automatically re-routed to healthy model hosts, minimizing impact to users and we are updating our monitors and failover mechanism to mitigate similar issues in the future.",
"name": "Incident with Copilot",
"timestamp": "Nov \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e23:41\u003c/var\u003e - Nov \u003cvar data-var='date'\u003e6\u003c/var\u003e, \u003cvar data-var='time'\u003e00:06\u003c/var\u003e UTC"
},
{
"code": "d5c8rxcwt7xw",
"impact": "minor",
"message": "On November 5, 2025, between 21:46 and 23:36 UTC, ghost text requests experienced errors from upstream model providers that resulted in 0.9% of users seeing elevated error rates.\u003cbr\u003e\u003cbr\u003eDuring the service disruption, users were automatically re-routed to healthy model hosts but may have experienced increased latency in response times as a result of re-routing.\u003cbr\u003e\u003cbr\u003eWe are updating our monitors and tuning our failover mechanism to more quickly mitigate issues like this in the future.",
"name": "Copilot Code Completions partially unavailable",
"timestamp": "Nov \u003cvar data-var='date'\u003e5\u003c/var\u003e, \u003cvar data-var='time'\u003e22:56\u003c/var\u003e - \u003cvar data-var='time'\u003e23:26\u003c/var\u003e UTC"
},
{
"code": "0xpyl7hs9hm6",
"impact": "none",
"message": "On November 4, 2025, GitHub Enterprise Importer experienced a period of degraded migration performance and elevated error rates between 18:04 UTC and 23:36 UTC. During this interval customers queueing and running migrations experienced prolonged queue times and slower processing.\n\nThe degradation was ultimately connected to higher than normal system load, once load was reduced error rates returned to normal. The investigation is ongoing to pinpoint the precise root cause and prevent future recurrence.\n\nLong-term work is planned to strengthen system resilience under high load and promote better visibility into migration status for customers.",
"name": "Incident With GitHub Enterprise Importer",
"timestamp": "Nov \u003cvar data-var='date'\u003e4\u003c/var\u003e, \u003cvar data-var='time'\u003e22:00\u003c/var\u003e - \u003cvar data-var='time'\u003e22:00\u003c/var\u003e UTC"
},
{
"code": "y8hlsmxtgf0w",
"impact": "minor",
"message": "On November 3, 2025, between 14:10 UTC and 19:20 UTC, GitHub Packages experienced degraded performance, resulting in failures for 0.5% of Nuget package download requests. The incident resulted from an unexpected change in usage patterns affecting rate limiting infrastructure in the Packages service.\u003cbr\u003e\u003cbr\u003eWe mitigated the issue by scaling up services and refining our rate limiting implementation to ensure more consistent and reliable service for all users. To prevent similar problems, we are enhancing our resilience to shifts in usage patterns, improving capacity planning, and implementing better monitoring to accelerate detection and mitigation in the future.",
"name": "Incident with Packages",
"timestamp": "Nov \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e14:33\u003c/var\u003e - \u003cvar data-var='time'\u003e19:20\u003c/var\u003e UTC"
},
{
"code": "xkvk1yhmqfdl",
"impact": "minor",
"message": "On November 1, 2025, between 2:30 UTC and 6:14 UTC, Actions workflows could not be triggered manually from the UI. This impacted all customers queueing workflows from the UI for most of the impact window. The issue was caused by a faulty code change in the UI, which was promptly reverted once the impact was identified. Detection was delayed due to an alerting gap for UI breaks in this area when all underlying APIs are still healthy. We are implementing enhanced alerting and additional automated tests to prevent similar regressions and reduce detection time in the future.",
"name": "Incident with using workflow_dispatch for Actions",
"timestamp": "Nov \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e04:43\u003c/var\u003e - \u003cvar data-var='time'\u003e06:14\u003c/var\u003e UTC"
}
],
"name": "November",
"year": 2025
},
{
"incidents": [
{
"code": "6hygvwpw2vr3",
"impact": "minor",
"message": "On October 30th we shipped a change that broke 3 links in the \"Solutions\" dropdown of the marketing navigation seen on https://github.com/home. We noticed internally the broken links and declared an incident so our users would know no other functionality was impacted. We were able to revert a change and are evaluating our testing and rollout processes to prevent future incidents like these.",
"name": "Disruption with some GitHub services",
"timestamp": "Oct \u003cvar data-var='date'\u003e30\u003c/var\u003e, \u003cvar data-var='time'\u003e22:47\u003c/var\u003e - \u003cvar data-var='time'\u003e23:00\u003c/var\u003e UTC"
},
{
"code": "4jxdz4m769gy",
"impact": "major",
"message": "On October 29th, 2025 between 14:07 UTC and 23:15 UTC, multiple GitHub services were degraded due to a broad outage in one of our service providers:\u003cbr\u003e\u003cbr\u003e- Users of Codespaces experienced failures connecting to new and existing Codespaces through VSCode Desktop or Web. On average the Codespace connection error rate was 90% and peaked at 100% across all regions throughout the incident period.\u003cbr\u003e- GitHub Actions larger hosted runners experienced degraded performance, with 0.5% of overall workflow runs and 9.8% of larger hosted runner jobs failing or not starting within 5 minutes. These recovered by 20:40 UTC.\u003cbr\u003e- The GitHub Enterprise Importer service was degraded, with some users experiencing migration failures during git push operations and most users experiencing delayed migration processing.\u003cbr\u003e- Initiation of new trials for GitHub Enterprise Cloud with Data Residency were also delayed during this time.\u003cbr\u003e- Copilot Metrics via the API could not access the downloadable link during this time. There were approximately 100 requests during the incident that would have failed the download. Recovery began around 20:25 UTC.\u003cbr\u003e\u003cbr\u003eWe were able to apply a number of mitigations to reduce impact over the course of the incident, but we did not achieve 100% recovery until our service provider’s incident was resolved.\u003cbr\u003e\u003cbr\u003eWe are working to reduce critical path dependencies on the service provider and gracefully degrade experiences where possible so that we are more resilient to future dependency outages.",
"name": "Experiencing connection issues across Actions, Codespaces, and possibly other services",
"timestamp": "Oct \u003cvar data-var='date'\u003e29\u003c/var\u003e, \u003cvar data-var='time'\u003e16:17\u003c/var\u003e - \u003cvar data-var='time'\u003e23:15\u003c/var\u003e UTC"
},
{
"code": "pch0flk719dj",
"impact": "minor",
"message": "A cloud resource used by the Copilot bing-search tool was deleted as part of a resource cleanup operation. Once this was discovered, the resource was recreated. Going forward, more effective monitoring will be put in place to catch this issue earlier.",
"name": "Disruption with Copilot Bing search tool",
"timestamp": "Oct \u003cvar data-var='date'\u003e29\u003c/var\u003e, \u003cvar data-var='time'\u003e21:34\u003c/var\u003e - \u003cvar data-var='time'\u003e21:49\u003c/var\u003e UTC"
},
{
"code": "jlhnszknd9pj",
"impact": "minor",
"message": "From October 28th at 16:03 UTC until 17:11 UTC, the Copilot service experienced degradation due to an infrastructure issue which impacted the Claude Haiku 4.5 model, leading to a spike in errors affecting 1% of users. No other models were impacted. The incident was caused due to an outage with an upstream provider. We are working to improve redundancy during future occurrences.",
"name": "Inconsistent results when using the Haiku 4.5 model",
"timestamp": "Oct \u003cvar data-var='date'\u003e28\u003c/var\u003e, \u003cvar data-var='time'\u003e16:39\u003c/var\u003e - \u003cvar data-var='time'\u003e17:11\u003c/var\u003e UTC"
},
{
"code": "t6fcxny6yf14",
"impact": "minor",
"message": "Between October 23, 2025 19:27:29 UTC and October 27, 2025 17:42:42 UTC, users experienced timeouts when viewing repository landing pages. We observed the timeouts for approximately 5,000 users across less than 1,000 repositories including forked repositories. The impact was limited to logged in users accessing repositories in organizations with more than 200,000 members. Forks of repositories from affected large organizations were also impacted. Git operations were functional throughout this period.\u003cbr\u003e\u003cbr\u003eThis was caused by feature flagged changes impacting organization membership. The changes caused unintended timeouts for organization membership count evaluations which led to repository landing pages not loading.\u003cbr\u003e\u003cbr\u003eThe flag was turned off and a fix addressing the timeouts was deployed, including additional optimizations to better support organizations of this size. We are reviewing related areas and will continue to monitor for similar performance regressions.",
"name": "Disruption with viewing some repository pages from large organizations",
"timestamp": "Oct \u003cvar data-var='date'\u003e27\u003c/var\u003e, \u003cvar data-var='time'\u003e16:25\u003c/var\u003e - \u003cvar data-var='time'\u003e17:51\u003c/var\u003e UTC"
},
{
"code": "jkll48jj78zv",
"impact": "none",
"message": "On UTC Oct 24 2:55 - 3:15 AM, githubstatus.com was unreachable due to service interruption with our status page provider. \nDuring this time, GitHub systems were not experiencing any outages or disruptions.\nWe are working our vendor to understand how to improve availability of githubstatus.com.",
"name": "githubstatus.com was unavailable UTC 2025 Oct 24 02:55 to 03:13",
"timestamp": "Oct \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e14:17\u003c/var\u003e - \u003cvar data-var='time'\u003e14:17\u003c/var\u003e UTC"
},
{
"code": "n7hf73qtpz2l",
"impact": "minor",
"message": "From Oct 22, 2025 15:00 UTC to Oct 24, 2025 14:30 UTC git operations via SSH saw periods of increased latency and failed requests, with failure rates ranging from 1.5% to a single spike of 15%. Git operations over http were not affected. This was due to resource exhaustion on our backend ssh servers. \u003cbr\u003e\u003cbr\u003eWe mitigated the incident by increasing the available resources for ssh connections. We are improving the observability and dynamic scalability of our backend to prevent issues like this in the future.",
"name": "git operations over ssh seeing increased latency on github.com",
"timestamp": "Oct \u003cvar data-var='date'\u003e24\u003c/var\u003e, \u003cvar data-var='time'\u003e09:31\u003c/var\u003e - \u003cvar data-var='time'\u003e10:10\u003c/var\u003e UTC"
},
{
"code": "8vql81b3xcgq",
"impact": "minor",
"message": "On October 23, 2025, between 15:54 UTC and 19:20 UTC, GitHub Actions larger hosted runners experienced degraded performance, with 1.4% of overall workflow runs and 29% of larger hosted runner jobs failing to start or timing out within 5 minutes.\u003cbr\u003e\u003cbr\u003eThe full set of contributing factors is still under investigation, but the customer impact was due to database performance degradation, triggered by routine database changes causing a load profile that triggered a bug in the underlying database platform used for larger runners.\u003cbr\u003e\u003cbr\u003eImpact was mitigated through a combination of scaling up the database and reducing load. We are working with partners to resolve the underlying bug and have paused similar database changes until it is resolved.",
"name": "Incident with Actions - Larger hosted runners",
"timestamp": "Oct \u003cvar data-var='date'\u003e23\u003c/var\u003e, \u003cvar data-var='time'\u003e16:33\u003c/var\u003e - \u003cvar data-var='time'\u003e20:25\u003c/var\u003e UTC"
},
{
"code": "dlvf3sfmz7dm",
"impact": "minor",
"message": "On October 22, 2025, between 14:06 UTC and 15:17 UTC, less than 0.5% of web users experienced intermittent slow page loads on GitHub.com. During this time, API requests showed increased latency, with up to 2% timing out. \u003cbr\u003e\u003cbr\u003eThe issue was caused by elevated loads on one of our databases caused by a poorly performing query, which impacted performance for a subset of requests.\u003cbr\u003e\u003cbr\u003eWe identified the source of the load and optimized the query to restore normal performance. We’ve added monitors for early detection for query performance, and we continue to monitor the system closely to ensure ongoing stability. \u003cbr\u003e",
"name": "Incident with API Requests",
"timestamp": "Oct \u003cvar data-var='date'\u003e22\u003c/var\u003e, \u003cvar data-var='time'\u003e14:29\u003c/var\u003e - \u003cvar data-var='time'\u003e15:53\u003c/var\u003e UTC"
},
{
"code": "v61nk2fpysnq",
"impact": "minor",
"message": "On October 21, 2025, between 13:30 and 17:30 UTC, GitHub Enterprise Cloud Organization SAML Single Sign-On experienced degraded performance. Customers may have been unable to successfully authenticate into their GitHub Organizations during this period. Organization SAML recorded a maximum of 0.4% of SSO requests failing during this timeframe.\u003cbr\u003e\u003cbr\u003eThis incident stemmed from a failure in a read replica database partition responsible for storing license usage information for GitHub Enterprise Cloud Organizations. This partition failure resulted in users from affected organizations, whose license usage information was stored on this partition, being unable to access SSO during the aforementioned window. A successful SSO requires an available license for the user who is accessing a GitHub Enterprise Cloud Organization backed by SSO.\u003cbr\u003eThe failing partition was subsequently taken out of service, thereby mitigating the issue. \u003cbr\u003e\u003cbr\u003eRemedial actions are currently underway to ensure that a read replica failure does not compromise the overall service availability.\u003cbr\u003e",
"name": "Disruption with some GitHub services",
"timestamp": "Oct \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e16:00\u003c/var\u003e - \u003cvar data-var='time'\u003e17:39\u003c/var\u003e UTC"
},
{
"code": "qqd6b1xb63tq",
"impact": "minor",
"message": "On October 21, 2025, between 07:55 UTC and 12:20 UTC, GitHub Actions experienced degraded performance. During this time, 2.11% workflow runs failed to start within 5 minutes, with an average delay of 8.2 minutes. The root cause was increased latency on a node in one of our Redis clusters, triggered by resource contention after a patching event became stuck. \u003cbr\u003e\u003cbr\u003eRecovery began once the patching process was unstuck and normal connectivity to the Redis cluster was restored at 11:45 UTC, but it took until 12:20 UTC to clear the backlog of queued work. We are implementing safeguards to prevent this failure mode and enhancing our monitoring to detect and address problems like this more quickly in the future.",
"name": "Incident with Actions",
"timestamp": "Oct \u003cvar data-var='date'\u003e21\u003c/var\u003e, \u003cvar data-var='time'\u003e09:12\u003c/var\u003e - \u003cvar data-var='time'\u003e12:28\u003c/var\u003e UTC"
},
{
"code": "9klytnsknx20",
"impact": "minor",
"message": "From October 20th at 14:10 UTC until 16:40 UTC, the Copilot service experienced degradation due to an infrastructure issue which impacted the Grok Code Fast 1 model, leading to a spike in errors affecting 30% of users. No other models were impacted. The incident was caused due to an outage with an upstream provider.",
"name": "Disruption with Grok Code Fast 1 in Copilot",
"timestamp": "Oct \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e14:46\u003c/var\u003e - \u003cvar data-var='time'\u003e16:40\u003c/var\u003e UTC"
},
{
"code": "krd9y2m82fbn",
"impact": "major",
"message": "On October 20, 2025, between 08:05 UTC and 10:50 UTC the Codespaces service was degraded, with users experiencing failures creating new codespaces and resuming existing ones. On average, the error rate for codespace creation was 39.5% and peaked at 71% of requests to the service during the incident window. Resume operations averaged 23.4% error rate with a peak of 46%. This was due to a cascading failure triggered by an outage in a 3rd-party dependency required to build devcontainer images.\u003cbr\u003e\u003cbr\u003eThe impact was mitigated when the 3rd-party dependency recovered.\u003cbr\u003e\u003cbr\u003eWe are investigating opportunities to make this dependency not a critical path for our container build process and working to improve our monitoring and alerting systems to reduce our time to detection of issues like this one in the future.",
"name": "Codespaces creation failling",
"timestamp": "Oct \u003cvar data-var='date'\u003e20\u003c/var\u003e, \u003cvar data-var='time'\u003e08:56\u003c/var\u003e - \u003cvar data-var='time'\u003e11:01\u003c/var\u003e UTC"
},
{
"code": "vs7qnzbydz2p",
"impact": "major",
"message": "On October 17th, 2025, between 12:51 UTC and 14:01 UTC, mobile push notifications failed to be delivered for a total duration of 70 minutes. This affected github.com and GitHub Enterprise Cloud in all regions. The disruption was related to an erroneous configuration change to cloud resources used for mobile push notification delivery.\u003cbr\u003e\u003cbr\u003eWe are reviewing our procedures and management of these cloud resources to prevent such an incident in the future.",
"name": "Disruption with push notifications",
"timestamp": "Oct \u003cvar data-var='date'\u003e17\u003c/var\u003e, \u003cvar data-var='time'\u003e13:11\u003c/var\u003e - \u003cvar data-var='time'\u003e14:12\u003c/var\u003e UTC"
},
{
"code": "g8rmr5p85tzv",
"impact": "minor",
"message": "On October 14th, 2025, between 18:26 UTC and 18:57 UTC a subset of unauthenticated requests to the commit endpoint for certain repositories received 503 errors. During the event, the average error rate was 3%, peaking at 3.5% of total requests.\u003cbr\u003e\u003cbr\u003eThis event was triggered by a recent configuration change and some traffic pattern shifts on the service. We were alerted of the issue immediately and made changes to the configuration in order to mitigate the problem. We are working on automatic mitigation solutions and better traffic handling in order to prevent issues like this in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Oct \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e18:26\u003c/var\u003e - \u003cvar data-var='time'\u003e18:57\u003c/var\u003e UTC"
},
{
"code": "fpbpfb33dw5s",
"impact": "minor",
"message": "On Oct 14th, 2025, between 13:34 UTC and 16:00 UTC the Copilot service was degraded for GPT-5 mini model. On average, 18% of the requests to GPT-5 mini failed due to an issue with our upstream provider.\u003cbr\u003e\u003cbr\u003eWe notified the upstream provider of the problem as soon as it was detected and mitigated the issue by failing over to other providers. The upstream provider has since resolved the issue.\u003cbr\u003e\u003cbr\u003eWe are working to improve our failover logic to mitigate similar upstream failures more quickly in the future.",
"name": "Disruption with GPT-5-mini in Copilot",
"timestamp": "Oct \u003cvar data-var='date'\u003e14\u003c/var\u003e, \u003cvar data-var='time'\u003e14:05\u003c/var\u003e - \u003cvar data-var='time'\u003e16:00\u003c/var\u003e UTC"
},
{
"code": "k7bhmjkblcwp",
"impact": "major",
"message": "On October 9th, 2025, between 14:35 UTC and 15:21 UTC, a network device in maintenance mode that was undergoing repairs was brought back into production before repairs were completed. Network traffic traversing this device experienced significant packet loss.\u003cbr\u003e\u003cbr\u003eAuthenticated users of the github.com UI experienced increased latency during the first 5 minutes of the incident. API users experienced up to 7.3% error rates, after which it stabilized to about 0.05% until mitigated. Actions service experienced 24% of runs being delayed for an average of 13 minutes. Large File Storage (LFS) requests experienced minimally increased error rate, with 0.038% of requests erroring.\u003cbr\u003e\u003cbr\u003eTo prevent similar issues, we are enhancing the validation process for device repairs of this category.",
"name": "Incident with Webhooks",
"timestamp": "Oct \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e14:45\u003c/var\u003e - \u003cvar data-var='time'\u003e16:40\u003c/var\u003e UTC"
},
{
"code": "kk58nfytx0c7",
"impact": "minor",
"message": "Between 13:39 UTC and 13:42 UTC on Oct 9, 2025, around 2.3% of REST API calls and 0.4% Web traffic were impacted due to the partial rollout of a new feature that had more impact on one of our primary databases than anticipated. When the feature was partially rolled out it performed an excessive number of writes per request which caused excessive latency for writes from other API and Web endpoints and resulted in 5xx errors to customers. \u003cbr\u003e\u003cbr\u003eThe issue was identified by our automatic alerting and reverted by turning down the percentage of traffic to the new feature, which led to recovery of the data cluster and services. \u003cbr\u003e\u003cbr\u003eWe are working to improve the way we roll out new features like this and move the specific writes from this incident to a storage solution more suited to this type of activity. We have also optimized this particular feature to avoid its rollout from having future impact on other areas of the site. We are also investigating how we can even more quickly identify issues like this.",
"name": "Multiple GitHub API endpoints are experiencing errors",
"timestamp": "Oct \u003cvar data-var='date'\u003e9\u003c/var\u003e, \u003cvar data-var='time'\u003e13:52\u003c/var\u003e - \u003cvar data-var='time'\u003e13:56\u003c/var\u003e UTC"
},
{
"code": "9f6pfjj0scrd",
"impact": "minor",
"message": "On October 7, 2025, between 7:48 PM UTC and October 8, 12:05 AM UTC (approximately 4 hours and 17 minutes), the audit log service was degraded, creating a backlog and delaying availability of new audit log events. The issue originated in a third-party dependency.\u003cbr\u003e\u003cbr\u003eWe mitigated the incident by working with the vendor to identify and resolve the issue. Write operations recovered first, followed by the processing of the accumulated backlog of audit log events.\u003cbr\u003e\u003cbr\u003eWe are working to improve our monitoring and alerting for audit log ingestion delays and strengthen our incident response procedures to reduce our time to detection and mitigation of issues like this one in the future.",
"name": "Disruption with some GitHub services",
"timestamp": "Oct \u003cvar data-var='date'\u003e7\u003c/var\u003e, \u003cvar data-var='time'\u003e19:48\u003c/var\u003e - Oct \u003cvar data-var='date'\u003e8\u003c/var\u003e, \u003cvar data-var='time'\u003e00:05\u003c/var\u003e UTC"
},
{
"code": "34wtrn4nngwk",
"impact": "minor",
"message": "\u003cp\u003eOn October 3rd, between approximately 10:00 PM and 11:30 Eastern, the Copilot service experienced degradation due to an issue with our upstream provider. Users encountered elevated error rates when using the following Claude models: Claude Sonnet 3.7, Claude Opus 4, Claude Opus 4.1, Claude Sonnet 4, and Claude Sonnet 4.5. No other models were impacted.\u003c/p\u003e\u003cp\u003eThe issue was mitigated by temporarily disabling affected endpoints while our provider resolved the upstream issue. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.\u003c/p\u003e",
"name": "Incident with Copilot",
"timestamp": "Oct \u003cvar data-var='date'\u003e3\u003c/var\u003e, \u003cvar data-var='time'\u003e02:41\u003c/var\u003e - \u003cvar data-var='time'\u003e03:47\u003c/var\u003e UTC"
},
{
"code": "l94jr9wnhs4r",
"impact": "minor",
"message": "Between October 1st, 2025 at 1 AM UTC and October 2nd, 2025 at 10:33 PM UTC, the Copilot service experienced a degradation of the Gemini 2.5 Pro model due to an issue with our upstream provider. Before 15:53 UTC on October 1st, users experienced higher error rates with large context requests while using Gemini 2.5 Pro. After 15:53 UTC and until 10:33 PM UTC on October 2nd, requests were restricted to smaller context windows when using Gemini 2.5. Pro. No other models were impacted.\u003cbr\u003e\u003cbr\u003eThe issue was resolved by a mitigation put in place by our provider. GitHub is collaborating with our provider to enhance communication and improve the ability to reproduce issues with the aim to reduce resolution time.",
"name": "Degraded Gemini 2.5 Pro experience in Copilot",
"timestamp": "Oct \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e16:43\u003c/var\u003e - Oct \u003cvar data-var='date'\u003e2\u003c/var\u003e, \u003cvar data-var='time'\u003e22:33\u003c/var\u003e UTC"
},
{
"code": "071h21gptcp0",
"impact": "minor",
"message": "On October 1, 2025 between 07:00 UTC and 17:20 UTC, Mac hosted runner capacity for Actions was degraded, leading to timed out jobs and long queue times. On average, the error rate was 46% and peaked at 96% of requests to the service. XL and Intel runners recovered by 10:10 UTC, with the other types taking longer to recover.\u003cbr\u003e\u003cbr\u003eThe degraded capacity was triggered by a scheduled event at 07:00 UTC that led to a permission failure on Mac runner hosts, blocking reimage operations. The permission issue was resolved by 9:41 UTC, but the recovery of available runners took longer than expected due to a combination of backoff logic slowing backend operations and some hosts needing state resets.\u003cbr\u003e\u003cbr\u003eWe deployed changes immediately following the incident to address the scheduled event and ensure that similar failures will not block critical operations in the future. We are also working to reduce the end-to-end time for self-healing of offline hosts for quicker full recovery of future capacity or host events.",
"name": "Degraded Performance for GitHub Actions MacOS Runners",
"timestamp": "Oct \u003cvar data-var='date'\u003e1\u003c/var\u003e, \u003cvar data-var='time'\u003e07:59\u003c/var\u003e - \u003cvar data-var='time'\u003e16:55\u003c/var\u003e UTC"
}
],
"name": "October",
"year": 2025
}
],
"start_time": "2025-10-01T00:00:00Z",
"time_zone": "UTC"
}
]
}