G

GitHub Outage History

Past incidents and downtime events

Complete history of GitHub outages, incidents, and service disruptions. Showing 50 most recent incidents.

October 2026(8 incidents)

criticalresolvedOct 7, 05:17 PM — Resolved Oct 7, 06:04 PM

Incident with Git Operations, Issues, Actions and Pull Requests

5 updates
resolvedOct 7, 06:04 PM

This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

monitoringOct 7, 06:00 PM

All systems have recovered. We have applied mitigations and do not anticipate another recurrence.

monitoringOct 7, 05:27 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingOct 7, 05:25 PM

We observed widespread impact across GitHub services from 16:52 UTC to 17:01 UTC, and are investigating the root cause. All systems have recovered. This incident is a reoccurrence of the incident from earlier in the day https://www.githubstatus.com/incidents/djlmxz2zd0j7. Durable mitigation is in progress. We will provide another update in 30 minutes.

investigatingOct 7, 05:17 PM

We are investigating reports of degraded performance for Actions, Git Operations, Issues and Pull Requests

criticalresolvedOct 7, 03:14 PM — Resolved Oct 7, 04:25 PM

Incident with Git Operations, Pull Requests and Actions

20 updates
resolvedOct 7, 04:25 PM

This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

monitoringOct 7, 04:23 PM

Git Operations, Pull Requests and Actions are fully recovered and stable. The root cause has been identified and mitigations are being applied to prevent a reoccurrence.

monitoringOct 7, 03:58 PM

We are seeing full recovery across all systems including Git Operations, Pull Requests and Actions. We are continuing to investigate the root cause of the disruption between 15:06 UTC and 15:16 UTC, and will share another update shortly.

monitoringOct 7, 03:56 PM

The degradation has been mitigated. We are monitoring to ensure stability.

monitoringOct 7, 03:49 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingOct 7, 03:49 PM

The degradation affecting Git Operations and Webhooks has been mitigated. We are monitoring to ensure stability.

investigatingOct 7, 03:47 PM

The degradation affecting Issues has been mitigated. We are monitoring to ensure stability.

investigatingOct 7, 03:41 PM

We are observing partial recovery across all systems. We are continuing to monitor and investigate the root cause of the disruption.

investigatingOct 7, 03:35 PM

Git Operations is experiencing degraded performance. We are continuing to investigate.

investigatingOct 7, 03:32 PM

Actions is operating normally.

investigatingOct 7, 03:31 PM

The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.

investigatingOct 7, 03:27 PM

Webhooks is experiencing degraded performance. We are continuing to investigate.

investigatingOct 7, 03:27 PM

Git Operations is operating normally.

investigatingOct 7, 03:25 PM

Pull Requests is experiencing degraded performance. We are continuing to investigate.

investigatingOct 7, 03:24 PM

We observed widespread impact across GitHub services between 15:06 UTC and 15:16 UTC, and are investigating the root cause. We are seeing all systems recover, and will provide another update shortly.

investigatingOct 7, 03:21 PM

Git Operations is experiencing degraded performance. We are continuing to investigate.

investigatingOct 7, 03:16 PM

Webhooks is experiencing degraded availability. We are continuing to investigate.

investigatingOct 7, 03:15 PM

Issues is experiencing degraded performance. We are continuing to investigate.

investigatingOct 7, 03:14 PM

Webhooks is experiencing degraded performance. We are continuing to investigate.

investigatingOct 7, 03:14 PM

We are investigating reports of degraded availability for Actions, Git Operations and Pull Requests

criticalresolvedOct 6, 07:57 PM — Resolved Oct 6, 11:40 PM

Several services are degraded

9 updates
resolvedOct 6, 11:40 PM

This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

investigatingOct 6, 10:40 PM

We are deploying a mitigation and plan to resume migrations shortly thereafter

investigatingOct 6, 09:29 PM

We are working towards a mitigation that will allow us to resume processing enterprise migrations.

investigatingOct 6, 08:57 PM

All services are functioning normally except for enterprise migrations, which are still paused. They will resume soon.

monitoringOct 6, 08:26 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingOct 6, 08:22 PM

A number of services including issues, pull requests, and webhooks, were briefly degraded, but have now recovered. Enterprise migrations are paused while we validate causes and mitigations.

investigatingOct 6, 08:12 PM

Pull Requests is operating normally.

investigatingOct 6, 08:04 PM

Webhooks is operating normally.

investigatingOct 6, 07:57 PM

We are investigating reports of degraded performance for Pull Requests and Webhooks

majorresolvedOct 5, 11:47 PM — Resolved Oct 6, 01:32 AM

Disruption with some GitHub services

3 updates
resolvedOct 6, 01:32 AM

This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

investigatingOct 6, 12:05 AM

Users may experience issues accessing organization and enterprise billing and licensing pages. We are deploying a mitigation and monitoring for recovery.

investigatingOct 5, 11:47 PM

We are investigating reports of impacted performance for some GitHub services.

criticalresolvedOct 5, 07:11 PM — Resolved Oct 5, 10:49 PM

Incident with Actions

12 updates
resolvedOct 5, 10:49 PM

On October 5, 2026, between 18:48 and 22:49 UTC, GitHub Actions and Hosted Runners experienced degraded performance. Over the full incident, 14.3% of workflow runs using GitHub-hosted runners and 26.5% of individual GitHub-hosted jobs did not start within five minutes. During the roughly 90-minute period of highest impact, 47.0% of workflow runs using GitHub-hosted runners failed and 71.9% of GitHub-hosted jobs did not start within five minutes, while existing Pages sites continued to serve normally. A portion of customers also encountered intermittent errors in repository lists, licensing and billing pages, and some Copilot features. Self-hosted runners were not affected. The incident was caused by a partial networking failure that disrupted connectivity between one of our on-premises datacenter regions and a subset of cloud-hosted regional database and storage services. We mitigated the incident by shifting database and storage traffic to healthy regions and endpoints and increasing capacity for affected compute services. GitHub Actions and Hosted Runners recovered by 21:54 UTC, and all affected services recovered by 22:49 UTC. We are partnering with our cloud provider to understand the network failure and identify improvements to network-path resilience, detection, and recovery for similar incidents.

investigatingOct 5, 10:40 PM

Pages is operating normally.

investigatingOct 5, 09:54 PM

Actions is operating normally.

investigatingOct 5, 09:32 PM

We have applied mitigations to address GitHub Actions failures. Queued jobs are clearing, new jobs are not delayed, and we expect full runner recovery shortly. We continue to monitor for full recovery and are working to mitigate ongoing issues affecting other services, including access to repository lists, licensing, and billing pages.

investigatingOct 5, 09:31 PM

Actions is experiencing degraded performance. We are continuing to investigate.

investigatingOct 5, 09:22 PM

Pages is experiencing degraded performance. We are continuing to investigate.

investigatingOct 5, 09:09 PM

In addition to ongoing issues with Actions and Hosted Runners, some customers are also unable to access repository list, licensing and billing pages within their GitHub accounts. We continue to investigate and are working on mitigation.

investigatingOct 5, 08:47 PM

Actions is experiencing degraded availability. We are continuing to investigate.

investigatingOct 5, 08:39 PM

We’re continuing to investigate job failures and delays affecting GitHub-hosted runner assignment and workflow start times. Our teams are actively working to mitigate the impact and will provide another update as we learn more.

investigatingOct 5, 07:50 PM

We’re still investigating delays in assigning GitHub-hosted runners, affecting workflow start times across multiple runner configurations. Our teams are working to mitigate the impact; we’ll share another update as we learn more.

investigatingOct 5, 07:15 PM

We’re investigating an issue causing delays when assigning GitHub-hosted runners to Actions jobs. Some workflows may take longer to start across runner configurations. Our teams are working to mitigate the issue, and we’ll provide an update as we learn more.

investigatingOct 5, 07:11 PM

We are investigating reports of degraded performance for Actions

minorresolvedOct 1, 02:47 PM — Resolved Oct 1, 05:56 PM

Actions Job Delays

10 updates
resolvedOct 1, 05:56 PM

This incident has been resolved. Thank you for your patience and understanding as we addressed this issue. A detailed root cause analysis will be shared as soon as it is available.

investigatingOct 1, 05:50 PM

GitHub Actions experienced degraded performance for some hosted runners due to throttling within an upstream Azure dependency. Service capacity has recovered, and we are continuing to monitor while working with Azure on the underlying condition.

investigatingOct 1, 04:48 PM

We are currently applying a mitigation and anticipate recovery within thirty minutes.

investigatingOct 1, 04:10 PM

We have identified an issue with our upstream provider which is causing Actions requests to 429 which is creating the delays. We have escalated to the owning team and are investigating how to mitigate the 429s.

investigatingOct 1, 03:29 PM

We are seeing a reoccurrence in run-start delays, and are continuing to investigate to issue. Customers will potentially experience delays of up to ten minutes.

investigatingOct 1, 03:20 PM

Actions is experiencing degraded performance. We are continuing to investigate.

monitoringOct 1, 03:10 PM

The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

investigatingOct 1, 03:02 PM

Run-start delays on Ubuntu runners have been resolved. We are investigating run-start delays on Windows runners and will share more information as it becomes available.

investigatingOct 1, 02:54 PM

We have identified the cause of increased Actions run start delays on Ubuntu runners and are actively deploying a fix. Customers may continue to experience intermittent delays while the mitigation rolls out and service metrics return to normal.

investigatingOct 1, 02:47 PM

We are investigating reports of degraded performance for Actions

minorresolvedOct 1, 01:37 PM — Resolved Oct 1, 01:57 PM

Elevated request latency

4 updates
resolvedOct 1, 01:57 PM

On October 1, 2026, between 13:04 UTC and 13:34 UTC, users experienced two periods of slow page loads and intermittent request failures on GitHub.com. On average, the error rate was 0.52% and peaked at 6.13%. An unusually high volume of incoming traffic placed additional load on our traffic-handling infrastructure, delaying requests before they reached our application servers. Performance returned to normal as the traffic subsided, and we have since applied additional traffic filtering to protect against similar traffic. We are improving alerting for delays in our traffic-handling infrastructure and strengthening our automated traffic protections to reduce our time to detection and mitigation.

monitoringOct 1, 01:51 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingOct 1, 01:39 PM

We are investigating recurrent periods of elevated latency affecting web requests. We’ll share updates as more information becomes available.

investigatingOct 1, 01:37 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedOct 1, 02:00 AM — Resolved Oct 1, 02:00 AM

[Retroactive] Actions workflow run failures after deployment gate approvals

1 update
resolvedOct 1, 11:34 AM

On October 1, around 02:00 UTC, an isolated infrastructure failure caused GitHub Actions to lose execution state for a small number of existing workflow runs. Affected runs may remain stuck, fail deployment approvals, or return errors when cancelled. Service connectivity has recovered, but the lost state cannot be restored by retrying an approval. If you're affected, you can trigger a new run or contact GitHub Support with links to your stuck runs so we can unblock them. Once cleared, select "Re-run all jobs." This preserves the workflow run ID but starts a new attempt, rebuilds artifacts, and requires fresh deployment approvals. Before rerunning, check whether any deployment steps already completed to avoid repeating changes.

September 2026(14 incidents)

criticalresolvedSep 28, 09:16 PM — Resolved Sep 28, 10:08 PM

Copilot Code Review is unable to complete reviews

4 updates
resolvedSep 28, 10:08 PM

On September 28, 2026, between 19:23 UTC and 22:08 UTC the Copilot Code Review service was degraded and pull request reviews did not complete in all environments. On average, approximately 70% of requested reviews did not complete. Reviews requested during this window were not automatically retried, and users needed to re-request a review from their pull request. This was due to a change in a dependent service that stopped delivering data Copilot Code Review needs to complete a review. Review completion is an inherently delayed signal, because each review job takes time to start and finish, so it took longer to detect the impact.We mitigated the incident by reverting the change.We are working to add a faster signal for reviews that fail to complete, add deployment validation that confirms a Copilot Code Review completes end to end, improve how we track internal dependencies between services, and make our deployment pipeline more reliable to reduce our time to detection and mitigation of issues like this one in the future.

investigatingSep 28, 10:08 PM

Copilot Code review has recovered, for reviews that were previously not successful you can re-request review from your PR.

investigatingSep 28, 09:23 PM

Revert of the impacted change is in flight, expect full recovery once the deployment is done.

investigatingSep 28, 09:16 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedSep 24, 04:51 PM — Resolved Sep 24, 08:41 PM

Disruption with billing information updates

7 updates
resolvedSep 24, 08:41 PM

Beginning on September 22, 2026 at 11:51 UTC, approximately 1,700 users experienced delays and errors when adding or updating billing address information because requests to the address validation service timed out or failed. Within the affected address validation flow, requests failed at an average rate of 69.3%, reaching 100% at peak. An operating system upgrade exposed a compatibility issue in an adapter used by our HTTP client library, causing requests to stall. A subsequent change to shorten request timeouts caused stalled requests to return connection errors and increased the failure rate. We mitigated the incident on September 24, 2026 at 20:24 UTC by changing the HTTP client used for address validation. We have added alerting to detect similar issues sooner and are extending the fix to other integrations.

monitoringSep 24, 08:24 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingSep 24, 08:10 PM

We have applied the mitigation and are seeing signs of recovery. We are continuing to monitor the system closely.

investigatingSep 24, 06:22 PM

We are continuing to work on mitigation. We will post another update in approximately one hour.

investigatingSep 24, 05:40 PM

We have identified the problem and are actively working on mitigation.

investigatingSep 24, 04:51 PM

We are investigating failures on billing information updates. Customers may be unable to create or update their billing information in the meantime.

investigatingSep 24, 04:51 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedSep 23, 10:11 AM — Resolved Sep 24, 04:55 AM

Incident across several services

16 updates
resolvedSep 24, 04:55 AM

Starting at 07:57 UTC on September 23, GitHub experienced elevated 500 and 404 responses across several application pages. This caused failures when installing GitHub Apps, creating organizations, and making some organization membership changes. Customers also experienced delayed label updates and stale search results in Projects. The infrastructure failure was isolated to our Azure Central US region. The API errors were mitigated by 10:58 UTC on September 23. Projects' processing continued to recover while an accumulated backlog was drained, and full service was restored at 04:55 UTC on September 24. The incident was caused by a failed planned maintenance operation on a primary database. Automated recovery initiated an emergency database failover, after which several replicas in the affected region were unable to resume replication correctly. This reduced available database capacity and caused the API errors and downstream Projects processing delays. We have mitigated the immediate failure mode. We are also improving maintenance safety checks, database failover handling, post-failover replica validation, and downstream processing resilience to reduce the likelihood and impact of similar incidents.

monitoringSep 24, 04:55 AM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingSep 24, 03:14 AM

We are continuing to process the backlog of issue label updates for Projects. Label changes may still be delayed. All other services are operating normally.

investigatingSep 24, 02:00 AM

We are continuing to process the backlog of issue label updates for Projects. Users may still see delays before label changes are reflected in Projects. All other services are operating normally.

investigatingSep 24, 12:16 AM

We've deployed a change intended to accelerate processing of the backlog of issue label updates in Projects. A sizable backlog still remains and we continue working through it. All other services are operating normally. We will provide another update within the next hour.

investigatingSep 23, 09:39 PM

Updates to issue labels may be delayed in being reflected in Projects. We are continuing to deploy a change that will accelerate processing of the backlog of label updates. All other services are available. We will provide another update within the next hour.

investigatingSep 23, 08:26 PM

We are preparing to deploy a change that will mitigate the impact.

investigatingSep 23, 06:42 PM

Continuing to investigate the lag that may be experienced in issue labels being accurately reflected in Projects. We are working on alternate solutions to process the backlog of label updates.

investigatingSep 23, 05:30 PM

We will post another update in approximately one hour to share our progress.

investigatingSep 23, 05:01 PM

Updates to issue labels may be delayed in being reflected in Projects by about ~10 minutes. We have added some capacity to work through the backlog more quickly, but it'll likely be a few hours to complete processing the full backlog of messages. All other services are available.

investigatingSep 23, 01:35 PM

Users may experience stale Project search results. We are working to increase indexing speed. All other services are available.

investigatingSep 23, 11:38 AM

We are seeing recovery for Projects. Users may experience stale search results for Projects while indexing catches up.

investigatingSep 23, 10:58 AM

The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.

investigatingSep 23, 10:57 AM

Database replicas have been restored. Org creation and the GitHub API are no longer degraded.

investigatingSep 23, 10:20 AM

Database replicas have detached. We're working to restore the database replicas. Users may experience issues beyond creating organizations and a degraded experience with the GitHub API and Projects.

investigatingSep 23, 10:11 AM

We are investigating reports of degraded performance for API Requests

minorresolvedSep 20, 10:13 PM — Resolved Sep 20, 11:22 PM

Incident with Pull Requests

4 updates
resolvedSep 20, 11:22 PM

On September 20, 2026, between 21:46 and 22:24 UTC the Pull Requests service was degraded and pull request merge and test-merge commits were created late, with delays reaching approximately four minutes at peak. Merge commits were delayed rather than lost. Because some Actions workflow runs start only after a pull request's merge commit is created, a subset of workflow runs for pull request events were also delayed. This was due to a routine repository maintenance job for an unusually large repository consuming nearly all of the memory on a single Git storage server, which left that server unable to serve the Git operations used to create merge commits. We mitigated the incident by removing the affected server from service at 22:18 UTC, after which the queued merge commits were created within six minutes. We have capped the memory a single repository maintenance job may consume so that one repository cannot exhaust a server, and we have improved monitoring and alerting on storage server health to reduce our time to detection and mitigation of issues like this one in the future.

monitoringSep 20, 10:32 PM

A git fileserver issue caused a brief delay in creating some merge commits - we've isolated the underlying server and already observed recovery.

monitoringSep 20, 10:27 PM

The degradation affecting Pull Requests has been mitigated. We are monitoring to ensure stability.

investigatingSep 20, 10:13 PM

We are investigating reports of degraded performance for Pull Requests

minorresolvedSep 17, 08:59 PM — Resolved Sep 17, 09:49 PM

Elevated rate of errors for OpenAI models provided by Copilot

3 updates
resolvedSep 17, 09:49 PM

Between 20:26 and 21:17 UTC on September 17, 2026, GitHub Copilot experienced degradation affecting several GPT models, including GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol, GPT-5.3-Codex, and GPT-6 Astra. Users encountered elevated error rates when using these models.The degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring and coordinated with the provider. Our automated model-warning system activated in-product warnings for affected models during the incident. Service returned to normal after the provider implemented a mitigation.

monitoringSep 17, 09:39 PM

We are experiencing degraded availability for GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol, GPT-6 Astra, GPT-5.3-Codex in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. The provider is working to mitigate the problem and we are monitoring recovery. We recommend choosing another model or selecting 'Auto' to continue using Copilot.

investigatingSep 17, 08:59 PM

We are investigating reports of degraded performance for Copilot AI Model Providers

majorresolvedSep 16, 07:20 AM — Resolved Sep 16, 05:48 PM

Degradation with Gemini 3.8 Flash

5 updates
resolvedSep 16, 05:48 PM

On September 16, 2026, between 04:40 and 11:45 UTC, the Gemini 3.8 Flash model in GitHub Copilot experienced degraded availability. Requests to this model failed at an average rate of 6.4%, and the impact was highest during peak traffic hours. Other Copilot models were not affected. Users could continue to work with a different model or with 'Auto'.The degradation was caused due to a capacity issue with an upstream model provider. Failure rates returned to normal as the daily traffic peak passed. We monitored the model until it was healthy and resolved the incident at 17:48 UTC. We are working to make our systems resilient to cover peak demand for all the Copilot models.

monitoringSep 16, 11:45 AM

The degradation affecting Copilot AI Model Providers has been mitigated. We are monitoring to ensure stability.

investigatingSep 16, 10:58 AM

Copilot AI Model Providers is experiencing degraded performance. We are continuing to investigate.

investigatingSep 16, 08:28 AM

Copilot AI Model Providers is experiencing degraded performance. We are continuing to investigate.

investigatingSep 16, 07:21 AM

We are investigating reports of degraded availability for Copilot AI Model Providers

minorresolvedSep 15, 07:11 PM — Resolved Sep 15, 08:00 PM

Disruption with some GitHub services

3 updates
resolvedSep 15, 08:00 PM

On September 15, 2026 between 15:30 and 20:00 UTC, some Copilot code reviews on pull requests failed to complete. The cause was increased latency in an internal caching service that GitHub Copilot Code Review relies on to coordinate its review jobs. This caused a timeout in lock acquisition, which interrupted the job. We reverted the change to the internal caching service and restored normal operation by 20:00 UTC.We sincerely apologize for the disruption.

monitoringSep 15, 07:48 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingSep 15, 07:11 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedSep 15, 09:47 AM — Resolved Sep 15, 11:17 AM

Disruption with some GitHub services

7 updates
resolvedSep 15, 11:17 AM

On September 15, 2026, between 05:45 and 09:50 UTC, the Claude Fable 5.1 model in GitHub Copilot experienced intermittently degraded availability, with an average error rate of 2.8%. During brief recurring 15 minute periods that recurred every ~45 minutes, availability for Claude Fable 5.1 dropped to a maximum of ~40% before recovering completely. Other Copilot models were not affected. Users could continue to work with a different model or with 'Auto'.The cause was an issue with an upstream model provider that intermittently rejected requests while overloaded. GitHub worked with the provider, who acknowledged and then resolved the underlying issue at 9:50 UTC, after which the model returned to constant normal availability. Once recovery was guaranteed, we resolved the incident at 11:17 UTC.To reduce the chance of recurrence and customer impact, GitHub is reviewing per-model availability alerting and automatic in-product fallback so that requests to a degraded model can shift to a healthy alternative more quickly.

investigatingSep 15, 11:17 AM

The issues with our upstream model provider have been resolved, and Claude Fable 5.1 is once again available in Copilot products and IDE surfaces.We will continue monitoring to ensure stability, but mitigation is complete.

investigatingSep 15, 11:09 AM

We continue to monitor intermittent errors affecting Claude Fable 5.1 in some Copilot products and integrated development environments. Customer-facing metrics have recovered, and we are awaiting confirmation from our upstream provider that the issue will not recur. Customers can select another model or Auto in the meantime.

investigatingSep 15, 10:32 AM

We continue to investigate intermittent errors affecting Claude Fable 5.1 in some Copilot products and integrated development environments. Customers can use another model or select Auto while we monitor the situation.

investigatingSep 15, 10:20 AM

Copilot AI Model Providers is experiencing degraded performance. We are continuing to investigate.

monitoringSep 15, 09:48 AM

We are investigating degraded availability for Claude Fable 5.1, affecting some Copilot products and integrated development environments. The issue is caused by a problem with our upstream model provider. Customers can use another model or select Auto while we investigate.

investigatingSep 15, 09:47 AM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedSep 14, 06:40 PM — Resolved Sep 14, 07:35 PM

Actions Larger Runner Jobs for some customers may be slow to start

3 updates
resolvedSep 14, 07:35 PM

On September 14, 2026, between 16:10 and 19:01 UTC, some customers using GitHub Actions larger runners experienced longer-than-normal wait times for jobs to start. During this period, 5.7% of larger-runner jobs were affected. A routine expansion of our compute capacity exposed a bug in how our provisioning system handled capacity records when selecting where to create runner virtual machines. This slowed the creation of new runners, leaving insufficient runner capacity to start affected jobs promptly. We restored normal provisioning by correcting the affected capacity records. We have fixed the underlying capacity-selection bug to prevent this failure from recurring. We have also added alerts for VM-record creation failures associated with this capacity issue.

monitoringSep 14, 07:01 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingSep 14, 06:40 PM

We are investigating reports of impacted performance for some GitHub services.

criticalresolvedSep 13, 09:16 AM — Resolved Sep 13, 10:44 AM

Incident with several GitHub Services

6 updates
resolvedSep 13, 10:44 AM

On September 13, 2026, between 08:43 and 10:44 UTC, GitHub experienced degraded availability across approximately 28 services, including Issues, Pull Requests, Actions, Codespaces, Pages, Notifications, Code Scanning, Git LFS, and new account signup. At peak, 8.8% of requests to create GitHub App installation access tokens failed. Token issuance for Actions workflows was also affected, impacting approximately 4% of workflows during the incident time frame. Creating issues through the web interface failed for about 96% of attempts, and signup failures were above 90%. The cause was an internal data-cleanup job that began writing to a shared database cluster at 07:33 UTC. That cluster stores permission data read on nearly every authenticated request. The safeguard that was pacing the background job watched only one health signal — how far the database replicas were lagging — and that signal stayed low the whole time. It did not account for the load building on the primary itself, so the job kept writing while the primary quietly ran toward its limit. When the primary ran out of available connections, requests that needed it could not complete. First, there was no quick timeout on these database calls, so request handlers waited on the stalled database instead of failing fast, and the shared request-handling capacity degraded into site-wide errors. Second, a retry loop around token creation kept re-sending the writes that were already failing, which held the database saturated rather than letting it recover. Monitoring declared the incident at 08:50 UTC, but due to the broad impact and amplification from token creation, it took time to identify the source of the load. First responders mitigated by shedding internal load and pausing the job, and all services recovered by 10:44 UTC. To prevent recurrence, we are rate-limiting background jobs against shared, customer-serving databases by default, and adding automatic pausing and paging on primary-server load rather than replication lag alone. We are also surfacing running background work directly alongside database health signals so responders can see and pause it without leaving those dashboards, bounding retries in the token-issuing path, and adding request-level timeouts so one unhealthy database cannot consume shared web server capacity. In addition, we are breaking apart this database cluster to remove the single point of failure. We will be moving various service-specific data, including the authorization data, out of this shared cluster in the next two weeks.

investigatingSep 13, 10:28 AM

Pull Requests is experiencing degraded performance. We are continuing to investigate.

investigatingSep 13, 10:26 AM

We have reduced load on this cluster with internal load-shedding and are seeing signs of recovery but continue to monitor

investigatingSep 13, 09:36 AM

We're seeing increased database replication delays on collab which is causing increased error rates in authorization endpoints and follow-on increased error rates across the system - we are investigating

investigatingSep 13, 09:25 AM

Actions is experiencing degraded performance. We are continuing to investigate.

investigatingSep 13, 09:16 AM

We are investigating reports of degraded availability for API Requests, Issues, Pages and Pull Requests

majorresolvedSep 4, 08:39 PM — Resolved Sep 4, 10:26 PM

Disruption with Copilot Code Review

5 updates
resolvedSep 4, 10:26 PM

On September 4, 2026, between 20:04 and 22:26 UTC, GitHub Copilot code review experienced an increased failure rate. Affected pull request reviews failed to complete or post review comments.The incident was caused by a change to the service’s authentication permissions that prevented it from submitting affected reviews to the GitHub API. We reverted the change and restored normal operation by 22:26 UTC.We apologize for the disruption.

monitoringSep 4, 10:25 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingSep 4, 09:54 PM

We are applying the mitigation and expect recovery within approximately 30 minutes.

investigatingSep 4, 08:57 PM

Some users may be experiencing failures when using Copilot code review. We have identified the root cause and are working on a mitigation.

investigatingSep 4, 08:39 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedSep 4, 10:02 PM — Resolved Sep 4, 10:23 PM

Degradation in repos contents API

2 updates
resolvedSep 4, 10:23 PM

On September 4, 2026, between approximately 21:45 and 22:07 UTC, some users experienced errors and elevated latency for repository operations. The incident was fully resolved at 22:23 UTC.The cause was a capacity change that spread one of our clusters across additional availability zones; our zone-aware traffic routing kept sending requests to the original zone for performance, overloading a small set of servers while the new capacity sat idle. We resolved the incident by reverting the change and letting traffic rebalance.We are improving per-zone capacity guarantees, cross-zone load-shedding, and pre-production testing of multi-zone changes to prevent recurrence.

investigatingSep 4, 10:02 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedSep 3, 02:17 PM — Resolved Sep 3, 05:11 PM

Incident with Grok Copilot AI Model Provider

5 updates
resolvedSep 3, 05:11 PM

Between 13:22 and 17:11 UTC on September 03, 2026, GitHub Copilot experienced degradation affecting several Grok models, including Grok 4.5 and Grok 4.6. Users encountered elevated error rates, but other models were not affected. The degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring, displayed in-product warnings for the affected models, and coordinated with the provider. Service returned to normal after the provider implemented a mitigation.

investigatingSep 3, 05:11 PM

The issues with our upstream model provider have been resolved, and Grok models are once again available in Copilot products and IDE surfaces.We will continue monitoring to ensure stability, but mitigation is complete.

investigatingSep 3, 02:22 PM

The Grok 4.5 model has degraded availability as well. We are working with the upstream provider to resolve the issue.

investigatingSep 3, 02:20 PM

We are experiencing degraded availability for the Grok 4.6 model in Copilot Chat, VS Code and other Copilot products. This is due to an issue with an upstream model provider. We are working with them to resolve the issue.

investigatingSep 3, 02:17 PM

We are investigating reports of degraded performance for Copilot AI Model Providers

minorresolvedSep 1, 03:00 PM — Resolved Sep 1, 04:01 PM

Delays in commit processing

4 updates
resolvedSep 1, 04:01 PM

On September 1, 2026, between approximately 14:01 and 16:01 UTC, updates in response to pushes were delayed, temporarily showing stale diffs. The median time to refresh a diff after a push rose from the normal level of about 3 seconds to over 2 minutes at the peak, and more than 140,000 customer accounts had at least one delayed refresh during the most affected 75 minutes. Pushing commits and opening pull requests continued to work normally. The incident was caused by a sharp, concentrated surge in push volume that saturated worker pools and job queueing infrastructure. Autoscaling did not increase capacity as intended, so the backlog did not clear on its own. The incident was mitigated by manually scaling the affected worker pools and increasing push-processing capacity. This allowed the system to process the backlog, after which refresh times returned to normal. To reduce the likelihood and impact of similar incidents, we are adding quotas and throttling earlier in the push path so a single concentrated source of load cannot saturate shared capacity, improving worker-pool autoscaling so capacity is added automatically, and improving monitors for background job processing so on-call is paged before customers experience delayed pull request updates.

investigatingSep 1, 04:00 PM

Time to update pull request diffs have improved to normal thresholds.

investigatingSep 1, 03:00 PM

Diffs in the PR view may be stale for several minutes. We are investigating and scaling up resources.

investigatingSep 1, 03:00 PM

We are investigating reports of degraded performance for Pull Requests

August 2026(27 incidents)

minorresolvedAug 31, 09:15 AM — Resolved Aug 31, 09:58 AM

Elevated rate of errors for OpenAI models provided by Copilot

6 updates
resolvedAug 31, 09:58 AM

Between 08:37 and 09:41 UTC on August 31, 2026, GitHub Copilot experienced degradation affecting several GPT models, including gpt-5.2, gpt-5.3-codex, gpt-5.4, gpt-5.4-mini, gpt-5.4-nano, and the gpt-5.6 family (Luna, Sol, and Terra). Users encountered elevated error rates and interrupted streaming responses. Other models were not affected.The degradation was caused by an issue with an upstream model provider. GitHub engineers detected the issue through automated monitoring, displayed in-product warnings for the affected models, and coordinated with the provider. Service returned to normal after the provider implemented a mitigation.

monitoringAug 31, 09:58 AM

The issues with our upstream model provider have been resolved, and gpt-5.3-codex, gpt-5.4-mini, gpt-5.4-nano, gpt-5.5, and the gpt-5.6 family of models are once again available in Copilot products and IDE surfaces.We will continue monitoring to ensure stability, but mitigation is complete.

monitoringAug 31, 09:51 AM

The degradation affecting Copilot AI Model Providers has been mitigated. We are monitoring to ensure stability.

investigatingAug 31, 09:48 AM

One of our model providers has confirmed an incident on their end. We have provided them details to help identify the issue. We are starting to see recovery.

investigatingAug 31, 09:22 AM

Copilot is experiencing a higher rate of errors for OpenAI models, including gpt-5.2, gpt-5.3-codex, gpt-5.4, gpt-5.4, and the gpt-5.6 family of models. Other models are not impacted.

investigatingAug 31, 09:15 AM

We are investigating reports of degraded performance for Copilot AI Model Providers

minorresolvedAug 26, 11:37 PM — Resolved Aug 27, 07:44 PM

Disruption with GitHub Billing

8 updates
resolvedAug 27, 07:44 PM

On August 26, 2026, between 20:40 UTC and 00:51 UTC on August 27, GitHub Billing experienced degraded performance affecting billing budget pages and GitHub Copilot CLI sessions. Affected customers encountered failed budget page loads or failures when starting or continuing CLI sessions. We confirmed this impact for a small number of customers (This was caused by a concentrated workload that created processing delays in our data storage layer. Automated retries increased the load and prolonged the degradation. We mitigated the incident by rebalancing traffic within our infrastructure. We are improving workload isolation, retry behavior, and detection of concentrated load to reduce the likelihood of recurrence and shorten our time to detect and mitigate similar incidents.

investigatingAug 27, 05:58 PM

No material change since the previous update. Service conditions remain stable following the mitigation, and we have not observed any further customer impact. We are actively monitoring the service while implementing targeted fixes to address the underlying root cause.

investigatingAug 27, 04:20 PM

Our mitigation continues to hold, and service conditions remain stable. We are continuing to investigate the concentrated workload responsible for the issue and are preparing additional preventative improvements. We have not identified a material change in customer impact since the previous update. We will provide another update as the investigation progresses.

investigatingAug 27, 02:49 PM

Our mitigation is still holding as we continue to investigate to find the root cause.

investigatingAug 27, 01:35 AM

We are continuing to monitor the mitigation that we have applied for the billing page disruption.

investigatingAug 27, 12:31 AM

We've applied a mitigation to unblock Copilot usage and have observed recovery for this particular impact. We're continuing to investigate and apply mitigations for the billing page disruption while monitoring to ensure Copilot remains recovered.

investigatingAug 26, 11:42 PM

We are currently investigating increased errors with billing services. Customers may observe failed billing budget page loads, and users of the Copilot CLI may observe failures starting or continuing sessions.

investigatingAug 26, 11:37 PM

We are investigating reports of impacted performance for some GitHub services.

criticalresolvedAug 27, 10:04 AM — Resolved Aug 27, 12:12 PM

Incident with Copilot AI Model Providers

5 updates
resolvedAug 27, 12:12 PM

On August 27th, 2026, between approximately 09:20 and 12:14 UTC, the Copilot service experienced a degradation of the Kimi K3 model due to an issue with our upstream provider. Users encountered elevated error rates when using Kimi K3. No other models were impacted. The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.

investigatingAug 27, 12:12 PM

The issues with our upstream model provider have been mitigated, and Kimi K3 is once again available in Copilot products and IDE surfaces.We will continue monitoring to ensure stability.

investigatingAug 27, 11:58 AM

Copilot AI Model Providers is experiencing degraded performance. We are continuing to investigate.

investigatingAug 27, 10:43 AM

We are experiencing degraded availability for the Kimi K3 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

investigatingAug 27, 10:04 AM

We are investigating reports of degraded availability for Copilot AI Model Providers

minorresolvedAug 26, 10:56 PM — Resolved Aug 27, 12:26 AM

Incident with Actions and Pull Requests

6 updates
resolvedAug 27, 12:26 AM

On August 26, 2026, from 21:55 UTC to 23:58 UTC, 2.6% of workflow runs triggered by pull request events were delayed, with the impact rising as high as 25% at its peak. Some users also experienced delays in pull request merge-commit generation, mergeability information, and merge-button availability. Actions and Pull Requests fully recovered by 23:58 UTC; the incident was resolved at 00:26 UTC after normal operation was confirmed. Background jobs that process pull request updates and generate merge commits were impacted by timeouts reaching a single partition of git data. This resulted in a backlog in pull request merge-commit processing, delaying pull request-triggered GitHub Actions workflows and some mergeability information. We reduced workload, shifted traffic away from affected infrastructure, and restored the affected service component to a healthy state. Together, these actions helped drain the backlog and restore normal operations. We are working to improve resource saturation detection and to eliminate customer impact in this scenario by isolating impact, placing better bounds on retries, and strengthening backpressure to make our systems more resilient under load.

monitoringAug 27, 12:26 AM

The degradation affecting Actions and Pull Requests has been mitigated. We are monitoring to ensure stability.

investigatingAug 27, 12:25 AM

We confirmed full recovery beginning at 23:58 UTC. Actions workflow runs and pull request merges are operating normally. We will now resolve the incident while continuing to monitor service health.

investigatingAug 27, 12:01 AM

We've applied mitigations and are seeing recovery in Actions workflow runs and blocked pull request merges. We're continuing to monitor for sustained health of merge commit creates before resolving.

investigatingAug 26, 10:57 PM

We are investigating elevated delays and timeouts affecting Actions workflow runs triggered by pull request events. 20% of actions runs have delayed starts of more than 5 minutes and up to 4% of runs failed to trigger. We are actively working on mitigation and will provide updates as we learn more.

investigatingAug 26, 10:56 PM

We are investigating reports of degraded performance for Actions and Pull Requests

criticalresolvedAug 26, 03:11 PM — Resolved Aug 26, 06:01 PM

Incident with Actions

11 updates
resolvedAug 26, 06:01 PM

On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. 3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs. Customers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27. Some runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued even after service recovery. In the backend, these runs had already failed and will automatically move to canceled state 24 hours after creation. As follow-up, we are fixing the root cause of this queued state and improving our ability to bulk-cancel affected runs. Several changes to improve the general scalability of this part of Actions were already complete and deploying to production. Rollout of those changes will be complete within the next 24 hours. Further work to improve scale, resiliency, and more graceful degradation of Actions workflows are in flight. We are also taking a repair item to accelerate clearing of stuck queued or waiting jobs in similar future cases.

monitoringAug 26, 06:00 PM

All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.

monitoringAug 26, 05:54 PM

The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

investigatingAug 26, 05:32 PM

We are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.

investigatingAug 26, 04:50 PM

We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.

investigatingAug 26, 04:49 PM

Pages is operating normally.

investigatingAug 26, 04:14 PM

We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.

investigatingAug 26, 03:48 PM

primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

investigatingAug 26, 03:23 PM

We've identified an issue with a database primary and are failing over to a replica immediately

investigatingAug 26, 03:12 PM

Pages is experiencing degraded performance. We are continuing to investigate.

investigatingAug 26, 03:11 PM

We are investigating reports of degraded availability for Actions

minorresolvedAug 26, 03:09 PM — Resolved Aug 26, 04:07 PM

Disruption with some GitHub services

2 updates
resolvedAug 26, 04:07 PM

Please refer to the combined summary in this related incident: https://www.githubstatus.com/incidents/y1t7p9fzrlj2

investigatingAug 26, 03:09 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedAug 24, 01:56 PM — Resolved Aug 24, 02:34 PM

Actions delays in starting runs

4 updates
resolvedAug 24, 02:34 PM

On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.

monitoringAug 24, 02:26 PM

The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

investigatingAug 24, 02:22 PM

Failures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery.

investigatingAug 24, 01:56 PM

We are investigating reports of degraded performance for Actions

majorresolvedAug 24, 07:12 AM — Resolved Aug 24, 07:58 AM

Elevated errors on Fable 5 due to upstream provider

3 updates
resolvedAug 24, 07:58 AM

On August 24th, 2026, between approximately 06:35 and 07:25 UTC, the Copilot service experienced a degradation of the Claude Fable 5 model due to an issue with our upstream provider. Users encountered elevated error rates when using Claude Fable 5, with requests sometimes failing mid-response. No other models were impacted.The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.

investigatingAug 24, 07:12 AM

We are experiencing degraded availability for the Fable model in Copilot products and IDE surfaces. This is due to an issue with the upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

investigatingAug 24, 07:12 AM

We are investigating reports of degraded availability for Copilot AI Model Providers

noneresolvedAug 21, 02:00 PM — Resolved Aug 21, 02:00 PM

Degraded Git Operations over SSH

1 update
resolvedAug 24, 07:19 AM

On August 21, 2026, between 14:00 and 14:07 UTC, dotcom Git operations over SSH were degraded. Successful Git operations over SSH fell by more than 95% for during the peak impact window, making clone, fetch, or push over SSH effectively unavailable to most users for approximately four minutes. Git operations over HTTPS were not affected. The incident was caused by a software defect in our load-balancing infrastructure that was triggered by a configuration change. The defect only occurred when connections passed through multiple layers of load balancers running the new configuration, which meant it was not detected during canary testing. We mitigated the incident by rolling back the configuration change. We are adding regression coverage for multi-layer load-balancer configurations and improving monitoring and alerting for Git operations over SSH to reduce our time to detection and mitigation of similar issues in the future.

criticalresolvedAug 20, 02:43 PM — Resolved Aug 21, 12:37 AM

Intermittent failures creating agent tasks

12 updates
resolvedAug 21, 12:37 AM

Between 13:57 UTC on August 20 and 00:37 UTC on August 21, 2026, some users of the Copilot Cloud Agent experienced delays of up to 60 to 90 minutes in seeing the status and results of their agent tasks. The agent tasks themselves continued to run and complete during this time; only the visibility of their status was delayed.The cause was a regional outage in a third-party cloud database service that Copilot uses to store agent task status. We failed over the affected database to a healthy region, added processing capacity to work through the backlog, and restored normal operation once the underlying service recovered. No task data was lost during the incident.To prevent repetition of similar incidents, we are removing the database configuration that made us vulnerable to this regional outage and improving our database failover procedures.

investigatingAug 20, 08:37 PM

We are seeing gradual recovery in Copilot Cloud Agent task status visibility as we deploy a fix for the root cause. Session output remains delayed by approximately one hour while remediation continues.

investigatingAug 20, 07:35 PM

We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility. Session output continues to be delayed by approximately 1 hour as our remediation steps take effect.

investigatingAug 20, 06:45 PM

We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed by approximately 1 hour. We have taken additional steps to accelerate the recovery and expect this to take effect within the next hour.

investigatingAug 20, 06:04 PM

We are continuing to observe gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed by approximately 1 hour. We have taken additional steps to accelerate the recovery and are continuing to monitor the impact.

investigatingAug 20, 05:32 PM

We are observing gradual recovery for Copilot Cloud Agent task status visibility, with session output delayed approximately 1 hour. We have taken additional steps to accelerate the recovery and are continuing to monitor the impact.

investigatingAug 20, 05:05 PM

We are seeing signs of recovery for Copilot Cloud Agent task status visibility, but this recovery is slower than anticipated. We are pursuing additional mitigating measures to accelerate recovery.

investigatingAug 20, 04:14 PM

Users are experiencing delays when starting tasks using Copilot Cloud Agent and are not be able to see the status of these tasks. Copilot Cloud Agent tasks are still being completed. We have identified the cause of the issue and are putting mitigations in place to return service to normal levels. We will provide another update about the expected recovery time shortly.

investigatingAug 20, 03:41 PM

We are experiencing issues with Copilot Cloud Agent tasks, resulting in newly started tasks not properly displaying on-going progress. These Copilot Cloud Agent tasks are still being completed correctly but lack proper visibility. We are actively investigating the issue and will provide updates as we learn more.

investigatingAug 20, 03:01 PM

We have identified the problematic component and are working to fail over to a healthy instance. Further updates will be provided as we perform mitigations.

investigatingAug 20, 02:51 PM

Users may experience delays when starting tasks using Copilot Cloud Agent. We are actively investigating the issue and will provide updates as we learn more.

investigatingAug 20, 02:43 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedAug 18, 07:40 AM — Resolved Aug 18, 11:42 AM

Intermittent failures in runner group and runner-related permissions pages

5 updates
resolvedAug 18, 11:42 AM

On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API. The issue was caused by failures in backend requests reading runner and runner group data. The failures were caused by an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents triggered by this operation. The impact was mitigated by completing the enablement of the new certificate in the backend system. We have added additional monitoring to this and other certificates. This service is also in the process of being replaced as part of our availability and scale work, bringing this authentication path and secret management in line with patterns across all GitHub services.

monitoringAug 18, 11:24 AM

We have applied a mitigation and are seeing recovery signals. We will continue monitoring recovery and providing updates.

monitoringAug 18, 10:41 AM

We have identified the source of a communication issue between Actions services and are working toward mitigation. Customers may experience failure to load runner groups and runner-related permissions issues when using Larger Runners.

monitoringAug 18, 07:40 AM

We are investigating reports of failure to load runner groups and runner-related permissions for customers using larger runners.

investigatingAug 18, 07:40 AM

We are investigating reports of impacted performance for some GitHub services.

majorresolvedAug 18, 09:36 AM — Resolved Aug 18, 10:23 AM

Incident with Actions

2 updates
resolvedAug 18, 10:23 AM

On August 18, 2026, between 05:02 UTC and 11:30 UTC, customers were unable to run jobs on Actions Larger Runners and were unable to view or manage Actions Runners and Runner Groups through the GitHub UI and API. These issues were caused by failures in backend requests resolving essential metadata for starting Larger Runner workflow runs and for reading runner and runner group data. The failures were caused by an expired authentication certificate unique to this service. The certificate had been rotated in KeyVault, but a step to enable use at runtime had been paused to prevent recurrence of previous incidents that had been triggered by this operation. We mitigated the issues by completing the enablement of the new certificate in the backend system. We have added additional monitoring to this and other certificates. The relevant service is also in the process of being replaced as part of our availability and scale work, bringing this authentication path and secret management in line with patterns across all GitHub services.

investigatingAug 18, 09:36 AM

We are investigating reports of impacted performance for some GitHub services.

criticalresolvedAug 17, 01:40 PM — Resolved Aug 17, 09:15 PM

Incident with GitHub.com

36 updates
resolvedAug 17, 09:15 PM

On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02. Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service. The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery. The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed. Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery. Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints. To prevent recurrence, our follow-up actions include: - Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity. - Auditing Istio request, concurrency, and scaling limits across affected services. - Reviewing retry limits and backoff behavior across gateways and clients. - Addressing the VS Code retry behavior that amplified Copilot token traffic. - Improving load-balancer capacity monitoring and regional failover safeguards.

investigatingAug 17, 08:45 PM

We are continuing to apply mitigations to address sporadic Copilot authentication failures in some applications. We expect full recovery within the next 30 minutes. Copilot usage via the GitHub CLI and GitHub App are unaffected.

investigatingAug 17, 08:22 PM

Issues is operating normally.

investigatingAug 17, 08:08 PM

We are continuing to investigate sporadic failures affecting Copilot authentication in some applications. Copilot usage via the GitHub CLI and GitHub App are unaffected.

investigatingAug 17, 07:13 PM

We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation.

investigatingAug 17, 07:01 PM

API Requests is operating normally.

investigatingAug 17, 06:48 PM

API Requests is experiencing degraded availability. We are continuing to investigate.

investigatingAug 17, 06:23 PM

The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability.

investigatingAug 17, 06:11 PM

We identified the problematic component and have taken corrective actions, but we are seeing residual impact in the form of sporadic authentication failures. We are continuing to apply additional mitigations and investigate the remaining impact.

investigatingAug 17, 05:36 PM

Issues is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 05:34 PM

We identified the problematic component and have taken corrective actions, but we are seeing residual impact across numerous services. We are continuing to apply additional mitigations and investigate the remaining impact.

investigatingAug 17, 05:30 PM

Git Operations is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 04:59 PM

The degradation affecting API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.

investigatingAug 17, 04:36 PM

We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues.

investigatingAug 17, 04:16 PM

We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are still working to identify the root cause and will continue to post updates as we learn more and perform mitigation.

investigatingAug 17, 03:42 PM

We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations and will post updates as we progress.

investigatingAug 17, 03:40 PM

Webhooks is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 03:21 PM

Git Operations is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 03:10 PM

Pages is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 03:01 PM

API Requests is experiencing degraded availability. We are continuing to investigate.

investigatingAug 17, 02:58 PM

Webhooks is experiencing degraded availability. We are continuing to investigate.

investigatingAug 17, 02:58 PM

We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations based on our investigation thus far and are monitoring for improvement.

investigatingAug 17, 02:58 PM

Actions is experiencing degraded availability. We are continuing to investigate.

investigatingAug 17, 02:54 PM

Pull Requests is experiencing degraded availability. We are continuing to investigate.

investigatingAug 17, 02:49 PM

Issues is experiencing degraded availability. We are continuing to investigate.

investigatingAug 17, 02:45 PM

Pull Requests is experiencing degraded availability. We are continuing to investigate.

investigatingAug 17, 02:31 PM

Copilot is experiencing degraded availability. We are continuing to investigate.

investigatingAug 17, 02:24 PM

We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. Investigations are on-going and we will continue to provide updates as we discover more information.

investigatingAug 17, 02:04 PM

We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. Investigations are on-going into the root cause, and updates will continue to be provided as we investigate.

investigatingAug 17, 01:58 PM

Pull Requests is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 01:46 PM

Issues is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 01:45 PM

We are seeing an approximate 20% error rate across numerous experiences including Pull Requests, Issues, and others. Investigations are currently under way and we will be posting updates as they become available

investigatingAug 17, 01:44 PM

Webhooks is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 01:42 PM

Actions is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 01:41 PM

API Requests is experiencing degraded performance. We are continuing to investigate.

investigatingAug 17, 01:40 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedAug 13, 04:21 PM — Resolved Aug 13, 06:27 PM

Disruption with GHEC Team Sync

4 updates
resolvedAug 13, 06:27 PM

On August 13, 2026, from 15:31:21 UTC to 18:27:55 UTC, GitHub Enterprise Cloud team synchronization was degraded for enterprises using personal accounts. Organization teams experienced delays of up to 3 to 13 hours (median 8 hours) when syncing with IdP groups, resulting in delayed access grants or removals for enterprise users across 2.8% of teams. A temporary change introduced to address a previous issue due to increased usage of this feature remained active after it was intended to be removed, causing synchronization delays during periods of high volume. We removed the temporary change and provisioned additional resources to handle the increased volume.

investigatingAug 13, 06:27 PM

We have deployed a mitigation. At this time GHEC Team Sync has recovered for enterprises with personal accounts. Teams syncing to IdP groups have returned to their normal cadence.

investigatingAug 13, 04:21 PM

GHEC Team Sync is currently degraded for enterprises with personal accounts, causing delays when syncing teams to IdP groups. We have identified the cause of the delays and are working on a mitigation. We will provide an update on our progress at 20:00 UTC.

investigatingAug 13, 04:21 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedAug 13, 02:43 PM — Resolved Aug 13, 03:47 PM

Errors with the Fable 5 Model in Copilot

5 updates
resolvedAug 13, 03:47 PM

On August 13th, 2026, between approximately 14:06 and 15:47 UTC, the Copilot service experienced a degradation of the Claude Fable 5 model due to an issue with our upstream provider. Users encountered elevated error rates, peaking at 43% and averaging 12%. Users who selected Auto or alternative models were unaffected.The issue was resolved by a mitigation put in place by our provider. GitHub is working with our provider to further improve the resiliency of the service to prevent similar incidents in the future.

investigatingAug 13, 03:47 PM

The issues with our upstream model provider have been resolved, and Fable 5 is once again available in Copilot products and IDE surfaces.We will continue monitoring to ensure stability, but mitigation is complete.

investigatingAug 13, 03:23 PM

We are seeing modest recovery, but are still experiencing degraded availability for the Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

investigatingAug 13, 02:50 PM

We are experiencing degraded availability for the Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

investigatingAug 13, 02:43 PM

We are investigating reports of degraded performance for Copilot AI Model Providers

minorresolvedAug 13, 02:45 PM — Resolved Aug 13, 03:36 PM

Incident with Webhooks

9 updates
resolvedAug 13, 03:36 PM

Between 14:24 and 14:53 UTC on 13 August 2026, a routine background job to delete an organization overwhelmed a key shared database, causing multiple GitHub services to briefly return elevated errors and slower responses. Most affected was the webhook management API, with smaller impact to Git operations, pull requests, issues, packages, sign-in, and Copilot. Impact cleared on its own at about 14:53 UTC once the job finished; we resolved the incident at 15:36 UTC. Affected users may have experienced a brief increase in errors and slower responses, primarily when creating, listing, or updating webhooks, with smaller impacts to pull requests, issues, packages, and Git operations. Failures peaked at about 1% for several minutes around 14:37 UTC. To prevent future incidents, we've already shipped an update that turns on the safer deletion path for organizations, along with caps on deletion holds on databases. Building on these changes, we're auditing all bulk deletion and cleanup jobs that write to shared databases to prevent similar issues in future.

monitoringAug 13, 03:33 PM

We have temporarily disabled a background job which caused the impact. At this time the impact is fully mitigated.

monitoringAug 13, 03:33 PM

The degradation affecting Git Operations, Issues, Packages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.

investigatingAug 13, 02:58 PM

We are currently investigating a brief degradation of service for Git operations (specifically pushes), issues, pull requests, package registry, and webhooks between 14:32 and 14:46 UTC. We have identified the source of the degradation and are investigating mitigation strategies to prevent recurrence.

investigatingAug 13, 02:56 PM

Packages is experiencing degraded performance. We are continuing to investigate.

investigatingAug 13, 02:46 PM

Git Operations is experiencing degraded performance. We are continuing to investigate.

investigatingAug 13, 02:46 PM

Issues is experiencing degraded performance. We are continuing to investigate.

investigatingAug 13, 02:46 PM

Pull Requests is experiencing degraded performance. We are continuing to investigate.

investigatingAug 13, 02:45 PM

We are investigating reports of degraded performance for Webhooks

majorresolvedAug 12, 09:39 PM — Resolved Aug 12, 10:56 PM

Disruption with Login and Release Asset downloads

4 updates
resolvedAug 12, 10:56 PM

On August 12 and 13, 2026, some anonymous (logged-out) requests to github.com experienced HTTP 5xx errors when loading pages like the sign-in page, and when downloading release assets, due to an unusual traffic pattern that repeatedly overloaded a part of our infrastructure that serves these types of requests. There were three windows of impact: (1) August 12 from 16:34 to 18:34 UTC, with an average error rate of 16.16% that peaked at 28.6%; (2) August 12 from 19:00 to 22:56 UTC, with an average error rate of 16.55% that peaked at 24.18%; and (3) August 13 from 06:19 to 08:05 UTC, with an average error rate of 2.01% that peaked at 7.49%.Requests from signed-in users were unaffected.We mitigated the incidents by applying traffic controls at our network edge that limited any requests matching the pattern identified previously, thereby preventing overload on our systems.Since these incidents occurred, we have tightened our monitoring systems to alert server-side errors that affect logged-out traffic. We are also working to further strengthen our edge protections and reduce the time to detect and mitigate similar incidents.

investigatingAug 12, 10:22 PM

We have identified the root cause and are working on mitigation. Errors on the login page and downloading release assets have decreased, but we are not fully mitigated. We will continue to provide updates.

investigatingAug 12, 09:43 PM

We are investigating issues with Login and when downloading Release Assets. We will continue to keep users updated on progress towards mitigation.

investigatingAug 12, 09:39 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedAug 12, 04:16 PM — Resolved Aug 12, 04:41 PM

Incident with Pull Requests and Issues

5 updates
resolvedAug 12, 04:41 PM

Between 16:03 and 16:29 UTC on August 12, some users encountered errors when viewing pull requests, issues, and search results. During this period, about 1.9% of Pull Request requests and 0.9% of Issues requests failed. During a database migration, two indexes were removed while application settings still referenced them, causing affected requests to fail. We detected the issue after the migration reached one database shard and before it progressed to the remaining shards. We restored service by disabling both settings. We are improving safeguards around database migrations and application configuration to prevent similar mismatches from causing errors.

monitoringAug 12, 04:38 PM

We identified the source of errors affecting Pull Requests, Issues, and Search on GitHub.com and have applied a mitigation. A database index hint was referencing an index that had been removed by a recent migration, causing query failures for some users. We disabled the problematic configuration and are seeing recovery across affected services. We are continuing to monitor to confirm full resolution.

monitoringAug 12, 04:35 PM

The degradation affecting Issues and Pull Requests has been mitigated. We are monitoring to ensure stability.

investigatingAug 12, 04:24 PM

We are investigating reports of errors affecting Pull Requests and Issues on GitHub.com. Some users may encounter 500 errors when loading pull request and issue pages. Our engineering teams are actively investigating the root cause, which appears to be related to a database infrastructure issue. We will provide an update as soon as we have more information.

investigatingAug 12, 04:16 PM

We are investigating reports of degraded performance for Issues and Pull Requests

minorresolvedAug 11, 02:50 PM — Resolved Aug 11, 08:06 PM

Incident with GraphQL API Requests

6 updates
resolvedAug 11, 08:06 PM

On August 11, 2026, between 14:00 UTC and 16:00 UTC the GraphQL API service was degraded and customers in saw higher than normal timeouts. On average, the timeout rate was 0.06% and peaked at 0.14% of requests routing to the service. This was due to increased utilization at one of our sites which caused resource contention across our dependencies, leading to an increase in timeouts for GraphQL requests. We mitigated the incident by increasing capacity to alleviate the capacity bottleneck. We are working to improve our monitoring so that we can proactively reduce the impact of high consumption requests in addition to scaling up; Additionally, we will improve our time to detection and mitigation of issues like this one in the future.

monitoringAug 11, 08:06 PM

We have identified and mitigated increased error rates affecting GraphQL API requests. A fix to increase service capacity has been deployed and error rates have returned to normal levels. We are resolving this incident.

monitoringAug 11, 04:49 PM

The degradation affecting API Requests has been mitigated. We are monitoring to ensure stability.

investigatingAug 11, 04:49 PM

We have returned to a healthy baseline on GraphQL API requests. We will continue to work on investigations into the errors seen during this incident.

investigatingAug 11, 03:28 PM

We are investigating reports of a small increase in error rates affecting GraphQL API requests. We are working on increasing capacity and continue to investigate the increased errors. We will provide another update when we have more information

investigatingAug 11, 02:50 PM

We are investigating reports of degraded performance for API Requests

minorresolvedAug 10, 08:27 PM — Resolved Aug 10, 09:50 PM

Disruption with Copilot for access to some models

6 updates
resolvedAug 10, 09:50 PM

On August 10, 2026, between 19:48 UTC and 20:49 UTC, GitHub Copilot users saw an incomplete list of available models. During this window, the service could return as few as one model instead of the full catalog. Requests that tried to use a model missing from that shortened list failed with a "model not found" error. Copilot requests that used an available model were not affected. This did not affect customers on data-residency (Proxima) environments.The issue was caused by a change to how model data was published, which our systems could not read back correctly and fell back to a limited default list.We mitigated the incident by 20:49 UTC and deployed a fix to prevent immediate recurrence by 21:50 UTC. We are adding validation and retry safeguards so that model data is verified before it is served.We apologize for the disruption.

monitoringAug 10, 09:50 PM

We have deployed and validated the fix to prevent immediate reoccurrence. We will be performing additional work to limit these kinds of failures in the future.

monitoringAug 10, 09:19 PM

The issue has been mitigated across all affected environments. We are currently deploying on a fix to prevent reoccurrence. We will provide another update once the fix has been deployed.

monitoringAug 10, 08:49 PM

The degradation has been mitigated. We are monitoring to ensure stability.

monitoringAug 10, 08:39 PM

We are currently investigating reports of some Copilot users experiencing issues accessing certain models. Affected users may see errors or degraded functionality when attempting to use specific models. We are actively working on a fix and will provide updates as we have more information.

investigatingAug 10, 08:27 PM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedAug 10, 06:02 PM — Resolved Aug 10, 06:46 PM

Disruption with creation of fine grained personal access tokens

5 updates
resolvedAug 10, 06:46 PM

On August 10, 2026, between 17:16 and 18:21 UTC, users were unable to create new fine-grained personal access tokens (FG PAT) through the GitHub website. When a user submitted the FG PAT creation form, they were returned to the FG PAT list without an error message and no FG PAT was created. Creating classic personal access tokens, as well as editing or deleting existing FG PAT were not affected.The cause was a change to how the website loads certain front-end JavaScript that was enabled for all users at 17:15 UTC; the change interacted with an issue in the token creation form's confirmation step that prevented it from running, so the final submission that actually creates the token never completed. Because the page still loaded and the server returned a normal response, the failure produced no error message. GitHub mitigated the incident by disabling the change at 18:21 UTC, at which point token creation recovered immediately, and the incident was resolved at 18:46 UTC.To reduce the chance of recurrence, GitHub is adding monitoring and alerting for anomalies in the FG PAT creation success rate and is removing the issue in the FG PAT creation form that prevented the confirmation step from running. GitHub is also adding automated detection of the issue so other areas of the GitHub front end do not repeat the problem.

monitoringAug 10, 06:22 PM

We identified the source of the issue affecting creation of fine-grained personal access tokens and have applied a mitigation. Users should now be able to create new fine-grained tokens successfully. We are continuing to monitor to confirm full recovery.

monitoringAug 10, 06:21 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingAug 10, 06:09 PM

We are investigating reports of users being unable to create fine-grained Personal Access Tokens. Attempting to create a new token redirects the user back to the token overview page without an error message, but the token was not created.

investigatingAug 10, 06:02 PM

We are investigating reports of impacted performance for some GitHub services.

criticalresolvedAug 6, 03:22 PM — Resolved Aug 7, 02:04 AM

Incident with Actions

24 updates
resolvedAug 7, 02:04 AM

On August 6, 2026, between 15:05 UTC and 00:14 UTC on August 7, GitHub Actions experienced degraded availability. During the incident, workflow runs failed or remained queued for an extended period of time. Customers using both GitHub-hosted and self-hosted runners were affected. At peak, 71% of workflow runs experienced infrastructure failures and 75% of the remaining workflow runs were delayed by more than 5 minutes. The incident was triggered by a routine deployment to an internal Actions service responsible for processing events and generating Actions jobs. The deployment exposed an existing capacity and concurrency weakness. As pods were replaced during the deployment, remaining capacity became saturated, causing services to crash and triggering a cascading impact across multiple clusters and downstream services. These services recovered at 17:00 after expanding capacity, throttling incoming webhook-triggered work to allow the system to recover, and increasing processing capacity for the backlog of affected events. As the incident progressed, a backlog of work accumulated across the systems responsible for assigning jobs to runners. Due to a latent bug in one of the services responsible for job assignment, runners were getting assigned jobs that were no longer valid and then getting stuck retrying those jobs, preventing them from picking up valid work. This second stage of impact was mitigated by deploying changes to prevent runners from repeatedly attempting to acquire invalid jobs. These mitigations allowed the accumulated queues to drain and Actions to recover to normal operation. Some Actions Runner Controller (ARC) runners remained stuck after the incident. A mitigation deployed during the incident inadvertently affected these runners, causing some to remain offline until they were manually recovered. We subsequently rolled back the change and are adding automatic recovery in upcoming Runner and ARC releases. Some jobs created during the incident were also left stuck unable to be retried or canceled. CLI and UI solutions for customers to address these were shared at https://github.com/orgs/community/discussions/204152#discussioncomment-17946043. To prevent recurrence, we are making improvements to deployment and capacity safeguards for the affected services, strengthening monitoring for the conditions that preceded the incident, improving the resiliency and recovery of queued work and runner assignment, and adding automatic recovery for self-hosted runners affected by similar failure conditions. We are also making additional improvements to reduce the risk of cascading failures and accelerate recovery during large-scale Actions disruptions.

monitoringAug 7, 02:03 AM

During the incident, some Actions Runner Controller (ARC) runner pods became stuck in an idle state. Affected users can delete those pods using kubectl or redeploy their Actions Runner Controller application. ARC will automatically create replacement runners.The next releases of Actions Runner and Actions Runner Controller will include an automatic recovery mechanism, preventing the need for these manual steps in the future.Some workflow-triggering events, including push and pull request events, were not processed during the incident and cannot be replayed automatically. Customers may need to repeat the triggering action by pushing a new commit, updating the pull request, or manually re-running the workflow where applicable.

monitoringAug 7, 12:59 AM

We’re investigating reports that some Actions Runner Controller runners are taking longer than expected to recover. We’ll provide an update as our investigation progresses.

monitoringAug 7, 12:06 AM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingAug 7, 12:05 AM

The degradation affecting Actions and Pages has been mitigated. We are monitoring to ensure stability.

investigatingAug 7, 12:01 AM

System-wide queues have been drained, and new jobs are being processed as expected. The fix for self-hosted runners not picking up jobs has been fully rolled out.Webhook-triggered Actions workflows have been restored to full throughput. GitHub Pages, Copilot code review, and Copilot coding agent are showing recovery. Migrations using GitHub Enterprise Importer remain paused as a precaution.We are monitoring all affected services for sustained recovery and will provide another update shortly.

investigatingAug 7, 12:01 AM

System-wide queues have been drained, and new jobs are being processed as expected. The fix for self-hosted runners not picking up jobs has been fully rolled out.Webhook-triggered Actions workflows have been restored to full throughput. GitHub Pages, Copilot code review, and Copilot coding agent are showing recovery. Migrations using GitHub Enterprise Importer remain paused as a precaution.We are monitoring all affected services for sustained recovery and will provide another update shortly.

investigatingAug 6, 11:13 PM

We have deployed fixes that address runners being assigned invalid jobs and are taking additional steps to clear the backlog of affected jobs. Job completion rates for running workflows have improved significantly, with success rates now at 99%. Global queues for hosted runner assignment are nearly burned down and concurrency queues for customers are being processed. Another change was deployed to accelerate processing the backlog of job requests.We are gradually restoring throughput for webhook-triggered Actions workflows and monitoring system stability. We have deployed a fix for self-hosted runners that were not picking up jobs and are enabling it incrementally.GitHub Pages, Copilot code review, and Copilot coding agent may still experience intermittent failures or delays. Migrations using GitHub Enterprise Importer remain paused.We continue to monitor recovery across all affected services and will provide another update as conditions improve.

investigatingAug 6, 10:18 PM

We continue to make progress on the issue affecting GitHub Actions. We have deployed a fix that addresses runners being assigned jobs that are no longer valid, and are seeing improvement in job completion rates. For workflow runs that are starting, success rates have increased significantly and are now at 97%. Standard and larger runners are now draining queued work. A change is also in progress to mitigate issues with existing self-hosted runners that are not picking up jobs.Webhook triggers remain throttled to support recovery. Many push and pull request events are not yet triggering new workflow runs, and we are working to safely restore full throughput.GitHub Pages, Copilot code review, and Copilot coding agent may still experience failures or delays. Migrations using GitHub Enterprise Importer remain paused.We are continuing to monitor recovery and will provide another update as conditions improve.

investigatingAug 6, 09:30 PM

We are continuing to work on an issue affecting GitHub Actions. Webhook triggers remain throttled to aid recovery, so many push and pull request events are not triggering new workflow runs.We identified runners being assigned jobs that are no longer valid and are deploying a change to address this issue. Both GitHub-hosted and self-hosted runners are affected.Copilot code review, Copilot coding agent, and GitHub Pages may experience failures or delays. Migrations using GitHub Enterprise Importer have been paused to support mitigation efforts.

investigatingAug 6, 08:34 PM

We are continuing to work on an issue affecting GitHub Actions. Webhook triggers are currently throttled to help with recovery and and we are processing approximately 15% of webhooks, so many events such as pushes and pull requests are not triggering workflow runs. Of jobs queued, approximately 65% are succeeding, improved from a low of 30 to 40% earlier in this incident.We have narrowed the remaining impact to runners that are stuck retrying jobs that are no longer available. Both GitHub-hosted and self-hosted runners are affected, and we are working to recover them.Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected.

investigatingAug 6, 07:43 PM

We are continuing to work on an issue affecting GitHub Actions. Capacity remains constrained and jobs may still be delayed or fail while it recovers gradually. Customers using self-hosted runners may see errors or rate limiting when runners register.  Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed. Our engineers remain actively engaged.

investigatingAug 6, 06:46 PM

We are continuing to work on an issue affecting multiple GitHub services. Workflow runs are still failing, and jobs may remain queued for an extended period before starting or may time out. Jobs using GitHub-hosted runners are particularly affected while capacity is constrained. Customers using self-hosted runners may see errors or rate limiting when runners register. Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed. Recovery is taking longer than we expected, and engineers remain actively engaged.

investigatingAug 6, 06:11 PM

We are continuing to work on an issue affecting multiple GitHub services. Workflow runs are still failing or delayed in starting, and some queued jobs may time out. Customers using self-hosted runners may see errors or rate limiting when runners register. Copilot code review, Copilot coding agent, hosted runners, and migrations using GitHub Enterprise Importer may also be affected. Webhook deliveries may be delayed.Engineers have applied further mitigations and are continuing to work towards full recovery.

investigatingAug 6, 05:40 PM

We are continuing to work on an issue affecting multiple GitHub services.Workflow runs are failing or delayed in starting, and some queued jobs may time out. Copilot code review, Copilot coding agent, hosted runners, and migrations using GitHub Enterprise Importer might also affected. Webhook deliveries may be delayed. Engineers have applied a number of mitigations and are rolling out a further fix across all affected systems now.

investigatingAug 6, 05:02 PM

We are continuing to work on the issue affecting GitHub Actions. Workflow runs are still failing or delayed in starting, and some queued jobs may time out. Some requests to the Actions API are returning errors. Customers running migrations with GitHub Enterprise Importer may see failures. Our engineers have applied several mitigations and are rolling out a further fix now.

investigatingAug 6, 04:33 PM

Actions and Pages are experiencing degraded availability. We are continuing to investigate.

investigatingAug 6, 04:27 PM

We are continuing to work on the issue affecting GitHub Actions. Some workflow runs are still delayed or failing to complete, and some requests to the Actions API are returning errors. Customers running migrations with GitHub Enterprise Importer may also see failures. Engineers are actively working towards full recovery.

investigatingAug 6, 04:27 PM

Pages is experiencing degraded performance. We are continuing to investigate.

investigatingAug 6, 04:19 PM

Pages is operating normally.

investigatingAug 6, 03:53 PM

Pages is experiencing degraded performance. We are continuing to investigate.

investigatingAug 6, 03:45 PM

We are investigating errors affecting GitHub Actions. Some workflow runs are failing to start or failing partway through, and some requests to the Actions REST API are returning errors. Some customers may also see unexpected rate limiting in their workflows. Engineers have identified the source of the disruption and are actively working on a mitigation

investigatingAug 6, 03:41 PM

Actions is experiencing degraded availability. We are continuing to investigate.

investigatingAug 6, 03:22 PM

We are investigating reports of degraded performance for Actions

minorresolvedAug 6, 03:03 PM — Resolved Aug 6, 04:22 PM

Incident with Pages - Deployment Lag

3 updates
resolvedAug 6, 04:22 PM

On August 6, 2026, at 07:00 UTC, a configuration change inadvertently reduced the capacity of the service that processes GitHub Pages deployments. As traffic increased over the following hours, latency in the deployment pipeline progressively increased. At 12:09 UTC, latency crossed the alerting threshold and the team began investigating. We reverted the invalid configuration and applied additional mitigations, including reducing status deployment processing to lower the load on our Redis cluster. Latency returned to normal levels at 15:40 UTC. Customer impact occurred from 11:34 to 15:32 UTC. During this period, we failed to process approximately 128,000 deployments. We have updated our alerts to detect elevated processing latency sooner and to notify us immediately when latency causes deployment processing failures. We've confirmed this incident was not fully captured by our availability metrics. In the coming days, we'll update how GitHub Pages availability is measured so incidents like this are accurately reflected going forward.

monitoringAug 6, 03:50 PM

The degradation affecting Pages has been mitigated. We are monitoring to ensure stability.

investigatingAug 6, 03:03 PM

We are investigating reports of degraded performance for Pages

minorresolvedAug 5, 11:38 AM — Resolved Aug 5, 01:00 PM

Some Copilot Cloud Agent jobs not starting

4 updates
resolvedAug 5, 01:00 PM

On August 5, 2026, between 11:02 and 11:54 UTC, the GitHub Copilot cloud agent service was degraded and new cloud agent jobs were delayed from starting. During this period 100% of newly submitted agent jobs were affected. The incident was limited to delay of cloud agent jobs. No jobs were lost and the queued backlog was processed by 13:00 UTC. This was due to an internal rate limit used to protect service availability that was enabled more broadly than intended delaying more traffic than expected. The service recovered when the rate limit window expired. We then tuned the control so it no longer affected unrelated coding agent traffic. We are working to improve the control's scoping and our monitoring and alerting to reduce our time to detection and mitigation of similar issues in the future.

monitoringAug 5, 12:10 PM

Copilot cloud agent jobs have recovered and the backlog of delayed jobs is being processed.

monitoringAug 5, 12:01 PM

The degradation has been mitigated. We are monitoring to ensure stability.

investigatingAug 5, 11:38 AM

We are investigating reports of impacted performance for some GitHub services.

minorresolvedAug 3, 09:53 AM — Resolved Aug 3, 11:25 AM

Incident with Copilot

5 updates
resolvedAug 3, 11:25 AM

On 2026-08-03, between 06:52 and 11:25 UTC, some GitHub Copilot users experienced errors when using chat and agent features. Requests to list the available models failed, and because every chat or agent interaction begins by retrieving the list of models, affected users saw their requests fail. On average about 3% of these model-listing requests failed during the incident (roughly 97% succeeded), but failures were significantly higher during peak-traffic periods, at times approaching 100% for the affected internal lookups. Approximately 4,066 users were affected in a single 60-minute window, concentrated among IDE-based clients. The underlying AI models themselves remained healthy throughout.The incident was caused by an increase in how often clients requested the model list, which pushed an internal user-authorization lookup past a rate limit; the rate-limited responses were surfaced to users as errors. We mitigated the impact by increasing how long Copilot caches that authorization lookup, which reduced load on the internal service, and we have additional capacity and rate-limit changes in progress. To prevent recurrence we are improving monitoring for this class of failure, adjusting cache and rate-limit settings, and coordinating with client teams on request patterns.

monitoringAug 3, 11:19 AM

The degradation affecting Copilot has been mitigated. We are monitoring to ensure stability.

investigatingAug 3, 10:35 AM

We are still seeing intermittent errors with Copilot, and are continuing to investigate and consider mitigations.

investigatingAug 3, 09:54 AM

We are experiencing degraded availability for chat & agent models in Copilot. Multiple models are impacted and customers may experience requests failing. We are investigating and will provide an update as soon as possible.

investigatingAug 3, 09:53 AM

We are investigating reports of degraded performance for Copilot

minorresolvedAug 1, 06:03 PM — Resolved Aug 1, 06:44 PM

Incident with Copilot AI Model Providers

6 updates
resolvedAug 1, 06:44 PM

On August 1, 2026, between 17:47 UTC and 18:20 UTC, users of the Fable 5 model in GitHub Copilot experienced increased request failures and latency. The average failure rate across all Copilot requests was 0.007%, while failures for Fable 5 peaked at 5.6%. Other models remained available. This was caused by degradation of an upstream model provider.The affected endpoint recovered, and we monitored the service until error rates and latency returned to normal levels. We are working to add endpoint redundancy to mitigate similar provider issues in the future.

monitoringAug 1, 06:23 PM

The issues with our upstream model provider have been resolved, and Fable 5 is once again available in Copilot products and IDE surfaces.We will continue monitoring to ensure stability, but mitigation is complete.

monitoringAug 1, 06:20 PM

The degradation affecting Copilot AI Model Providers has been mitigated. We are monitoring to ensure stability.

investigatingAug 1, 06:20 PM

We are experiencing degraded availability for the Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

investigatingAug 1, 06:03 PM

We are seeing increased error rates from specific upstream AI Model Providers

investigatingAug 1, 06:03 PM

We are investigating reports of degraded performance for Copilot AI Model Providers

minorresolvedAug 1, 11:16 AM — Resolved Aug 1, 12:30 PM

Degraded availability GPT 5.6 Luna

5 updates
resolvedAug 1, 12:30 PM

On August 1st, 2026, the GPT-5.6 Luna model in GitHub Copilot experienced degraded availability in intermittent time intervals between ~08:05 UTC and ~16:30 UTC. Specifically the timeframes observed were 10:00-10:20 UTC, 10:45-11:50 UTC, 13:00-14:25 UTC, and 16:00-16:30 UTC. During this time, requests to GPT-5.6 Luna in Copilot chat and IDE surfaces frequently failed or timed out. This was caused by an issue with an upstream model provider. Other Copilot models were not affected, and users could continue working by selecting another model or 'Auto'. Availability for GPT-5.6 Luna fully recovered once the provider resolved their outage at 16:30 UTC.

investigatingAug 1, 12:29 PM

The issues with our upstream model provider have been resolved, and GPT-5.6 Luna is once again available in Copilot products and IDE surfaces.We will continue monitoring to ensure stability, but mitigation is complete.

investigatingAug 1, 12:13 PM

We keep working with our upstream model provider, and are observing recovery. We continue monitoring to ensure stability.

investigatingAug 1, 11:20 AM

We are experiencing degraded availability for the GPT-5.6 Luna model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot

investigatingAug 1, 11:16 AM

We are investigating reports of degraded performance for Copilot AI Model Providers

July 2026(1 incident)

minorresolvedJul 30, 09:07 AM — Resolved Jul 30, 10:12 AM

Copilot model Claude Fable 5 experiencing elevated errors

4 updates
resolvedJul 30, 10:12 AM

On July 30, 2026, the Claude Fable 5 model in GitHub Copilot experienced degraded availability for approximately 73 minutes, from 08:33 to 09:46 UTC. During this time, requests to Claude Fable 5 in Copilot chat and IDE surfaces frequently failed or timed out. This was caused by an issue with an upstream model provider. Other Copilot models were not affected, and users could continue working by selecting another model or 'Auto'. Availability for Claude Fable 5 fully recovered once the provider resolved their outage at 09:46 UTC, and we confirmed resolution at 10:12 UTC.

investigatingJul 30, 10:11 AM

The issues with our upstream model provider have been resolved, and Claude Fable 5 is once again available in Copilot products and IDE surfaces.We will continue monitoring to ensure stability, but mitigation is complete.

investigatingJul 30, 09:17 AM

We are experiencing degraded availability for the Claude Fable 5 model in Copilot products and IDE surfaces. This is due to an issue with an upstream model provider. While we work with them to resolve the issue, we recommend choosing another model or selecting 'Auto' to continue using Copilot.

investigatingJul 30, 09:07 AM

We are investigating reports of degraded performance for Copilot AI Model Providers

📡 Tired of checking GitHub status manually?

Better Stack monitors uptime every 30 seconds and alerts you instantly when GitHub goes down.

Start Free Monitoring →