F

Fly.io Outage History

Past incidents and downtime events

Complete history of Fly.io outages, incidents, and service disruptions. Showing 50 most recent incidents.

August 2026(2 incidents)

criticalresolvedAug 3, 03:16 PM — Resolved Aug 3, 06:50 PM

Inbound connection failure to Fly Apps

1 update
resolvedAug 3, 03:16 PM

A BGP misconfiguration while provisioning new edge capacity caused most traffic from Europe endpoints to be dropped, between 14:50 UTC and 15:02 UTC. The misconfiguration has been fixed and we are implementing safeguards against this kind of issue in the future.

minorresolvedAug 1, 05:01 PM — Resolved Aug 1, 06:30 PM

Managed Postgres v2 control plane issues in iad

3 updates
resolvedAug 1, 06:30 PM

This incident has been resolved.

monitoringAug 1, 05:26 PM

A fix has been implemented and we are monitoring the results.

investigatingAug 1, 05:01 PM

We are investigating an issue with the control plane for Managed Postgres v2 in the IAD region. Creating new v2 clusters in the IAD region may fail at this time. Existing clusters continue to run, but may experience lagging backups.

July 2026(17 incidents)

noneresolvedJul 31, 04:44 PM — Resolved Jul 31, 08:11 PM

Capacity issues in CDG

3 updates
resolvedJul 31, 08:11 PM

This incident has been resolved.

monitoringJul 31, 07:02 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 31, 04:44 PM

Creating machines in CDG region may fail at this time with an "no capacity available in cdg" message. Existing apps continue to run.

minorresolvedJul 31, 01:34 PM — Resolved Jul 31, 02:51 PM

Outbound email issues

3 updates
resolvedJul 31, 02:51 PM

This incident has been resolved.

monitoringJul 31, 02:38 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 31, 01:34 PM

Our dashboard is failing to send outbound email. Emails such as new account verification or password reset may fail to send at this time.

minorresolvedJul 29, 12:41 PM — Resolved Jul 29, 01:26 PM

Increased API latency

2 updates
resolvedJul 29, 01:26 PM

This incident has been resolved.

investigatingJul 29, 12:41 PM

We are investigating some database issues causing high latency on some API endpoints and dashboard operations. You may experience intermittent "503 service unavailable" errors at this time. Currently deployed apps continue to run.

criticalresolvedJul 20, 07:10 AM — Resolved Jul 20, 05:09 PM

High number of 5XX on the Machines API and dashboard

8 updates
resolvedJul 20, 05:09 PM

This incident has been resolved.

monitoringJul 20, 01:34 PM

We are still working on fixing degraded Managed Postgres clusters.

monitoringJul 20, 09:14 AM

Some Managed Postgres v1 clusters are degraded. We are working on fixing them. Managed Postgres v2 is unaffected.

monitoringJul 20, 08:02 AM

A fix has been implemented and we are monitoring the results.

identifiedJul 20, 07:52 AM

We've identified an internal service providing authentication to our Machines API has failed, our team is currently looking at our options for restoring this service. Existing Machines/Apps will continue to run as normal. Thank you for your patience.

identifiedJul 20, 07:51 AM

We've identified an internal service providing authentication to our Machines API has failed, our team is currently looking at our options for restoring this service. Existing Machines/Apps will continue to run as normal. Thank you for your patience.

investigatingJul 20, 07:48 AM

We are continuing to investigate this issue.

investigatingJul 20, 07:10 AM

Existing machines are unaffected. We are investigating the issue.

minorresolvedJul 19, 03:48 AM — Resolved Jul 19, 04:52 AM

Egress IPv6 issues in BOM

2 updates
resolvedJul 19, 04:52 AM

This incident has been resolved.

identifiedJul 19, 03:48 AM

We have identified an upstream issue that is preventing egress IPv6 addresses in BOM from reaching parts of the internet, and we're currently working with an upstream provider to resolve this issue. Normal IPv6 addresses remain unaffected.

noneresolvedJul 17, 06:30 PM — Resolved Jul 17, 06:30 PM

Brief Flycast / MPG interruption in YYZ

1 update
resolvedJul 17, 07:13 PM

A bad deployment momentarily caused issues with Flycast connectivity, and, by extension, MPG, in our YYZ region. The deployment was immediately rolled back and connectivity was restored shortly after.

minorresolvedJul 16, 11:34 AM — Resolved Jul 16, 01:44 PM

Edge proxy issues

3 updates
resolvedJul 16, 01:44 PM

This incident has been resolved.

monitoringJul 16, 12:27 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 16, 11:34 AM

We are investigating increased connection latency and "connection reset" errors from our edge proxy. Apps continue to run, but requests may experience increased connection latency or fail at this time.

minorresolvedJul 16, 11:54 AM — Resolved Jul 16, 12:17 PM

App creation failing

2 updates
resolvedJul 16, 12:17 PM

This incident has been resolved.

identifiedJul 16, 11:54 AM

An issue with our Machines API is causing app creations to fail in some cases. We are working on a fix.

majorresolvedJul 14, 02:51 PM — Resolved Jul 14, 04:59 PM

Partial outage in SJC

3 updates
resolvedJul 14, 04:59 PM

This incident has been resolved.

monitoringJul 14, 03:53 PM

A fix has been implemented and we are monitoring the results. Apps should be reachable at this point in time.

identifiedJul 14, 02:51 PM

A subset of hosts in SJC are currently offline. Some apps may be unreachable at this time.

minorresolvedJul 13, 10:25 PM — Resolved Jul 13, 10:58 PM

Some DFW hosts offline

3 updates
resolvedJul 13, 10:58 PM

This incident has been resolved.

monitoringJul 13, 10:36 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 13, 10:25 PM

A subset of hosts in DFW are currently offline, and we're investigating the issue.

minorresolvedJul 9, 03:21 PM — Resolved Jul 9, 04:18 PM

Delays starting Depot Builders in IAD

3 updates
resolvedJul 9, 04:18 PM

This incident has been resolved.

monitoringJul 9, 03:43 PM

A fix has been implemented and we are seeing improvements in builder performance, latency, and error rates. We are continuing to monitor for a full recovery. Customers still seeing issues can trigger a fly-hosted builder based deploy with `fly deploy --depot=false`.

identifiedJul 9, 03:21 PM

We have identified an issue causing delays or failures when starting depoting builders located in the IAD region. Customers with builders in IAD may see delays or timeouts starting builds during `fly deploy`. We are working on a fix. In the meantime customers can trigger a fly-hosted builder based deploy with `fly deploy --depot=false`. You can also change your builder region away from IAD via the `settings` tab of your Fly dashboard. We recommend DFW or ORD as alternate builder regions at this time. Please note, your builder region can differ from the region your machines are in.

minorresolvedJul 9, 03:33 PM — Resolved Jul 9, 04:04 PM

Registry performance issues

3 updates
resolvedJul 9, 04:04 PM

This incident has been resolved.

monitoringJul 9, 03:55 PM

A fix has been implemented and we are monitoring the results.

identifiedJul 9, 03:33 PM

The fly.io registry is currently experiencing capacity constraints that reduced performance and may lead to temporary high latency or failed pushes. We are currently working to add capacity and restore service.

majorresolvedJul 3, 05:00 PM — Resolved Jul 3, 06:11 PM

Partial Outage in ORD

4 updates
resolvedJul 3, 06:11 PM

This incident has been resolved.

monitoringJul 3, 05:49 PM

A fix has been implemented and we are monitoring the results.

identifiedJul 3, 05:02 PM

We've identified the issue as a networking hardware failure impacting a subset of hosts at one of our Upstream providers in ORD. We are working with our provider to restore connectivity.

investigatingJul 3, 05:00 PM

We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly. Deploys with machines on these hosts may fail at this time. Some Managed Postgres clusters in ORD region may be unavailable or see connectivity issues at this time.

majorresolvedJul 3, 12:11 AM — Resolved Jul 3, 05:59 AM

Partial outage in ORD

7 updates
resolvedJul 3, 05:59 AM

This incident has been resolved.

monitoringJul 3, 04:41 AM

Customer workloads are now starting and we're monitoring the affected hosts. Affected Managed Postgres instances will be investigated.

identifiedJul 3, 04:21 AM

Power restoration is ongoing and we're making sure the hosts are healthy before starting customer workloads to avoid issues. Customer impact remains and updates to come.

identifiedJul 3, 02:59 AM

Power restoration work in a subset of ORD is still in progress and impact remains ongoing for a subset of hosts and some Managed Postgres clusters.

identifiedJul 3, 12:54 AM

Our provider has advised us their facilities team is working on restoring the power, we'll provide another update as soon as we learn more.

identifiedJul 3, 12:29 AM

We've identified and reported power issues with one of our upstream providers in ORD. We're waiting for an update from our upstream for a resolution. Some Managed Postgres clusters in ORD will be unavailable due to placement.

investigatingJul 3, 12:11 AM

We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly. Deploys with machines on these hosts may fail at this time. Some Managed Postgres clusters in ORD region may be unavailable at this time.

criticalresolvedJul 2, 09:38 PM — Resolved Jul 2, 11:18 PM

Errors issuing new SSL certificates

4 updates
resolvedJul 2, 11:18 PM

This incident has been resolved.

monitoringJul 2, 10:50 PM

A fix has been implemented upstream and certificates are being issued successfully. We will continue to monitor.

identifiedJul 2, 10:26 PM

The issue has been identified and we are awaiting a fix.

investigatingJul 2, 09:38 PM

We are currently investigating errors when issuing new SSL certificates for hostnames.

minorresolvedJul 1, 01:06 PM — Resolved Jul 1, 01:57 PM

Static Egress IPv6 issues in NRT

3 updates
resolvedJul 1, 01:57 PM

This incident has been resolved.

monitoringJul 1, 01:44 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 1, 01:06 PM

We are investigating issues with static egress IPv6 addresses in NRT region. Apps using static egress IPs may experience connectivity failures to some destinations.

majorresolvedJul 1, 06:14 AM — Resolved Jul 1, 07:50 AM

Elevated API Errors

5 updates
resolvedJul 1, 07:50 AM

This incident has been resolved.

monitoringJul 1, 07:26 AM

Background jobs have caught up and the API is fully operational. We are continuing to monitor service health.

identifiedJul 1, 07:05 AM

A fix has been put in place, and we are no longer seeing elevated API errors. Some dashboard actions will be delayed while background processing catches up.

identifiedJul 1, 06:27 AM

The cause of the errors has been identified and we are working on a fix.

investigatingJul 1, 06:14 AM

We are investigating elevated errors with our GraphQL API and background job processing

June 2026(22 incidents)

majorresolvedJun 29, 01:03 AM — Resolved Jun 30, 10:19 PM

Delayed Metrics

12 updates
resolvedJun 30, 10:19 PM

This incident has been resolved.

identifiedJun 30, 07:58 PM

Almost all metrics have caught up aside from a small handful in `sin` and `syd`. We're continuing to monitor and expect these to complete in the next few hours.

identifiedJun 30, 04:50 PM

Backlogged metrics are still being processed. We're bringing extra processing capacity online to speed up the process.

identifiedJun 30, 05:17 AM

Backlogged metrics are still being processed and ingestion delays persist for some customers

identifiedJun 30, 12:58 AM

Backlogged metrics are still being processed and ingestion delays persist for some customers, but we’re continuing to see gradual improvement as the backlog continues to drain.

identifiedJun 29, 08:02 PM

Backlogged metrics data is still being processed, continue to monitor the situation while the system catches up.

identifiedJun 29, 03:56 PM

Metric ingestion is still delayed for some customers. We’re seeing gradual improvement and continue to monitor the situation while the system catches up.

identifiedJun 29, 08:43 AM

We are still in the process of increasing the metrics cluster throughput to catch up with metric backlog.

identifiedJun 29, 05:34 AM

We are continuing to see delayed metrics due to resource contention on a subset of metrics ingestion hosts, and we are working to rebalance ingestion traffic and reduce the backlog.

identifiedJun 29, 02:58 AM

We are continuing to see delayed metric exports from a number of hosts. Users will see delayed or missing metrics for machines on impacted hosts at this time.

identifiedJun 29, 01:46 AM

We've identified an issue on multiple hosts causing delayed metric ingestion into our hosted fly-metrics.net dashboards. We are working on a fix.

investigatingJun 29, 01:03 AM

We are investigating issues with customer facing metrics in the fly-metrics.net dashboard. Users may see delayed or missing metrics at this time.

minorresolvedJun 30, 01:45 PM — Resolved Jun 30, 02:03 PM

Egress IP issues in SIN and NRT

3 updates
resolvedJun 30, 02:03 PM

This incident has been resolved.

monitoringJun 30, 01:49 PM

A fix has been implemented and we are monitoring the results.

identifiedJun 30, 01:45 PM

We are aware of egress IP issues in SIN and NRT and are working on a fix. Some machines in SIN and NRT using egress IPs may temporarily lose connectivity or otherwise see degraded performance.

majorresolvedJun 28, 08:10 AM — Resolved Jun 28, 09:20 PM

Metrics currently experiencing issues

2 updates
resolvedJun 28, 09:20 PM

This incident has been resolved.

investigatingJun 28, 08:10 AM

We are currently investigating an issue with our metrics cluster.

majorresolvedJun 26, 10:17 PM — Resolved Jun 26, 11:48 PM

IPv6 Connectivity Issues in EWR

3 updates
resolvedJun 26, 11:48 PM

This incident has been resolved.

monitoringJun 26, 11:14 PM

A fix has been implemented and we are monitoring the results.

investigatingJun 26, 10:17 PM

One of our upstream providers is experiencing IPv6 network connectivity problems in EWR. Apps with machines on affected hosts may have impacted connectivity to certain IPv6 destinations while they investigate and resolve this issue.

minorresolvedJun 25, 01:41 PM — Resolved Jun 25, 03:38 PM

Deploys defaulting to Fly-hosted Builders

3 updates
resolvedJun 25, 03:38 PM

This incident has been resolved.

monitoringJun 25, 03:08 PM

We are seeing improvements in Depot builder provision times and are switching the default deploy strategy back to them. We will continue to monitor builder performance closely. Users with a preference can trigger a Fly builder based deploy with `fly deploy --depot=false` or use `fly deploy --depot=true` to force a depot-based deployment.

investigatingJun 25, 01:41 PM

We are investigating delays provisioning Depot backed builders for deploys. We have switched the default `fly deploy` strategy to use fly hosted builders at this time. Users can still trigger a depot based deploy with `fly deploy --depot=true`

minorresolvedJun 25, 12:11 PM — Resolved Jun 25, 03:18 PM

Elevated control plane latency

4 updates
resolvedJun 25, 03:18 PM

This incident has been resolved.

monitoringJun 25, 01:48 PM

A fix has been implemented and we are monitoring the results.

identifiedJun 25, 01:20 PM

The issue has been identified and a fix is being implemented.

investigatingJun 25, 12:11 PM

We're addressing elevated control plane latency and saturation affecting the BOM and NRT regions. Apps with machines in this region might experience longer response times and possible timeouts (502 errors).

minorresolvedJun 24, 01:42 AM — Resolved Jun 24, 06:57 AM

Degraded networking in North America

4 updates
resolvedJun 24, 06:57 AM

This incident has been resolved.

identifiedJun 24, 03:42 AM

Some 6PN Private Networking traffic remains impacted into and out of our LAX region, pending upstream resolution.

identifiedJun 24, 02:04 AM

Most networking is largely healthy between primary North American regions. Some Machines may see ongoing packet loss and higher latency communicating with other Machines on certain routes. We're continuing to monitor the backbone health upstream.

investigatingJun 24, 01:42 AM

We are currently investigating degraded network performance between sites in NA due to an upstream incident

majorresolvedJun 22, 08:15 PM — Resolved Jun 22, 09:56 PM

Network issues in SIN, NRT

2 updates
resolvedJun 22, 09:56 PM

This incident has been resolved.

identifiedJun 22, 08:15 PM

Our upstream provider is continuing to experience network issues in SIN and NRT regions. Apps running in those regions may be unreachable or experience high packet loss at this time.

majorresolvedJun 22, 05:10 PM — Resolved Jun 22, 07:48 PM

SIN, NRT network issues

5 updates
resolvedJun 22, 07:48 PM

This incident has been resolved.

identifiedJun 22, 06:33 PM

We are continuing to experience network issues with an upstream provider in SIN and NRT regions.

monitoringJun 22, 05:44 PM

A fix has been implemented and we are monitoring the results.

identifiedJun 22, 05:24 PM

Our upstream provider has identified the issue and is working on a fix. Apps running in NRT (Tokyo) region may also have issues reaching certain destinations at this time.

investigatingJun 22, 05:10 PM

We are investigating an upstream network issue in the SIN (Singapore) region. Apps may be unreachable or have higher packet loss.

minorresolvedJun 19, 06:05 PM — Resolved Jun 19, 09:53 PM

Log search unavailable

4 updates
resolvedJun 19, 09:53 PM

Most queued historical logs have been ingested and should now be available through log search. Log ingestion rates have returned to normal levels.

monitoringJun 19, 06:19 PM

We've applied a fix for this issue. Historical logs are currently backfilling. We will post an update once logs have finished backfilling and current logs are being ingested normally.

investigatingJun 19, 06:10 PM

Log search is available; however, new app logs since ~1 hour ago are missing and new logs are not being ingested. We are continuing to investigate.

investigatingJun 19, 06:05 PM

We are investigating an issue causing application log search to be unavailable. This is affecting the Fly Metrics log search panels, and historical application logs initially returned from the `fly logs` command. Streaming logs using `fly logs`, the Live Logs page in the dashboard, and Fly Log Shipper services continue to work as expected.

majorresolvedJun 17, 03:21 AM — Resolved Jun 17, 04:55 AM

Network Issues in SIN

5 updates
resolvedJun 17, 04:55 AM

This incident has been resolved.

monitoringJun 17, 04:38 AM

Network connectivity in SIN has been fully restored. We're continuing to monitor.

identifiedJun 17, 04:22 AM

We are seeing recovery of network connectivity between SIN and most destinations. We're continuing to work with our upstream provider to resolve the remaining issues.

identifiedJun 17, 03:47 AM

Some machines in SIN are unreachable. A few Managed Postgres clusters may fail to fail-over or update. We are in the process of fixing this with our upstream provider.

investigatingJun 17, 03:21 AM

We are currently investigating network connectivity issues in the SIN region. Hosted apps may be unavailable.

criticalresolvedJun 15, 03:03 PM — Resolved Jun 15, 04:41 PM

Macaroon Auth + Machines API Issues

9 updates
resolvedJun 15, 04:41 PM

This incident has been resolved and we are seeing all platform functions operate normally.

monitoringJun 15, 04:15 PM

A fix has been implemented and we are monitoring the results.

identifiedJun 15, 04:15 PM

We have deployed another change and are seeing wider improvements in platform stability across all regions. Performance is trending to normal, though users may still see some degradation at this time. We are continuing to closely monitor to ensure full, stable recovery. We will provide another update in 15m.

identifiedJun 15, 04:00 PM

We are seeing elevated cluster errors with Managed Postgres clusters as the MPG control plane recovers from the API outage. MPG Users may see elevated rates of failing or slow connections, as well as increased primary/replica failovers. The managed postgres team is addressing any degraded clusters. We will provide a further update within 15m.

identifiedJun 15, 03:57 PM

We continue to seeing degraded performance and increased errors with the Machines API and other platform features at this time. We are continuing to work on fully restoring service.

identifiedJun 15, 03:41 PM

An initial fix has been deployed and we are starting to see platform features recover. Users may still see degraded performance and intermittent failures at this time. We are continuing to address the issue to ensure a full stable recovery.

identifiedJun 15, 03:30 PM

We have identified the cause of the issue and are working on deploying a fix. Impacted features remain unavailable or degraded at this time. Already running customer applications/machines remain available. MPG clusters remain generally reachable and healthy, however new clusters cannot be provisioned and failovers may not complete. We will provide another update within 15 minutes.

identifiedJun 15, 03:13 PM

We are continuing to address this issue. Platform authentication with macaroon based tokens is currently failing. Platform features that authenticate with macaroons including Machines API operations, Dashboard logins, some flyctl commands, fly-metrics.net Grafana, and deployments are failing at this time. Existing, running customer applications and machines remain reachable and running. We will provide another update within 15 minutes

investigatingJun 15, 03:03 PM

We are investigating issues with Macaroon based authentication. This is impacting parts of the Machines API, Fly.io Dashboard, some flyctl operations and other platform features that rely on this.

noneresolvedJun 15, 06:20 AM — Resolved Jun 15, 06:39 AM

MPG cluster provisioning is broken

2 updates
resolvedJun 15, 06:39 AM

This incident has been resolved.

investigatingJun 15, 06:20 AM

New MPG cluster provisioning is broken. Existing MPG clusters are not affected. Newly created organizations may see errors while SSHing into their machines. We are investigating the issue.

majorresolvedJun 12, 02:35 AM — Resolved Jun 12, 04:30 AM

Elevated Sprites error rates in SIN

3 updates
resolvedJun 12, 04:30 AM

This incident has been resolved.

monitoringJun 12, 03:24 AM

A fix has been implemented and we are seeing error rates for sprites in SIN normalize. We are continuing to monitor to ensure full recovery.

investigatingJun 12, 02:35 AM

We are investigating elevated 500 / internal server error rates with Sprites in the SIN region. Users may see increased errors when accessing sprites located in this region, or for requests to the Sprites API originating from SIN

minorresolvedJun 11, 01:16 AM — Resolved Jun 11, 11:12 AM

Increased network latency in North America

8 updates
resolvedJun 11, 02:12 PM

This incident has been resolved.

monitoringJun 11, 10:09 AM

Our upstream paths have been fixed. We are monitoring the results.

identifiedJun 11, 09:35 AM

We are working with our upstream network provider to address periodic loss of connectivity over transit in ord

identifiedJun 11, 07:58 AM

We are seeing a recurrence of elevated latency across some hosts in ORD impacting the MPG control plane and a subset of clusters there. We are working to address this.

monitoringJun 11, 05:24 AM

The network between our regions has been performing well for the majority of traffic. We're still continuing to monitor a few impacted routes in North America that may be seeing elevated latency and packet loss.

monitoringJun 11, 03:30 AM

Impacted NA backbones have been sidestepped where possible, and we're continuing to monitor network health.

monitoringJun 11, 01:59 AM

Managed Postgres in ORD has returned to normal operation. We continue to see slightly elevated latencies and loss over transits in North America. We are working with our upstream network providers to improve performance.

investigatingJun 11, 01:16 AM

We are investigating elevated network instability across some hosts in ORD. Apps and managed postgres clusters on impacted hosts may see elevated latency or networking errors at this time.

majorresolvedJun 10, 07:39 PM — Resolved Jun 10, 07:52 PM

Ingress Traffic issues in GRU

2 updates
resolvedJun 10, 07:52 PM

This incident has been resolved.

identifiedJun 10, 07:39 PM

Some of our edge nodes in GRU has suffered an error that crashed some of the critical services. We're currently working to bring them back online. Some traffic entering through GRU (i.e. users connecting from around GRU) may be temporarily affected: connections may see increased latency or be occasionally dropped.

majorresolvedJun 9, 06:20 PM — Resolved Jun 9, 06:42 PM

Emergency maintenance of Petsem causing some control plane errors

3 updates
resolvedJun 9, 06:42 PM

This incident has been resolved.

monitoringJun 9, 06:28 PM

The maintenance has been completed and control plane functions should recover to normal. We're monitoring for any further complications.

identifiedJun 9, 06:20 PM

We're performing an emergency maintenance on Petsem, our secrets management service. Some control plane write operations may temporarily fail, for example, creating new apps or secrets. Existing apps and machines should keep functioning without issues.

majorresolvedJun 9, 02:57 PM — Resolved Jun 9, 04:15 PM

Managed Postgres Control Plane Issues in IAD

4 updates
resolvedJun 9, 09:31 PM

This incident has been resolved.

monitoringJun 9, 03:19 PM

An initial fix has been implemented and connectivity to all impacted clusters has been restored. We are continuing to monitor to ensure stable recovery.

identifiedJun 9, 03:06 PM

We are continuing to address this issue. Some clusters in IAD are unavailable at this time, some users may have seen unexpected cluter restarts. We are working on restoring normal performance for all clusters in IAD

investigatingJun 9, 02:57 PM

We are investigating MPG control plane instability in a subset of the IAD region. A small number of clusters in the region may have seen unexpected failovers or connection issues over the past 30m.

minorresolvedJun 9, 09:16 AM — Resolved Jun 9, 10:01 AM

egress ips are broken in ORD

3 updates
resolvedJun 9, 10:01 AM

This incident has been resolved.

monitoringJun 9, 09:37 AM

A fix has been implemented and we are monitoring the results.

investigatingJun 9, 09:16 AM

Egress ips are broken in most of ORD, we are currently investigating this issue

minorresolvedJun 8, 10:31 AM — Resolved Jun 8, 12:30 PM

Capacity issues in ARN region

2 updates
resolvedJun 8, 12:30 PM

This incident has been resolved.

investigatingJun 8, 10:31 AM

The ARN region is low on available host capacity. Creating new machines, or starting currently stopped/suspended machines, may fail at this time. We are working on provisioning new host capacity in the region. Please consider using nearby regions if possible.

noneresolvedJun 4, 01:44 PM — Resolved Jun 4, 05:03 PM

Consul cluster degradation

3 updates
resolvedJun 4, 05:03 PM

We have restored the degraded Consul cluster. All affected functionality is now working correctly: Unmanaged Postgres and LiteFS with dynamic leases.

monitoringJun 4, 04:02 PM

We have restored the degraded Consul cluster and are monitoring for stability. All affected functionality should now be working correctly: Unmanaged Postgres and LiteFS with dynamic leases.

identifiedJun 4, 01:44 PM

One of our Consul clusters is in degraded state due to a failed node. This can cause issues with LiteFS primary node selection, Unmanaged Postgres (14.x and older *only*), and creation of new Unmanaged Postgres clusters. Impact is limited to these legacy products and does not affect deployments, running Fly applications in general, or Managed Postgres clusters.

minorresolvedJun 1, 06:33 PM — Resolved Jun 1, 06:43 PM

Issues with flyctl ssh console and Machines OIDC

3 updates
resolvedJun 1, 06:43 PM

This incident has been resolved.

monitoringJun 1, 06:35 PM

A fix has been implemented and we are monitoring the results.

investigatingJun 1, 06:33 PM

We're currently investigating an issue affecting flyctl ssh console functionality and machines' OIDC tokens.

May 2026(9 incidents)

majorresolvedMay 31, 06:34 AM — Resolved May 31, 10:52 AM

IPv6 outage for some machines in ORD

3 updates
resolvedMay 31, 10:52 AM

This incident has been resolved.

monitoringMay 31, 07:03 AM

A fix has been implemented and we are monitoring the results.

investigatingMay 31, 06:34 AM

We're working with our upstream providers to investigate an IPv6 networking failure in ORD.

minorresolvedMay 30, 01:29 PM — Resolved May 30, 01:37 PM

Private networking issues in SYD

3 updates
resolvedMay 30, 01:37 PM

This incident has been resolved.

monitoringMay 30, 01:32 PM

A fix has been implemented and we are monitoring the results.

identifiedMay 30, 01:29 PM

Due to an upstream provider issue, Private Networking (6PN) is currently degraded in SYD region. Communication between Machines in SYD region and Machines in other regions may fail at this time. Newly created Machines in SYD may fail to sync to other regions (may not show up in Machines API List endpoint, or state may be incorrect). Additionally, TLS certificate resolution and Machines API authentication may currently be degraded in the SYD region. We are working with our upstream providers to resolve this issue.

minorresolvedMay 30, 02:42 AM — Resolved May 30, 04:21 AM

Elevated deployment errors

4 updates
resolvedMay 30, 04:21 AM

This incident has been resolved.

monitoringMay 30, 03:37 AM

A fix has been implemented and we are monitoring the results.

identifiedMay 30, 03:13 AM

We identified the issue and are working on a fix.

investigatingMay 30, 02:42 AM

We're investigating an increase in deployment errors affecting some users. At this time, creating or updating Machines may erroneously fail with the message: "We require your billing information, please add it at https://fly.io/dashboard//billing".

minorresolvedMay 29, 07:09 PM — Resolved May 29, 07:45 PM

Networking issues in ORD

3 updates
resolvedMay 29, 07:45 PM

This incident has been resolved.

monitoringMay 29, 07:21 PM

A fix has been implemented and we are monitoring the results.

identifiedMay 29, 07:09 PM

We are aware of increased latency and connection drops for clients located near Chicago (ORD) and are currently working on a fix.

minorresolvedMay 29, 06:59 AM — Resolved May 29, 04:50 PM

Networking issues in ORD

3 updates
resolvedMay 29, 04:50 PM

This incident has been resolved.

monitoringMay 29, 10:22 AM

A fix has been implemented and we are monitoring the results.

investigatingMay 29, 06:59 AM

We are currently investigating increased latency and dropped connections in ORD (Chicago).

minorresolvedMay 29, 12:42 AM — Resolved May 29, 01:27 AM

Networking issues in ORD

3 updates
resolvedMay 29, 01:27 AM

This incident has been resolved.

monitoringMay 29, 01:00 AM

A fix has been implemented and we are monitoring the results.

investigatingMay 29, 12:42 AM

We are currently investigating increased latency and dropped connections in ORD (Chicago).

minorresolvedMay 28, 09:08 PM — Resolved May 28, 11:07 PM

Increased latency in SJC

6 updates
resolvedMay 28, 11:07 PM

This incident has been resolved.

monitoringMay 28, 10:34 PM

We are continuing to monitor for any further issues.

monitoringMay 28, 10:33 PM

A fix has been implemented and we are monitoring the results.

identifiedMay 28, 10:25 PM

We're still seeing connection issues originating from SJC/LAX/West Coast US and are still investigating.

monitoringMay 28, 09:56 PM

We have implemented a mitigation for the issue and are monitoring for stability. This issue happened on the edge hosts in SJC, which means that any traffic near/around SJC would have been affected during the incident. We will provide a more detailed write-up on our Infra Log (https://fly.io/infra-log/) later once we have a better picture of the entire incident.

investigatingMay 28, 09:08 PM

We are currently investigating increased latency in SJC.

majorresolvedMay 27, 11:54 AM — Resolved May 27, 12:52 PM

Private networking issues in SYD

3 updates
resolvedMay 27, 12:52 PM

This incident has been resolved.

monitoringMay 27, 12:13 PM

A fix has been implemented and we are monitoring the results.

identifiedMay 27, 11:54 AM

Due to an upstream provider issue, Private Networking (6PN) is currently degraded in SYD region. Communication between Machines in SYD region and Machines in other regions may fail at this time. Newly created Machines in SYD may fail to sync to other regions (may not show up in Machines API List endpoint, or state may be incorrect). Additionally, TLS certificate resolution and Machines API authentication may currently be degraded in the SYD region. We are working with our upstream providers to resolve this issue.

minorresolvedMay 27, 03:50 AM — Resolved May 27, 04:36 AM

Elevated GraphQL API Latency

3 updates
resolvedMay 27, 04:36 AM

This incident has been resolved.

monitoringMay 27, 04:07 AM

A fix has been implemented and we are monitoring the results.

investigatingMay 27, 03:50 AM

We are investigating elevated API Latency. Users may see delays or errors creating apps, as well as on some dashboard pages.

📡 Tired of checking Fly.io status manually?

Better Stack monitors uptime every 30 seconds and alerts you instantly when Fly.io goes down.

Start Free Monitoring →