F

Fly.io Outage History

Past incidents and downtime events

Complete history of Fly.io outages, incidents, and service disruptions. Showing 50 most recent incidents.

September 2026(5 incidents)

minorresolvedSep 15, 04:08 AM — Resolved Sep 15, 05:34 AM

Depot builder failures

4 updates
resolvedSep 15, 05:34 AM

This incident has been resolved.

monitoringSep 15, 04:51 AM

A fix has been implemented and we are monitoring builds. Standard flyctl builds should be working again for customers in all regions.

identifiedSep 15, 04:45 AM

We have identified an issue with deploying via Depot for users connecting through our SYD and JNB regions. Affected customers in these regions can deploy successfully using the --depot=false or --buildkit arguments to flyctl.

investigatingSep 15, 04:08 AM

We are investigating reports of Depot builds failing for some customers. Affected customers can deploy successfully using the --depot=false or --buildkit arguments to flyctl.

minorresolvedSep 12, 09:22 PM — Resolved Sep 12, 10:12 PM

Network issues in US West Coast

3 updates
resolvedSep 12, 10:12 PM

This incident has been resolved.

monitoringSep 12, 09:57 PM

Private networking between Fly Machines is resolved, and most outbound connections are healthy. We're continuing to monitor the network, and some issues will still be expected from clients physically located in US West until upstream transit issues are resolved.

investigatingSep 12, 09:22 PM

We are investigating upstream network issues from US West Coast (SJC, LAX). Apps hosted in US West regions may experience higher latency or packet loss, and requests from clients physically located in US West may experience higher latency.

minorresolvedSep 2, 11:03 PM — Resolved Sep 3, 12:41 AM

Sprites API Partial Outage

3 updates
resolvedSep 3, 12:41 AM

This incident has been resolved.

monitoringSep 2, 11:33 PM

Error rates have decreased. We are continuing to monitor the API health.

investigatingSep 2, 11:03 PM

We're aware of a problem affecting a subset of Sprites users. We are investigating the source of the issue.

minorresolvedSep 2, 02:18 PM — Resolved Sep 2, 03:57 PM

Upstream network issues

5 updates
resolvedSep 2, 03:57 PM

This incident has been resolved.

monitoringSep 2, 03:33 PM

A fix has been implemented and we are monitoring the results.

identifiedSep 2, 02:56 PM

We're updating the affected region list to also include SJC since this seems to be a wider upstream issue in US West Coast.

identifiedSep 2, 02:48 PM

We have put in some temporary mitigations along with our providers. However since the root cause of this issue lies within a bigger upstream transit provider, you may continue to see some elevated latency and connection issues in/around affected regions. We're still working closely with them to resolve the root cause.

identifiedSep 2, 02:18 PM

We have observed an upstream network issue in LAX. Connections to some destinations may see elevated latency and packet loss. We're working with our upstream to resolve this issue.

majorresolvedSep 2, 07:21 AM — Resolved Sep 2, 07:44 AM

API background job queue failure

3 updates
resolvedSep 2, 07:44 AM

This incident has been resolved.

monitoringSep 2, 07:31 AM

A fix has been implemented and we are monitoring the results.

investigatingSep 2, 07:21 AM

We are investigating an issue with the background job runner for our API. Actions that require a background job, such as creating apps, assigning IP addresses, or creating/renewing certificates, may fail at this time.

August 2026(21 incidents)

noneresolvedAug 31, 03:30 PM — Resolved Aug 31, 03:30 PM

HTTP/2 traffic disruptions

1 update
resolvedAug 31, 03:42 PM

A configuration update caused temporary failures for incoming HTTP/2 traffic for Fly Machines located on a subset of hosts for a few minutes. This incident has since been resolved. Managed Postgres depends on HTTP/2 and some control plane ops may have been affected as well. However, downstream Postgres connections were unlikely to have been affected by this since they do not use the HTTP2 handler.

minorresolvedAug 31, 06:26 AM — Resolved Aug 31, 08:10 AM

Packet loss in ORD

3 updates
resolvedAug 31, 08:10 AM

This incident has been resolved.

monitoringAug 31, 07:39 AM

Packet loss in ORD is improving and impacted services are recovering; we’re continuing to monitor for intermittent issues

investigatingAug 31, 06:26 AM

Due to an upstream provider, we are seeing ~50% packet loss on a subset of hosts in ORD. Some MPG clusters in ORD are slow to replicate as a result.

noneresolvedAug 30, 10:45 PM — Resolved Aug 30, 10:45 PM

Sprite deletion jobs failing

1 update
resolvedAug 30, 10:45 PM

We saw Sprite deletion jobs failing between 21:18 and 22:05 UTC. This issue has been resolved.

minorresolvedAug 28, 09:56 PM — Resolved Aug 28, 10:36 PM

Networking Issues in GRU

5 updates
resolvedAug 28, 10:36 PM

This incident has been resolved.

identifiedAug 28, 10:36 PM

Our upstream provider has implemented a fix. Network performance in GRU has normalized.

identifiedAug 28, 10:12 PM

We are seeing a recurrance in networking issues in GRU. Some apps in the region may experience increased latency or packet loss. We are working with our upstream networking provider to resolve.

monitoringAug 28, 10:00 PM

Networking performance in GRU has normalized and we are no longer seeing issues. We are continuing to monitor to ensure a full recovery.

investigatingAug 28, 09:56 PM

We are investigating networking issues impacting some hosts in GRU (São Paulo, Brazil) region. Some apps in GRU may experience increased latency or packet loss.

minorresolvedAug 28, 08:09 AM — Resolved Aug 28, 10:22 AM

Increased packet loss

2 updates
resolvedAug 28, 10:22 AM

This incident has been resolved.

investigatingAug 28, 08:09 AM

We are currently investigating this issue.

minorresolvedAug 26, 06:14 PM — Resolved Aug 26, 06:43 PM

WireGuard gateway issues

4 updates
resolvedAug 26, 06:43 PM

This incident has been resolved.

monitoringAug 26, 06:33 PM

Our testing and monitoring indicates gateways should be back to normal; if you are still having problem using `flyctl ssh console`, try restarting the `flyctl` agent by `flyctl agent restart`.

monitoringAug 26, 06:27 PM

A fix has been implemented and we are monitoring the results.

investigatingAug 26, 06:14 PM

We are investigating issues with our WireGuard gateways. Some CLI commands like `flyctl ssh console` or `flyctl proxy` may not work at this time. Apps continue to run.

minorresolvedAug 24, 10:23 AM — Resolved Aug 24, 01:24 PM

Metrics in some regions are lagging behind

3 updates
resolvedAug 24, 01:24 PM

This is now resolved

monitoringAug 24, 12:47 PM

All hosts have caught up with metrics and we're monitoring the situation

investigatingAug 24, 10:23 AM

We are currently experiencing some metrics lag on servers in some regions. We are provisioning more metric processing instances to accommodate the backlog and catch up.

majorresolvedAug 23, 01:28 AM — Resolved Aug 23, 02:10 AM

Network Issues in LAX Region

3 updates
resolvedAug 23, 02:10 AM

This incident has been resolved.

monitoringAug 23, 02:04 AM

Upstream networking issues have resolved.

investigatingAug 23, 01:28 AM

We are investigating network issues in the Los Angeles region. Apps may experience higher latency or be unreachable at this time.

minorresolvedAug 20, 07:30 PM — Resolved Aug 20, 07:30 PM

Temporary DNS resolution failure

1 update
resolvedAug 20, 08:07 PM

A BGP configuration error caused our Anycast DNS to route to some nodes without the proper DNS infrastructure. The issue was temporary and was resolved as soon as we removed that node from BGP.

noneresolvedAug 20, 01:54 PM — Resolved Aug 20, 02:16 PM

Oauth/Macaroon Errors from flyctl

3 updates
resolvedAug 20, 02:16 PM

This incident has been resolved.

monitoringAug 20, 02:05 PM

A fix has been deployed and this error should no longer be occurring. We're monitoring to ensure full recovery.

identifiedAug 20, 01:54 PM

We have identified an issue causing authentication errors for some operations from `flyctl`. These operations are failing with an error like: `This endpoint no longer accepts legacy OAuth tokens (starting with `fo1_`). Please use a macaroon token (starting with `fm2_`) instead. We have identified the issue and are rolling out a fix

majorresolvedAug 20, 07:25 AM — Resolved Aug 20, 07:57 AM

MPG (v1) partially down in ORD

4 updates
resolvedAug 20, 07:57 AM

This incident has been resolved.

monitoringAug 20, 07:32 AM

A fix has been implemented and we are monitoring the results.

investigatingAug 20, 07:26 AM

We are continuing to investigate this issue.

investigatingAug 20, 07:25 AM

We had an issue with the ord-0 Fly Kubernetes cluster, and many MPG clusters are failing to restart. Our MPG team is actively working on it.

noneresolvedAug 18, 08:00 PM — Resolved Aug 19, 12:00 AM

6PN Networking issue in YYZ

1 update
resolvedAug 19, 01:15 AM

6PN networking issues between some machines in YYZ during a rollout which was rolled back once we noticed errors. During this time some machines were unable to talk to internal resources like other DBs, other apps or MPG clusters.

noneresolvedAug 17, 01:19 PM — Resolved Aug 17, 03:41 PM

No capacity in ARN

2 updates
resolvedAug 17, 03:41 PM

The capacity issue in the ARN region has been resolved.

investigatingAug 17, 01:19 PM

New machines may fail to create in ARN because we lack capacity.

majorresolvedAug 14, 07:30 PM — Resolved Aug 14, 08:33 PM

Secrets service outage

3 updates
resolvedAug 14, 08:33 PM

This incident has been resolved.

monitoringAug 14, 08:08 PM

We have failed over the secrets database to a replica, and the Machines API appears healthy now. We are monitoring for any further issues.

identifiedAug 14, 07:30 PM

We are working to recover our secrets service after a failed deployment. Apps continue to run, but it is not possible to create new apps or update secrets at this time.

minorresolvedAug 13, 05:45 PM — Resolved Aug 13, 07:41 PM

IPv6 Networking Issues

6 updates
resolvedAug 13, 07:41 PM

This incident has been resolved.

monitoringAug 13, 06:37 PM

We are continuing to monitor for any further issues.

monitoringAug 13, 06:36 PM

A fix has been implemented and we are monitoring the results.

identifiedAug 13, 06:14 PM

The issue has been identified and a fix is being implemented.

investigatingAug 13, 05:51 PM

We are continuing to investigate this issue.

investigatingAug 13, 05:45 PM

We are currently investigating degraded ipv6 networking on a subset of hosts

majorresolvedAug 9, 02:50 AM — Resolved Aug 9, 07:10 AM

Increased app-not-found errors

6 updates
resolvedAug 9, 07:10 AM

This incident has been resolved.

monitoringAug 9, 06:42 AM

A fix has been implemented and we are monitoring the results

identifiedAug 9, 06:03 AM

We’ve deployed an additional mitigation to further reduce Corrosion retry pressure and are seeing improvement; we’re continuing to monitor while remaining affected nodes catch up.

identifiedAug 9, 04:58 AM

We’ve applied a mitigation to reduce the impact from Corrosion batch insertion retries and are continuing to monitor while affected nodes catch up.

identifiedAug 9, 03:50 AM

We have identified the issue as failed insertions in a subset of Corrosion batches. These failures trigger retries, which can cause timeouts for other batches.

investigatingAug 9, 02:50 AM

We are currently investigating app-not-found errors returned by Machines API calls made shortly after creating new applications.

minorresolvedAug 5, 08:26 PM — Resolved Aug 5, 08:41 PM

MPG IAD data plane degraded for new clusters

2 updates
resolvedAug 5, 08:41 PM

Etcd is stable. Services are back to normal.

identifiedAug 5, 08:26 PM

High CPU pressure on a shared etcd instance is causing lags on MPG creation in the IAD region

majorresolvedAug 4, 07:03 PM — Resolved Aug 5, 03:22 AM

MPG creation is failing in GRU due to lack of capacity

3 updates
resolvedAug 5, 03:22 AM

This incident has been resolved.

monitoringAug 4, 08:59 PM

We tweaked hosts to allow for more machine allocation. We'll be monitoring the region over the next hours.

identifiedAug 4, 07:03 PM

New MPG clusters may fail to create in GRU because we lack capacity.

minorresolvedAug 4, 02:53 PM — Resolved Aug 4, 11:58 PM

Certificate issuance delays

5 updates
resolvedAug 4, 11:58 PM

This incident has been resolved.

monitoringAug 4, 08:36 PM

A fix has been implemented and we are monitoring the results.

identifiedAug 4, 06:26 PM

We have the size of TLS certificate issuance backlog under control, but are still seeing some remaining issues and are currently working to clean up the edge cases.

identifiedAug 4, 05:02 PM

We believe we have identified the issue and are releasing a fix.

investigatingAug 4, 02:53 PM

We're currently investigating issues related to certificate issuance. Certificates may be delayed for new custom domains.

criticalresolvedAug 3, 03:16 PM — Resolved Aug 3, 06:50 PM

Inbound connection failure to Fly Apps

1 update
resolvedAug 3, 03:16 PM

A BGP misconfiguration while provisioning new edge capacity caused most traffic from Europe endpoints to be dropped, between 14:50 UTC and 15:02 UTC. The misconfiguration has been fixed and we are implementing safeguards against this kind of issue in the future.

minorresolvedAug 1, 05:01 PM — Resolved Aug 1, 06:30 PM

Managed Postgres v2 control plane issues in iad

3 updates
resolvedAug 1, 06:30 PM

This incident has been resolved.

monitoringAug 1, 05:26 PM

A fix has been implemented and we are monitoring the results.

investigatingAug 1, 05:01 PM

We are investigating an issue with the control plane for Managed Postgres v2 in the IAD region. Creating new v2 clusters in the IAD region may fail at this time. Existing clusters continue to run, but may experience lagging backups.

July 2026(17 incidents)

noneresolvedJul 31, 04:44 PM — Resolved Jul 31, 08:11 PM

Capacity issues in CDG

3 updates
resolvedJul 31, 08:11 PM

This incident has been resolved.

monitoringJul 31, 07:02 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 31, 04:44 PM

Creating machines in CDG region may fail at this time with an "no capacity available in cdg" message. Existing apps continue to run.

minorresolvedJul 31, 01:34 PM — Resolved Jul 31, 02:51 PM

Outbound email issues

3 updates
resolvedJul 31, 02:51 PM

This incident has been resolved.

monitoringJul 31, 02:38 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 31, 01:34 PM

Our dashboard is failing to send outbound email. Emails such as new account verification or password reset may fail to send at this time.

minorresolvedJul 29, 12:41 PM — Resolved Jul 29, 01:26 PM

Increased API latency

2 updates
resolvedJul 29, 01:26 PM

This incident has been resolved.

investigatingJul 29, 12:41 PM

We are investigating some database issues causing high latency on some API endpoints and dashboard operations. You may experience intermittent "503 service unavailable" errors at this time. Currently deployed apps continue to run.

criticalresolvedJul 20, 07:10 AM — Resolved Jul 20, 05:09 PM

High number of 5XX on the Machines API and dashboard

8 updates
resolvedJul 20, 05:09 PM

This incident has been resolved.

monitoringJul 20, 01:34 PM

We are still working on fixing degraded Managed Postgres clusters.

monitoringJul 20, 09:14 AM

Some Managed Postgres v1 clusters are degraded. We are working on fixing them. Managed Postgres v2 is unaffected.

monitoringJul 20, 08:02 AM

A fix has been implemented and we are monitoring the results.

identifiedJul 20, 07:52 AM

We've identified an internal service providing authentication to our Machines API has failed, our team is currently looking at our options for restoring this service. Existing Machines/Apps will continue to run as normal. Thank you for your patience.

identifiedJul 20, 07:51 AM

We've identified an internal service providing authentication to our Machines API has failed, our team is currently looking at our options for restoring this service. Existing Machines/Apps will continue to run as normal. Thank you for your patience.

investigatingJul 20, 07:48 AM

We are continuing to investigate this issue.

investigatingJul 20, 07:10 AM

Existing machines are unaffected. We are investigating the issue.

minorresolvedJul 19, 03:48 AM — Resolved Jul 19, 04:52 AM

Egress IPv6 issues in BOM

2 updates
resolvedJul 19, 04:52 AM

This incident has been resolved.

identifiedJul 19, 03:48 AM

We have identified an upstream issue that is preventing egress IPv6 addresses in BOM from reaching parts of the internet, and we're currently working with an upstream provider to resolve this issue. Normal IPv6 addresses remain unaffected.

noneresolvedJul 17, 06:30 PM — Resolved Jul 17, 06:30 PM

Brief Flycast / MPG interruption in YYZ

1 update
resolvedJul 17, 07:13 PM

A bad deployment momentarily caused issues with Flycast connectivity, and, by extension, MPG, in our YYZ region. The deployment was immediately rolled back and connectivity was restored shortly after.

minorresolvedJul 16, 11:34 AM — Resolved Jul 16, 01:44 PM

Edge proxy issues

3 updates
resolvedJul 16, 01:44 PM

This incident has been resolved.

monitoringJul 16, 12:27 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 16, 11:34 AM

We are investigating increased connection latency and "connection reset" errors from our edge proxy. Apps continue to run, but requests may experience increased connection latency or fail at this time.

minorresolvedJul 16, 11:54 AM — Resolved Jul 16, 12:17 PM

App creation failing

2 updates
resolvedJul 16, 12:17 PM

This incident has been resolved.

identifiedJul 16, 11:54 AM

An issue with our Machines API is causing app creations to fail in some cases. We are working on a fix.

majorresolvedJul 14, 02:51 PM — Resolved Jul 14, 04:59 PM

Partial outage in SJC

3 updates
resolvedJul 14, 04:59 PM

This incident has been resolved.

monitoringJul 14, 03:53 PM

A fix has been implemented and we are monitoring the results. Apps should be reachable at this point in time.

identifiedJul 14, 02:51 PM

A subset of hosts in SJC are currently offline. Some apps may be unreachable at this time.

minorresolvedJul 13, 10:25 PM — Resolved Jul 13, 10:58 PM

Some DFW hosts offline

3 updates
resolvedJul 13, 10:58 PM

This incident has been resolved.

monitoringJul 13, 10:36 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 13, 10:25 PM

A subset of hosts in DFW are currently offline, and we're investigating the issue.

minorresolvedJul 9, 03:21 PM — Resolved Jul 9, 04:18 PM

Delays starting Depot Builders in IAD

3 updates
resolvedJul 9, 04:18 PM

This incident has been resolved.

monitoringJul 9, 03:43 PM

A fix has been implemented and we are seeing improvements in builder performance, latency, and error rates. We are continuing to monitor for a full recovery. Customers still seeing issues can trigger a fly-hosted builder based deploy with `fly deploy --depot=false`.

identifiedJul 9, 03:21 PM

We have identified an issue causing delays or failures when starting depoting builders located in the IAD region. Customers with builders in IAD may see delays or timeouts starting builds during `fly deploy`. We are working on a fix. In the meantime customers can trigger a fly-hosted builder based deploy with `fly deploy --depot=false`. You can also change your builder region away from IAD via the `settings` tab of your Fly dashboard. We recommend DFW or ORD as alternate builder regions at this time. Please note, your builder region can differ from the region your machines are in.

minorresolvedJul 9, 03:33 PM — Resolved Jul 9, 04:04 PM

Registry performance issues

3 updates
resolvedJul 9, 04:04 PM

This incident has been resolved.

monitoringJul 9, 03:55 PM

A fix has been implemented and we are monitoring the results.

identifiedJul 9, 03:33 PM

The fly.io registry is currently experiencing capacity constraints that reduced performance and may lead to temporary high latency or failed pushes. We are currently working to add capacity and restore service.

majorresolvedJul 3, 05:00 PM — Resolved Jul 3, 06:11 PM

Partial Outage in ORD

4 updates
resolvedJul 3, 06:11 PM

This incident has been resolved.

monitoringJul 3, 05:49 PM

A fix has been implemented and we are monitoring the results.

identifiedJul 3, 05:02 PM

We've identified the issue as a networking hardware failure impacting a subset of hosts at one of our Upstream providers in ORD. We are working with our provider to restore connectivity.

investigatingJul 3, 05:00 PM

We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly. Deploys with machines on these hosts may fail at this time. Some Managed Postgres clusters in ORD region may be unavailable or see connectivity issues at this time.

majorresolvedJul 3, 12:11 AM — Resolved Jul 3, 05:59 AM

Partial outage in ORD

7 updates
resolvedJul 3, 05:59 AM

This incident has been resolved.

monitoringJul 3, 04:41 AM

Customer workloads are now starting and we're monitoring the affected hosts. Affected Managed Postgres instances will be investigated.

identifiedJul 3, 04:21 AM

Power restoration is ongoing and we're making sure the hosts are healthy before starting customer workloads to avoid issues. Customer impact remains and updates to come.

identifiedJul 3, 02:59 AM

Power restoration work in a subset of ORD is still in progress and impact remains ongoing for a subset of hosts and some Managed Postgres clusters.

identifiedJul 3, 12:54 AM

Our provider has advised us their facilities team is working on restoring the power, we'll provide another update as soon as we learn more.

identifiedJul 3, 12:29 AM

We've identified and reported power issues with one of our upstream providers in ORD. We're waiting for an update from our upstream for a resolution. Some Managed Postgres clusters in ORD will be unavailable due to placement.

investigatingJul 3, 12:11 AM

We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly. Deploys with machines on these hosts may fail at this time. Some Managed Postgres clusters in ORD region may be unavailable at this time.

criticalresolvedJul 2, 09:38 PM — Resolved Jul 2, 11:18 PM

Errors issuing new SSL certificates

4 updates
resolvedJul 2, 11:18 PM

This incident has been resolved.

monitoringJul 2, 10:50 PM

A fix has been implemented upstream and certificates are being issued successfully. We will continue to monitor.

identifiedJul 2, 10:26 PM

The issue has been identified and we are awaiting a fix.

investigatingJul 2, 09:38 PM

We are currently investigating errors when issuing new SSL certificates for hostnames.

minorresolvedJul 1, 01:06 PM — Resolved Jul 1, 01:57 PM

Static Egress IPv6 issues in NRT

3 updates
resolvedJul 1, 01:57 PM

This incident has been resolved.

monitoringJul 1, 01:44 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 1, 01:06 PM

We are investigating issues with static egress IPv6 addresses in NRT region. Apps using static egress IPs may experience connectivity failures to some destinations.

majorresolvedJul 1, 06:14 AM — Resolved Jul 1, 07:50 AM

Elevated API Errors

5 updates
resolvedJul 1, 07:50 AM

This incident has been resolved.

monitoringJul 1, 07:26 AM

Background jobs have caught up and the API is fully operational. We are continuing to monitor service health.

identifiedJul 1, 07:05 AM

A fix has been put in place, and we are no longer seeing elevated API errors. Some dashboard actions will be delayed while background processing catches up.

identifiedJul 1, 06:27 AM

The cause of the errors has been identified and we are working on a fix.

investigatingJul 1, 06:14 AM

We are investigating elevated errors with our GraphQL API and background job processing

June 2026(7 incidents)

majorresolvedJun 29, 01:03 AM — Resolved Jun 30, 10:19 PM

Delayed Metrics

12 updates
resolvedJun 30, 10:19 PM

This incident has been resolved.

identifiedJun 30, 07:58 PM

Almost all metrics have caught up aside from a small handful in `sin` and `syd`. We're continuing to monitor and expect these to complete in the next few hours.

identifiedJun 30, 04:50 PM

Backlogged metrics are still being processed. We're bringing extra processing capacity online to speed up the process.

identifiedJun 30, 05:17 AM

Backlogged metrics are still being processed and ingestion delays persist for some customers

identifiedJun 30, 12:58 AM

Backlogged metrics are still being processed and ingestion delays persist for some customers, but we’re continuing to see gradual improvement as the backlog continues to drain.

identifiedJun 29, 08:02 PM

Backlogged metrics data is still being processed, continue to monitor the situation while the system catches up.

identifiedJun 29, 03:56 PM

Metric ingestion is still delayed for some customers. We’re seeing gradual improvement and continue to monitor the situation while the system catches up.

identifiedJun 29, 08:43 AM

We are still in the process of increasing the metrics cluster throughput to catch up with metric backlog.

identifiedJun 29, 05:34 AM

We are continuing to see delayed metrics due to resource contention on a subset of metrics ingestion hosts, and we are working to rebalance ingestion traffic and reduce the backlog.

identifiedJun 29, 02:58 AM

We are continuing to see delayed metric exports from a number of hosts. Users will see delayed or missing metrics for machines on impacted hosts at this time.

identifiedJun 29, 01:46 AM

We've identified an issue on multiple hosts causing delayed metric ingestion into our hosted fly-metrics.net dashboards. We are working on a fix.

investigatingJun 29, 01:03 AM

We are investigating issues with customer facing metrics in the fly-metrics.net dashboard. Users may see delayed or missing metrics at this time.

minorresolvedJun 30, 01:45 PM — Resolved Jun 30, 02:03 PM

Egress IP issues in SIN and NRT

3 updates
resolvedJun 30, 02:03 PM

This incident has been resolved.

monitoringJun 30, 01:49 PM

A fix has been implemented and we are monitoring the results.

identifiedJun 30, 01:45 PM

We are aware of egress IP issues in SIN and NRT and are working on a fix. Some machines in SIN and NRT using egress IPs may temporarily lose connectivity or otherwise see degraded performance.

majorresolvedJun 28, 08:10 AM — Resolved Jun 28, 09:20 PM

Metrics currently experiencing issues

2 updates
resolvedJun 28, 09:20 PM

This incident has been resolved.

investigatingJun 28, 08:10 AM

We are currently investigating an issue with our metrics cluster.

majorresolvedJun 26, 10:17 PM — Resolved Jun 26, 11:48 PM

IPv6 Connectivity Issues in EWR

3 updates
resolvedJun 26, 11:48 PM

This incident has been resolved.

monitoringJun 26, 11:14 PM

A fix has been implemented and we are monitoring the results.

investigatingJun 26, 10:17 PM

One of our upstream providers is experiencing IPv6 network connectivity problems in EWR. Apps with machines on affected hosts may have impacted connectivity to certain IPv6 destinations while they investigate and resolve this issue.

minorresolvedJun 25, 01:41 PM — Resolved Jun 25, 03:38 PM

Deploys defaulting to Fly-hosted Builders

3 updates
resolvedJun 25, 03:38 PM

This incident has been resolved.

monitoringJun 25, 03:08 PM

We are seeing improvements in Depot builder provision times and are switching the default deploy strategy back to them. We will continue to monitor builder performance closely. Users with a preference can trigger a Fly builder based deploy with `fly deploy --depot=false` or use `fly deploy --depot=true` to force a depot-based deployment.

investigatingJun 25, 01:41 PM

We are investigating delays provisioning Depot backed builders for deploys. We have switched the default `fly deploy` strategy to use fly hosted builders at this time. Users can still trigger a depot based deploy with `fly deploy --depot=true`

minorresolvedJun 25, 12:11 PM — Resolved Jun 25, 03:18 PM

Elevated control plane latency

4 updates
resolvedJun 25, 03:18 PM

This incident has been resolved.

monitoringJun 25, 01:48 PM

A fix has been implemented and we are monitoring the results.

identifiedJun 25, 01:20 PM

The issue has been identified and a fix is being implemented.

investigatingJun 25, 12:11 PM

We're addressing elevated control plane latency and saturation affecting the BOM and NRT regions. Apps with machines in this region might experience longer response times and possible timeouts (502 errors).

minorresolvedJun 24, 01:42 AM — Resolved Jun 24, 06:57 AM

Degraded networking in North America

4 updates
resolvedJun 24, 06:57 AM

This incident has been resolved.

identifiedJun 24, 03:42 AM

Some 6PN Private Networking traffic remains impacted into and out of our LAX region, pending upstream resolution.

identifiedJun 24, 02:04 AM

Most networking is largely healthy between primary North American regions. Some Machines may see ongoing packet loss and higher latency communicating with other Machines on certain routes. We're continuing to monitor the backbone health upstream.

investigatingJun 24, 01:42 AM

We are currently investigating degraded network performance between sites in NA due to an upstream incident

📡 Tired of checking Fly.io status manually?

Better Stack monitors uptime every 30 seconds and alerts you instantly when Fly.io goes down.

Start Free Monitoring →