Grafana Cloud Outage History
Past incidents and downtime events
Complete history of Grafana Cloud outages, incidents, and service disruptions. Showing 50 most recent incidents.
July 2026(19 incidents)
Fleet Managment Interface 404's
1 update
Between 9:45 and 12:00 UTC, we experienced an issue affecting the Fleet Management interface. During this time, the Remote Configuration tab for production stacks returned a 404 error when accessed through the UI. Backend systems were not affected. Configuration-as-code workflows and other non-UI methods of interacting with Fleet Management continued to operate normally throughout the incident. This issue has since been resolved.
Mimir Write Performance Degradation
2 updates
This incident has been resolved. Thank you for your patience.
From 19:24 until 19:40 UTC, backend performance degradation impacted ingestion. We are currently monitoring.
Mimir Partial Write Outage
3 updates
This incident has been resolved.
The outage is recovered as of 16:56 UTC, and we're continuing to monitor.
We're investigating backend degradation which has resulted in a partial write outage beginning around 16:40 UTC. This degradation has also impacted reads and rule evaluation. Issue has been identified and we are working on mitigation.
Issue with Dashboard Views Being Registered
3 updates
We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.
We've identified the cause of the issue impacting dashboard view and error counts in the Dashboards and Folder list. Our team is currently working on implementing a fix and validating the solution. We will provide another update within the next 24 hours, or sooner if we have additional information to share.
This is related to https://status.grafana.com/incidents/rhrk2ck6ly0y which was resolved by mistake. For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly. This is ultimately causing inaccurate or missing view counts. We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.
IRM Mobile App forcing some users to logout
5 updates
This incident has been resolved.
A new iOS mobile app version (v2.39.5) has been released, which includes a fix for this issue. We are monitoring the rollout and its impact to ensure the issue has been fully resolved. If you continue to experience any problems after updating to the latest version, please let us know.
A fix has been submitted for review and will be released as Grafana Mobile v2.39.5 once it is approved. If your app is currently on v2.39.3, we recommend not upgrading to v2.39.4 and instead waiting for v2.39.5 to become available. Users who have already been signed out by this issue can log back into the app and continue using it. We will provide another update once v2.39.5 is available.
Our investigation has narrowed the impact to version 2.39.4 of the Grafana mobile app on iOS. Users should avoid upgrading to this version until an updated release is available. As a precaution, we also recommend ensuring you have alternative notification methods configured (such as SMS, email, or phone calls) if you rely on mobile push notifications for alerting. We are actively working on a resolution and will provide another update as soon as more information is available.
We are aware of an issue in the latest mobile app release that is causing some users to be signed out and asked to log in again. We have reproduced the behavior and are investigating and working on a fix. We will share another update as soon as we have more information.
Issue with Dashboard Views Being Registered
4 updates
At this stage, we are considering the incident resolved.
We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.
We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.
For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly This is ultimately causing inaccurate or missing view counts.
Network Degredation in prod-us-central-0
1 update
From approximately 20:24 UTC - 20:47 UTC a network issue in prod-us-central-0 isolated part of our infrastructure in one availability zone, temporarily cutting off connectivity to a subset of backend services. This caused elevated query latency and some delayed metric evaluations, along with related write-path errors in a logging subsystem. All systems recovered automatically once network connectivity was restored. Impacted customers may have had queries or rule evaluations fail during the incident, but ingestion was not impacted.
PDC Degraded performance
8 updates
This incident has been resolved. Thank you for your patience.
We are now seeing recovery of PDC traffic to the affected customers. Our team will continue to monitor across shifts.
We are continuing to see disruptions across multiple deployments, ranging from degraded performance to full outages for some customers. Additional resources have been engaged to mitigate this issue. We will post updates as they become available.
We are continuing to monitor for any further issues.
We are seeing recovery in some deployments, while others not just yet. We are continuing to monitor this incident.
A fix has been applied and we are seeing recovery. We will continue to monitor this.
We are still currently investigating the issue.
We are currently facing performance degradation on PDC service hosted on Multiple clusters. Our Engineering Team is currently working on fixing the issue, we do apologize for any inconvenience.
Delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1
3 updates
This incident has been resolved.
A fix has been implemented. We are currently monitoring the results.
We are observing delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1. As of yet we have not noticed any customer impact however we are currently observing the cell.
Loki read path in prod-eu-west-2 was down
1 update
Between 5:49 and 5:54 UTC, the read path in prod-eu-west-2 was down. This has completely recovered by 5:57 UTC. Recording rules may have failed to evaluate during this period which may result in gaps.
Grafana rulers crash-looping on prometheus
3 updates
This incident has been resolved.
We are in the process of rolling out the fix. Our engineers are monitoring the progress.
We are investigating issues with the grafana-ruler service on prod-us-east-2 and prod-us-west-0 which are causing periodic crash conditions. A code fix is currently being deployed to mitigate this
Loki Billing Data Lost
1 update
Between 13:57 and 14:20 UTC (23 minutes), a data gap occurred affecting billing usage data for Loki. Data for this window was not recorded and cannot be recovered.
Cannot access dashboard set up as a home page
2 updates
We have found that the issues are related to the grafana.unifiedHomepage feature rollout active 16:50 UTC yesterday, to 10:20 UTC today. We have rolled this back and systems are now working as expected.
We are currently investigating an issue affecting some customers who are unable to load or access their dashboards when configured as their home page. We will update once we have more information on this.
Delayed usage and billing data
1 update
We identified an issue with the internal job that calculates month-to-date usage and cost data, which caused usage attribution and billing dashboards to display stale information. The root cause has been identified and resolved. Your billing dashboard may show a sudden jump in usage. This is expected. It reflects several days of accumulated usage that hadn't been showing up while the issue was ongoing, not a sudden change in your actual usage. This issue does not affect your end-of-month invoice. Your bill will be calculated based on actual usage, not the numbers displayed during this period.
Grafana Cloud IRM alert groups failing to produce alerts in us-east-3 region.
3 updates
We haven't noticed any further issues in this region for alert group processing since yesterday. This incident i fully resolved.
The work on bringing services back to healthy state is completed and alerts should be created correctly, without a delay now. We're keeping the incident open and monitoring for any potential hiccups that might occur.
Some IRM alert groups in us-east-3 region may not be able to produce alerts and could be sluggish/misbehaving in general. The issue started around 01:00 UTC on July 3rd. The issue has been identified and the cause of the issue has been fixed. Our team is actively working on getting the service back to healthy state.
Partial Write Outage for Grafana Cloud Logs in prod-eu-north-0.
1 update
Grafana Cloud Logs in prod-eu-north-0 experienced a 10-minute partial write outage between 13:45 and 13:54 UTC. Impacted users may have experienced 5xx errors during this time.
Some Queries Failing
3 updates
This incident has been resolved. Thank you for your patience.
We are continuing to work on rolling back a PR responsible for this behavior. Once we have more information, we will share it here. Thank you for your patience.
Queries with drop __error__, including log volume histogram queries (the queries that generate the histogram visualization in Grafana), are failing due to a bug with series limit checks. The Root cause has been identified, and we are working on a fix.
Loki and Frontend Observability - Major Outage in prod-us-central-0 region
3 updates
This incident has been resolved by restarting the affected services.
The AWS Logs integration in the same region is affected as well. We will provide further updates as our investigation progresses.
We are currently investigating a major outage in Loki writes and Frontend Observability in the prod-us-central-0 region. Our Engineering team is investigating this and we will provide further updates as our investigation progresses.
Elevated Loki Query Bytes Reporting
2 updates
We’ve implemented a fix and can confirm the issue is fully resolved as of 20:25 UTC. Thank you for your patience.
We are investigating an issue where some customers may see higher Loki query byte usage reported than was actually consumed. This affects usage reporting only; there is no impact to query execution or service availability. The issue began at approximately 13:20 UTC and is ongoing. We expect the issue to be resolved soon and will provide another update as more information becomes available.
June 2026(23 incidents)
Mimir read errors and high latency in prod-eu-west-0
5 updates
Since the mitigation has been applied, we have not seen the errors return. At this point, we are considering the incident resolved.
We've identified a possible cause, and a mitigation is in place to prevent further occurrences.
The errors and latency have now recovered, we continue investigating the root cause.
The errors are recovering, and we are still looking into the root cause of this.
We are currently investigating an issue with Mimir in prod-eu-west-0 we are seeing read errors and high latency. This incident is currently ongoing. The errors are recovering but we are currently looking into the route cause of this.
Confluent API Outage
2 updates
This incident has been resolved.
We are investigating an issue affecting Confluent metrics ingestion across all regions. Due to an elevated error rate on the Confluent side, some metrics may not be ingested, resulting in potential data loss. We are actively investigating the issue and will provide updates as more information becomes available.
Rule evaluation error on cluster prod-gb-south-0
3 updates
This incident has been resolved.
A fix has been applied and we are currently monitoring results.
We are currently investigating Rule Evaluation errors on the cluster prod-gb-south-0 which is leading to error codes showing within the stacks. We are looking into the issue and will update accordingly.
K6 - Test run metrics processing is delayed
6 updates
This incident has been resolved.
We've improved the metric ingestion delay time and are working on additional fixes to bring it down to expected range. Customers can currently expect a delay of 2 to 5 minutes before their test run metrics show up (after starting a test).
Update: Changed incident title to "Test run metrics processing is delayed" We have found the issue and are working on deploying the fix.
We are experiencing intermittent delays with secondary metrics processing for k6 Cloud test runs due to heavy load. We don't expect any data loss or impact on user runs, but results may take longer time to appear in UI.
A small update: The issue is isolated to new test runs, and users can go see the metrics of all the previous test runs
We’re currently investigating an issue causing metrics not to appear during test runs. . Our team is actively working to identify the cause. Thank you for your patience.
Potential Issues Loading Grafana for Users in India
9 updates
This incident has been resolved. Error rates continued to remain near 0 and operations are performing as expected.
Error rates have remained near zero, and we continue to monitor.
We are continuing to monitor for further issues.
We have deployed additional mitigations that should help with remaining errors. We are continuing to monitor error rates.
We’ve verified and begun to implement a fix that will improve loading errors. We are continuing to roll this out to all regions and monitor for efficacy.
We're actively monitoring this issue and working with our 3rd party provider. The next update will be sent on Monday unless there's new information to share.
Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr We are continuing to work with our CSP on this investigation. Impacted users may receive intermittent error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in.
Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr Impacted users may receive intermittent error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in. We continue to work with our CSP on this investigation.
Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr Impacted users may receive error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in. We are currently investigating this issue from our end, and will provide updates as they are available.
Elevated Logs Query Usage — Unexpected Billing Impact
1 update
Starting June 20, 2026 at approximately 20:00 UTC, some Grafana Cloud customers experienced unexpectedly elevated logs query usage. The issue persisted until it was resolved on June 22, 2026 at approximately 16:00 UTC. Our engineering team identified and mitigated the issue. Systems have since stabilized and are operating normally. Grafana Labs is reviewing affected accounts for appropriate remediation.
Rule Evaluation Outage in prod-us-central-0
3 updates
This incident has been resolved. Thank you for your patience.
We’re continuing to track progress post-mitigation. While we don’t have new information to share yet, our team remains actively engaged.
We had an outage affecting rule evaluations between 15:16-15:59 UTC in the prod-us-central-0 region. Our team quickly identified the issue and has since mitigated. The engineering team is monitoring.
Issues with actions in the Grafana IRM mobile app
3 updates
This incident has been resolved. Thank you for your patience.
We've verified a fix in our staging environment to restore functionality to the mobile app. The fix is currently being deployed to production. Thanks for your patience as we continue to roll this out and monitor the resolution.
We're noticing an uptick in users being unable to respond to actions on the mobile app (acknowledging and silencing alerts, for example). Users working in the web UI should not be affected. Ingestion and notification delivery are working as expected. We have a fix in place and are in the process of deploying.
Degraded k6 cloud UI performance
5 updates
This incident has been resolved. Thank you for your patience.
We are continuing to monitor for any further issues.
The root cause of the issue has been identified and a fix has been successfully deployed. We are observing widespread improvements across all systems. Our team is currently monitoring the environment to ensure performance remains stable.
We are continuing to investigate this issue.
We’re currently investigating an issue resulting in degraded k6 cloud UI performance and API response time. Our team is actively working to rectify this issue.
Loki data source-managed alert rules not visible in the Grafana Cloud Alerting UI
6 updates
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
We are continuing to deploy the fix and monitor recovery efforts. As part of the rollout, we identified an issue that required adjustments to our deployment plan, which has extended the timeline for mitigation. Work remains actively underway, and we will share additional updates as progress continues.
Deployment of the fix is still in progress. We are continuing to monitor the rollout and validate recovery across affected systems. We will share further updates as they become available.
Our Engineering Team has implemented a fix which is now being rolled out. We will continue to monitor the situation and update as soon as we have more information.
We have identified an issue where alert rules and alerts managed directly in a Loki data source (data source-managed alerting) are not displayed in the Grafana Cloud Alerting UI. Rules created via Prometheus/Mimir data sources and Grafana-managed alert rules are not affected. Impact is limited to visibility and management in the UI. Affected alert rules continue to evaluate and send notifications normally — there is no impact to alert delivery. Workaround: Loki alert rules can still be viewed and managed directly through the Loki ruler API (for example, using cortextool against /loki/api/v1/rules). A fix has been identified and is in progress. We will provide a further update once it has been rolled out.
Frontend Observability - Suspected commit feature not working as expected
2 updates
A fix has been deployed and the issue after monitoring as been fixed.
We’re currently investigating an issue affecting Frontend Observability product. The "Suspected commit" feature is not currently working as expected. Ingestion and querying is unaffected by this. Our team has identified the cause and is actively working on a fix. Thank you for your patience.
Brief Rule Evaluation Failures in prod-eu-west-3
7 updates
This incident has been resolved. Thank you for your patience.
The incident has been mitigated, and services are operating normally. We continue to monitor the service to ensure full stability.
The incident has been mitigated, and services are operating normally. We are currently monitor the service to ensure full stability.
We’re making ongoing progress on the investigation alongside our upstream provider.
We are continuing to investigate this issue.
Intermittent spikes in rule evaluations continuing.
From 00:20:00 to 00:27:00 and again 00:32:00 to 00:38:00 there were brief spikes in rule evaluation failures. Engineers are investigating.
Brief Loki Prod-012-eu-west-2 Disruption
1 update
Our team had discovered a read issue around 19:35-20:08 UTC. Impact at the time would have provided errors similar to context deadline exceeded (DatasourceError response). This has since been resolved, and should not have caused any data loss, only a short query disruption.
Grafana Dashboards page not displaying when set to ‘View by Folders’
2 updates
This incident has been resolved.
We’re currently investigating an issue affecting The Grafana Dashboards page. When set to view by folders, is currently experiencing an issue where no dashboards are shown. Our team is working on fixing the problem. In the meantime, switching to ‘View as list’ allows access to dashboards as usual”.
Investigating Issues with Data Source-Managed Alerting
3 updates
We continue to observe a continued period of recovery. At this time, we are considering this issue resolved. No further updates.
Our team has implemented a fix and we are currently monitoring the results of this.
We are currently investigating an issue affecting data source-managed alerting management functionality in Grafana Cloud. Customers may experience problems viewing, creating, updating, or managing alerts through Grafana when using data source-managed alerting. This issue is limited to alert management functionality within Grafana. Alert evaluation and backend alerting services continue to operate normally. Direct alerting APIs for Mimir and Loki remain fully operational and are unaffected. Grafana-managed alerting is not impacted. We identified this issue at approximately 20:45 UTC and are actively working on a resolution. We will provide additional updates as more information becomes available. Workaround: Customers can continue to use the direct Mimir and Loki alerting APIs while we work to restore normal functionality.
IRM Degraded Performance
8 updates
This incident has been resolved.
We've released a fix to the IRM app that should restore service for affected customers with issues related to labels. Thanks for your patience while investigating. We're continuing to monitor as we confirm the resolution in place.
We are continuing to work on a fix for this. To further clarify, this issue is not about accessing IRM or alert ingestion/notification/delivery, but rather with handling labels.
The degraded performance is about labels, and we have seen this degradation in more regions.
We are continuing to work on a fix for this issue.
The issue has been identified and a fix is being implemented.
We are continuing to investigate this issue.
We are experiencing access issues in IRM as there are elevated 500 API responses in prod-us-central-0.
Permissions Issues with IRM
8 updates
This incident has been resolved.
Continuing to monitor progress. Most customers affected should have all services restored, with a few remaining customers receiving updates as the rollout finishes out. Thanks again for your patience.
A fix has been released to prod and rolling out across the fleet for IRM, restoring access to affected customers. Thanks for your patience through this work. We're continuing to monitor to confirm we've returned to a steady state.
We've identified an earlier regression in one of our recent code changes that was affecting resolution of our previous fix. We're deploying this change now and applying a hot fix in the interim to restore access quickly
A fix is being deployed now, and we are monitoring the progress.
We've identified an issue with RBAC, and are working on a fix to restore permission services for those affected.
Our engineering team is still investigating this issue. We do not have any new information to share at this time, but will continue to provide timely updates.
We are currently investigating an issue impacting permissions for IRM. As a result, users are not currently getting paged. We will provide updates as they become available.
Silences not Working as Expected
2 updates
This incident has been resolved.
We have identified an issue causing Silences to not work as expected in the Cloud (Mimir) Alertmanager. Grafana Alertmanager is working ok, this is only affecting Data source-managed alerts.
Grafana Assistant Skills Page Blank
4 updates
This incident has been resolved.
The issue has been identified, and we are working on a fix.
We are continuing to investigate this issue.
We are currently investigating an issue affecting the Skills page of Grafana Assistant. Impacted deployments will encounter a blank screen when attempting to access this page. At this time, we have observed partial impact in the us-east-0 and us-central-0 regions, and will provide an update here if the scope of impact expands.
K6 Test Runs Degraded Performance
3 updates
This incident has been resolved.
We have applied a fix, and are monitoring the results.
We are currently investigating an issue causing k6 test runs to take longer than expected to complete, or to time out within Grafana Cloud.
Synthetic Scripted/Browser checks failure
3 updates
This incident has been resolved.
We are in the process of deploying a fix for this issue.
We’re currently investigating an issue affecting Synthetic Monitoring where updates for Scripted/Browser checks might fail. Our team is actively working to identify the cause. Thank you for your patience.
tempo prod-25 write-path-down
2 updates
This incident has been resolved.
Between 21:20 and 22:40 UTC, writes to tempo-prod-25 failed due to an outage. tempo-prod-24 was also affected during an overlapping window from 22:32 to 22:40 UTC."
Alert manager unavailable in prod-us-central-0
3 updates
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
Starting at 18:30 UTC, we noticed alert manager unavailability limited to prod-us-central-0 which affects grafana-managed and datasource-managed alerting, causing disruption to updating alertmanager config and limited disruption to alert sending. We have identified the cause and are in the process of remediation.
May 2026(7 incidents)
Grafana Loki Log Query Issues
3 updates
This incident has been resolved.
We have identified the cause of this incident and a fix has been applied. Normal functions are returning. We are currently monitoring the recovery process.
We’re currently investigating an issue affecting Loki queries in Grafana. We have had reports from customers showing the logs are not loading or showing missing logs. Our team is actively working to identify the cause. Thank you for your patience.
Prometheus Datasource Errors/Outage in prod-us-east-0
4 updates
This incident has been resolved. Thank you for your patience.
We are seeing recovery across affected Prometheus datasources, and error rates have significantly improved. The service is recovering without any required customer action, and our team continues to monitor stability while we investigate the underlying cause. We’ll provide another update as we learn more.
We continue to investigate an issue affecting Prometheus datasources causing intermittent timeouts and unexpected errors, primarily impacting alert rule evaluations. Our team is actively working to identify the cause. Thank you for your patience.
We’re currently investigating an issue affecting Prometheus datasources causing 500 internal or Unexpected errors. Our team is actively working to identify the cause. Thank you for your patience.
Grafana K6 metrics processing and test runs degradation
4 updates
This incident has been resolved.
We've stabilized the system and test runs no longer result in timeout. There is a small delay (a few minutes) in processing metrics at the end of the test run, but most users shouldn't be too negatively impacted by that. We expected the delay/lag to also resolve within the next 30-60 minutes.
We have identified that test runs are getting timed out as a result of the issue This issue first occurred on May 05/15/2026 at 8:00PM UTC.
We’re currently investigating an issue that is resulting in degraded performance in metrics processing and test run metrics may take longer than usual to show up. Our team is actively working to identify the cause. Thank you for your patience.
Intermittent Errors and High latency Writing to Cloud Metrics, Cloud Logs and Cloud Traces
7 updates
We continue to observe an extended period of recovery and we're marking the incident as resolved at this point in time.
We continue to see signs of recovery and improved stability across impacted services. Our teams continue to closely monitor the situation while working with the cloud provider.
We continue to see signs of recovery and improved stability across impacted services. Our teams continue to closely monitor the situation while working with the cloud provider.
We are seeing signs of recovery and improved stability across impacted services over the past hour. Our teams continue to closely monitor the situation while working with the cloud provider.
We have identified expanded impact affecting Grafana Cloud Logs and Grafana Cloud Traces in addition to Cloud Metrics, causing intermittent errors and increased latency when writing data. Our teams continue working on a fix and investigating the issue with the cloud provider’s support team.
We’re continuing to investigate the issue causing intermittent errors and high latency when writing to Cloud Metrics. We are in contact with the cloud provider’s support team, and they are investigating the issue alongside us.
We’re currently investigating an issue causing intermittent errors and high latency when writing to Cloud Metrics. Our team is actively working to identify the cause. Thank you for your patience.
"Failed to Load Dashboard" Errors
5 updates
This incident has been resolved. Thank you for your patience.
The fix is currently being rolled out to all impacted environments.
Our teams continue working on a fix for this issue. We do not have additional information to share at this time, but we will continue to provide updates as progress is made.
We are continuing to work on a fix for this issue. While we do not have additional updates to share at this time, our teams remain actively engaged and we will provide further updates as soon as they become available.
Customers on Grafana Cloud may see an error on dashboard panels with "Failed to load dashboard ... json unmarshal number ...". We have identified the issue and are working to deploy out the fix.
SSL/TLS Connectivity Issues
2 updates
This incident has been resolved. Thank you for your patience.
We are currently investigating reports of service disruption affecting a subset of customers. Customers may experience intermittent connectivity issues, degraded performance, or SSL/TLS certificate validation errors when accessing affected services. Our engineering teams are actively working to identify the scope of impact and restore full functionality as quickly as possible. We will continue to provide updates as more information becomes available.
Cloud Metrics -High Write Latency and Errors in prod-us-central-7
2 updates
We have continued to observe stability. This incident is now being considered as resolved. Thank you for your patience.
From approximately 20:40-21:00 UTc, we experienced an issue affecting Grafana Cloud Metrics in prod-us-central-7. Affected users may have experienced high latency and/or errors during ingestion and rule evaluation. Our team has identified the cause and mitigated. We are currently monitoring for long-term stability.
March 2026(1 incident)
Complete outage in prod-me-central-1
15 updates
Following our ongoing communications regarding the complete outage in prod-me-central-1, we are now closing this incident. As noted in the latest AWS update, the Middle East (UAE) region (ME-CENTRAL-1) has suffered significant damage and restoration is expected to take several months. We strongly recommend all affected customers migrate workloads to an alternate Grafana Cloud region as soon as possible. If you have not already done so, please follow the migration steps outlined in our previous updates: 1. Create a Grafana Cloud stack in an alternate region 2. Update clients to send telemetry to the new region, if using Grafana Alloy then you can use Fleet Management https://grafana.com/docs/grafana-cloud/send-data/fleet-management/introduction/ 3. If your instance remains available and you have not configured your dashboards as code, then you may be able to use `grafanactl` to migrate dashboards https://grafana.com/docs/grafana/latest/as-code/observability-as-code/grafana-cli/grafanacli-workflows/ https://grafana.github.io/grafanactl/ For further details, please refer to the AWS incident communication directly: https://health.aws.amazon.com/health/status Please reach out to our Support team if you need any assistance with the above - https://grafana.com/profile/org#support We will continue to monitor the situation and update the incident once circumstances change.
The TLS certificates serving prod-me-central-1 endpoints expire on May 30, 2026. Replacement certificates have been imported, but the ongoing AWS regional incident is preventing them from propagating to all load balancer nodes, so customers may see certificate errors after that date until AWS restores normal operation. We do not have any additional updates to share at this time. Our team is actively monitoring the situation and will provide further information as it becomes available. In the meantime, please continue to refer to the AWS Status Page for the most detailed and up-to-date information.
AWS UAE - prod-me-central-1: Public Probe checks might suffer degraded experience. We recommend migrating checks from the UAE probe to the next nearest probe suitable for your use case.
We do not have any additional updates to share at this time. Our team is actively monitoring the situation and will provide further information as it becomes available. In the meantime, please continue to refer to the AWS Status Page for the most detailed and up-to-date information.
We are continuing to investigate this issue.
We have not received any further updates from AWS at this time. However, we are actively monitoring the outage and will provide additional information as it becomes available. Also, please continue to refer to the AWS status page for more detailed updates. https://health.aws.amazon.com/health/status All the guidance previously included about stack migration is still relevant. Please reach out to our Support team if you have any questions.
We are actively monitoring the situation, but at this time there are no new updates to share. The next update will be provided once we have more information to share. Please reach out to our Support team if you have any questions.
We are continuing to investigate this issue.
Please continue to refer to the AWS status page for more detailed updates specific to AWS. https://health.aws.amazon.com/health/status AWS are recommending that affected customers move workloads to alternate regions, and we are recommending the same. Customers who are impacted and who cannot wait for a restoration of service are asked to: 1. Create a Grafana Cloud stack in an alternate region 2. Update clients to send telemetry to the new region, if using Grafana Alloy then you can use Fleet Management https://grafana.com/docs/grafana-cloud/send-data/fleet-management/introduction/ 3. If your instance remains available and you have not configured your dashboards as code, then you may be able to use `grafanactl` to migrate dashboards https://grafana.com/docs/grafana/latest/as-code/observability-as-code/grafana-cli/grafanacli-workflows/ https://grafana.github.io/grafanactl/ We are continuing to work with our CSP at this time, and will provide updates as they are available.
AWS are recommending that affected customers move workloads to alternate regions https://health.aws.amazon.com/health/status and we are recommending the same. Customers who are impacted and who cannot wait for a restoration of service are asked to: 1. Create a Grafana Cloud stack in an alternate region 2. Update clients to send telemetry to the new region, if using Grafana Alloy then you can use Fleet Management https://grafana.com/docs/grafana-cloud/send-data/fleet-management/introduction/ 3. If your instance remains available and you have not configured your dashboards as code, then you may be able to use `grafanactl` to migrate dashboards https://grafana.com/docs/grafana/latest/as-code/observability-as-code/grafana-cli/grafanacli-workflows/ https://grafana.github.io/grafanactl/ We will provide updates when we have them, but we do not have an expected resolution time at this point.
Customers are recommended to configure a new blank stack in an alternative Grafana Cloud region and to reconfigure their clients (such as Grafana Alloy) to send telemetry to that region, Fleet Management can be used for this purpose https://grafana.com/docs/grafana-cloud/send-data/fleet-management/introduction/
We are updating this incident to reflect a complete outage in prod-me-central-1, due to an on-going AWS UAE data center issue. We will provide further updates accordingly.
We are observing write and read outage errors across all databases (metrics, logs, traces) in prod-me-central-1, due to an on-going AWS UAE data center issue. We will provide further updates accordingly.
We are observing write and read outage errors across all databases (metrics, logs, traces) in prod-me-central-1, due to an on-going AWS UAE data center issue. We will provide further updates accordingly.
We are seeing elevated write and read path errors in prod-me-central-1, due to an on-going AWS UAE data center issue. We will provide further updates accordingly.
📡 Tired of checking Grafana Cloud status manually?
Better Stack monitors uptime every 30 seconds and alerts you instantly when Grafana Cloud goes down.