Grafana Cloud Outage History
Past incidents and downtime events
Complete history of Grafana Cloud outages, incidents, and service disruptions. Showing 50 most recent incidents.
September 2026(15 incidents)
Increased Execution Time for Browser Checks in Synthetic Monitoring
4 updates
Between September 9 and September 15, 2026, some browser checks in Synthetic Monitoring experienced longer than usual execution times, with impact varying by location. This has been resolved and check execution times have returned to baseline. Thank you for your patience.
A fix has been deployed and CPU throttling is returning to normal. Check execution times should return to baseline shortly. We will continue monitoring recovery. Next update by 11:56 UTC.
A fix has been merged and is now being rolled out across affected locations. We are monitoring deployment progress. Next update by 11:49 UTC.
Since September 9, 2026, some browser checks in Synthetic Monitoring may take longer to complete than usual, with impact varying by location. We have identified the cause and are actively working on a fix. Next update by 10:49 UTC.
Intermittent Metric Write Errors in GCP US Central (prod-us-central-0)
1 update
Between 15:31 and 20:43 UTC on September 14, 2026 (approximately 5 hours 12 minutes), some customers in GCP US Central (prod-us-central-0) may have experienced intermittent errors or delays when writing metrics. This has been fully resolved and all services are operating normally.
CloudWatch request timeouts in prod-us-central-0
1 update
Some customers experienced timeouts affecting CloudWatch traffic in prod-us-central-0 for approximately one hour. Service has since been restored. This incident did not affect all customers. We are publishing this update retrospectively and apologise for the disruption.
Elevated Latency Managing Cloud Provider Integrations in GCP US Central
1 update
Between 03:30 UTC on September 14, 2026 and 10:50 UTC on September 17, 2026 (approximately 3 days), customers managing cloud provider integrations in GCP US Central may have experienced elevated latency when creating, viewing, or removing integrations. This has been fully resolved and all services are operating normally.
Failures Submitting Self-Serve Configuration Changes for Grafana Cloud Logs in AWS Sweden (prod-eu-north-0)
1 update
Between approximately 23:00 UTC on September 12, 2026 and 09:12 UTC on September 14, 2026, some Grafana Cloud Logs customers in AWS Sweden (prod-eu-north-0) may have intermittently experienced errors when submitting configuration changes through the Cloud Portal. This has been resolved, and configuration change requests are being accepted normally.
Investigating elevated database load in AWS Germany
3 updates
This incident has been resolved.
A fix has been implemented and we have been seeing recovery to the affected region. We will continue to monitor.
We are investigating elevated database load affecting infrastructure supporting a subset of Grafana Cloud instances in the AWS Germany region (prod-eu-west-4). Affected users may experience slow responses when accessing Grafana. We will provide an update as more information becomes available.
Issues with Geomap Tiles
2 updates
This incident has been resolved.
Some geomap tiles are displaying the “api key required". The issue has been identified, and our team is working on a fix. As a workaround, you can open the Geomap panel, navigate to the Map layers section, and change the Basemap layer from CARTO, or "Default base layer," to OpenStreetMap. The watermark should disappear immediately. Alternatively, if you would prefer to retain a dark-styled map similar to the current CARTO appearance, you can select MapLibre as the basemap and configure the URL as: https://tiles.openfreemap.org/styles/dark This option also renders without the watermark, does not require an API key, and provides a visual style similar to the dark CARTO basemap.
High Latency in prod-us-east-2
3 updates
This incident has been resolved.
We have observed latency return to normal levels, and we will continue to monitor this incident.
We are currently investigating an issue causing high latency for deployments in the prod-us-east-2 cluster.
US Central Region Instability
3 updates
The US Central region is fully recovered now and the issue is now resolved.
The underlying cloud provider issue is not yet fully resolved, but our systems have recovered and affected Grafana Cloud stacks are loading normally again. Some customers may still see brief intermittent slowness. We are continuing to monitor until our provider confirms resolution.
We’re currently investigating an issue affecting the us central region due to ongoing cloud provider incident. Multiple components impacted. Our team is actively working on this. Thank you for your patience.
Degradation of Hosted Grafana in US Central Region
3 updates
The US Central region is now fully operational and this incident is now resolved.
We are currently observing minimal impact and continuing to monitor the situation.
We are aware of an issue causing some Grafana Cloud stacks in the US region to be unavailable or fail to load. The cause is reduced compute capacity in one of our cloud provider's availability zones, which is preventing affected stacks from starting. Other regions are unaffected. We are working with our cloud provider to restore capacity.
Alert rule creation, deletion, and update degradation in prod-us-east-2
2 updates
The issue has been mitigated and there's a code fix pending.
We’re currently investigating an issue affecting alert rule creation, deletion, and updates in prod-us-east-2. Our team is actively working to identify the cause. Thank you for your patience.
Tempo and Mimir Read and Write Failures
4 updates
This incident has been resolved.
Tempo recovered.
Cluster Metrics are healthy for the last 15 minutes
The issue has been identified and a fix is being implemented.
High Latency in prod-ap-south-1
3 updates
This incident has been resolved.
A fix has been applied and we are seeing signs of recovery. We will continue to monitor the recovery.
We are currently investigating an issue impacting deployments in prod-ap-south-1. Impacted customers may be experiencing increased latency and issues with ingestion.
Elevated Latency in Grafana Cloud Logs (prod-ap-south-1)
1 update
We observed elevated latency in Grafana Cloud Logs between 08:00 and 11:00 UTC for a subset of deployments in prod-ap-south-1. The issue is now resolved.
Investigating issues in US Central (prod-us-central-0, prod-us-central-5)
3 updates
This incident has been resolved.
We have identified the cause as an issue in our cloud provider’s US Central region. A mitigation is in place and recovery is in progress. Customers in prod-us-central-0 and prod-us-central-5 may still see errors or delays with metrics, logs, or Grafana while services come back. We will update this page as recovery continues.
We are investigating reports of errors affecting Grafana Cloud in US Central (prod-us-central-0, prod-us-central-5). Customers in this region may see timeouts or failed requests when writing or querying logs and metrics. We are looking into this and will post another update as soon as we have more information.
August 2026(14 incidents)
Partial Logs Write Outage
3 updates
This incident has been resolved.
This incident also had impact on Frontend Observability for the same region. A fix has been deployed, and things have stablized. We will continue to monitor on our end.
We’ve identified the cause of an issue impacting Log writes. Our team is currently implementing a fix.
Some Grafana UI features may be unavailable or reverting to legacy behaviour
3 updates
This incident has been resolved.
a resolution is rolling out to resolve the issues
Some Grafana UI features may be unavailable or reverting to legacy behaviour
Elevated error rates affecting metrics writes in prod-us-central-0
3 updates
This incident has been resolved.
A fix has been applied and metrics write error rates in prod-us-central-0 have returned to normal levels. Metrics ingestion is operating as expected. We are continuing to monitor the affected systems to confirm full recovery before marking this incident resolved.
We are investigating elevated error rates affecting metrics writes in prod-us-central-0. Customers in this region may see failed or delayed metrics ingestion, gaps in recent data on dashboards and queries, and alert rules that depend on the affected series may not evaluate as expected.
Mimir Writes Incident in prod-us-central-0
1 update
Mimir writes in the prod-us-central-0 region had elevated error rates for approximately 15 minutes from 1:05 to 1:20 UTC. The issue has been resolved and we are monitoring.
Incident Management unavailable in US Central
4 updates
This incident has been resolved. Incident Management in US Central is operating as normal. Customers can create, view, and query Incidents, and the Incident public API is fully available. Grafana OnCall was not affected at any point during this incident.
We are continuing to monitor for any further issues.
A fix has been applied and Incident Management in US Central is recovering. The Incident public API is responding normally, and customers can create, view, and query Incidents again. We are continuing to monitor the recovery to confirm the service is fully stable.
We are investigating an outage affecting Grafana Incident in the US Central region (prod-us-central-0). Customers are unable to create, view, or query Incidents, and requests to the Incident public API are timing out. All customers in this region are affected. Grafana OnCall is not impacted — alert ingestion, alert group creation, and notification delivery are all operating normally.
Cloud Logs read path outage on eu-west-2
1 update
We observed an issue which impacted the following service: reads of Cloud Logs in the eu-west-2 region. Affected services would include Explore, dashboards, alert evaluation when querying Logs. The time of impact lasted from approximately 22h15 and 22h23 UTC, on loki-prod-012. This incident has since been resolved.
K6 Test Outage
2 updates
This incident has been resolved.
Telemetry push APIs went down at 15:06 UTC, causing tests to lose telemetry data, like metrics, logs, traces, and browser screenshots. This issue has now been mitigated, and we are monitoring to determine root cause as well as full resolution.
Metrics: Elevated Error Rates Reads/Writes
1 update
Between 12:20 and 13:15 UTC, we experienced elevated error rates for metrics reads and writes due to a networking issue. Impact was limited to a subset of deployments in the prod-us-central-0 region. A fix has been applied and error rates have since recovered.
Alerting expressions pipeline failing when recovery settings
1 update
Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.
Some Cloud Test Runs Terminated
1 update
Between approximately 13:40 and 15:05 UTC today, a subset of cloud test runs were unexpectedly terminated and marked Aborted (by system). This affected runs that were in progress during three short windows in that period; test results and metrics for completed runs were not impacted. We have identified the cause and no further occurrences have been observed since 15:05 UTC. A fix is in place. The platform is currently operating normally and test runs are executing as expected
o11y requests too large for nods in the prod-eu-west-2 region.
2 updates
This incident has been resolved.
We are currently investigating an issue with critical o11y requests being too large for nods in the prod-eu-west-2 region. This is leading to issues in scaling up services, and may may lead increased errors in requests.
Some Grafana Instances Unavailable
5 updates
We are observing a continued period of stability and at this point, we are marking the incident resolved.
We have rolled out a change to the affected regions to restore services.
We have identified the cause: our cloud provider has exhausted compute capacity in the affected region. We are working directly with them and are moving affected workloads to alternative capacity to restore service.
We are continuing to investigate this issue.
We are currently investigating an issue that causing some Grafana deployments in prod-us-central-0 and prod-us-central-3 to become unavailable.
Logs latency increase within prod-eu-west-3
3 updates
This incident has been resolved.
There's still a minor latency increase in this cell however we have identified that it is improving although not back to normal as of yet. We will continue to look into this.
We have identified minor latency increases in write endpoints within the prod-eu-west-3. Our team is currently monitoring and investigating this.
K6 - Cloud test-run issues
3 updates
This incident has been resolved.
A critical piece of infrastructure behaved poorly after a reboot. It has now been properly recovered
We are currently investigating an issue which is resulting in some test runs to abort and metrics data to be incomplete for a subset of tests
July 2026(21 incidents)
Partial Read Outage for Loki in prod-us-east-4
1 update
We are investigating a partial read outage affecting Loki in prod-us-east-4. Between 19:41 UTC and 19:50 UTC, a significant portion of read queries may have failed or returned errors. The issue has been identified and service has been restored. We are continuing to investigate the underlying cause and will provide additional information as it becomes available.
Degraded Performance: Stack Provisioning Failures within certain reigons (PDC Setup)
2 updates
He have identified the cause and applied a fix for this issue and all effected stacks within the affected regions are not working as expected.
We are currently investigating an issue impacting stack provisioning. Attempting to set up PDC on newly created stacks will currently fail across several regions. We are currently looking into the cause of this.
PDC Authentication Issues
5 updates
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
We are continuing to investigate this issue.
We are continuing to investigate this issue.
We are currently investigating an issue impacting PDC authentication. We will provide additional updates as they become available.
Issues with Billing/Usage Dashboard Metrics and Panels.
4 updates
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
We have identified the issue, and are working on deploying a fix.
We are investigating an outage for the Billing / Usage dashboard metrics and panels. This appears to be partially affecting organizations using Grafana Cloud. We are working on identifying and resolving the issue
IRM Performance Degradation in EU Region
3 updates
The issue affecting Grafana IRM in the EU region has been resolved. The IRM UI, public API, and alert notification processing have been restored and are operating normally.
We have identified the underlying issue affecting Grafana IRM in our EU region and are actively working to restore normal service. Users may continue to experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notifications while mitigation efforts are underway. Our engineering team is working to restore full functionality as quickly as possible. We will provide another update as more information becomes available.
We are currently investigating an issue affecting Grafana IRM in our EU region. Users may experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notifications. Our engineering team is actively investigating the issue and working to restore normal service. We will provide another update as more information becomes available.
Grafana Cloud non-billing usage metrics gaps in select regions
1 update
Impact period: April 1 – July 29, 2026 Summary: During this period, some customers in a subset of regions may have experienced gaps in select ruler/recording rule metrics within their Grafana Cloud instances. Not all customers or environments in the listed regions were affected. Affected regions: A subset of environments across the following regions may have been impacted: ap-south-1 au-southeast-1 ca-east-0 eu-central-0 eu-west-2 eu-west-7 us-west-0 Affected values: grafanacloud_instance_queries_per_second grafanacloud_instance_rule_config_last_reload_successful grafanacloud_instance_rule_evaluations_total:rate5m grafanacloud_instance_rule_evaluation_failures_total:rate5m grafanacloud_instance_rule_group_interval_seconds grafanacloud_instance_rule_group_last_duration_seconds grafanacloud_instance_rule_group_iterations_total:rate5m grafanacloud_instance_rule_group_iterations_missed_total:rate5m grafanacloud_instance_rule_group_last_evaluation_timestamp_seconds grafanacloud_instance_rule_group_rules grafanacloud_instance_ruler_queries_failed_total:rate5m grafanacloud_instance_ruler_queries_zero_fetched_series_total:rate5m grafanacloud_instance_ruler_notifications_sent_total:rate5m grafanacloud_instance_ruler_notifications_errors_total:rate5m grafanacloud_instance_ruler_notifications_queue_capacity grafanacloud_instance_ruler_notifications_queue_length grafanacloud_instance_ruler_notifications_latency_seconds:99quantile grafanacloud_instance_ruler_notifications_latency_seconds:50quantile Current status: This issue has been resolved. No further action is required from customers. If you continue to notice gaps in the metrics described above, please reach out to support and reference this incident.
Partial OTLP Write Outage in prod-us-east-3
2 updates
The issue affecting OTLP ingestion in the prod-us-east-3 region has been resolved. Between 13:30 UTC and 17:45 UTC, some customers experienced intermittent failures when writing telemetry to the OTLP endpoint. Logs were confirmed to be affected, and metrics and traces may also have experienced intermittent ingestion failures. Our investigation has concluded, and the affected services have recovered.
We are investigating an issue affecting OTLP ingestion in the prod-us-east-3 region. Customers may experience intermittent failures when writing telemetry to the OTLP endpoint. Based on current information, logs are confirmed to be affected, and metrics and traces may also be impacted. Our investigation indicates intermittent failures began around 08:00 UTC, with the primary period of impact occurring between 15:30 UTC and 17:45 UTC. We are continuing to investigate the scope and root cause of this issue and will provide additional updates as more information becomes available.
Write Outage
3 updates
This incident has been resolved.
Healthy as of 17:00 UTC. We are continuing to monitor.
We are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.
Errors creating new Slack integration for Grafana IRM
4 updates
This incident has been resolved.
We have verified the fix, and we are starting to roll it out
We have identified the root cause of the problem, and we are working on a fix
We are investigating a possible error creating Slack integrations. Will update the status soon.
Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23
2 updates
Writes are looking stable in the last 6h
status is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor
Metrics Write Path Errors
3 updates
This incident has been resolved.
Error rates have dropped, and we are monitoring this issue for any recurrence.
We are currently investigating an elevated rate of error and latency in the impacted write path.
Stacks using SCIM user provisioning are currently unable to log into Grafana
4 updates
This incident has been resolved.
We have applied a fix and are currently awaiting feedback from affected customers.
The issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.
We are currently facing an issue where Stacks using SCIM user provisioning are currently unable to log into Grafana. We are currently investigating this issue and working on a fix.
K6 - Cloud output test-runs are failing to fetch the script logs
3 updates
This incident has been resolved.
We have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.
We are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.
Cloud Log Exporter Unavailable
3 updates
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
Cloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.
PDC Issues
4 updates
This incident has been resolved.
A fix has been implemented, and we are observing recovery. We will continue to monitor the results.
We are continuing to investigate this issue.
We are currently investigating an issue that is causing issues with PDC in the prod-eu-west-2 region. We will provide another update in 1-2 hours.
Issue with Dashboard Views Being Registered
8 updates
A fix has been implemented and dashboard view and error counts in the Dashboards and Folder list are updating as expected. Thank you for your patience while we worked to address this issue.
We are currently working on a fix related to this incident. There are no new updates to share at this time.
We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.
We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.
We've identified the cause of the issue impacting dashboard view and error counts in the Dashboards and Folder list. Our team is currently working on implementing a fix and validating the solution. We will provide another update within the next 24 hours, or sooner if we have additional information to share.
This is related to https://status.grafana.com/incidents/rhrk2ck6ly0y which was resolved by mistake. For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly. This is ultimately causing inaccurate or missing view counts. We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.
Grafana Cloud login issues for users with role of None
4 updates
Users with the None role authenticating via Grafana.com should be able to access their stacks again
We are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.
We’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.
We’re currently investigating an issue with user logins with the None role in Grafana Cloud. Our team is actively working to identify the cause. Thank you for your patience.
Adaptive Metrics aggregation delay in eu-west-0 region.
3 updates
The incident is now fully resolved.
Services are fully recovered now. We're monitoring to be sure the issue won't re-occur.
We're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the appropriate fix is already deployed. We're seeing services recovering. More updated to come soon.
Delayed Aggregated Metrics (prod-us-central-0)
3 updates
This incident has been resolved.
We have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete. Thank you for your patience.
We are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region. Our engineering team has identified the issue and is actively working on a mitigation. We will provide further updates as the investigation progresses.
Partial Outage in prod-eu-west-2
5 updates
This incident has been resolved. Thank you for your patience.
We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
We are continuing to investigate this issue.
We are investigating a broader issue affecting multiple Grafana Cloud products in prod-eu-west-2. This appears to be caused by a third-party provider issue rather than a Loki-specific problem. Impact is currently inconsistent: some components are affected while others continue to function normally. We have also seen some impact to Tempo write paths. We are continuing to investigate and will share another update as soon as we have more information.
We are investigating an issue affecting Loki queries in prod-eu-west-2. We first observed this behavior at approximately 21:57 UTC. Affected users may see elevated query errors, timeouts, or intermittent failures when running Loki queries in this region. The issue appears to be improving, but it is not fully resolved yet. We are continuing to investigate and will share another update as soon as we have more information.
Some Reports of Grafana Not Loading.
4 updates
This incident has been resolved. Thank you for your patience.
The fix is in the process of being rolled out.
We believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.
We are currently investigating an issue impacting a small subset of stacks in the impacted regions. We will provide more details as they become available.
📡 Tired of checking Grafana Cloud status manually?
Better Stack monitors uptime every 30 seconds and alerts you instantly when Grafana Cloud goes down.