Grafana Cloud Outage History
50 incidents reported. Data sourced from the official Grafana Cloud status page.
50
Total Incidents
17
Major/Critical
27
Minor
50
Resolved
September 2026
Increased Execution Time for Browser Checks in Synthetic Monitoring
minorSep 15, 10:00 AM→Sep 15, 11:17 AMresolved
Sep 15, 11:17 AM
resolved — Between September 9 and September 15, 2026, some browser checks in Synthetic Monitoring experienced longer than usual execution times, with impact varying by location. This has been resolved and check...
Sep 15, 10:58 AM
monitoring — A fix has been deployed and CPU throttling is returning to normal. Check execution times should return to baseline shortly. We will continue monitoring recovery. Next update by 11:56 UTC.
Sep 15, 10:51 AM
identified — A fix has been merged and is now being rolled out across affected locations. We are monitoring deployment progress. Next update by 11:49 UTC.
+1 more updates
Intermittent Metric Write Errors in GCP US Central (prod-us-central-0)
minorSep 14, 03:30 PM→Sep 14, 03:30 PMresolved
Sep 15, 07:49 AM
resolved — Between 15:31 and 20:43 UTC on September 14, 2026 (approximately 5 hours 12 minutes), some customers in GCP US Central (prod-us-central-0) may have experienced intermittent errors or delays when writi...
CloudWatch request timeouts in prod-us-central-0
noneSep 14, 12:00 PM→Sep 14, 12:00 PMresolved
Sep 14, 01:04 PM
resolved — Some customers experienced timeouts affecting CloudWatch traffic in prod-us-central-0 for approximately one hour. Service has since been restored.
This incident did not affect all customers. We are p...
Elevated Latency Managing Cloud Provider Integrations in GCP US Central
minorSep 14, 03:30 AM→Sep 14, 03:30 AMresolved
Sep 17, 12:34 PM
resolved — Between 03:30 UTC on September 14, 2026 and 10:50 UTC on September 17, 2026 (approximately 3 days), customers managing cloud provider integrations in GCP US Central may have experienced elevated laten...
Failures Submitting Self-Serve Configuration Changes for Grafana Cloud Logs in AWS Sweden (prod-eu-north-0)
minorSep 12, 11:00 PM→Sep 12, 11:00 PMresolved
Sep 14, 10:18 AM
resolved — Between approximately 23:00 UTC on September 12, 2026 and 09:12 UTC on September 14, 2026, some Grafana Cloud Logs customers in AWS Sweden (prod-eu-north-0) may have intermittently experienced errors ...
Investigating elevated database load in AWS Germany
minorSep 9, 11:27 AM→Sep 9, 01:51 PMresolved
Sep 9, 01:51 PM
resolved — This incident has been resolved.
Sep 9, 12:44 PM
monitoring — A fix has been implemented and we have been seeing recovery to the affected region. We will continue to monitor.
Sep 9, 11:27 AM
investigating — We are investigating elevated database load affecting infrastructure supporting a subset of Grafana Cloud instances in the AWS Germany region (prod-eu-west-4).
Affected users may experience slow resp...
Issues with Geomap Tiles
minorSep 8, 06:05 PM→Sep 8, 10:41 PMresolved
Sep 8, 10:41 PM
resolved — This incident has been resolved.
Sep 8, 06:05 PM
identified — Some geomap tiles are displaying the “api key required". The issue has been identified, and our team is working on a fix.
As a workaround, you can open the Geomap panel, navigate to the Map layers se...
High Latency in prod-us-east-2
minorSep 8, 04:14 PM→Sep 8, 08:34 PMresolved
Sep 8, 08:34 PM
resolved — This incident has been resolved.
Sep 8, 05:04 PM
monitoring — We have observed latency return to normal levels, and we will continue to monitor this incident.
Sep 8, 04:14 PM
investigating — We are currently investigating an issue causing high latency for deployments in the prod-us-east-2 cluster.
US Central Region Instability
minorSep 4, 04:45 PM→Sep 4, 11:00 PMresolved
Sep 4, 11:00 PM
resolved — The US Central region is fully recovered now and the issue is now resolved.
Sep 4, 06:58 PM
monitoring — The underlying cloud provider issue is not yet fully resolved, but our systems have recovered and affected Grafana Cloud stacks are loading normally again.
Some customers may still see brief intermit...
Sep 4, 04:45 PM
identified — We’re currently investigating an issue affecting the us central region due to ongoing cloud provider incident. Multiple components impacted. Our team is actively working on this. Thank you for your pa...
Degradation of Hosted Grafana in US Central Region
majorSep 4, 05:03 PM→Sep 4, 08:42 PMresolved
Sep 4, 08:42 PM
resolved — The US Central region is now fully operational and this incident is now resolved.
Sep 4, 06:24 PM
monitoring — We are currently observing minimal impact and continuing to monitor the situation.
Sep 4, 05:03 PM
identified — We are aware of an issue causing some Grafana Cloud stacks in the US region to be unavailable or fail to load. The cause is reduced compute capacity in one of our cloud provider's availability zones, ...
Alert rule creation, deletion, and update degradation in prod-us-east-2
minorSep 4, 01:36 PM→Sep 4, 04:32 PMresolved
Sep 4, 04:32 PM
resolved — The issue has been mitigated and there's a code fix pending.
Sep 4, 01:36 PM
investigating — We’re currently investigating an issue affecting alert rule creation, deletion, and updates in prod-us-east-2. Our team is actively working to identify the cause. Thank you for your patience.
Tempo and Mimir Read and Write Failures
minorSep 3, 10:13 PM→Sep 4, 12:55 AMresolved
Sep 4, 12:55 AM
resolved — This incident has been resolved.
Sep 3, 10:48 PM
monitoring — Tempo recovered.
Sep 3, 10:31 PM
identified — Cluster Metrics are healthy for the last 15 minutes
+1 more updates
High Latency in prod-ap-south-1
majorSep 3, 09:32 AM→Sep 3, 12:49 PMresolved
Sep 3, 12:49 PM
resolved — This incident has been resolved.
Sep 3, 11:37 AM
monitoring — A fix has been applied and we are seeing signs of recovery. We will continue to monitor the recovery.
Sep 3, 09:32 AM
investigating — We are currently investigating an issue impacting deployments in prod-ap-south-1. Impacted customers may be experiencing increased latency and issues with ingestion.
Elevated Latency in Grafana Cloud Logs (prod-ap-south-1)
noneSep 2, 02:14 PM→Sep 2, 02:14 PMresolved
Sep 2, 02:14 PM
resolved — We observed elevated latency in Grafana Cloud Logs between 08:00 and 11:00 UTC for a subset of deployments in prod-ap-south-1. The issue is now resolved.
Investigating issues in US Central (prod-us-central-0, prod-us-central-5)
minorSep 1, 03:25 PM→Sep 1, 08:20 PMresolved
Sep 1, 08:20 PM
resolved — This incident has been resolved.
Sep 1, 04:14 PM
identified — We have identified the cause as an issue in our cloud provider’s US Central region. A mitigation is in place and recovery is in progress. Customers in prod-us-central-0 and prod-us-central-5 may still...
Sep 1, 03:25 PM
investigating — We are investigating reports of errors affecting Grafana Cloud in US Central (prod-us-central-0, prod-us-central-5). Customers in this region may see timeouts or failed requests when writing or queryi...
August 2026
Partial Logs Write Outage
majorAug 28, 05:12 PM→Aug 28, 06:42 PMresolved
Aug 28, 06:42 PM
resolved — This incident has been resolved.
Aug 28, 05:55 PM
monitoring — This incident also had impact on Frontend Observability for the same region.
A fix has been deployed, and things have stablized. We will continue to monitor on our end.
Aug 28, 05:12 PM
identified — We’ve identified the cause of an issue impacting Log writes. Our team is currently implementing a fix.
Some Grafana UI features may be unavailable or reverting to legacy behaviour
minorAug 28, 12:58 AM→Aug 28, 02:40 AMresolved
Aug 28, 02:40 AM
resolved — This incident has been resolved.
Aug 28, 01:08 AM
identified — a resolution is rolling out to resolve the issues
Aug 28, 12:58 AM
identified — Some Grafana UI features may be unavailable or reverting to legacy behaviour
Elevated error rates affecting metrics writes in prod-us-central-0
minorAug 27, 12:09 PM→Aug 27, 01:06 PMresolved
Aug 27, 01:06 PM
resolved — This incident has been resolved.
Aug 27, 12:37 PM
monitoring — A fix has been applied and metrics write error rates in prod-us-central-0 have returned to normal levels. Metrics ingestion is operating as expected. We are continuing to monitor the affected systems ...
Aug 27, 12:09 PM
investigating — We are investigating elevated error rates affecting metrics writes in prod-us-central-0. Customers in this region may see failed or delayed metrics ingestion, gaps in recent data on dashboards and que...
Mimir Writes Incident in prod-us-central-0
noneAug 27, 01:54 AM→Aug 27, 01:54 AMresolved
Aug 27, 01:54 AM
resolved — Mimir writes in the prod-us-central-0 region had elevated error rates for approximately 15 minutes from 1:05 to 1:20 UTC. The issue has been resolved and we are monitoring.
Incident Management unavailable in US Central
minorAug 25, 10:39 AM→Aug 25, 11:02 AMresolved
Aug 25, 11:02 AM
resolved — This incident has been resolved. Incident Management in US Central is operating as normal. Customers can create, view, and query Incidents, and the Incident public API is fully available. Grafana OnCa...
Aug 25, 11:01 AM
monitoring — We are continuing to monitor for any further issues.
Aug 25, 10:49 AM
monitoring — A fix has been applied and Incident Management in US Central is recovering. The Incident public API is responding normally, and customers can create, view, and query Incidents again. We are continuing...
+1 more updates
Cloud Logs read path outage on eu-west-2
criticalAug 18, 10:00 PM→Aug 18, 10:00 PMresolved
Aug 18, 10:53 PM
resolved — We observed an issue which impacted the following service: reads of Cloud Logs in the eu-west-2 region.
Affected services would include Explore, dashboards, alert evaluation when querying Logs.
The...
K6 Test Outage
criticalAug 13, 03:25 PM→Aug 13, 04:47 PMresolved
Aug 13, 04:47 PM
resolved — This incident has been resolved.
Aug 13, 03:25 PM
monitoring — Telemetry push APIs went down at 15:06 UTC, causing tests to lose telemetry data, like metrics, logs, traces, and browser screenshots.
This issue has now been mitigated, and we are monitoring to dete...
Metrics: Elevated Error Rates Reads/Writes
minorAug 10, 01:30 PM→Aug 10, 01:30 PMresolved
Aug 10, 01:30 PM
resolved — Between 12:20 and 13:15 UTC, we experienced elevated error rates for metrics reads and writes due to a networking issue. Impact was limited to a subset of deployments in the prod-us-central-0 region. ...
Alerting expressions pipeline failing when recovery settings
minorAug 8, 07:48 PM→Aug 8, 07:48 PMresolved
Aug 8, 07:48 PM
resolved — Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.
Some Cloud Test Runs Terminated
noneAug 5, 05:57 PM→Aug 5, 05:57 PMresolved
Aug 5, 05:57 PM
resolved — Between approximately 13:40 and 15:05 UTC today, a subset of cloud test runs were unexpectedly terminated and marked Aborted (by system). This affected runs that were in progress during three short wi...
o11y requests too large for nods in the prod-eu-west-2 region.
minorAug 5, 11:41 AM→Aug 5, 11:44 AMresolved
Aug 5, 11:44 AM
resolved — This incident has been resolved.
Aug 5, 11:41 AM
investigating — We are currently investigating an issue with critical o11y requests being too large for nods in the prod-eu-west-2 region. This is leading to issues in scaling up services, and may may lead increased ...
Some Grafana Instances Unavailable
criticalAug 4, 07:11 PM→Aug 5, 01:32 AMresolved
Aug 5, 01:32 AM
resolved — We are observing a continued period of stability and at this point, we are marking the incident resolved.
Aug 5, 12:54 AM
monitoring — We have rolled out a change to the affected regions to restore services.
Aug 4, 10:58 PM
investigating — We have identified the cause: our cloud provider has exhausted compute capacity in the affected region. We are working directly with them and are moving affected workloads to alternative capacity to r...
+2 more updates
Logs latency increase within prod-eu-west-3
minorAug 4, 10:34 AM→Aug 4, 01:49 PMresolved
Aug 4, 01:49 PM
resolved — This incident has been resolved.
Aug 4, 12:46 PM
investigating — There's still a minor latency increase in this cell however we have identified that it is improving although not back to normal as of yet. We will continue to look into this.
Aug 4, 10:34 AM
investigating — We have identified minor latency increases in write endpoints within the prod-eu-west-3. Our team is currently monitoring and investigating this.
K6 - Cloud test-run issues
minorAug 1, 07:25 AM→Aug 1, 08:31 AMresolved
Aug 1, 08:31 AM
resolved — This incident has been resolved.
Aug 1, 08:06 AM
monitoring — A critical piece of infrastructure behaved poorly after a reboot. It has now been properly recovered
Aug 1, 07:25 AM
investigating — We are currently investigating an issue which is resulting in some test runs to abort and metrics data to be incomplete for a subset of tests
July 2026
Partial Read Outage for Loki in prod-us-east-4
noneJul 31, 07:30 PM→Jul 31, 07:30 PMresolved
Jul 31, 08:33 PM
resolved — We are investigating a partial read outage affecting Loki in prod-us-east-4. Between 19:41 UTC and 19:50 UTC, a significant portion of read queries may have failed or returned errors.
The issue has b...
Degraded Performance: Stack Provisioning Failures within certain reigons (PDC Setup)
minorJul 31, 10:51 AM→Jul 31, 12:13 PMresolved
Jul 31, 12:13 PM
resolved — He have identified the cause and applied a fix for this issue and all effected stacks within the affected regions are not working as expected.
Jul 31, 10:51 AM
investigating — We are currently investigating an issue impacting stack provisioning. Attempting to set up PDC on newly created stacks will currently fail across several regions. We are currently looking into the cau...
PDC Authentication Issues
majorJul 30, 02:45 PM→Jul 30, 05:40 PMresolved
Jul 30, 05:40 PM
resolved — This incident has been resolved.
Jul 30, 03:56 PM
monitoring — A fix has been implemented and we are monitoring the results.
Jul 30, 03:11 PM
investigating — We are continuing to investigate this issue.
+2 more updates
Issues with Billing/Usage Dashboard Metrics and Panels.
majorJul 30, 01:15 PM→Jul 30, 03:17 PMresolved
Jul 30, 03:17 PM
resolved — This incident has been resolved.
Jul 30, 02:34 PM
monitoring — A fix has been implemented and we are monitoring the results.
Jul 30, 01:50 PM
identified — We have identified the issue, and are working on deploying a fix.
+1 more updates
IRM Performance Degradation in EU Region
minorJul 29, 09:46 PM→Jul 30, 12:39 AMresolved
Jul 30, 12:39 AM
resolved — The issue affecting Grafana IRM in the EU region has been resolved. The IRM UI, public API, and alert notification processing have been restored and are operating normally.
Jul 29, 11:07 PM
identified — We have identified the underlying issue affecting Grafana IRM in our EU region and are actively working to restore normal service.
Users may continue to experience a degraded or unresponsive IRM UI, ...
Jul 29, 09:46 PM
investigating — We are currently investigating an issue affecting Grafana IRM in our EU region. Users may experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notificati...
Grafana Cloud non-billing usage metrics gaps in select regions
minorJul 29, 06:49 PM→Jul 29, 06:49 PMresolved
Jul 29, 06:49 PM
resolved — Impact period: April 1 – July 29, 2026
Summary:
During this period, some customers in a subset of regions may have experienced gaps in select ruler/recording rule metrics within their Grafana Cloud i...
Partial OTLP Write Outage in prod-us-east-3
majorJul 28, 07:40 PM→Jul 28, 07:43 PMresolved
Jul 28, 07:43 PM
resolved — The issue affecting OTLP ingestion in the prod-us-east-3 region has been resolved. Between 13:30 UTC and 17:45 UTC, some customers experienced intermittent failures when writing telemetry to the OTLP ...
Jul 28, 07:40 PM
investigating — We are investigating an issue affecting OTLP ingestion in the prod-us-east-3 region. Customers may experience intermittent failures when writing telemetry to the OTLP endpoint. Based on current inform...
Write Outage
criticalJul 24, 05:02 PM→Jul 24, 06:20 PMresolved
Jul 24, 06:20 PM
resolved — This incident has been resolved.
Jul 24, 05:18 PM
monitoring — Healthy as of 17:00 UTC. We are continuing to monitor.
Jul 24, 05:02 PM
investigating — We are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.
Errors creating new Slack integration for Grafana IRM
majorJul 24, 09:17 AM→Jul 24, 02:03 PMresolved
Jul 24, 02:03 PM
resolved — This incident has been resolved.
Jul 24, 10:12 AM
identified — We have verified the fix, and we are starting to roll it out
Jul 24, 09:49 AM
identified — We have identified the root cause of the problem, and we are working on a fix
+1 more updates
Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23
noneJul 23, 11:00 PM→Jul 24, 05:32 AMresolved
Jul 24, 05:32 AM
resolved — Writes are looking stable in the last 6h
Jul 23, 11:00 PM
monitoring — status is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor
Metrics Write Path Errors
majorJul 23, 05:14 PM→Jul 23, 07:47 PMresolved
Jul 23, 07:47 PM
resolved — This incident has been resolved.
Jul 23, 05:37 PM
monitoring — Error rates have dropped, and we are monitoring this issue for any recurrence.
Jul 23, 05:14 PM
investigating — We are currently investigating an elevated rate of error and latency in the impacted write path.
Stacks using SCIM user provisioning are currently unable to log into Grafana
majorJul 23, 10:13 AM→Jul 23, 03:37 PMresolved
Jul 23, 03:37 PM
resolved — This incident has been resolved.
Jul 23, 01:23 PM
identified — We have applied a fix and are currently awaiting feedback from affected customers.
Jul 23, 11:26 AM
identified — The issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.
+1 more updates
K6 - Cloud output test-runs are failing to fetch the script logs
minorJul 22, 10:27 AM→Jul 22, 11:25 AMresolved
Jul 22, 11:25 AM
resolved — This incident has been resolved.
Jul 22, 10:52 AM
monitoring — We have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.
Jul 22, 10:27 AM
investigating — We are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.
Cloud Log Exporter Unavailable
majorJul 21, 06:43 PM→Jul 21, 09:06 PMresolved
Jul 21, 09:06 PM
resolved — This incident has been resolved.
Jul 21, 08:04 PM
monitoring — A fix has been implemented and we are monitoring the results.
Jul 21, 06:43 PM
investigating — Cloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.
PDC Issues
majorJul 20, 01:48 PM→Jul 20, 03:03 PMresolved
Jul 20, 03:03 PM
resolved — This incident has been resolved.
Jul 20, 02:23 PM
monitoring — A fix has been implemented, and we are observing recovery. We will continue to monitor the results.
Jul 20, 02:13 PM
investigating — We are continuing to investigate this issue.
+1 more updates
Issue with Dashboard Views Being Registered
minorJul 14, 08:10 AM→Jul 20, 10:15 AMresolved
Jul 20, 10:15 AM
resolved — A fix has been implemented and dashboard view and error counts in the Dashboards and Folder list are updating as expected. Thank you for your patience while we worked to address this issue.
Jul 20, 08:17 AM
identified — We are currently working on a fix related to this incident. There are no new updates to share at this time.
Jul 16, 09:27 PM
identified — We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
+5 more updates
Grafana Cloud login issues for users with role of None
minorJul 17, 06:22 PM→Jul 18, 01:15 PMresolved
Jul 18, 01:15 PM
resolved — Users with the None role authenticating via Grafana.com should be able to access their stacks again
Jul 17, 09:33 PM
identified — We are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.
Jul 17, 07:28 PM
identified — We’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.
+1 more updates
Adaptive Metrics aggregation delay in eu-west-0 region.
minorJul 16, 12:56 PM→Jul 16, 02:58 PMresolved
Jul 16, 02:58 PM
resolved — The incident is now fully resolved.
Jul 16, 01:45 PM
monitoring — Services are fully recovered now. We're monitoring to be sure the issue won't re-occur.
Jul 16, 12:56 PM
identified — We're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the app...
Delayed Aggregated Metrics (prod-us-central-0)
minorJul 15, 07:19 PM→Jul 16, 03:54 AMresolved
Jul 16, 03:54 AM
resolved — This incident has been resolved.
Jul 15, 09:23 PM
monitoring — We have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete.
Thank you for your patience.
Jul 15, 07:19 PM
identified — We are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region.
Our engineering team has identified the issue and is actively ...
Partial Outage in prod-eu-west-2
majorJul 15, 10:24 PM→Jul 15, 11:36 PMresolved
Jul 15, 11:36 PM
resolved — This incident has been resolved. Thank you for your patience.
Jul 15, 10:53 PM
monitoring — We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
Jul 15, 10:38 PM
investigating — We are continuing to investigate this issue.
+2 more updates
Some Reports of Grafana Not Loading.
majorJul 15, 06:07 PM→Jul 15, 09:00 PMresolved
Jul 15, 09:00 PM
resolved — This incident has been resolved. Thank you for your patience.
Jul 15, 08:50 PM
identified — The fix is in the process of being rolled out.
Jul 15, 06:47 PM
identified — We believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.
+1 more updates
Related Incident Histories
Get Grafana Cloud Outage Alerts
Be the first to know when Grafana Cloud go down.