Grafana Cloud Outage History

50 incidents reported. Data sourced from the official Grafana Cloud status page.

50
Total Incidents
17
Major/Critical
27
Minor
50
Resolved

September 2026

Increased Execution Time for Browser Checks in Synthetic Monitoring

minor
Sep 15, 10:00 AMSep 15, 11:17 AMresolved
Sep 15, 11:17 AM
resolvedBetween September 9 and September 15, 2026, some browser checks in Synthetic Monitoring experienced longer than usual execution times, with impact varying by location. This has been resolved and check...
Sep 15, 10:58 AM
monitoringA fix has been deployed and CPU throttling is returning to normal. Check execution times should return to baseline shortly. We will continue monitoring recovery. Next update by 11:56 UTC.
Sep 15, 10:51 AM
identifiedA fix has been merged and is now being rolled out across affected locations. We are monitoring deployment progress. Next update by 11:49 UTC.
+1 more updates

Intermittent Metric Write Errors in GCP US Central (prod-us-central-0)

minor
Sep 14, 03:30 PMSep 14, 03:30 PMresolved
Sep 15, 07:49 AM
resolvedBetween 15:31 and 20:43 UTC on September 14, 2026 (approximately 5 hours 12 minutes), some customers in GCP US Central (prod-us-central-0) may have experienced intermittent errors or delays when writi...

CloudWatch request timeouts in prod-us-central-0

none
Sep 14, 12:00 PMSep 14, 12:00 PMresolved
Sep 14, 01:04 PM
resolvedSome customers experienced timeouts affecting CloudWatch traffic in prod-us-central-0 for approximately one hour. Service has since been restored. This incident did not affect all customers. We are p...

Elevated Latency Managing Cloud Provider Integrations in GCP US Central

minor
Sep 14, 03:30 AMSep 14, 03:30 AMresolved
Sep 17, 12:34 PM
resolvedBetween 03:30 UTC on September 14, 2026 and 10:50 UTC on September 17, 2026 (approximately 3 days), customers managing cloud provider integrations in GCP US Central may have experienced elevated laten...

Failures Submitting Self-Serve Configuration Changes for Grafana Cloud Logs in AWS Sweden (prod-eu-north-0)

minor
Sep 12, 11:00 PMSep 12, 11:00 PMresolved
Sep 14, 10:18 AM
resolvedBetween approximately 23:00 UTC on September 12, 2026 and 09:12 UTC on September 14, 2026, some Grafana Cloud Logs customers in AWS Sweden (prod-eu-north-0) may have intermittently experienced errors ...

Investigating elevated database load in AWS Germany

minor
Sep 9, 11:27 AMSep 9, 01:51 PMresolved
Sep 9, 01:51 PM
resolvedThis incident has been resolved.
Sep 9, 12:44 PM
monitoringA fix has been implemented and we have been seeing recovery to the affected region. We will continue to monitor.
Sep 9, 11:27 AM
investigatingWe are investigating elevated database load affecting infrastructure supporting a subset of Grafana Cloud instances in the AWS Germany region (prod-eu-west-4). Affected users may experience slow resp...

Issues with Geomap Tiles

minor
Sep 8, 06:05 PMSep 8, 10:41 PMresolved
Sep 8, 10:41 PM
resolvedThis incident has been resolved.
Sep 8, 06:05 PM
identifiedSome geomap tiles are displaying the “api key required". The issue has been identified, and our team is working on a fix. As a workaround, you can open the Geomap panel, navigate to the Map layers se...

High Latency in prod-us-east-2

minor
Sep 8, 04:14 PMSep 8, 08:34 PMresolved
Sep 8, 08:34 PM
resolvedThis incident has been resolved.
Sep 8, 05:04 PM
monitoringWe have observed latency return to normal levels, and we will continue to monitor this incident.
Sep 8, 04:14 PM
investigatingWe are currently investigating an issue causing high latency for deployments in the prod-us-east-2 cluster.

US Central Region Instability

minor
Sep 4, 04:45 PMSep 4, 11:00 PMresolved
Sep 4, 11:00 PM
resolvedThe US Central region is fully recovered now and the issue is now resolved.
Sep 4, 06:58 PM
monitoringThe underlying cloud provider issue is not yet fully resolved, but our systems have recovered and affected Grafana Cloud stacks are loading normally again. Some customers may still see brief intermit...
Sep 4, 04:45 PM
identifiedWe’re currently investigating an issue affecting the us central region due to ongoing cloud provider incident. Multiple components impacted. Our team is actively working on this. Thank you for your pa...

Degradation of Hosted Grafana in US Central Region

major
Sep 4, 05:03 PMSep 4, 08:42 PMresolved
Sep 4, 08:42 PM
resolvedThe US Central region is now fully operational and this incident is now resolved.
Sep 4, 06:24 PM
monitoringWe are currently observing minimal impact and continuing to monitor the situation.
Sep 4, 05:03 PM
identifiedWe are aware of an issue causing some Grafana Cloud stacks in the US region to be unavailable or fail to load. The cause is reduced compute capacity in one of our cloud provider's availability zones, ...

Alert rule creation, deletion, and update degradation in prod-us-east-2

minor
Sep 4, 01:36 PMSep 4, 04:32 PMresolved
Sep 4, 04:32 PM
resolvedThe issue has been mitigated and there's a code fix pending.
Sep 4, 01:36 PM
investigatingWe’re currently investigating an issue affecting alert rule creation, deletion, and updates in prod-us-east-2. Our team is actively working to identify the cause. Thank you for your patience.

Tempo and Mimir Read and Write Failures

minor
Sep 3, 10:13 PMSep 4, 12:55 AMresolved
Sep 4, 12:55 AM
resolvedThis incident has been resolved.
Sep 3, 10:48 PM
monitoringTempo recovered.
Sep 3, 10:31 PM
identifiedCluster Metrics are healthy for the last 15 minutes
+1 more updates

High Latency in prod-ap-south-1

major
Sep 3, 09:32 AMSep 3, 12:49 PMresolved
Sep 3, 12:49 PM
resolvedThis incident has been resolved.
Sep 3, 11:37 AM
monitoringA fix has been applied and we are seeing signs of recovery. We will continue to monitor the recovery.
Sep 3, 09:32 AM
investigatingWe are currently investigating an issue impacting deployments in prod-ap-south-1. Impacted customers may be experiencing increased latency and issues with ingestion.

Elevated Latency in Grafana Cloud Logs (prod-ap-south-1)

none
Sep 2, 02:14 PMSep 2, 02:14 PMresolved
Sep 2, 02:14 PM
resolvedWe observed elevated latency in Grafana Cloud Logs between 08:00 and 11:00 UTC for a subset of deployments in prod-ap-south-1. The issue is now resolved.

Investigating issues in US Central (prod-us-central-0, prod-us-central-5)

minor
Sep 1, 03:25 PMSep 1, 08:20 PMresolved
Sep 1, 08:20 PM
resolvedThis incident has been resolved.
Sep 1, 04:14 PM
identifiedWe have identified the cause as an issue in our cloud provider’s US Central region. A mitigation is in place and recovery is in progress. Customers in prod-us-central-0 and prod-us-central-5 may still...
Sep 1, 03:25 PM
investigatingWe are investigating reports of errors affecting Grafana Cloud in US Central (prod-us-central-0, prod-us-central-5). Customers in this region may see timeouts or failed requests when writing or queryi...

August 2026

Partial Logs Write Outage

major
Aug 28, 05:12 PMAug 28, 06:42 PMresolved
Aug 28, 06:42 PM
resolvedThis incident has been resolved.
Aug 28, 05:55 PM
monitoringThis incident also had impact on Frontend Observability for the same region. A fix has been deployed, and things have stablized. We will continue to monitor on our end.
Aug 28, 05:12 PM
identifiedWe’ve identified the cause of an issue impacting Log writes. Our team is currently implementing a fix.

Some Grafana UI features may be unavailable or reverting to legacy behaviour

minor
Aug 28, 12:58 AMAug 28, 02:40 AMresolved
Aug 28, 02:40 AM
resolvedThis incident has been resolved.
Aug 28, 01:08 AM
identifieda resolution is rolling out to resolve the issues
Aug 28, 12:58 AM
identifiedSome Grafana UI features may be unavailable or reverting to legacy behaviour

Elevated error rates affecting metrics writes in prod-us-central-0

minor
Aug 27, 12:09 PMAug 27, 01:06 PMresolved
Aug 27, 01:06 PM
resolvedThis incident has been resolved.
Aug 27, 12:37 PM
monitoringA fix has been applied and metrics write error rates in prod-us-central-0 have returned to normal levels. Metrics ingestion is operating as expected. We are continuing to monitor the affected systems ...
Aug 27, 12:09 PM
investigatingWe are investigating elevated error rates affecting metrics writes in prod-us-central-0. Customers in this region may see failed or delayed metrics ingestion, gaps in recent data on dashboards and que...

Mimir Writes Incident in prod-us-central-0

none
Aug 27, 01:54 AMAug 27, 01:54 AMresolved
Aug 27, 01:54 AM
resolvedMimir writes in the prod-us-central-0 region had elevated error rates for approximately 15 minutes from 1:05 to 1:20 UTC. The issue has been resolved and we are monitoring.

Incident Management unavailable in US Central

minor
Aug 25, 10:39 AMAug 25, 11:02 AMresolved
Aug 25, 11:02 AM
resolvedThis incident has been resolved. Incident Management in US Central is operating as normal. Customers can create, view, and query Incidents, and the Incident public API is fully available. Grafana OnCa...
Aug 25, 11:01 AM
monitoringWe are continuing to monitor for any further issues.
Aug 25, 10:49 AM
monitoringA fix has been applied and Incident Management in US Central is recovering. The Incident public API is responding normally, and customers can create, view, and query Incidents again. We are continuing...
+1 more updates

Cloud Logs read path outage on eu-west-2

critical
Aug 18, 10:00 PMAug 18, 10:00 PMresolved
Aug 18, 10:53 PM
resolvedWe observed an issue which impacted the following service: reads of Cloud Logs in the eu-west-2 region. Affected services would include Explore, dashboards, alert evaluation when querying Logs. The...

K6 Test Outage

critical
Aug 13, 03:25 PMAug 13, 04:47 PMresolved
Aug 13, 04:47 PM
resolvedThis incident has been resolved.
Aug 13, 03:25 PM
monitoringTelemetry push APIs went down at 15:06 UTC, causing tests to lose telemetry data, like metrics, logs, traces, and browser screenshots. This issue has now been mitigated, and we are monitoring to dete...

Metrics: Elevated Error Rates Reads/Writes

minor
Aug 10, 01:30 PMAug 10, 01:30 PMresolved
Aug 10, 01:30 PM
resolvedBetween 12:20 and 13:15 UTC, we experienced elevated error rates for metrics reads and writes due to a networking issue. Impact was limited to a subset of deployments in the prod-us-central-0 region. ...

Alerting expressions pipeline failing when recovery settings

minor
Aug 8, 07:48 PMAug 8, 07:48 PMresolved
Aug 8, 07:48 PM
resolvedDue to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.

Some Cloud Test Runs Terminated

none
Aug 5, 05:57 PMAug 5, 05:57 PMresolved
Aug 5, 05:57 PM
resolvedBetween approximately 13:40 and 15:05 UTC today, a subset of cloud test runs were unexpectedly terminated and marked Aborted (by system). This affected runs that were in progress during three short wi...

o11y requests too large for nods in the prod-eu-west-2 region.

minor
Aug 5, 11:41 AMAug 5, 11:44 AMresolved
Aug 5, 11:44 AM
resolvedThis incident has been resolved.
Aug 5, 11:41 AM
investigatingWe are currently investigating an issue with critical o11y requests being too large for nods in the prod-eu-west-2 region. This is leading to issues in scaling up services, and may may lead increased ...

Some Grafana Instances Unavailable

critical
Aug 4, 07:11 PMAug 5, 01:32 AMresolved
Aug 5, 01:32 AM
resolvedWe are observing a continued period of stability and at this point, we are marking the incident resolved.
Aug 5, 12:54 AM
monitoringWe have rolled out a change to the affected regions to restore services.
Aug 4, 10:58 PM
investigatingWe have identified the cause: our cloud provider has exhausted compute capacity in the affected region. We are working directly with them and are moving affected workloads to alternative capacity to r...
+2 more updates

Logs latency increase within prod-eu-west-3

minor
Aug 4, 10:34 AMAug 4, 01:49 PMresolved
Aug 4, 01:49 PM
resolvedThis incident has been resolved.
Aug 4, 12:46 PM
investigatingThere's still a minor latency increase in this cell however we have identified that it is improving although not back to normal as of yet. We will continue to look into this.
Aug 4, 10:34 AM
investigatingWe have identified minor latency increases in write endpoints within the prod-eu-west-3. Our team is currently monitoring and investigating this.

K6 - Cloud test-run issues

minor
Aug 1, 07:25 AMAug 1, 08:31 AMresolved
Aug 1, 08:31 AM
resolvedThis incident has been resolved.
Aug 1, 08:06 AM
monitoringA critical piece of infrastructure behaved poorly after a reboot. It has now been properly recovered
Aug 1, 07:25 AM
investigatingWe are currently investigating an issue which is resulting in some test runs to abort and metrics data to be incomplete for a subset of tests

July 2026

Partial Read Outage for Loki in prod-us-east-4

none
Jul 31, 07:30 PMJul 31, 07:30 PMresolved
Jul 31, 08:33 PM
resolvedWe are investigating a partial read outage affecting Loki in prod-us-east-4. Between 19:41 UTC and 19:50 UTC, a significant portion of read queries may have failed or returned errors. The issue has b...

Degraded Performance: Stack Provisioning Failures within certain reigons (PDC Setup)

minor
Jul 31, 10:51 AMJul 31, 12:13 PMresolved
Jul 31, 12:13 PM
resolvedHe have identified the cause and applied a fix for this issue and all effected stacks within the affected regions are not working as expected.
Jul 31, 10:51 AM
investigatingWe are currently investigating an issue impacting stack provisioning. Attempting to set up PDC on newly created stacks will currently fail across several regions. We are currently looking into the cau...

PDC Authentication Issues

major
Jul 30, 02:45 PMJul 30, 05:40 PMresolved
Jul 30, 05:40 PM
resolvedThis incident has been resolved.
Jul 30, 03:56 PM
monitoringA fix has been implemented and we are monitoring the results.
Jul 30, 03:11 PM
investigatingWe are continuing to investigate this issue.
+2 more updates

Issues with Billing/Usage Dashboard Metrics and Panels.

major
Jul 30, 01:15 PMJul 30, 03:17 PMresolved
Jul 30, 03:17 PM
resolvedThis incident has been resolved.
Jul 30, 02:34 PM
monitoringA fix has been implemented and we are monitoring the results.
Jul 30, 01:50 PM
identifiedWe have identified the issue, and are working on deploying a fix.
+1 more updates

IRM Performance Degradation in EU Region

minor
Jul 29, 09:46 PMJul 30, 12:39 AMresolved
Jul 30, 12:39 AM
resolvedThe issue affecting Grafana IRM in the EU region has been resolved. The IRM UI, public API, and alert notification processing have been restored and are operating normally.
Jul 29, 11:07 PM
identifiedWe have identified the underlying issue affecting Grafana IRM in our EU region and are actively working to restore normal service. Users may continue to experience a degraded or unresponsive IRM UI, ...
Jul 29, 09:46 PM
investigatingWe are currently investigating an issue affecting Grafana IRM in our EU region. Users may experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notificati...

Grafana Cloud non-billing usage metrics gaps in select regions

minor
Jul 29, 06:49 PMJul 29, 06:49 PMresolved
Jul 29, 06:49 PM
resolvedImpact period: April 1 – July 29, 2026 Summary: During this period, some customers in a subset of regions may have experienced gaps in select ruler/recording rule metrics within their Grafana Cloud i...

Partial OTLP Write Outage in prod-us-east-3

major
Jul 28, 07:40 PMJul 28, 07:43 PMresolved
Jul 28, 07:43 PM
resolvedThe issue affecting OTLP ingestion in the prod-us-east-3 region has been resolved. Between 13:30 UTC and 17:45 UTC, some customers experienced intermittent failures when writing telemetry to the OTLP ...
Jul 28, 07:40 PM
investigatingWe are investigating an issue affecting OTLP ingestion in the prod-us-east-3 region. Customers may experience intermittent failures when writing telemetry to the OTLP endpoint. Based on current inform...

Write Outage

critical
Jul 24, 05:02 PMJul 24, 06:20 PMresolved
Jul 24, 06:20 PM
resolvedThis incident has been resolved.
Jul 24, 05:18 PM
monitoringHealthy as of 17:00 UTC. We are continuing to monitor.
Jul 24, 05:02 PM
investigatingWe are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.

Errors creating new Slack integration for Grafana IRM

major
Jul 24, 09:17 AMJul 24, 02:03 PMresolved
Jul 24, 02:03 PM
resolvedThis incident has been resolved.
Jul 24, 10:12 AM
identifiedWe have verified the fix, and we are starting to roll it out
Jul 24, 09:49 AM
identifiedWe have identified the root cause of the problem, and we are working on a fix
+1 more updates

Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23

none
Jul 23, 11:00 PMJul 24, 05:32 AMresolved
Jul 24, 05:32 AM
resolvedWrites are looking stable in the last 6h
Jul 23, 11:00 PM
monitoringstatus is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor

Metrics Write Path Errors

major
Jul 23, 05:14 PMJul 23, 07:47 PMresolved
Jul 23, 07:47 PM
resolvedThis incident has been resolved.
Jul 23, 05:37 PM
monitoringError rates have dropped, and we are monitoring this issue for any recurrence.
Jul 23, 05:14 PM
investigatingWe are currently investigating an elevated rate of error and latency in the impacted write path.

Stacks using SCIM user provisioning are currently unable to log into Grafana

major
Jul 23, 10:13 AMJul 23, 03:37 PMresolved
Jul 23, 03:37 PM
resolvedThis incident has been resolved.
Jul 23, 01:23 PM
identifiedWe have applied a fix and are currently awaiting feedback from affected customers.
Jul 23, 11:26 AM
identifiedThe issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.
+1 more updates

K6 - Cloud output test-runs are failing to fetch the script logs

minor
Jul 22, 10:27 AMJul 22, 11:25 AMresolved
Jul 22, 11:25 AM
resolvedThis incident has been resolved.
Jul 22, 10:52 AM
monitoringWe have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.
Jul 22, 10:27 AM
investigatingWe are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.

Cloud Log Exporter Unavailable

major
Jul 21, 06:43 PMJul 21, 09:06 PMresolved
Jul 21, 09:06 PM
resolvedThis incident has been resolved.
Jul 21, 08:04 PM
monitoringA fix has been implemented and we are monitoring the results.
Jul 21, 06:43 PM
investigatingCloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.

PDC Issues

major
Jul 20, 01:48 PMJul 20, 03:03 PMresolved
Jul 20, 03:03 PM
resolvedThis incident has been resolved.
Jul 20, 02:23 PM
monitoringA fix has been implemented, and we are observing recovery. We will continue to monitor the results.
Jul 20, 02:13 PM
investigatingWe are continuing to investigate this issue.
+1 more updates

Issue with Dashboard Views Being Registered

minor
Jul 14, 08:10 AMJul 20, 10:15 AMresolved
Jul 20, 10:15 AM
resolvedA fix has been implemented and dashboard view and error counts in the Dashboards and Folder list are updating as expected. Thank you for your patience while we worked to address this issue.
Jul 20, 08:17 AM
identifiedWe are currently working on a fix related to this incident. There are no new updates to share at this time.
Jul 16, 09:27 PM
identifiedWe continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
+5 more updates

Grafana Cloud login issues for users with role of None

minor
Jul 17, 06:22 PMJul 18, 01:15 PMresolved
Jul 18, 01:15 PM
resolvedUsers with the None role authenticating via Grafana.com should be able to access their stacks again
Jul 17, 09:33 PM
identifiedWe are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.
Jul 17, 07:28 PM
identifiedWe’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.
+1 more updates

Adaptive Metrics aggregation delay in eu-west-0 region.

minor
Jul 16, 12:56 PMJul 16, 02:58 PMresolved
Jul 16, 02:58 PM
resolvedThe incident is now fully resolved.
Jul 16, 01:45 PM
monitoringServices are fully recovered now. We're monitoring to be sure the issue won't re-occur.
Jul 16, 12:56 PM
identifiedWe're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the app...

Delayed Aggregated Metrics (prod-us-central-0)

minor
Jul 15, 07:19 PMJul 16, 03:54 AMresolved
Jul 16, 03:54 AM
resolvedThis incident has been resolved.
Jul 15, 09:23 PM
monitoringWe have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete. Thank you for your patience.
Jul 15, 07:19 PM
identifiedWe are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region. Our engineering team has identified the issue and is actively ...

Partial Outage in prod-eu-west-2

major
Jul 15, 10:24 PMJul 15, 11:36 PMresolved
Jul 15, 11:36 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jul 15, 10:53 PM
monitoringWe’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
Jul 15, 10:38 PM
investigatingWe are continuing to investigate this issue.
+2 more updates

Some Reports of Grafana Not Loading.

major
Jul 15, 06:07 PMJul 15, 09:00 PMresolved
Jul 15, 09:00 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jul 15, 08:50 PM
identifiedThe fix is in the process of being rolled out.
Jul 15, 06:47 PM
identifiedWe believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.
+1 more updates

Get Grafana Cloud Outage Alerts

Be the first to know when Grafana Cloud go down.