Grafana Cloud Outage History

50 incidents reported. Data sourced from the official Grafana Cloud status page.

50
Total Incidents
22
Major/Critical
23
Minor
50
Resolved

August 2026

K6 - Cloud test-run issues

minor
Aug 1, 07:25 AMAug 1, 08:31 AMresolved
Aug 1, 08:31 AM
resolvedThis incident has been resolved.
Aug 1, 08:06 AM
monitoringA critical piece of infrastructure behaved poorly after a reboot. It has now been properly recovered
Aug 1, 07:25 AM
investigatingWe are currently investigating an issue which is resulting in some test runs to abort and metrics data to be incomplete for a subset of tests

July 2026

Partial Read Outage for Loki in prod-us-east-4

none
Jul 31, 07:30 PMJul 31, 07:30 PMresolved
Jul 31, 08:33 PM
resolvedWe are investigating a partial read outage affecting Loki in prod-us-east-4. Between 19:41 UTC and 19:50 UTC, a significant portion of read queries may have failed or returned errors. The issue has b...

Degraded Performance: Stack Provisioning Failures within certain reigons (PDC Setup)

minor
Jul 31, 10:51 AMJul 31, 12:13 PMresolved
Jul 31, 12:13 PM
resolvedHe have identified the cause and applied a fix for this issue and all effected stacks within the affected regions are not working as expected.
Jul 31, 10:51 AM
investigatingWe are currently investigating an issue impacting stack provisioning. Attempting to set up PDC on newly created stacks will currently fail across several regions. We are currently looking into the cau...

PDC Authentication Issues

major
Jul 30, 02:45 PMJul 30, 05:40 PMresolved
Jul 30, 05:40 PM
resolvedThis incident has been resolved.
Jul 30, 03:56 PM
monitoringA fix has been implemented and we are monitoring the results.
Jul 30, 03:11 PM
investigatingWe are continuing to investigate this issue.
+2 more updates

Issues with Billing/Usage Dashboard Metrics and Panels.

major
Jul 30, 01:15 PMJul 30, 03:17 PMresolved
Jul 30, 03:17 PM
resolvedThis incident has been resolved.
Jul 30, 02:34 PM
monitoringA fix has been implemented and we are monitoring the results.
Jul 30, 01:50 PM
identifiedWe have identified the issue, and are working on deploying a fix.
+1 more updates

IRM Performance Degradation in EU Region

minor
Jul 29, 09:46 PMJul 30, 12:39 AMresolved
Jul 30, 12:39 AM
resolvedThe issue affecting Grafana IRM in the EU region has been resolved. The IRM UI, public API, and alert notification processing have been restored and are operating normally.
Jul 29, 11:07 PM
identifiedWe have identified the underlying issue affecting Grafana IRM in our EU region and are actively working to restore normal service. Users may continue to experience a degraded or unresponsive IRM UI, ...
Jul 29, 09:46 PM
investigatingWe are currently investigating an issue affecting Grafana IRM in our EU region. Users may experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notificati...

Grafana Cloud non-billing usage metrics gaps in select regions

minor
Jul 29, 06:49 PMJul 29, 06:49 PMresolved
Jul 29, 06:49 PM
resolvedImpact period: April 1 – July 29, 2026 Summary: During this period, some customers in a subset of regions may have experienced gaps in select ruler/recording rule metrics within their Grafana Cloud i...

Partial OTLP Write Outage in prod-us-east-3

major
Jul 28, 07:40 PMJul 28, 07:43 PMresolved
Jul 28, 07:43 PM
resolvedThe issue affecting OTLP ingestion in the prod-us-east-3 region has been resolved. Between 13:30 UTC and 17:45 UTC, some customers experienced intermittent failures when writing telemetry to the OTLP ...
Jul 28, 07:40 PM
investigatingWe are investigating an issue affecting OTLP ingestion in the prod-us-east-3 region. Customers may experience intermittent failures when writing telemetry to the OTLP endpoint. Based on current inform...

Write Outage

critical
Jul 24, 05:02 PMJul 24, 06:20 PMresolved
Jul 24, 06:20 PM
resolvedThis incident has been resolved.
Jul 24, 05:18 PM
monitoringHealthy as of 17:00 UTC. We are continuing to monitor.
Jul 24, 05:02 PM
investigatingWe are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.

Errors creating new Slack integration for Grafana IRM

major
Jul 24, 09:17 AMJul 24, 02:03 PMresolved
Jul 24, 02:03 PM
resolvedThis incident has been resolved.
Jul 24, 10:12 AM
identifiedWe have verified the fix, and we are starting to roll it out
Jul 24, 09:49 AM
identifiedWe have identified the root cause of the problem, and we are working on a fix
+1 more updates

Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23

none
Jul 23, 11:00 PMJul 24, 05:32 AMresolved
Jul 24, 05:32 AM
resolvedWrites are looking stable in the last 6h
Jul 23, 11:00 PM
monitoringstatus is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor

Metrics Write Path Errors

major
Jul 23, 05:14 PMJul 23, 07:47 PMresolved
Jul 23, 07:47 PM
resolvedThis incident has been resolved.
Jul 23, 05:37 PM
monitoringError rates have dropped, and we are monitoring this issue for any recurrence.
Jul 23, 05:14 PM
investigatingWe are currently investigating an elevated rate of error and latency in the impacted write path.

Stacks using SCIM user provisioning are currently unable to log into Grafana

major
Jul 23, 10:13 AMJul 23, 03:37 PMresolved
Jul 23, 03:37 PM
resolvedThis incident has been resolved.
Jul 23, 01:23 PM
identifiedWe have applied a fix and are currently awaiting feedback from affected customers.
Jul 23, 11:26 AM
identifiedThe issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.
+1 more updates

K6 - Cloud output test-runs are failing to fetch the script logs

minor
Jul 22, 10:27 AMJul 22, 11:25 AMresolved
Jul 22, 11:25 AM
resolvedThis incident has been resolved.
Jul 22, 10:52 AM
monitoringWe have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.
Jul 22, 10:27 AM
investigatingWe are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.

Cloud Log Exporter Unavailable

major
Jul 21, 06:43 PMJul 21, 09:06 PMresolved
Jul 21, 09:06 PM
resolvedThis incident has been resolved.
Jul 21, 08:04 PM
monitoringA fix has been implemented and we are monitoring the results.
Jul 21, 06:43 PM
investigatingCloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.

PDC Issues

major
Jul 20, 01:48 PMJul 20, 03:03 PMresolved
Jul 20, 03:03 PM
resolvedThis incident has been resolved.
Jul 20, 02:23 PM
monitoringA fix has been implemented, and we are observing recovery. We will continue to monitor the results.
Jul 20, 02:13 PM
investigatingWe are continuing to investigate this issue.
+1 more updates

Issue with Dashboard Views Being Registered

minor
Jul 14, 08:10 AMJul 20, 10:15 AMresolved
Jul 20, 10:15 AM
resolvedA fix has been implemented and dashboard view and error counts in the Dashboards and Folder list are updating as expected. Thank you for your patience while we worked to address this issue.
Jul 20, 08:17 AM
identifiedWe are currently working on a fix related to this incident. There are no new updates to share at this time.
Jul 16, 09:27 PM
identifiedWe continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
+5 more updates

Grafana Cloud login issues for users with role of None

minor
Jul 17, 06:22 PMJul 18, 01:15 PMresolved
Jul 18, 01:15 PM
resolvedUsers with the None role authenticating via Grafana.com should be able to access their stacks again
Jul 17, 09:33 PM
identifiedWe are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.
Jul 17, 07:28 PM
identifiedWe’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.
+1 more updates

Adaptive Metrics aggregation delay in eu-west-0 region.

minor
Jul 16, 12:56 PMJul 16, 02:58 PMresolved
Jul 16, 02:58 PM
resolvedThe incident is now fully resolved.
Jul 16, 01:45 PM
monitoringServices are fully recovered now. We're monitoring to be sure the issue won't re-occur.
Jul 16, 12:56 PM
identifiedWe're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the app...

Delayed Aggregated Metrics (prod-us-central-0)

minor
Jul 15, 07:19 PMJul 16, 03:54 AMresolved
Jul 16, 03:54 AM
resolvedThis incident has been resolved.
Jul 15, 09:23 PM
monitoringWe have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete. Thank you for your patience.
Jul 15, 07:19 PM
identifiedWe are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region. Our engineering team has identified the issue and is actively ...

Partial Outage in prod-eu-west-2

major
Jul 15, 10:24 PMJul 15, 11:36 PMresolved
Jul 15, 11:36 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jul 15, 10:53 PM
monitoringWe’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
Jul 15, 10:38 PM
investigatingWe are continuing to investigate this issue.
+2 more updates

Some Reports of Grafana Not Loading.

major
Jul 15, 06:07 PMJul 15, 09:00 PMresolved
Jul 15, 09:00 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jul 15, 08:50 PM
identifiedThe fix is in the process of being rolled out.
Jul 15, 06:47 PM
identifiedWe believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.
+1 more updates

Fleet Managment Interface 404's

minor
Jul 15, 01:54 PMJul 15, 01:54 PMresolved
Jul 15, 01:54 PM
resolvedBetween 9:45 and 12:00 UTC, we experienced an issue affecting the Fleet Management interface. During this time, the Remote Configuration tab for production stacks returned a 404 error when accessed th...

Mimir Write Performance Degradation

minor
Jul 14, 08:28 PMJul 14, 10:00 PMresolved
Jul 14, 10:00 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jul 14, 08:28 PM
monitoringFrom 19:24 until 19:40 UTC, backend performance degradation impacted ingestion.  We are currently monitoring.

Mimir Partial Write Outage

major
Jul 14, 05:09 PMJul 14, 07:03 PMresolved
Jul 14, 07:03 PM
resolvedThis incident has been resolved.
Jul 14, 05:23 PM
monitoringThe outage is recovered as of 16:56 UTC, and we're continuing to monitor.
Jul 14, 05:09 PM
identifiedWe're investigating backend degradation which has resulted in a partial write outage beginning around 16:40 UTC.  This degradation has also impacted reads and rule evaluation.  Issue has been identifi...

IRM Mobile App forcing some users to logout

minor
Jul 10, 08:27 PMJul 13, 01:00 PMresolved
Jul 13, 01:00 PM
resolvedThis incident has been resolved.
Jul 11, 03:58 AM
monitoringA new iOS mobile app version (v2.39.5) has been released, which includes a fix for this issue. We are monitoring the rollout and its impact to ensure the issue has been fully resolved. If you continue...
Jul 10, 10:17 PM
identifiedA fix has been submitted for review and will be released as Grafana Mobile v2.39.5 once it is approved. If your app is currently on v2.39.3, we recommend not upgrading to v2.39.4 and instead waiting ...
+2 more updates

Issue with Dashboard Views Being Registered

minor
Jul 10, 03:57 PMJul 13, 05:23 AMresolved
Jul 13, 05:23 AM
resolvedAt this stage, we are considering the incident resolved.
Jul 13, 05:21 AM
investigatingWe are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soo...
Jul 10, 09:40 PM
investigatingWe are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soo...
+1 more updates

Network Degredation in prod-us-central-0

major
Jul 10, 09:31 PMJul 10, 09:31 PMresolved
Jul 10, 09:31 PM
resolvedFrom approximately 20:24 UTC - 20:47 UTC a network issue in prod-us-central-0 isolated part of our infrastructure in one availability zone, temporarily cutting off connectivity to a subset of backend ...

PDC Degraded performance

critical
Jul 9, 11:37 AMJul 9, 07:54 PMresolved
Jul 9, 07:54 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jul 9, 05:52 PM
monitoringWe are now seeing recovery of PDC traffic to the affected customers. Our team will continue to monitor across shifts.
Jul 9, 04:08 PM
identifiedWe are continuing to see disruptions across multiple deployments, ranging from degraded performance to full outages for some customers. Additional resources have been engaged to mitigate this issue. ...
+5 more updates

Delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1

minor
Jul 9, 10:28 AMJul 9, 01:07 PMresolved
Jul 9, 01:07 PM
resolvedThis incident has been resolved.
Jul 9, 11:18 AM
monitoringA fix has been implemented. We are currently monitoring the results.
Jul 9, 10:28 AM
investigatingWe are observing delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1. As of yet we have not noticed any customer impact however we are currently observing the cell.

Loki read path in prod-eu-west-2 was down

none
Jul 9, 07:37 AMJul 9, 07:37 AMresolved
Jul 9, 07:37 AM
resolvedBetween 5:49 and 5:54 UTC, the read path in prod-eu-west-2 was down. This has completely recovered by 5:57 UTC. Recording rules may have failed to evaluate during this period which may result in gaps.

Grafana rulers crash-looping on prometheus

minor
Jul 8, 12:17 PMJul 8, 05:23 PMresolved
Jul 8, 05:23 PM
resolvedThis incident has been resolved.
Jul 8, 02:23 PM
monitoringWe are in the process of rolling out the fix. Our engineers are monitoring the progress.
Jul 8, 12:17 PM
investigatingWe are investigating issues with the grafana-ruler service on prod-us-east-2 and prod-us-west-0 which are causing periodic crash conditions. A code fix is currently being deployed to mitigate this

Loki Billing Data Lost

minor
Jul 8, 03:33 PMJul 8, 03:33 PMresolved
Jul 8, 03:33 PM
resolvedBetween 13:57 and 14:20 UTC (23 minutes), a data gap occurred affecting billing usage data for Loki. Data for this window was not recorded and cannot be recovered.

Cannot access dashboard set up as a home page

minor
Jul 8, 10:28 AMJul 8, 11:00 AMresolved
Jul 8, 11:00 AM
resolvedWe have found that the issues are related to the grafana.unifiedHomepage feature rollout active 16:50 UTC yesterday, to 10:20 UTC today. We have rolled this back and systems are now working as expecte...
Jul 8, 10:28 AM
investigatingWe are currently investigating an issue affecting some customers who are unable to load or access their dashboards when configured as their home page. We will update once we have more information on t...

Delayed usage and billing data

none
Jul 7, 01:44 PMJul 7, 01:44 PMresolved
Jul 7, 01:44 PM
resolvedWe identified an issue with the internal job that calculates month-to-date usage and cost data, which caused usage attribution and billing dashboards to display stale information. The root cause has b...

Grafana Cloud IRM alert groups failing to produce alerts in us-east-3 region.

major
Jul 5, 11:37 AMJul 6, 07:37 AMresolved
Jul 6, 07:37 AM
resolvedWe haven't noticed any further issues in this region for alert group processing since yesterday. This incident i fully resolved.
Jul 5, 01:38 PM
monitoringThe work on bringing services back to healthy state is completed and alerts should be created correctly, without a delay now. We're keeping the incident open and monitoring for any potential hiccups t...
Jul 5, 11:37 AM
identifiedSome IRM alert groups in us-east-3 region may not be able to produce alerts and could be sluggish/misbehaving in general. The issue started around 01:00 UTC on July 3rd. The issue has been identified ...

Partial Write Outage for Grafana Cloud Logs in prod-eu-north-0.

minor
Jul 3, 02:27 PMJul 3, 02:27 PMresolved
Jul 3, 02:27 PM
resolvedGrafana Cloud Logs in prod-eu-north-0 experienced a 10-minute partial write outage between 13:45 and 13:54 UTC. Impacted users may have experienced 5xx errors during this time.

Some Queries Failing

minor
Jul 2, 06:10 PMJul 2, 09:35 PMresolved
Jul 2, 09:35 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jul 2, 08:16 PM
identifiedWe are continuing to work on rolling back a PR responsible for this behavior. Once we have more information, we will share it here. Thank you for your patience.
Jul 2, 06:10 PM
identifiedQueries with drop __error__, including log volume histogram queries (the queries that generate the histogram visualization in Grafana), are failing due to a bug with series limit checks. The Root cau...

Loki and Frontend Observability - Major Outage in prod-us-central-0 region

critical
Jul 2, 04:55 AMJul 2, 06:49 AMresolved
Jul 2, 06:49 AM
resolvedThis incident has been resolved by restarting the affected services.
Jul 2, 04:57 AM
investigatingThe AWS Logs integration in the same region is affected as well. We will provide further updates as our investigation progresses.
Jul 2, 04:55 AM
investigatingWe are currently investigating a major outage in Loki writes and Frontend Observability in the prod-us-central-0 region. Our Engineering team is investigating this and we will provide further updates ...

Elevated Loki Query Bytes Reporting

minor
Jul 1, 07:43 PMJul 1, 10:00 PMresolved
Jul 1, 10:00 PM
resolvedWe’ve implemented a fix and can confirm the issue is fully resolved as of 20:25 UTC. Thank you for your patience.
Jul 1, 07:43 PM
investigatingWe are investigating an issue where some customers may see higher Loki query byte usage reported than was actually consumed. This affects usage reporting only; there is no impact to query execution or...

June 2026

Mimir read errors and high latency in prod-eu-west-0

minor
Jun 29, 11:37 AMJun 30, 08:20 AMresolved
Jun 30, 08:20 AM
resolvedSince the mitigation has been applied, we have not seen the errors return. At this point, we are considering the incident resolved.
Jun 29, 05:47 PM
identifiedWe've identified a possible cause, and a mitigation is in place to prevent further occurrences.
Jun 29, 03:37 PM
investigatingThe errors and latency have now recovered, we continue investigating the root cause.
+2 more updates

Confluent API Outage

major
Jun 29, 12:37 PMJun 29, 02:27 PMresolved
Jun 29, 02:27 PM
resolvedThis incident has been resolved.
Jun 29, 12:37 PM
investigatingWe are investigating an issue affecting Confluent metrics ingestion across all regions. Due to an elevated error rate on the Confluent side, some metrics may not be ingested, resulting in potential da...

Rule evaluation error on cluster prod-gb-south-0

minor
Jun 29, 10:34 AMJun 29, 01:12 PMresolved
Jun 29, 01:12 PM
resolvedThis incident has been resolved.
Jun 29, 11:14 AM
monitoringA fix has been applied and we are currently monitoring results.
Jun 29, 10:34 AM
investigatingWe are currently investigating Rule Evaluation errors on the cluster prod-gb-south-0 which is leading to error codes showing within the stacks. We are looking into the issue and will update accordingl...

K6 - Test run metrics processing is delayed

minor
Jun 26, 01:08 PMJun 28, 02:35 AMresolved
Jun 28, 02:35 AM
resolvedThis incident has been resolved.
Jun 26, 11:30 PM
investigatingWe've improved the metric ingestion delay time and are working on additional fixes to bring it down to expected range. Customers can currently expect a delay of 2 to 5 minutes before their test run me...
Jun 26, 04:33 PM
investigatingUpdate: Changed incident title to "Test run metrics processing is delayed" We have found the issue and are working on deploying the fix.
+3 more updates

Potential Issues Loading Grafana for Users in India

major
Jun 18, 03:18 PMJun 24, 12:46 PMresolved
Jun 24, 12:46 PM
resolvedThis incident has been resolved. Error rates continued to remain near 0 and operations are performing as expected.
Jun 23, 05:54 PM
monitoringError rates have remained near zero, and we continue to monitor.
Jun 23, 01:15 PM
monitoringWe are continuing to monitor for further issues.
+6 more updates

Elevated Logs Query Usage — Unexpected Billing Impact

none
Jun 20, 08:00 PMJun 20, 08:00 PMresolved
Jun 22, 10:10 PM
resolvedStarting June 20, 2026 at approximately 20:00 UTC, some Grafana Cloud customers experienced unexpectedly elevated logs query usage. The issue persisted until it was resolved on June 22, 2026 at approx...

Rule Evaluation Outage in prod-us-central-0

major
Jun 19, 04:39 PMJun 19, 07:19 PMresolved
Jun 19, 07:19 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jun 19, 05:46 PM
monitoringWe’re continuing to track progress post-mitigation. While we don’t have new information to share yet, our team remains actively engaged.
Jun 19, 04:39 PM
monitoringWe had an outage affecting rule evaluations between 15:16-15:59 UTC in the prod-us-central-0 region. Our team quickly identified the issue and has since mitigated. The engineering team is monitoring...

Issues with actions in the Grafana IRM mobile app

major
Jun 18, 05:01 PMJun 18, 06:50 PMresolved
Jun 18, 06:50 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jun 18, 06:09 PM
monitoringWe've verified a fix in our staging environment to restore functionality to the mobile app. The fix is currently being deployed to production. Thanks for your patience as we continue to roll this out ...
Jun 18, 05:01 PM
identifiedWe're noticing an uptick in users being unable to respond to actions on the mobile app (acknowledging and silencing alerts, for example). Users working in the web UI should not be affected. Ingestion ...

Degraded k6 cloud UI performance

critical
Jun 18, 11:25 AMJun 18, 04:21 PMresolved
Jun 18, 04:21 PM
resolvedThis incident has been resolved. Thank you for your patience.
Jun 18, 02:04 PM
monitoringWe are continuing to monitor for any further issues.
Jun 18, 01:15 PM
monitoringThe root cause of the issue has been identified and a fix has been successfully deployed. We are observing widespread improvements across all systems. Our team is currently monitoring the environment ...
+2 more updates

Loki data source-managed alert rules not visible in the Grafana Cloud Alerting UI

major
Jun 17, 08:17 PMJun 18, 02:08 PMresolved
Jun 18, 02:08 PM
resolvedThis incident has been resolved.
Jun 18, 08:04 AM
monitoringA fix has been implemented and we are monitoring the results.
Jun 17, 11:22 PM
identifiedWe are continuing to deploy the fix and monitor recovery efforts. As part of the rollout, we identified an issue that required adjustments to our deployment plan, which has extended the timeline for m...
+3 more updates

Get Grafana Cloud Outage Alerts

Be the first to know when Grafana Cloud go down.