Grafana Cloud Outage History
50 incidents reported. Data sourced from the official Grafana Cloud status page.
50
Total Incidents
22
Major/Critical
23
Minor
50
Resolved
August 2026
K6 - Cloud test-run issues
minorAug 1, 07:25 AM→Aug 1, 08:31 AMresolved
Aug 1, 08:31 AM
resolved — This incident has been resolved.
Aug 1, 08:06 AM
monitoring — A critical piece of infrastructure behaved poorly after a reboot. It has now been properly recovered
Aug 1, 07:25 AM
investigating — We are currently investigating an issue which is resulting in some test runs to abort and metrics data to be incomplete for a subset of tests
July 2026
Partial Read Outage for Loki in prod-us-east-4
noneJul 31, 07:30 PM→Jul 31, 07:30 PMresolved
Jul 31, 08:33 PM
resolved — We are investigating a partial read outage affecting Loki in prod-us-east-4. Between 19:41 UTC and 19:50 UTC, a significant portion of read queries may have failed or returned errors.
The issue has b...
Degraded Performance: Stack Provisioning Failures within certain reigons (PDC Setup)
minorJul 31, 10:51 AM→Jul 31, 12:13 PMresolved
Jul 31, 12:13 PM
resolved — He have identified the cause and applied a fix for this issue and all effected stacks within the affected regions are not working as expected.
Jul 31, 10:51 AM
investigating — We are currently investigating an issue impacting stack provisioning. Attempting to set up PDC on newly created stacks will currently fail across several regions. We are currently looking into the cau...
PDC Authentication Issues
majorJul 30, 02:45 PM→Jul 30, 05:40 PMresolved
Jul 30, 05:40 PM
resolved — This incident has been resolved.
Jul 30, 03:56 PM
monitoring — A fix has been implemented and we are monitoring the results.
Jul 30, 03:11 PM
investigating — We are continuing to investigate this issue.
+2 more updates
Issues with Billing/Usage Dashboard Metrics and Panels.
majorJul 30, 01:15 PM→Jul 30, 03:17 PMresolved
Jul 30, 03:17 PM
resolved — This incident has been resolved.
Jul 30, 02:34 PM
monitoring — A fix has been implemented and we are monitoring the results.
Jul 30, 01:50 PM
identified — We have identified the issue, and are working on deploying a fix.
+1 more updates
IRM Performance Degradation in EU Region
minorJul 29, 09:46 PM→Jul 30, 12:39 AMresolved
Jul 30, 12:39 AM
resolved — The issue affecting Grafana IRM in the EU region has been resolved. The IRM UI, public API, and alert notification processing have been restored and are operating normally.
Jul 29, 11:07 PM
identified — We have identified the underlying issue affecting Grafana IRM in our EU region and are actively working to restore normal service.
Users may continue to experience a degraded or unresponsive IRM UI, ...
Jul 29, 09:46 PM
investigating — We are currently investigating an issue affecting Grafana IRM in our EU region. Users may experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notificati...
Grafana Cloud non-billing usage metrics gaps in select regions
minorJul 29, 06:49 PM→Jul 29, 06:49 PMresolved
Jul 29, 06:49 PM
resolved — Impact period: April 1 – July 29, 2026
Summary:
During this period, some customers in a subset of regions may have experienced gaps in select ruler/recording rule metrics within their Grafana Cloud i...
Partial OTLP Write Outage in prod-us-east-3
majorJul 28, 07:40 PM→Jul 28, 07:43 PMresolved
Jul 28, 07:43 PM
resolved — The issue affecting OTLP ingestion in the prod-us-east-3 region has been resolved. Between 13:30 UTC and 17:45 UTC, some customers experienced intermittent failures when writing telemetry to the OTLP ...
Jul 28, 07:40 PM
investigating — We are investigating an issue affecting OTLP ingestion in the prod-us-east-3 region. Customers may experience intermittent failures when writing telemetry to the OTLP endpoint. Based on current inform...
Write Outage
criticalJul 24, 05:02 PM→Jul 24, 06:20 PMresolved
Jul 24, 06:20 PM
resolved — This incident has been resolved.
Jul 24, 05:18 PM
monitoring — Healthy as of 17:00 UTC. We are continuing to monitor.
Jul 24, 05:02 PM
investigating — We are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.
Errors creating new Slack integration for Grafana IRM
majorJul 24, 09:17 AM→Jul 24, 02:03 PMresolved
Jul 24, 02:03 PM
resolved — This incident has been resolved.
Jul 24, 10:12 AM
identified — We have verified the fix, and we are starting to roll it out
Jul 24, 09:49 AM
identified — We have identified the root cause of the problem, and we are working on a fix
+1 more updates
Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23
noneJul 23, 11:00 PM→Jul 24, 05:32 AMresolved
Jul 24, 05:32 AM
resolved — Writes are looking stable in the last 6h
Jul 23, 11:00 PM
monitoring — status is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor
Metrics Write Path Errors
majorJul 23, 05:14 PM→Jul 23, 07:47 PMresolved
Jul 23, 07:47 PM
resolved — This incident has been resolved.
Jul 23, 05:37 PM
monitoring — Error rates have dropped, and we are monitoring this issue for any recurrence.
Jul 23, 05:14 PM
investigating — We are currently investigating an elevated rate of error and latency in the impacted write path.
Stacks using SCIM user provisioning are currently unable to log into Grafana
majorJul 23, 10:13 AM→Jul 23, 03:37 PMresolved
Jul 23, 03:37 PM
resolved — This incident has been resolved.
Jul 23, 01:23 PM
identified — We have applied a fix and are currently awaiting feedback from affected customers.
Jul 23, 11:26 AM
identified — The issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.
+1 more updates
K6 - Cloud output test-runs are failing to fetch the script logs
minorJul 22, 10:27 AM→Jul 22, 11:25 AMresolved
Jul 22, 11:25 AM
resolved — This incident has been resolved.
Jul 22, 10:52 AM
monitoring — We have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.
Jul 22, 10:27 AM
investigating — We are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.
Cloud Log Exporter Unavailable
majorJul 21, 06:43 PM→Jul 21, 09:06 PMresolved
Jul 21, 09:06 PM
resolved — This incident has been resolved.
Jul 21, 08:04 PM
monitoring — A fix has been implemented and we are monitoring the results.
Jul 21, 06:43 PM
investigating — Cloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.
PDC Issues
majorJul 20, 01:48 PM→Jul 20, 03:03 PMresolved
Jul 20, 03:03 PM
resolved — This incident has been resolved.
Jul 20, 02:23 PM
monitoring — A fix has been implemented, and we are observing recovery. We will continue to monitor the results.
Jul 20, 02:13 PM
investigating — We are continuing to investigate this issue.
+1 more updates
Issue with Dashboard Views Being Registered
minorJul 14, 08:10 AM→Jul 20, 10:15 AMresolved
Jul 20, 10:15 AM
resolved — A fix has been implemented and dashboard view and error counts in the Dashboards and Folder list are updating as expected. Thank you for your patience while we worked to address this issue.
Jul 20, 08:17 AM
identified — We are currently working on a fix related to this incident. There are no new updates to share at this time.
Jul 16, 09:27 PM
identified — We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.
+5 more updates
Grafana Cloud login issues for users with role of None
minorJul 17, 06:22 PM→Jul 18, 01:15 PMresolved
Jul 18, 01:15 PM
resolved — Users with the None role authenticating via Grafana.com should be able to access their stacks again
Jul 17, 09:33 PM
identified — We are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.
Jul 17, 07:28 PM
identified — We’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.
+1 more updates
Adaptive Metrics aggregation delay in eu-west-0 region.
minorJul 16, 12:56 PM→Jul 16, 02:58 PMresolved
Jul 16, 02:58 PM
resolved — The incident is now fully resolved.
Jul 16, 01:45 PM
monitoring — Services are fully recovered now. We're monitoring to be sure the issue won't re-occur.
Jul 16, 12:56 PM
identified — We're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the app...
Delayed Aggregated Metrics (prod-us-central-0)
minorJul 15, 07:19 PM→Jul 16, 03:54 AMresolved
Jul 16, 03:54 AM
resolved — This incident has been resolved.
Jul 15, 09:23 PM
monitoring — We have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete.
Thank you for your patience.
Jul 15, 07:19 PM
identified — We are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region.
Our engineering team has identified the issue and is actively ...
Partial Outage in prod-eu-west-2
majorJul 15, 10:24 PM→Jul 15, 11:36 PMresolved
Jul 15, 11:36 PM
resolved — This incident has been resolved. Thank you for your patience.
Jul 15, 10:53 PM
monitoring — We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.
Jul 15, 10:38 PM
investigating — We are continuing to investigate this issue.
+2 more updates
Some Reports of Grafana Not Loading.
majorJul 15, 06:07 PM→Jul 15, 09:00 PMresolved
Jul 15, 09:00 PM
resolved — This incident has been resolved. Thank you for your patience.
Jul 15, 08:50 PM
identified — The fix is in the process of being rolled out.
Jul 15, 06:47 PM
identified — We believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.
+1 more updates
Fleet Managment Interface 404's
minorJul 15, 01:54 PM→Jul 15, 01:54 PMresolved
Jul 15, 01:54 PM
resolved — Between 9:45 and 12:00 UTC, we experienced an issue affecting the Fleet Management interface. During this time, the Remote Configuration tab for production stacks returned a 404 error when accessed th...
Mimir Write Performance Degradation
minorJul 14, 08:28 PM→Jul 14, 10:00 PMresolved
Jul 14, 10:00 PM
resolved — This incident has been resolved. Thank you for your patience.
Jul 14, 08:28 PM
monitoring — From 19:24 until 19:40 UTC, backend performance degradation impacted ingestion. We are currently monitoring.
Mimir Partial Write Outage
majorJul 14, 05:09 PM→Jul 14, 07:03 PMresolved
Jul 14, 07:03 PM
resolved — This incident has been resolved.
Jul 14, 05:23 PM
monitoring — The outage is recovered as of 16:56 UTC, and we're continuing to monitor.
Jul 14, 05:09 PM
identified — We're investigating backend degradation which has resulted in a partial write outage beginning around 16:40 UTC. This degradation has also impacted reads and rule evaluation. Issue has been identifi...
IRM Mobile App forcing some users to logout
minorJul 10, 08:27 PM→Jul 13, 01:00 PMresolved
Jul 13, 01:00 PM
resolved — This incident has been resolved.
Jul 11, 03:58 AM
monitoring — A new iOS mobile app version (v2.39.5) has been released, which includes a fix for this issue. We are monitoring the rollout and its impact to ensure the issue has been fully resolved. If you continue...
Jul 10, 10:17 PM
identified — A fix has been submitted for review and will be released as Grafana Mobile v2.39.5 once it is approved.
If your app is currently on v2.39.3, we recommend not upgrading to v2.39.4 and instead waiting ...
+2 more updates
Issue with Dashboard Views Being Registered
minorJul 10, 03:57 PM→Jul 13, 05:23 AMresolved
Jul 13, 05:23 AM
resolved — At this stage, we are considering the incident resolved.
Jul 13, 05:21 AM
investigating — We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soo...
Jul 10, 09:40 PM
investigating — We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soo...
+1 more updates
Network Degredation in prod-us-central-0
majorJul 10, 09:31 PM→Jul 10, 09:31 PMresolved
Jul 10, 09:31 PM
resolved — From approximately 20:24 UTC - 20:47 UTC a network issue in prod-us-central-0 isolated part of our infrastructure in one availability zone, temporarily cutting off connectivity to a subset of backend ...
PDC Degraded performance
criticalJul 9, 11:37 AM→Jul 9, 07:54 PMresolved
Jul 9, 07:54 PM
resolved — This incident has been resolved. Thank you for your patience.
Jul 9, 05:52 PM
monitoring — We are now seeing recovery of PDC traffic to the affected customers. Our team will continue to monitor across shifts.
Jul 9, 04:08 PM
identified — We are continuing to see disruptions across multiple deployments, ranging from degraded performance to full outages for some customers.
Additional resources have been engaged to mitigate this issue. ...
+5 more updates
Delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1
minorJul 9, 10:28 AM→Jul 9, 01:07 PMresolved
Jul 9, 01:07 PM
resolved — This incident has been resolved.
Jul 9, 11:18 AM
monitoring — A fix has been implemented. We are currently monitoring the results.
Jul 9, 10:28 AM
investigating — We are observing delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1. As of yet we have not noticed any customer impact however we are currently observing the cell.
Loki read path in prod-eu-west-2 was down
noneJul 9, 07:37 AM→Jul 9, 07:37 AMresolved
Jul 9, 07:37 AM
resolved — Between 5:49 and 5:54 UTC, the read path in prod-eu-west-2 was down. This has completely recovered by 5:57 UTC. Recording rules may have failed to evaluate during this period which may result in gaps.
Grafana rulers crash-looping on prometheus
minorJul 8, 12:17 PM→Jul 8, 05:23 PMresolved
Jul 8, 05:23 PM
resolved — This incident has been resolved.
Jul 8, 02:23 PM
monitoring — We are in the process of rolling out the fix. Our engineers are monitoring the progress.
Jul 8, 12:17 PM
investigating — We are investigating issues with the grafana-ruler service on prod-us-east-2 and prod-us-west-0 which are causing periodic crash conditions. A code fix is currently being deployed to mitigate this
Loki Billing Data Lost
minorJul 8, 03:33 PM→Jul 8, 03:33 PMresolved
Jul 8, 03:33 PM
resolved — Between 13:57 and 14:20 UTC (23 minutes), a data gap occurred affecting billing usage data for Loki. Data for this window was not recorded and cannot be recovered.
Cannot access dashboard set up as a home page
minorJul 8, 10:28 AM→Jul 8, 11:00 AMresolved
Jul 8, 11:00 AM
resolved — We have found that the issues are related to the grafana.unifiedHomepage feature rollout active 16:50 UTC yesterday, to 10:20 UTC today. We have rolled this back and systems are now working as expecte...
Jul 8, 10:28 AM
investigating — We are currently investigating an issue affecting some customers who are unable to load or access their dashboards when configured as their home page. We will update once we have more information on t...
Delayed usage and billing data
noneJul 7, 01:44 PM→Jul 7, 01:44 PMresolved
Jul 7, 01:44 PM
resolved — We identified an issue with the internal job that calculates month-to-date usage and cost data, which caused usage attribution and billing dashboards to display stale information. The root cause has b...
Grafana Cloud IRM alert groups failing to produce alerts in us-east-3 region.
majorJul 5, 11:37 AM→Jul 6, 07:37 AMresolved
Jul 6, 07:37 AM
resolved — We haven't noticed any further issues in this region for alert group processing since yesterday. This incident i fully resolved.
Jul 5, 01:38 PM
monitoring — The work on bringing services back to healthy state is completed and alerts should be created correctly, without a delay now. We're keeping the incident open and monitoring for any potential hiccups t...
Jul 5, 11:37 AM
identified — Some IRM alert groups in us-east-3 region may not be able to produce alerts and could be sluggish/misbehaving in general. The issue started around 01:00 UTC on July 3rd. The issue has been identified ...
Partial Write Outage for Grafana Cloud Logs in prod-eu-north-0.
minorJul 3, 02:27 PM→Jul 3, 02:27 PMresolved
Jul 3, 02:27 PM
resolved — Grafana Cloud Logs in prod-eu-north-0 experienced a 10-minute partial write outage between 13:45 and 13:54 UTC. Impacted users may have experienced 5xx errors during this time.
Some Queries Failing
minorJul 2, 06:10 PM→Jul 2, 09:35 PMresolved
Jul 2, 09:35 PM
resolved — This incident has been resolved. Thank you for your patience.
Jul 2, 08:16 PM
identified — We are continuing to work on rolling back a PR responsible for this behavior. Once we have more information, we will share it here.
Thank you for your patience.
Jul 2, 06:10 PM
identified — Queries with drop __error__, including log volume histogram queries (the queries that generate the histogram visualization in Grafana), are failing due to a bug with series limit checks.
The Root cau...
Loki and Frontend Observability - Major Outage in prod-us-central-0 region
criticalJul 2, 04:55 AM→Jul 2, 06:49 AMresolved
Jul 2, 06:49 AM
resolved — This incident has been resolved by restarting the affected services.
Jul 2, 04:57 AM
investigating — The AWS Logs integration in the same region is affected as well. We will provide further updates as our investigation progresses.
Jul 2, 04:55 AM
investigating — We are currently investigating a major outage in Loki writes and Frontend Observability in the prod-us-central-0 region. Our Engineering team is investigating this and we will provide further updates ...
Elevated Loki Query Bytes Reporting
minorJul 1, 07:43 PM→Jul 1, 10:00 PMresolved
Jul 1, 10:00 PM
resolved — We’ve implemented a fix and can confirm the issue is fully resolved as of 20:25 UTC.
Thank you for your patience.
Jul 1, 07:43 PM
investigating — We are investigating an issue where some customers may see higher Loki query byte usage reported than was actually consumed. This affects usage reporting only; there is no impact to query execution or...
June 2026
Mimir read errors and high latency in prod-eu-west-0
minorJun 29, 11:37 AM→Jun 30, 08:20 AMresolved
Jun 30, 08:20 AM
resolved — Since the mitigation has been applied, we have not seen the errors return. At this point, we are considering the incident resolved.
Jun 29, 05:47 PM
identified — We've identified a possible cause, and a mitigation is in place to prevent further occurrences.
Jun 29, 03:37 PM
investigating — The errors and latency have now recovered, we continue investigating the root cause.
+2 more updates
Confluent API Outage
majorJun 29, 12:37 PM→Jun 29, 02:27 PMresolved
Jun 29, 02:27 PM
resolved — This incident has been resolved.
Jun 29, 12:37 PM
investigating — We are investigating an issue affecting Confluent metrics ingestion across all regions. Due to an elevated error rate on the Confluent side, some metrics may not be ingested, resulting in potential da...
Rule evaluation error on cluster prod-gb-south-0
minorJun 29, 10:34 AM→Jun 29, 01:12 PMresolved
Jun 29, 01:12 PM
resolved — This incident has been resolved.
Jun 29, 11:14 AM
monitoring — A fix has been applied and we are currently monitoring results.
Jun 29, 10:34 AM
investigating — We are currently investigating Rule Evaluation errors on the cluster prod-gb-south-0 which is leading to error codes showing within the stacks. We are looking into the issue and will update accordingl...
K6 - Test run metrics processing is delayed
minorJun 26, 01:08 PM→Jun 28, 02:35 AMresolved
Jun 28, 02:35 AM
resolved — This incident has been resolved.
Jun 26, 11:30 PM
investigating — We've improved the metric ingestion delay time and are working on additional fixes to bring it down to expected range. Customers can currently expect a delay of 2 to 5 minutes before their test run me...
Jun 26, 04:33 PM
investigating — Update: Changed incident title to "Test run metrics processing is delayed"
We have found the issue and are working on deploying the fix.
+3 more updates
Potential Issues Loading Grafana for Users in India
majorJun 18, 03:18 PM→Jun 24, 12:46 PMresolved
Jun 24, 12:46 PM
resolved — This incident has been resolved. Error rates continued to remain near 0 and operations are performing as expected.
Jun 23, 05:54 PM
monitoring — Error rates have remained near zero, and we continue to monitor.
Jun 23, 01:15 PM
monitoring — We are continuing to monitor for further issues.
+6 more updates
Elevated Logs Query Usage — Unexpected Billing Impact
noneJun 20, 08:00 PM→Jun 20, 08:00 PMresolved
Jun 22, 10:10 PM
resolved — Starting June 20, 2026 at approximately 20:00 UTC, some Grafana Cloud customers experienced unexpectedly elevated logs query usage. The issue persisted until it was resolved on June 22, 2026 at approx...
Rule Evaluation Outage in prod-us-central-0
majorJun 19, 04:39 PM→Jun 19, 07:19 PMresolved
Jun 19, 07:19 PM
resolved — This incident has been resolved. Thank you for your patience.
Jun 19, 05:46 PM
monitoring — We’re continuing to track progress post-mitigation. While we don’t have new information to share yet, our team remains actively engaged.
Jun 19, 04:39 PM
monitoring — We had an outage affecting rule evaluations between 15:16-15:59 UTC in the prod-us-central-0 region.
Our team quickly identified the issue and has since mitigated. The engineering team is monitoring...
Issues with actions in the Grafana IRM mobile app
majorJun 18, 05:01 PM→Jun 18, 06:50 PMresolved
Jun 18, 06:50 PM
resolved — This incident has been resolved. Thank you for your patience.
Jun 18, 06:09 PM
monitoring — We've verified a fix in our staging environment to restore functionality to the mobile app. The fix is currently being deployed to production. Thanks for your patience as we continue to roll this out ...
Jun 18, 05:01 PM
identified — We're noticing an uptick in users being unable to respond to actions on the mobile app (acknowledging and silencing alerts, for example). Users working in the web UI should not be affected. Ingestion ...
Degraded k6 cloud UI performance
criticalJun 18, 11:25 AM→Jun 18, 04:21 PMresolved
Jun 18, 04:21 PM
resolved — This incident has been resolved. Thank you for your patience.
Jun 18, 02:04 PM
monitoring — We are continuing to monitor for any further issues.
Jun 18, 01:15 PM
monitoring — The root cause of the issue has been identified and a fix has been successfully deployed. We are observing widespread improvements across all systems. Our team is currently monitoring the environment ...
+2 more updates
Loki data source-managed alert rules not visible in the Grafana Cloud Alerting UI
majorJun 17, 08:17 PM→Jun 18, 02:08 PMresolved
Jun 18, 02:08 PM
resolved — This incident has been resolved.
Jun 18, 08:04 AM
monitoring — A fix has been implemented and we are monitoring the results.
Jun 17, 11:22 PM
identified — We are continuing to deploy the fix and monitor recovery efforts. As part of the rollout, we identified an issue that required adjustments to our deployment plan, which has extended the timeline for m...
+3 more updates
Related Incident Histories
Get Grafana Cloud Outage Alerts
Be the first to know when Grafana Cloud go down.