G

Grafana Cloud Outage History

Past incidents and downtime events

Complete history of Grafana Cloud outages, incidents, and service disruptions. Showing 50 most recent incidents.

August 2026(1 incident)

minorresolvedAug 1, 07:25 AM — Resolved Aug 1, 08:31 AM

K6 - Cloud test-run issues

3 updates
resolvedAug 1, 08:31 AM

This incident has been resolved.

monitoringAug 1, 08:06 AM

A critical piece of infrastructure behaved poorly after a reboot. It has now been properly recovered

investigatingAug 1, 07:25 AM

We are currently investigating an issue which is resulting in some test runs to abort and metrics data to be incomplete for a subset of tests

July 2026(39 incidents)

noneresolvedJul 31, 07:30 PM — Resolved Jul 31, 07:30 PM

Partial Read Outage for Loki in prod-us-east-4

1 update
resolvedJul 31, 08:33 PM

We are investigating a partial read outage affecting Loki in prod-us-east-4. Between 19:41 UTC and 19:50 UTC, a significant portion of read queries may have failed or returned errors. The issue has been identified and service has been restored. We are continuing to investigate the underlying cause and will provide additional information as it becomes available.

minorresolvedJul 31, 10:51 AM — Resolved Jul 31, 12:13 PM

Degraded Performance: Stack Provisioning Failures within certain reigons (PDC Setup)

2 updates
resolvedJul 31, 12:13 PM

He have identified the cause and applied a fix for this issue and all effected stacks within the affected regions are not working as expected.

investigatingJul 31, 10:51 AM

We are currently investigating an issue impacting stack provisioning. Attempting to set up PDC on newly created stacks will currently fail across several regions. We are currently looking into the cause of this.

majorresolvedJul 30, 02:45 PM — Resolved Jul 30, 05:40 PM

PDC Authentication Issues

5 updates
resolvedJul 30, 05:40 PM

This incident has been resolved.

monitoringJul 30, 03:56 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 30, 03:11 PM

We are continuing to investigate this issue.

investigatingJul 30, 02:51 PM

We are continuing to investigate this issue.

investigatingJul 30, 02:45 PM

We are currently investigating an issue impacting PDC authentication. We will provide additional updates as they become available.

majorresolvedJul 30, 01:15 PM — Resolved Jul 30, 03:17 PM

Issues with Billing/Usage Dashboard Metrics and Panels.

4 updates
resolvedJul 30, 03:17 PM

This incident has been resolved.

monitoringJul 30, 02:34 PM

A fix has been implemented and we are monitoring the results.

identifiedJul 30, 01:50 PM

We have identified the issue, and are working on deploying a fix.

investigatingJul 30, 01:15 PM

We are investigating an outage for the Billing / Usage dashboard metrics and panels. This appears to be partially affecting organizations using Grafana Cloud. We are working on identifying and resolving the issue

minorresolvedJul 29, 09:46 PM — Resolved Jul 30, 12:39 AM

IRM Performance Degradation in EU Region

3 updates
resolvedJul 30, 12:39 AM

The issue affecting Grafana IRM in the EU region has been resolved. The IRM UI, public API, and alert notification processing have been restored and are operating normally.

identifiedJul 29, 11:07 PM

We have identified the underlying issue affecting Grafana IRM in our EU region and are actively working to restore normal service. Users may continue to experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notifications while mitigation efforts are underway. Our engineering team is working to restore full functionality as quickly as possible. We will provide another update as more information becomes available.

investigatingJul 29, 09:46 PM

We are currently investigating an issue affecting Grafana IRM in our EU region. Users may experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notifications. Our engineering team is actively investigating the issue and working to restore normal service. We will provide another update as more information becomes available.

minorresolvedJul 29, 06:49 PM — Resolved Jul 29, 06:49 PM

Grafana Cloud non-billing usage metrics gaps in select regions

1 update
resolvedJul 29, 06:49 PM

Impact period: April 1 – July 29, 2026 Summary: During this period, some customers in a subset of regions may have experienced gaps in select ruler/recording rule metrics within their Grafana Cloud instances. Not all customers or environments in the listed regions were affected. Affected regions: A subset of environments across the following regions may have been impacted: ap-south-1 au-southeast-1 ca-east-0 eu-central-0 eu-west-2 eu-west-7 us-west-0 Affected values: grafanacloud_instance_queries_per_second grafanacloud_instance_rule_config_last_reload_successful grafanacloud_instance_rule_evaluations_total:rate5m grafanacloud_instance_rule_evaluation_failures_total:rate5m grafanacloud_instance_rule_group_interval_seconds grafanacloud_instance_rule_group_last_duration_seconds grafanacloud_instance_rule_group_iterations_total:rate5m grafanacloud_instance_rule_group_iterations_missed_total:rate5m grafanacloud_instance_rule_group_last_evaluation_timestamp_seconds grafanacloud_instance_rule_group_rules grafanacloud_instance_ruler_queries_failed_total:rate5m grafanacloud_instance_ruler_queries_zero_fetched_series_total:rate5m grafanacloud_instance_ruler_notifications_sent_total:rate5m grafanacloud_instance_ruler_notifications_errors_total:rate5m grafanacloud_instance_ruler_notifications_queue_capacity grafanacloud_instance_ruler_notifications_queue_length grafanacloud_instance_ruler_notifications_latency_seconds:99quantile grafanacloud_instance_ruler_notifications_latency_seconds:50quantile Current status: This issue has been resolved. No further action is required from customers. If you continue to notice gaps in the metrics described above, please reach out to support and reference this incident.

majorresolvedJul 28, 07:40 PM — Resolved Jul 28, 07:43 PM

Partial OTLP Write Outage in prod-us-east-3

2 updates
resolvedJul 28, 07:43 PM

The issue affecting OTLP ingestion in the prod-us-east-3 region has been resolved. Between 13:30 UTC and 17:45 UTC, some customers experienced intermittent failures when writing telemetry to the OTLP endpoint. Logs were confirmed to be affected, and metrics and traces may also have experienced intermittent ingestion failures. Our investigation has concluded, and the affected services have recovered.

investigatingJul 28, 07:40 PM

We are investigating an issue affecting OTLP ingestion in the prod-us-east-3 region. Customers may experience intermittent failures when writing telemetry to the OTLP endpoint. Based on current information, logs are confirmed to be affected, and metrics and traces may also be impacted. Our investigation indicates intermittent failures began around 08:00 UTC, with the primary period of impact occurring between 15:30 UTC and 17:45 UTC. We are continuing to investigate the scope and root cause of this issue and will provide additional updates as more information becomes available.

criticalresolvedJul 24, 05:02 PM — Resolved Jul 24, 06:20 PM

Write Outage

3 updates
resolvedJul 24, 06:20 PM

This incident has been resolved.

monitoringJul 24, 05:18 PM

Healthy as of 17:00 UTC. We are continuing to monitor.

investigatingJul 24, 05:02 PM

We are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.

majorresolvedJul 24, 09:17 AM — Resolved Jul 24, 02:03 PM

Errors creating new Slack integration for Grafana IRM

4 updates
resolvedJul 24, 02:03 PM

This incident has been resolved.

identifiedJul 24, 10:12 AM

We have verified the fix, and we are starting to roll it out

identifiedJul 24, 09:49 AM

We have identified the root cause of the problem, and we are working on a fix

investigatingJul 24, 09:17 AM

We are investigating a possible error creating Slack integrations. Will update the status soon.

noneresolvedJul 23, 11:00 PM — Resolved Jul 24, 05:32 AM

Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23

2 updates
resolvedJul 24, 05:32 AM

Writes are looking stable in the last 6h

monitoringJul 23, 11:00 PM

status is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor

majorresolvedJul 23, 05:14 PM — Resolved Jul 23, 07:47 PM

Metrics Write Path Errors

3 updates
resolvedJul 23, 07:47 PM

This incident has been resolved.

monitoringJul 23, 05:37 PM

Error rates have dropped, and we are monitoring this issue for any recurrence.

investigatingJul 23, 05:14 PM

We are currently investigating an elevated rate of error and latency in the impacted write path.

majorresolvedJul 23, 10:13 AM — Resolved Jul 23, 03:37 PM

Stacks using SCIM user provisioning are currently unable to log into Grafana

4 updates
resolvedJul 23, 03:37 PM

This incident has been resolved.

identifiedJul 23, 01:23 PM

We have applied a fix and are currently awaiting feedback from affected customers.

identifiedJul 23, 11:26 AM

The issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.

investigatingJul 23, 10:13 AM

We are currently facing an issue where Stacks using SCIM user provisioning are currently unable to log into Grafana. We are currently investigating this issue and working on a fix.

minorresolvedJul 22, 10:27 AM — Resolved Jul 22, 11:25 AM

K6 - Cloud output test-runs are failing to fetch the script logs

3 updates
resolvedJul 22, 11:25 AM

This incident has been resolved.

monitoringJul 22, 10:52 AM

We have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.

investigatingJul 22, 10:27 AM

We are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.

majorresolvedJul 21, 06:43 PM — Resolved Jul 21, 09:06 PM

Cloud Log Exporter Unavailable

3 updates
resolvedJul 21, 09:06 PM

This incident has been resolved.

monitoringJul 21, 08:04 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 21, 06:43 PM

Cloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.

majorresolvedJul 20, 01:48 PM — Resolved Jul 20, 03:03 PM

PDC Issues

4 updates
resolvedJul 20, 03:03 PM

This incident has been resolved.

monitoringJul 20, 02:23 PM

A fix has been implemented, and we are observing recovery. We will continue to monitor the results.

investigatingJul 20, 02:13 PM

We are continuing to investigate this issue.

investigatingJul 20, 01:48 PM

We are currently investigating an issue that is causing issues with PDC in the prod-eu-west-2 region. We will provide another update in 1-2 hours.

minorresolvedJul 14, 08:10 AM — Resolved Jul 20, 10:15 AM

Issue with Dashboard Views Being Registered

8 updates
resolvedJul 20, 10:15 AM

A fix has been implemented and dashboard view and error counts in the Dashboards and Folder list are updating as expected. Thank you for your patience while we worked to address this issue.

identifiedJul 20, 08:17 AM

We are currently working on a fix related to this incident. There are no new updates to share at this time.

identifiedJul 16, 09:27 PM

We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.

identifiedJul 16, 01:34 PM

We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.

identifiedJul 15, 08:52 PM

We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.

identifiedJul 15, 12:05 PM

We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.

identifiedJul 14, 02:27 PM

We've identified the cause of the issue impacting dashboard view and error counts in the Dashboards and Folder list. Our team is currently working on implementing a fix and validating the solution. We will provide another update within the next 24 hours, or sooner if we have additional information to share.

investigatingJul 14, 08:10 AM

This is related to https://status.grafana.com/incidents/rhrk2ck6ly0y which was resolved by mistake. For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly. This is ultimately causing inaccurate or missing view counts. We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.

minorresolvedJul 17, 06:22 PM — Resolved Jul 18, 01:15 PM

Grafana Cloud login issues for users with role of None

4 updates
resolvedJul 18, 01:15 PM

Users with the None role authenticating via Grafana.com should be able to access their stacks again

identifiedJul 17, 09:33 PM

We are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.

identifiedJul 17, 07:28 PM

We’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.

investigatingJul 17, 06:22 PM

We’re currently investigating an issue with user logins with the None role in Grafana Cloud. Our team is actively working to identify the cause. Thank you for your patience.

minorresolvedJul 16, 12:56 PM — Resolved Jul 16, 02:58 PM

Adaptive Metrics aggregation delay in eu-west-0 region.

3 updates
resolvedJul 16, 02:58 PM

The incident is now fully resolved.

monitoringJul 16, 01:45 PM

Services are fully recovered now. We're monitoring to be sure the issue won't re-occur.

identifiedJul 16, 12:56 PM

We're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the appropriate fix is already deployed. We're seeing services recovering. More updated to come soon.

minorresolvedJul 15, 07:19 PM — Resolved Jul 16, 03:54 AM

Delayed Aggregated Metrics (prod-us-central-0)

3 updates
resolvedJul 16, 03:54 AM

This incident has been resolved.

monitoringJul 15, 09:23 PM

We have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete. Thank you for your patience.

identifiedJul 15, 07:19 PM

We are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region. Our engineering team has identified the issue and is actively working on a mitigation. We will provide further updates as the investigation progresses.

majorresolvedJul 15, 10:24 PM — Resolved Jul 15, 11:36 PM

Partial Outage in prod-eu-west-2

5 updates
resolvedJul 15, 11:36 PM

This incident has been resolved. Thank you for your patience.

monitoringJul 15, 10:53 PM

We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.

investigatingJul 15, 10:38 PM

We are continuing to investigate this issue.

investigatingJul 15, 10:35 PM

We are investigating a broader issue affecting multiple Grafana Cloud products in prod-eu-west-2. This appears to be caused by a third-party provider issue rather than a Loki-specific problem. Impact is currently inconsistent: some components are affected while others continue to function normally. We have also seen some impact to Tempo write paths. We are continuing to investigate and will share another update as soon as we have more information.

investigatingJul 15, 10:24 PM

We are investigating an issue affecting Loki queries in prod-eu-west-2. We first observed this behavior at approximately 21:57 UTC. Affected users may see elevated query errors, timeouts, or intermittent failures when running Loki queries in this region. The issue appears to be improving, but it is not fully resolved yet. We are continuing to investigate and will share another update as soon as we have more information.

majorresolvedJul 15, 06:07 PM — Resolved Jul 15, 09:00 PM

Some Reports of Grafana Not Loading.

4 updates
resolvedJul 15, 09:00 PM

This incident has been resolved. Thank you for your patience.

identifiedJul 15, 08:50 PM

The fix is in the process of being rolled out.

identifiedJul 15, 06:47 PM

We believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.

investigatingJul 15, 06:07 PM

We are currently investigating an issue impacting a small subset of stacks in the impacted regions. We will provide more details as they become available.

minorresolvedJul 15, 01:54 PM — Resolved Jul 15, 01:54 PM

Fleet Managment Interface 404's

1 update
resolvedJul 15, 01:54 PM

Between 9:45 and 12:00 UTC, we experienced an issue affecting the Fleet Management interface. During this time, the Remote Configuration tab for production stacks returned a 404 error when accessed through the UI. Backend systems were not affected. Configuration-as-code workflows and other non-UI methods of interacting with Fleet Management continued to operate normally throughout the incident. This issue has since been resolved.

minorresolvedJul 14, 08:28 PM — Resolved Jul 14, 10:00 PM

Mimir Write Performance Degradation

2 updates
resolvedJul 14, 10:00 PM

This incident has been resolved. Thank you for your patience.

monitoringJul 14, 08:28 PM

From 19:24 until 19:40 UTC, backend performance degradation impacted ingestion.  We are currently monitoring.

majorresolvedJul 14, 05:09 PM — Resolved Jul 14, 07:03 PM

Mimir Partial Write Outage

3 updates
resolvedJul 14, 07:03 PM

This incident has been resolved.

monitoringJul 14, 05:23 PM

The outage is recovered as of 16:56 UTC, and we're continuing to monitor.

identifiedJul 14, 05:09 PM

We're investigating backend degradation which has resulted in a partial write outage beginning around 16:40 UTC.  This degradation has also impacted reads and rule evaluation.  Issue has been identified and we are working on mitigation.

minorresolvedJul 10, 08:27 PM — Resolved Jul 13, 01:00 PM

IRM Mobile App forcing some users to logout

5 updates
resolvedJul 13, 01:00 PM

This incident has been resolved.

monitoringJul 11, 03:58 AM

A new iOS mobile app version (v2.39.5) has been released, which includes a fix for this issue. We are monitoring the rollout and its impact to ensure the issue has been fully resolved. If you continue to experience any problems after updating to the latest version, please let us know.

identifiedJul 10, 10:17 PM

A fix has been submitted for review and will be released as Grafana Mobile v2.39.5 once it is approved. If your app is currently on v2.39.3, we recommend not upgrading to v2.39.4 and instead waiting for v2.39.5 to become available. Users who have already been signed out by this issue can log back into the app and continue using it. We will provide another update once v2.39.5 is available.

investigatingJul 10, 08:37 PM

Our investigation has narrowed the impact to version 2.39.4 of the Grafana mobile app on iOS. Users should avoid upgrading to this version until an updated release is available. As a precaution, we also recommend ensuring you have alternative notification methods configured (such as SMS, email, or phone calls) if you rely on mobile push notifications for alerting. We are actively working on a resolution and will provide another update as soon as more information is available.

investigatingJul 10, 08:27 PM

We are aware of an issue in the latest mobile app release that is causing some users to be signed out and asked to log in again. We have reproduced the behavior and are investigating and working on a fix. We will share another update as soon as we have more information.

minorresolvedJul 10, 03:57 PM — Resolved Jul 13, 05:23 AM

Issue with Dashboard Views Being Registered

4 updates
resolvedJul 13, 05:23 AM

At this stage, we are considering the incident resolved.

investigatingJul 13, 05:21 AM

We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.

investigatingJul 10, 09:40 PM

We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.

investigatingJul 10, 03:57 PM

For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly This is ultimately causing inaccurate or missing view counts.

majorresolvedJul 10, 09:31 PM — Resolved Jul 10, 09:31 PM

Network Degredation in prod-us-central-0

1 update
resolvedJul 10, 09:31 PM

From approximately 20:24 UTC - 20:47 UTC a network issue in prod-us-central-0 isolated part of our infrastructure in one availability zone, temporarily cutting off connectivity to a subset of backend services. This caused elevated query latency and some delayed metric evaluations, along with related write-path errors in a logging subsystem. All systems recovered automatically once network connectivity was restored. Impacted customers may have had queries or rule evaluations fail during the incident, but ingestion was not impacted.

criticalresolvedJul 9, 11:37 AM — Resolved Jul 9, 07:54 PM

PDC Degraded performance

8 updates
resolvedJul 9, 07:54 PM

This incident has been resolved. Thank you for your patience.

monitoringJul 9, 05:52 PM

We are now seeing recovery of PDC traffic to the affected customers. Our team will continue to monitor across shifts.

identifiedJul 9, 04:08 PM

We are continuing to see disruptions across multiple deployments, ranging from degraded performance to full outages for some customers. Additional resources have been engaged to mitigate this issue. We will post updates as they become available.

monitoringJul 9, 02:44 PM

We are continuing to monitor for any further issues.

monitoringJul 9, 02:10 PM

We are seeing recovery in some deployments, while others not just yet. We are continuing to monitor this incident.

monitoringJul 9, 01:04 PM

A fix has been applied and we are seeing recovery. We will continue to monitor this.

investigatingJul 9, 12:36 PM

We are still currently investigating the issue.

investigatingJul 9, 11:37 AM

We are currently facing performance degradation on PDC service hosted on Multiple clusters. Our Engineering Team is currently working on fixing the issue, we do apologize for any inconvenience.

minorresolvedJul 9, 10:28 AM — Resolved Jul 9, 01:07 PM

Delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1

3 updates
resolvedJul 9, 01:07 PM

This incident has been resolved.

monitoringJul 9, 11:18 AM

A fix has been implemented. We are currently monitoring the results.

investigatingJul 9, 10:28 AM

We are observing delayed ingestion and recording rule evaluation failures for Mimir in prod-ap-south-1. As of yet we have not noticed any customer impact however we are currently observing the cell.

noneresolvedJul 9, 07:37 AM — Resolved Jul 9, 07:37 AM

Loki read path in prod-eu-west-2 was down

1 update
resolvedJul 9, 07:37 AM

Between 5:49 and 5:54 UTC, the read path in prod-eu-west-2 was down. This has completely recovered by 5:57 UTC. Recording rules may have failed to evaluate during this period which may result in gaps.

minorresolvedJul 8, 12:17 PM — Resolved Jul 8, 05:23 PM

Grafana rulers crash-looping on prometheus

3 updates
resolvedJul 8, 05:23 PM

This incident has been resolved.

monitoringJul 8, 02:23 PM

We are in the process of rolling out the fix. Our engineers are monitoring the progress.

investigatingJul 8, 12:17 PM

We are investigating issues with the grafana-ruler service on prod-us-east-2 and prod-us-west-0 which are causing periodic crash conditions. A code fix is currently being deployed to mitigate this

minorresolvedJul 8, 03:33 PM — Resolved Jul 8, 03:33 PM

Loki Billing Data Lost

1 update
resolvedJul 8, 03:33 PM

Between 13:57 and 14:20 UTC (23 minutes), a data gap occurred affecting billing usage data for Loki. Data for this window was not recorded and cannot be recovered.

minorresolvedJul 8, 10:28 AM — Resolved Jul 8, 11:00 AM

Cannot access dashboard set up as a home page

2 updates
resolvedJul 8, 11:00 AM

We have found that the issues are related to the grafana.unifiedHomepage feature rollout active 16:50 UTC yesterday, to 10:20 UTC today. We have rolled this back and systems are now working as expected.

investigatingJul 8, 10:28 AM

We are currently investigating an issue affecting some customers who are unable to load or access their dashboards when configured as their home page. We will update once we have more information on this.

noneresolvedJul 7, 01:44 PM — Resolved Jul 7, 01:44 PM

Delayed usage and billing data

1 update
resolvedJul 7, 01:44 PM

We identified an issue with the internal job that calculates month-to-date usage and cost data, which caused usage attribution and billing dashboards to display stale information. The root cause has been identified and resolved. Your billing dashboard may show a sudden jump in usage. This is expected. It reflects several days of accumulated usage that hadn't been showing up while the issue was ongoing, not a sudden change in your actual usage. This issue does not affect your end-of-month invoice. Your bill will be calculated based on actual usage, not the numbers displayed during this period.

majorresolvedJul 5, 11:37 AM — Resolved Jul 6, 07:37 AM

Grafana Cloud IRM alert groups failing to produce alerts in us-east-3 region.

3 updates
resolvedJul 6, 07:37 AM

We haven't noticed any further issues in this region for alert group processing since yesterday. This incident i fully resolved.

monitoringJul 5, 01:38 PM

The work on bringing services back to healthy state is completed and alerts should be created correctly, without a delay now. We're keeping the incident open and monitoring for any potential hiccups that might occur.

identifiedJul 5, 11:37 AM

Some IRM alert groups in us-east-3 region may not be able to produce alerts and could be sluggish/misbehaving in general. The issue started around 01:00 UTC on July 3rd. The issue has been identified and the cause of the issue has been fixed. Our team is actively working on getting the service back to healthy state.

minorresolvedJul 3, 02:27 PM — Resolved Jul 3, 02:27 PM

Partial Write Outage for Grafana Cloud Logs in prod-eu-north-0.

1 update
resolvedJul 3, 02:27 PM

Grafana Cloud Logs in prod-eu-north-0 experienced a 10-minute partial write outage between 13:45 and 13:54 UTC. Impacted users may have experienced 5xx errors during this time.

minorresolvedJul 2, 06:10 PM — Resolved Jul 2, 09:35 PM

Some Queries Failing

3 updates
resolvedJul 2, 09:35 PM

This incident has been resolved. Thank you for your patience.

identifiedJul 2, 08:16 PM

We are continuing to work on rolling back a PR responsible for this behavior. Once we have more information, we will share it here. Thank you for your patience.

identifiedJul 2, 06:10 PM

Queries with drop __error__, including log volume histogram queries (the queries that generate the histogram visualization in Grafana), are failing due to a bug with series limit checks. The Root cause has been identified, and we are working on a fix.

criticalresolvedJul 2, 04:55 AM — Resolved Jul 2, 06:49 AM

Loki and Frontend Observability - Major Outage in prod-us-central-0 region

3 updates
resolvedJul 2, 06:49 AM

This incident has been resolved by restarting the affected services.

investigatingJul 2, 04:57 AM

The AWS Logs integration in the same region is affected as well. We will provide further updates as our investigation progresses.

investigatingJul 2, 04:55 AM

We are currently investigating a major outage in Loki writes and Frontend Observability in the prod-us-central-0 region. Our Engineering team is investigating this and we will provide further updates as our investigation progresses.

minorresolvedJul 1, 07:43 PM — Resolved Jul 1, 10:00 PM

Elevated Loki Query Bytes Reporting

2 updates
resolvedJul 1, 10:00 PM

We’ve implemented a fix and can confirm the issue is fully resolved as of 20:25 UTC. Thank you for your patience.

investigatingJul 1, 07:43 PM

We are investigating an issue where some customers may see higher Loki query byte usage reported than was actually consumed. This affects usage reporting only; there is no impact to query execution or service availability. The issue began at approximately 13:20 UTC and is ongoing. We expect the issue to be resolved soon and will provide another update as more information becomes available.

June 2026(10 incidents)

minorresolvedJun 29, 11:37 AM — Resolved Jun 30, 08:20 AM

Mimir read errors and high latency in prod-eu-west-0

5 updates
resolvedJun 30, 08:20 AM

Since the mitigation has been applied, we have not seen the errors return. At this point, we are considering the incident resolved.

identifiedJun 29, 05:47 PM

We've identified a possible cause, and a mitigation is in place to prevent further occurrences.

investigatingJun 29, 03:37 PM

The errors and latency have now recovered, we continue investigating the root cause.

investigatingJun 29, 01:52 PM

The errors are recovering, and we are still looking into the root cause of this.

investigatingJun 29, 11:37 AM

We are currently investigating an issue with Mimir in prod-eu-west-0 we are seeing read errors and high latency. This incident is currently ongoing. The errors are recovering but we are currently looking into the route cause of this.

majorresolvedJun 29, 12:37 PM — Resolved Jun 29, 02:27 PM

Confluent API Outage

2 updates
resolvedJun 29, 02:27 PM

This incident has been resolved.

investigatingJun 29, 12:37 PM

We are investigating an issue affecting Confluent metrics ingestion across all regions. Due to an elevated error rate on the Confluent side, some metrics may not be ingested, resulting in potential data loss. We are actively investigating the issue and will provide updates as more information becomes available.

minorresolvedJun 29, 10:34 AM — Resolved Jun 29, 01:12 PM

Rule evaluation error on cluster prod-gb-south-0

3 updates
resolvedJun 29, 01:12 PM

This incident has been resolved.

monitoringJun 29, 11:14 AM

A fix has been applied and we are currently monitoring results.

investigatingJun 29, 10:34 AM

We are currently investigating Rule Evaluation errors on the cluster prod-gb-south-0 which is leading to error codes showing within the stacks. We are looking into the issue and will update accordingly.

minorresolvedJun 26, 01:08 PM — Resolved Jun 28, 02:35 AM

K6 - Test run metrics processing is delayed

6 updates
resolvedJun 28, 02:35 AM

This incident has been resolved.

investigatingJun 26, 11:30 PM

We've improved the metric ingestion delay time and are working on additional fixes to bring it down to expected range. Customers can currently expect a delay of 2 to 5 minutes before their test run metrics show up (after starting a test).

investigatingJun 26, 04:33 PM

Update: Changed incident title to "Test run metrics processing is delayed" We have found the issue and are working on deploying the fix.

investigatingJun 26, 02:09 PM

We are experiencing intermittent delays with secondary metrics processing for k6 Cloud test runs due to heavy load. We don't expect any data loss or impact on user runs, but results may take longer time to appear in UI.

investigatingJun 26, 01:19 PM

A small update: The issue is isolated to new test runs, and users can go see the metrics of all the previous test runs

investigatingJun 26, 01:08 PM

We’re currently investigating an issue causing metrics not to appear during test runs. . Our team is actively working to identify the cause. Thank you for your patience.

majorresolvedJun 18, 03:18 PM — Resolved Jun 24, 12:46 PM

Potential Issues Loading Grafana for Users in India

9 updates
resolvedJun 24, 12:46 PM

This incident has been resolved. Error rates continued to remain near 0 and operations are performing as expected.

monitoringJun 23, 05:54 PM

Error rates have remained near zero, and we continue to monitor.

monitoringJun 23, 01:15 PM

We are continuing to monitor for further issues.

monitoringJun 23, 09:14 AM

We have deployed additional mitigations that should help with remaining errors. We are continuing to monitor error rates.

monitoringJun 22, 05:16 PM

We’ve verified and begun to implement a fix that will improve loading errors. We are continuing to roll this out to all regions and monitor for efficacy.

monitoringJun 19, 07:07 PM

We're actively monitoring this issue and working with our 3rd party provider. The next update will be sent on Monday unless there's new information to share.

monitoringJun 19, 02:33 AM

Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr We are continuing to work with our CSP on this investigation. Impacted users may receive intermittent error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in.

monitoringJun 18, 04:57 PM

Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr Impacted users may receive intermittent error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in. We continue to work with our CSP on this investigation.

investigatingJun 18, 03:18 PM

Due to the linked GCP outage below, users located in India may have trouble loading parts of Grafana. https://status.cloud.google.com/incidents/5fGQt4VbkDnr3Yp8PXPr Impacted users may receive error messages such as "Error Loading" or "Failed to load Assets". To be clear, it does not matter the region the stack is located, but the geography where the user is physically in. We are currently investigating this issue from our end, and will provide updates as they are available.

noneresolvedJun 20, 08:00 PM — Resolved Jun 20, 08:00 PM

Elevated Logs Query Usage — Unexpected Billing Impact

1 update
resolvedJun 22, 10:10 PM

Starting June 20, 2026 at approximately 20:00 UTC, some Grafana Cloud customers experienced unexpectedly elevated logs query usage. The issue persisted until it was resolved on June 22, 2026 at approximately 16:00 UTC. Our engineering team identified and mitigated the issue. Systems have since stabilized and are operating normally. Grafana Labs is reviewing affected accounts for appropriate remediation.

majorresolvedJun 19, 04:39 PM — Resolved Jun 19, 07:19 PM

Rule Evaluation Outage in prod-us-central-0

3 updates
resolvedJun 19, 07:19 PM

This incident has been resolved. Thank you for your patience.

monitoringJun 19, 05:46 PM

We’re continuing to track progress post-mitigation. While we don’t have new information to share yet, our team remains actively engaged.

monitoringJun 19, 04:39 PM

We had an outage affecting rule evaluations between 15:16-15:59 UTC in the prod-us-central-0 region. Our team quickly identified the issue and has since mitigated. The engineering team is monitoring.

majorresolvedJun 18, 05:01 PM — Resolved Jun 18, 06:50 PM

Issues with actions in the Grafana IRM mobile app

3 updates
resolvedJun 18, 06:50 PM

This incident has been resolved. Thank you for your patience.

monitoringJun 18, 06:09 PM

We've verified a fix in our staging environment to restore functionality to the mobile app. The fix is currently being deployed to production. Thanks for your patience as we continue to roll this out and monitor the resolution.

identifiedJun 18, 05:01 PM

We're noticing an uptick in users being unable to respond to actions on the mobile app (acknowledging and silencing alerts, for example). Users working in the web UI should not be affected. Ingestion and notification delivery are working as expected. We have a fix in place and are in the process of deploying.

criticalresolvedJun 18, 11:25 AM — Resolved Jun 18, 04:21 PM

Degraded k6 cloud UI performance

5 updates
resolvedJun 18, 04:21 PM

This incident has been resolved. Thank you for your patience.

monitoringJun 18, 02:04 PM

We are continuing to monitor for any further issues.

monitoringJun 18, 01:15 PM

The root cause of the issue has been identified and a fix has been successfully deployed. We are observing widespread improvements across all systems. Our team is currently monitoring the environment to ensure performance remains stable.

investigatingJun 18, 11:45 AM

We are continuing to investigate this issue.

investigatingJun 18, 11:25 AM

We’re currently investigating an issue resulting in degraded k6 cloud UI performance and API response time. Our team is actively working to rectify this issue.

majorresolvedJun 17, 08:17 PM — Resolved Jun 18, 02:08 PM

Loki data source-managed alert rules not visible in the Grafana Cloud Alerting UI

6 updates
resolvedJun 18, 02:08 PM

This incident has been resolved.

monitoringJun 18, 08:04 AM

A fix has been implemented and we are monitoring the results.

identifiedJun 17, 11:22 PM

We are continuing to deploy the fix and monitor recovery efforts. As part of the rollout, we identified an issue that required adjustments to our deployment plan, which has extended the timeline for mitigation. Work remains actively underway, and we will share additional updates as progress continues.

identifiedJun 17, 09:43 PM

Deployment of the fix is still in progress. We are continuing to monitor the rollout and validate recovery across affected systems. We will share further updates as they become available.

identifiedJun 17, 08:55 PM

Our Engineering Team has implemented a fix which is now being rolled out. We will continue to monitor the situation and update as soon as we have more information.

identifiedJun 17, 08:17 PM

We have identified an issue where alert rules and alerts managed directly in a Loki data source (data source-managed alerting) are not displayed in the Grafana Cloud Alerting UI. Rules created via Prometheus/Mimir data sources and Grafana-managed alert rules are not affected. Impact is limited to visibility and management in the UI. Affected alert rules continue to evaluate and send notifications normally — there is no impact to alert delivery. Workaround: Loki alert rules can still be viewed and managed directly through the Loki ruler API (for example, using cortextool against /loki/api/v1/rules). A fix has been identified and is in progress. We will provide a further update once it has been rolled out.

📡 Tired of checking Grafana Cloud status manually?

Better Stack monitors uptime every 30 seconds and alerts you instantly when Grafana Cloud goes down.

Start Free Monitoring →