G

Grafana Cloud Outage History

Past incidents and downtime events

Complete history of Grafana Cloud outages, incidents, and service disruptions. Showing 50 most recent incidents.

September 2026(15 incidents)

minorresolvedSep 15, 10:00 AM — Resolved Sep 15, 11:17 AM

Increased Execution Time for Browser Checks in Synthetic Monitoring

4 updates
resolvedSep 15, 11:17 AM

Between September 9 and September 15, 2026, some browser checks in Synthetic Monitoring experienced longer than usual execution times, with impact varying by location. This has been resolved and check execution times have returned to baseline. Thank you for your patience.

monitoringSep 15, 10:58 AM

A fix has been deployed and CPU throttling is returning to normal. Check execution times should return to baseline shortly. We will continue monitoring recovery. Next update by 11:56 UTC.

identifiedSep 15, 10:51 AM

A fix has been merged and is now being rolled out across affected locations. We are monitoring deployment progress. Next update by 11:49 UTC.

identifiedSep 15, 10:00 AM

Since September 9, 2026, some browser checks in Synthetic Monitoring may take longer to complete than usual, with impact varying by location. We have identified the cause and are actively working on a fix. Next update by 10:49 UTC.

minorresolvedSep 14, 03:30 PM — Resolved Sep 14, 03:30 PM

Intermittent Metric Write Errors in GCP US Central (prod-us-central-0)

1 update
resolvedSep 15, 07:49 AM

Between 15:31 and 20:43 UTC on September 14, 2026 (approximately 5 hours 12 minutes), some customers in GCP US Central (prod-us-central-0) may have experienced intermittent errors or delays when writing metrics. This has been fully resolved and all services are operating normally.

noneresolvedSep 14, 12:00 PM — Resolved Sep 14, 12:00 PM

CloudWatch request timeouts in prod-us-central-0

1 update
resolvedSep 14, 01:04 PM

Some customers experienced timeouts affecting CloudWatch traffic in prod-us-central-0 for approximately one hour. Service has since been restored. This incident did not affect all customers. We are publishing this update retrospectively and apologise for the disruption.

minorresolvedSep 14, 03:30 AM — Resolved Sep 14, 03:30 AM

Elevated Latency Managing Cloud Provider Integrations in GCP US Central

1 update
resolvedSep 17, 12:34 PM

Between 03:30 UTC on September 14, 2026 and 10:50 UTC on September 17, 2026 (approximately 3 days), customers managing cloud provider integrations in GCP US Central may have experienced elevated latency when creating, viewing, or removing integrations. This has been fully resolved and all services are operating normally.

minorresolvedSep 12, 11:00 PM — Resolved Sep 12, 11:00 PM

Failures Submitting Self-Serve Configuration Changes for Grafana Cloud Logs in AWS Sweden (prod-eu-north-0)

1 update
resolvedSep 14, 10:18 AM

Between approximately 23:00 UTC on September 12, 2026 and 09:12 UTC on September 14, 2026, some Grafana Cloud Logs customers in AWS Sweden (prod-eu-north-0) may have intermittently experienced errors when submitting configuration changes through the Cloud Portal. This has been resolved, and configuration change requests are being accepted normally.

minorresolvedSep 9, 11:27 AM — Resolved Sep 9, 01:51 PM

Investigating elevated database load in AWS Germany

3 updates
resolvedSep 9, 01:51 PM

This incident has been resolved.

monitoringSep 9, 12:44 PM

A fix has been implemented and we have been seeing recovery to the affected region. We will continue to monitor.

investigatingSep 9, 11:27 AM

We are investigating elevated database load affecting infrastructure supporting a subset of Grafana Cloud instances in the AWS Germany region (prod-eu-west-4). Affected users may experience slow responses when accessing Grafana. We will provide an update as more information becomes available.

minorresolvedSep 8, 06:05 PM — Resolved Sep 8, 10:41 PM

Issues with Geomap Tiles

2 updates
resolvedSep 8, 10:41 PM

This incident has been resolved.

identifiedSep 8, 06:05 PM

Some geomap tiles are displaying the “api key required". The issue has been identified, and our team is working on a fix. As a workaround, you can open the Geomap panel, navigate to the Map layers section, and change the Basemap layer from CARTO, or "Default base layer," to OpenStreetMap. The watermark should disappear immediately. Alternatively, if you would prefer to retain a dark-styled map similar to the current CARTO appearance, you can select MapLibre as the basemap and configure the URL as: https://tiles.openfreemap.org/styles/dark This option also renders without the watermark, does not require an API key, and provides a visual style similar to the dark CARTO basemap.

minorresolvedSep 8, 04:14 PM — Resolved Sep 8, 08:34 PM

High Latency in prod-us-east-2

3 updates
resolvedSep 8, 08:34 PM

This incident has been resolved.

monitoringSep 8, 05:04 PM

We have observed latency return to normal levels, and we will continue to monitor this incident.

investigatingSep 8, 04:14 PM

We are currently investigating an issue causing high latency for deployments in the prod-us-east-2 cluster.

minorresolvedSep 4, 04:45 PM — Resolved Sep 4, 11:00 PM

US Central Region Instability

3 updates
resolvedSep 4, 11:00 PM

The US Central region is fully recovered now and the issue is now resolved.

monitoringSep 4, 06:58 PM

The underlying cloud provider issue is not yet fully resolved, but our systems have recovered and affected Grafana Cloud stacks are loading normally again. Some customers may still see brief intermittent slowness. We are continuing to monitor until our provider confirms resolution.

identifiedSep 4, 04:45 PM

We’re currently investigating an issue affecting the us central region due to ongoing cloud provider incident. Multiple components impacted. Our team is actively working on this. Thank you for your patience.

majorresolvedSep 4, 05:03 PM — Resolved Sep 4, 08:42 PM

Degradation of Hosted Grafana in US Central Region

3 updates
resolvedSep 4, 08:42 PM

The US Central region is now fully operational and this incident is now resolved.

monitoringSep 4, 06:24 PM

We are currently observing minimal impact and continuing to monitor the situation.

identifiedSep 4, 05:03 PM

We are aware of an issue causing some Grafana Cloud stacks in the US region to be unavailable or fail to load. The cause is reduced compute capacity in one of our cloud provider's availability zones, which is preventing affected stacks from starting. Other regions are unaffected. We are working with our cloud provider to restore capacity.

minorresolvedSep 4, 01:36 PM — Resolved Sep 4, 04:32 PM

Alert rule creation, deletion, and update degradation in prod-us-east-2

2 updates
resolvedSep 4, 04:32 PM

The issue has been mitigated and there's a code fix pending.

investigatingSep 4, 01:36 PM

We’re currently investigating an issue affecting alert rule creation, deletion, and updates in prod-us-east-2. Our team is actively working to identify the cause. Thank you for your patience.

minorresolvedSep 3, 10:13 PM — Resolved Sep 4, 12:55 AM

Tempo and Mimir Read and Write Failures

4 updates
resolvedSep 4, 12:55 AM

This incident has been resolved.

monitoringSep 3, 10:48 PM

Tempo recovered.

identifiedSep 3, 10:31 PM

Cluster Metrics are healthy for the last 15 minutes

identifiedSep 3, 10:13 PM

The issue has been identified and a fix is being implemented.

majorresolvedSep 3, 09:32 AM — Resolved Sep 3, 12:49 PM

High Latency in prod-ap-south-1

3 updates
resolvedSep 3, 12:49 PM

This incident has been resolved.

monitoringSep 3, 11:37 AM

A fix has been applied and we are seeing signs of recovery. We will continue to monitor the recovery.

investigatingSep 3, 09:32 AM

We are currently investigating an issue impacting deployments in prod-ap-south-1. Impacted customers may be experiencing increased latency and issues with ingestion.

noneresolvedSep 2, 02:14 PM — Resolved Sep 2, 02:14 PM

Elevated Latency in Grafana Cloud Logs (prod-ap-south-1)

1 update
resolvedSep 2, 02:14 PM

We observed elevated latency in Grafana Cloud Logs between 08:00 and 11:00 UTC for a subset of deployments in prod-ap-south-1. The issue is now resolved.

minorresolvedSep 1, 03:25 PM — Resolved Sep 1, 08:20 PM

Investigating issues in US Central (prod-us-central-0, prod-us-central-5)

3 updates
resolvedSep 1, 08:20 PM

This incident has been resolved.

identifiedSep 1, 04:14 PM

We have identified the cause as an issue in our cloud provider’s US Central region. A mitigation is in place and recovery is in progress. Customers in prod-us-central-0 and prod-us-central-5 may still see errors or delays with metrics, logs, or Grafana while services come back. We will update this page as recovery continues.

investigatingSep 1, 03:25 PM

We are investigating reports of errors affecting Grafana Cloud in US Central (prod-us-central-0, prod-us-central-5). Customers in this region may see timeouts or failed requests when writing or querying logs and metrics. We are looking into this and will post another update as soon as we have more information.

August 2026(14 incidents)

majorresolvedAug 28, 05:12 PM — Resolved Aug 28, 06:42 PM

Partial Logs Write Outage

3 updates
resolvedAug 28, 06:42 PM

This incident has been resolved.

monitoringAug 28, 05:55 PM

This incident also had impact on Frontend Observability for the same region. A fix has been deployed, and things have stablized. We will continue to monitor on our end.

identifiedAug 28, 05:12 PM

We’ve identified the cause of an issue impacting Log writes. Our team is currently implementing a fix.

minorresolvedAug 28, 12:58 AM — Resolved Aug 28, 02:40 AM

Some Grafana UI features may be unavailable or reverting to legacy behaviour

3 updates
resolvedAug 28, 02:40 AM

This incident has been resolved.

identifiedAug 28, 01:08 AM

a resolution is rolling out to resolve the issues

identifiedAug 28, 12:58 AM

Some Grafana UI features may be unavailable or reverting to legacy behaviour

minorresolvedAug 27, 12:09 PM — Resolved Aug 27, 01:06 PM

Elevated error rates affecting metrics writes in prod-us-central-0

3 updates
resolvedAug 27, 01:06 PM

This incident has been resolved.

monitoringAug 27, 12:37 PM

A fix has been applied and metrics write error rates in prod-us-central-0 have returned to normal levels. Metrics ingestion is operating as expected. We are continuing to monitor the affected systems to confirm full recovery before marking this incident resolved.

investigatingAug 27, 12:09 PM

We are investigating elevated error rates affecting metrics writes in prod-us-central-0. Customers in this region may see failed or delayed metrics ingestion, gaps in recent data on dashboards and queries, and alert rules that depend on the affected series may not evaluate as expected.

noneresolvedAug 27, 01:54 AM — Resolved Aug 27, 01:54 AM

Mimir Writes Incident in prod-us-central-0

1 update
resolvedAug 27, 01:54 AM

Mimir writes in the prod-us-central-0 region had elevated error rates for approximately 15 minutes from 1:05 to 1:20 UTC. The issue has been resolved and we are monitoring.

minorresolvedAug 25, 10:39 AM — Resolved Aug 25, 11:02 AM

Incident Management unavailable in US Central

4 updates
resolvedAug 25, 11:02 AM

This incident has been resolved. Incident Management in US Central is operating as normal. Customers can create, view, and query Incidents, and the Incident public API is fully available. Grafana OnCall was not affected at any point during this incident.

monitoringAug 25, 11:01 AM

We are continuing to monitor for any further issues.

monitoringAug 25, 10:49 AM

A fix has been applied and Incident Management in US Central is recovering. The Incident public API is responding normally, and customers can create, view, and query Incidents again. We are continuing to monitor the recovery to confirm the service is fully stable.

investigatingAug 25, 10:39 AM

We are investigating an outage affecting Grafana Incident in the US Central region (prod-us-central-0). Customers are unable to create, view, or query Incidents, and requests to the Incident public API are timing out. All customers in this region are affected. Grafana OnCall is not impacted — alert ingestion, alert group creation, and notification delivery are all operating normally.

criticalresolvedAug 18, 10:00 PM — Resolved Aug 18, 10:00 PM

Cloud Logs read path outage on eu-west-2

1 update
resolvedAug 18, 10:53 PM

We observed an issue which impacted the following service: reads of Cloud Logs in the eu-west-2 region. Affected services would include Explore, dashboards, alert evaluation when querying Logs. The time of impact lasted from approximately 22h15 and 22h23 UTC, on loki-prod-012. This incident has since been resolved.

criticalresolvedAug 13, 03:25 PM — Resolved Aug 13, 04:47 PM

K6 Test Outage

2 updates
resolvedAug 13, 04:47 PM

This incident has been resolved.

monitoringAug 13, 03:25 PM

Telemetry push APIs went down at 15:06 UTC, causing tests to lose telemetry data, like metrics, logs, traces, and browser screenshots. This issue has now been mitigated, and we are monitoring to determine root cause as well as full resolution.

minorresolvedAug 10, 01:30 PM — Resolved Aug 10, 01:30 PM

Metrics: Elevated Error Rates Reads/Writes

1 update
resolvedAug 10, 01:30 PM

Between 12:20 and 13:15 UTC, we experienced elevated error rates for metrics reads and writes due to a networking issue. Impact was limited to a subset of deployments in the prod-us-central-0 region. A fix has been applied and error rates have since recovered.

minorresolvedAug 8, 07:48 PM — Resolved Aug 8, 07:48 PM

Alerting expressions pipeline failing when recovery settings

1 update
resolvedAug 8, 07:48 PM

Due to a software bug the evaluation of some alert rules (primarily ones that have recovery threshold setting) were failing to be evaluated starting 14:00 UTC to 19:30 UTC today.

noneresolvedAug 5, 05:57 PM — Resolved Aug 5, 05:57 PM

Some Cloud Test Runs Terminated

1 update
resolvedAug 5, 05:57 PM

Between approximately 13:40 and 15:05 UTC today, a subset of cloud test runs were unexpectedly terminated and marked Aborted (by system). This affected runs that were in progress during three short windows in that period; test results and metrics for completed runs were not impacted. We have identified the cause and no further occurrences have been observed since 15:05 UTC. A fix is in place. The platform is currently operating normally and test runs are executing as expected

minorresolvedAug 5, 11:41 AM — Resolved Aug 5, 11:44 AM

o11y requests too large for nods in the prod-eu-west-2 region.

2 updates
resolvedAug 5, 11:44 AM

This incident has been resolved.

investigatingAug 5, 11:41 AM

We are currently investigating an issue with critical o11y requests being too large for nods in the prod-eu-west-2 region. This is leading to issues in scaling up services, and may may lead increased errors in requests.

criticalresolvedAug 4, 07:11 PM — Resolved Aug 5, 01:32 AM

Some Grafana Instances Unavailable

5 updates
resolvedAug 5, 01:32 AM

We are observing a continued period of stability and at this point, we are marking the incident resolved.

monitoringAug 5, 12:54 AM

We have rolled out a change to the affected regions to restore services.

investigatingAug 4, 10:58 PM

We have identified the cause: our cloud provider has exhausted compute capacity in the affected region. We are working directly with them and are moving affected workloads to alternative capacity to restore service.

investigatingAug 4, 08:12 PM

We are continuing to investigate this issue.

investigatingAug 4, 07:11 PM

We are currently investigating an issue that causing some Grafana deployments in prod-us-central-0 and prod-us-central-3 to become unavailable.

minorresolvedAug 4, 10:34 AM — Resolved Aug 4, 01:49 PM

Logs latency increase within prod-eu-west-3

3 updates
resolvedAug 4, 01:49 PM

This incident has been resolved.

investigatingAug 4, 12:46 PM

There's still a minor latency increase in this cell however we have identified that it is improving although not back to normal as of yet. We will continue to look into this.

investigatingAug 4, 10:34 AM

We have identified minor latency increases in write endpoints within the prod-eu-west-3. Our team is currently monitoring and investigating this.

minorresolvedAug 1, 07:25 AM — Resolved Aug 1, 08:31 AM

K6 - Cloud test-run issues

3 updates
resolvedAug 1, 08:31 AM

This incident has been resolved.

monitoringAug 1, 08:06 AM

A critical piece of infrastructure behaved poorly after a reboot. It has now been properly recovered

investigatingAug 1, 07:25 AM

We are currently investigating an issue which is resulting in some test runs to abort and metrics data to be incomplete for a subset of tests

July 2026(21 incidents)

noneresolvedJul 31, 07:30 PM — Resolved Jul 31, 07:30 PM

Partial Read Outage for Loki in prod-us-east-4

1 update
resolvedJul 31, 08:33 PM

We are investigating a partial read outage affecting Loki in prod-us-east-4. Between 19:41 UTC and 19:50 UTC, a significant portion of read queries may have failed or returned errors. The issue has been identified and service has been restored. We are continuing to investigate the underlying cause and will provide additional information as it becomes available.

minorresolvedJul 31, 10:51 AM — Resolved Jul 31, 12:13 PM

Degraded Performance: Stack Provisioning Failures within certain reigons (PDC Setup)

2 updates
resolvedJul 31, 12:13 PM

He have identified the cause and applied a fix for this issue and all effected stacks within the affected regions are not working as expected.

investigatingJul 31, 10:51 AM

We are currently investigating an issue impacting stack provisioning. Attempting to set up PDC on newly created stacks will currently fail across several regions. We are currently looking into the cause of this.

majorresolvedJul 30, 02:45 PM — Resolved Jul 30, 05:40 PM

PDC Authentication Issues

5 updates
resolvedJul 30, 05:40 PM

This incident has been resolved.

monitoringJul 30, 03:56 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 30, 03:11 PM

We are continuing to investigate this issue.

investigatingJul 30, 02:51 PM

We are continuing to investigate this issue.

investigatingJul 30, 02:45 PM

We are currently investigating an issue impacting PDC authentication. We will provide additional updates as they become available.

majorresolvedJul 30, 01:15 PM — Resolved Jul 30, 03:17 PM

Issues with Billing/Usage Dashboard Metrics and Panels.

4 updates
resolvedJul 30, 03:17 PM

This incident has been resolved.

monitoringJul 30, 02:34 PM

A fix has been implemented and we are monitoring the results.

identifiedJul 30, 01:50 PM

We have identified the issue, and are working on deploying a fix.

investigatingJul 30, 01:15 PM

We are investigating an outage for the Billing / Usage dashboard metrics and panels. This appears to be partially affecting organizations using Grafana Cloud. We are working on identifying and resolving the issue

minorresolvedJul 29, 09:46 PM — Resolved Jul 30, 12:39 AM

IRM Performance Degradation in EU Region

3 updates
resolvedJul 30, 12:39 AM

The issue affecting Grafana IRM in the EU region has been resolved. The IRM UI, public API, and alert notification processing have been restored and are operating normally.

identifiedJul 29, 11:07 PM

We have identified the underlying issue affecting Grafana IRM in our EU region and are actively working to restore normal service. Users may continue to experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notifications while mitigation efforts are underway. Our engineering team is working to restore full functionality as quickly as possible. We will provide another update as more information becomes available.

investigatingJul 29, 09:46 PM

We are currently investigating an issue affecting Grafana IRM in our EU region. Users may experience a degraded or unresponsive IRM UI, reduced availability of the public API, and delays in notifications. Our engineering team is actively investigating the issue and working to restore normal service. We will provide another update as more information becomes available.

minorresolvedJul 29, 06:49 PM — Resolved Jul 29, 06:49 PM

Grafana Cloud non-billing usage metrics gaps in select regions

1 update
resolvedJul 29, 06:49 PM

Impact period: April 1 – July 29, 2026 Summary: During this period, some customers in a subset of regions may have experienced gaps in select ruler/recording rule metrics within their Grafana Cloud instances. Not all customers or environments in the listed regions were affected. Affected regions: A subset of environments across the following regions may have been impacted: ap-south-1 au-southeast-1 ca-east-0 eu-central-0 eu-west-2 eu-west-7 us-west-0 Affected values: grafanacloud_instance_queries_per_second grafanacloud_instance_rule_config_last_reload_successful grafanacloud_instance_rule_evaluations_total:rate5m grafanacloud_instance_rule_evaluation_failures_total:rate5m grafanacloud_instance_rule_group_interval_seconds grafanacloud_instance_rule_group_last_duration_seconds grafanacloud_instance_rule_group_iterations_total:rate5m grafanacloud_instance_rule_group_iterations_missed_total:rate5m grafanacloud_instance_rule_group_last_evaluation_timestamp_seconds grafanacloud_instance_rule_group_rules grafanacloud_instance_ruler_queries_failed_total:rate5m grafanacloud_instance_ruler_queries_zero_fetched_series_total:rate5m grafanacloud_instance_ruler_notifications_sent_total:rate5m grafanacloud_instance_ruler_notifications_errors_total:rate5m grafanacloud_instance_ruler_notifications_queue_capacity grafanacloud_instance_ruler_notifications_queue_length grafanacloud_instance_ruler_notifications_latency_seconds:99quantile grafanacloud_instance_ruler_notifications_latency_seconds:50quantile Current status: This issue has been resolved. No further action is required from customers. If you continue to notice gaps in the metrics described above, please reach out to support and reference this incident.

majorresolvedJul 28, 07:40 PM — Resolved Jul 28, 07:43 PM

Partial OTLP Write Outage in prod-us-east-3

2 updates
resolvedJul 28, 07:43 PM

The issue affecting OTLP ingestion in the prod-us-east-3 region has been resolved. Between 13:30 UTC and 17:45 UTC, some customers experienced intermittent failures when writing telemetry to the OTLP endpoint. Logs were confirmed to be affected, and metrics and traces may also have experienced intermittent ingestion failures. Our investigation has concluded, and the affected services have recovered.

investigatingJul 28, 07:40 PM

We are investigating an issue affecting OTLP ingestion in the prod-us-east-3 region. Customers may experience intermittent failures when writing telemetry to the OTLP endpoint. Based on current information, logs are confirmed to be affected, and metrics and traces may also be impacted. Our investigation indicates intermittent failures began around 08:00 UTC, with the primary period of impact occurring between 15:30 UTC and 17:45 UTC. We are continuing to investigate the scope and root cause of this issue and will provide additional updates as more information becomes available.

criticalresolvedJul 24, 05:02 PM — Resolved Jul 24, 06:20 PM

Write Outage

3 updates
resolvedJul 24, 06:20 PM

This incident has been resolved.

monitoringJul 24, 05:18 PM

Healthy as of 17:00 UTC. We are continuing to monitor.

investigatingJul 24, 05:02 PM

We are currently investigating a Write outage ongoing, starting at 16:30 UTC for the impacted region.

majorresolvedJul 24, 09:17 AM — Resolved Jul 24, 02:03 PM

Errors creating new Slack integration for Grafana IRM

4 updates
resolvedJul 24, 02:03 PM

This incident has been resolved.

identifiedJul 24, 10:12 AM

We have verified the fix, and we are starting to roll it out

identifiedJul 24, 09:49 AM

We have identified the root cause of the problem, and we are working on a fix

investigatingJul 24, 09:17 AM

We are investigating a possible error creating Slack integrations. Will update the status soon.

noneresolvedJul 23, 11:00 PM — Resolved Jul 24, 05:32 AM

Intermittent write outage from 21:20-21:26, 21:54-21:56, and 22:22-22:23

2 updates
resolvedJul 24, 05:32 AM

Writes are looking stable in the last 6h

monitoringJul 23, 11:00 PM

status is health in prod-us-east-3.loki-prod-042 and we're continuing to monitor

majorresolvedJul 23, 05:14 PM — Resolved Jul 23, 07:47 PM

Metrics Write Path Errors

3 updates
resolvedJul 23, 07:47 PM

This incident has been resolved.

monitoringJul 23, 05:37 PM

Error rates have dropped, and we are monitoring this issue for any recurrence.

investigatingJul 23, 05:14 PM

We are currently investigating an elevated rate of error and latency in the impacted write path.

majorresolvedJul 23, 10:13 AM — Resolved Jul 23, 03:37 PM

Stacks using SCIM user provisioning are currently unable to log into Grafana

4 updates
resolvedJul 23, 03:37 PM

This incident has been resolved.

identifiedJul 23, 01:23 PM

We have applied a fix and are currently awaiting feedback from affected customers.

identifiedJul 23, 11:26 AM

The issue has been identified and we are currently working on the mitigations now. The scope is much more limited than initially thought only specific SCIM configurations are impacted.

investigatingJul 23, 10:13 AM

We are currently facing an issue where Stacks using SCIM user provisioning are currently unable to log into Grafana. We are currently investigating this issue and working on a fix.

minorresolvedJul 22, 10:27 AM — Resolved Jul 22, 11:25 AM

K6 - Cloud output test-runs are failing to fetch the script logs

3 updates
resolvedJul 22, 11:25 AM

This incident has been resolved.

monitoringJul 22, 10:52 AM

We have identified the cause and applied a fix which has has resulted in significant recovery. We will continue to monitor this before resolving.

investigatingJul 22, 10:27 AM

We are currently investigating an issue cloud output test-runs which is resulting in failure to fetch the script logs. We are working on identifying the cause and working on a fix.

majorresolvedJul 21, 06:43 PM — Resolved Jul 21, 09:06 PM

Cloud Log Exporter Unavailable

3 updates
resolvedJul 21, 09:06 PM

This incident has been resolved.

monitoringJul 21, 08:04 PM

A fix has been implemented and we are monitoring the results.

investigatingJul 21, 06:43 PM

Cloud Log Exporter is temporarily not available in the marked regions. We are currently investigating this issue and will update as soon as we have more info to share.

majorresolvedJul 20, 01:48 PM — Resolved Jul 20, 03:03 PM

PDC Issues

4 updates
resolvedJul 20, 03:03 PM

This incident has been resolved.

monitoringJul 20, 02:23 PM

A fix has been implemented, and we are observing recovery. We will continue to monitor the results.

investigatingJul 20, 02:13 PM

We are continuing to investigate this issue.

investigatingJul 20, 01:48 PM

We are currently investigating an issue that is causing issues with PDC in the prod-eu-west-2 region. We will provide another update in 1-2 hours.

minorresolvedJul 14, 08:10 AM — Resolved Jul 20, 10:15 AM

Issue with Dashboard Views Being Registered

8 updates
resolvedJul 20, 10:15 AM

A fix has been implemented and dashboard view and error counts in the Dashboards and Folder list are updating as expected. Thank you for your patience while we worked to address this issue.

identifiedJul 20, 08:17 AM

We are currently working on a fix related to this incident. There are no new updates to share at this time.

identifiedJul 16, 09:27 PM

We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.

identifiedJul 16, 01:34 PM

We continue to work on resolving the issue impacting dashboard view and error counts. There are no new updates to share at this time. Our next update will be provided within 24 hours.

identifiedJul 15, 08:52 PM

We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.

identifiedJul 15, 12:05 PM

We continue to work on resolving this issue. We'll provide another update within the next 24 hours, or sooner if additional information becomes available.

identifiedJul 14, 02:27 PM

We've identified the cause of the issue impacting dashboard view and error counts in the Dashboards and Folder list. Our team is currently working on implementing a fix and validating the solution. We will provide another update within the next 24 hours, or sooner if we have additional information to share.

investigatingJul 14, 08:10 AM

This is related to https://status.grafana.com/incidents/rhrk2ck6ly0y which was resolved by mistake. For some dashboards, the Views/Error counts in the Dashboards/Folder list renders as `-` and never updates, even after dashboards are viewed repeatedly. This is ultimately causing inaccurate or missing view counts. We are continuing to investigate the underlying cause of this issue. At this time, the scope and impact remain unchanged, and we have no new information to share. We will provide another update as soon as more information becomes available.

minorresolvedJul 17, 06:22 PM — Resolved Jul 18, 01:15 PM

Grafana Cloud login issues for users with role of None

4 updates
resolvedJul 18, 01:15 PM

Users with the None role authenticating via Grafana.com should be able to access their stacks again

identifiedJul 17, 09:33 PM

We are in the process of implementing a fix and are monitoring the rollout. Thank you for your patience.

identifiedJul 17, 07:28 PM

We’ve identified the cause of the issue impacting users with the role of None from logging in. Our team is currently implementing a fix.

investigatingJul 17, 06:22 PM

We’re currently investigating an issue with user logins with the None role in Grafana Cloud. Our team is actively working to identify the cause. Thank you for your patience.

minorresolvedJul 16, 12:56 PM — Resolved Jul 16, 02:58 PM

Adaptive Metrics aggregation delay in eu-west-0 region.

3 updates
resolvedJul 16, 02:58 PM

The incident is now fully resolved.

monitoringJul 16, 01:45 PM

Services are fully recovered now. We're monitoring to be sure the issue won't re-occur.

identifiedJul 16, 12:56 PM

We're currently facing an issue with Adaptive Metrics aggregation delay in the eu-west-0 region (GCP Belgium). The issue started at around 11:50 UTC, but we were able to identify the issue and the appropriate fix is already deployed. We're seeing services recovering. More updated to come soon.

minorresolvedJul 15, 07:19 PM — Resolved Jul 16, 03:54 AM

Delayed Aggregated Metrics (prod-us-central-0)

3 updates
resolvedJul 16, 03:54 AM

This incident has been resolved.

monitoringJul 15, 09:23 PM

We have applied the mitigation and are monitoring as the backlog catches up. We will post another update once that is complete. Thank you for your patience.

identifiedJul 15, 07:19 PM

We are currently investigating a delay in aggregated metric results for tenants using Adaptive Metrics in the prod-us-central-0 region. Our engineering team has identified the issue and is actively working on a mitigation. We will provide further updates as the investigation progresses.

majorresolvedJul 15, 10:24 PM — Resolved Jul 15, 11:36 PM

Partial Outage in prod-eu-west-2

5 updates
resolvedJul 15, 11:36 PM

This incident has been resolved. Thank you for your patience.

monitoringJul 15, 10:53 PM

We’ve implemented a fix and are monitoring the results to confirm the issue is fully resolved. Services may start to recover during this time.

investigatingJul 15, 10:38 PM

We are continuing to investigate this issue.

investigatingJul 15, 10:35 PM

We are investigating a broader issue affecting multiple Grafana Cloud products in prod-eu-west-2. This appears to be caused by a third-party provider issue rather than a Loki-specific problem. Impact is currently inconsistent: some components are affected while others continue to function normally. We have also seen some impact to Tempo write paths. We are continuing to investigate and will share another update as soon as we have more information.

investigatingJul 15, 10:24 PM

We are investigating an issue affecting Loki queries in prod-eu-west-2. We first observed this behavior at approximately 21:57 UTC. Affected users may see elevated query errors, timeouts, or intermittent failures when running Loki queries in this region. The issue appears to be improving, but it is not fully resolved yet. We are continuing to investigate and will share another update as soon as we have more information.

majorresolvedJul 15, 06:07 PM — Resolved Jul 15, 09:00 PM

Some Reports of Grafana Not Loading.

4 updates
resolvedJul 15, 09:00 PM

This incident has been resolved. Thank you for your patience.

identifiedJul 15, 08:50 PM

The fix is in the process of being rolled out.

identifiedJul 15, 06:47 PM

We believe we have found the cause and are working on remediation. It is also worth mentioning that only stacks on the "slow" release channel are impacted.

investigatingJul 15, 06:07 PM

We are currently investigating an issue impacting a small subset of stacks in the impacted regions. We will provide more details as they become available.

📡 Tired of checking Grafana Cloud status manually?

Better Stack monitors uptime every 30 seconds and alerts you instantly when Grafana Cloud goes down.

Start Free Monitoring →