Groq Outage History
Past incidents and downtime events
Complete history of Groq outages, incidents, and service disruptions. Showing 25 most recent incidents.
July 2026(1 incident)
Data Center Failure Impacting Capacity
5 updates
The **capacity issue has been resolved**. All services are operating at normal capacity. The incident was caused by a **power loss issue that led to a cooling system failure at one of our US Central data centers**. We apologize for any disruption and **appreciate your patience**.
We have **redistributed and allocated more capacity to production models** in **the US Central region.** Users should **see performance return to normal**. We’re now **monitoring** the infrastructure to ensure stability.
A power loss issue at approximately 22:20pm UTC led to a subsequent **cooling system failure** at one of our **US Central** data centers that is causing **reduced capacity**, primarily for the `openai/gpt-oss-20b` model. The team is **working on restoring capacity**.
We have **identified a cooling system failure** at one of our **US Central** data centers that is causing **reduced capacity**. The team is **working on restoring capacity**. Users **continue to see elevated latencies for specific models**.
We are currently investigating a potential issue at one of our US Central data centers that is impacting system capacity. Users **may experience higher latencies** due to reduced resources. Our engineering team is working to identify the root cause and remediate the problem.
March 2026(1 incident)
openai/gpt-oss-120b Performance Issue
3 updates
The issues affecting openai/gpt-oss-120b have been **resolved**. The model is operating normally. Actions were taken to cancel billing plans and restrict verification status of organizations engaged in coordinated abuse. We apologize for the disruption and thank you for your patience.
We have implemented a fix for openai/gpt-oss-120b and performance is improving. We’re now monitoring the model to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have identified the problem affecting the openai/gpt-oss-120b model. The issue has been traced to malicious traffic patterns. A fix is in progress, we are working to block the source of these requests. Users may still see degraded performance from requests made to openai/gpt-oss-120b as we identify and block the orgs from which this traffic is originating.
February 2026(3 incidents)
meta-llama/llama-4-scout-17b-16e-instruct Degraded Performance
5 updates
This incident has been resolved. Thank you for your patience.
We have implemented a new fix for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound and performance has improved. We’re continuing monitor the models to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have implemented a fix for meta-llama/llama-4-scout-17b-16e-instruct however performance is still degraded. We’re continuing to investigate the models to determine the cause of the issues.
We are continuing to investigating an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing elevated error rates or slow response time. Other models remain operational. Our team is continuing to analyze logs and infrastructure to identify the cause.
We are currently investigating reports of an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
meta-llama/llama-4-scout-17b-16e-instruct Degraded Performance
3 updates
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct have been **resolved**. The model is once again operating normally. We apologize for the disruption and thank you for your patience.
meta-llama/llama-4-scout-17b-16e-instruct and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are currently investigating reports of an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
meta-llama/llama-4-scout-17b-16e-instruct Degraded Performance
5 updates
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct have been fully **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We’re continuing to **monitor** the scout model to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have **implemented a fix** for meta-llama/llama-4-scout-17b-16e-instruct and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We have **identified the problem** affecting the meta-llama/llama-4-scout-17b-16e-instruct. A fix is in progress, users may still see delayed responses and errors from Scout until the fix completes.
We are currently investigating reports of an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
January 2026(2 incidents)
meta-llama/llama-4-scout-17b-16e-instruct Degraded Performance
4 updates
This incident has been resolved. Thank you for your patience.
We have implemented a new fix for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound and performance has improved. We’re continuing monitor the models to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have implemented a fix for meta-llama/llama-4-scout-17b-16e-instruct however performance is still degraded. We’re continuing to investigate the models to determine the cause of the issues.
We are currently investigating an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing elevated error rates or slow response time. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Data Center Failure Impacting Model Latency - SYD
5 updates
This incident is now resolved.
This issue is now resolved. All services are operating normally. We will continue to monitor for stability and resolve this incident shortly. Thank you for your patience.
The issue impacting our SYD data center has been fixed. Users should gradually see performance return to normal as we continue to recover all models. We will continue to monitor.
The team is still working to resolve the issue at our SYD Data Center. Customers in the area may still be experiencing some model latency. We will provide further updates at they are available.
We have **identified an issue in our Sydney, AUS data center** that may be causing some latency for customers in this region. The issue was traced to **a network issue**. Users may still see latency until the fix completes.
December 2025(4 incidents)
Model Performance or Availability Issue: llama-3.3-70b-versatile
3 updates
The issues affecting **llama-3.3-70b-versatile** have been **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for **llama-3.3-70b-versatile** and performance is improving. The team is continuing to work the issue and we hope to have it resolved shortly.
We’ve identified an issue causing 503s for the `llama-3.3-70b-versatile` production model. We’ve started mitigation. Error rates are improving, but some requests may still fail while we continue tuning and monitoring.
Data Center Failure Impacting Capacity - DMM1
4 updates
**The data center capacity issue has been resolved. All services are operating at normal capacity. The incident was caused by a failed ARP entry on the network gateway device caused the Salam network link to go down in the DMM1 data center, which has been addressed. We apologize for the disruption and appreciate your patience.**
We have restored capacity in DMM 1 and services are beginning to recover. Users should gradually see network performance return to normal. We’re now monitoring the infrastructure to ensure stability.
We have identified a failure in our DMM1 infrastructure that is causing reduced capacity and network latency. The team has isolated the cause to the Salam link confirmed down and the Mobily link operating at full capacity, resulting in significant packet loss and is working on restoring capacity. Users continue to see Network latency.
We are currently investigating a potential network issue at our DMM1 that is impacting system capacity. Users may experience slower response times or intermittent failures due to reduced resources. Our engineering team is working to identify the root cause and remediate the problem.
meta-llama/llama-4-scout-17b-16e-instruct Degraded Performance
5 updates
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct , meta-llama/llama-4-maverick-17b-128e-instruct, & groq/compound have been **resolved**. The models are operating normally. We apologize for the disruption and thank you for your patience.
We have implemented a new fix for meta-llama/llama-4-scout-17b-16e-instruct and performance and have observed improvement over the last 15 minutes. We’re continuing monitor to ensure stability persists. If all remains normal, we will resolve the incident in the next update. In addition to meta-llama/llama-4-scout-17b-16e-instruct, we realized there also was some impact to groq/compound, which has also recovered.
We are no longer seeing issues on meta-llama/llama-4-maverick-17b-128e-instruct. However, the team is still troubleshooting errors and latency issues with meta-llama/llama-4-scout-17b-16e-instruct.
We are now seeing similar degradation on meta-llama/llama-4-maverick-17b-128e-instruct. The team is still investigating the issue and working to resolve.
We are currently investigating reports of an issue with our meta-llama/llama-4-scout-17b-16e-instruct & groq/compound models. Users may be experiencing elevated error rates and or slow responses. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
meta-llama/llama-4-scout-17b-16e-instruct & groq/compound Degraded Performance
5 updates
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct & groq/compound have been **resolved**. Both of the models are operating normally. We apologize for the disruption and thank you for your patience.
We have **implemented a new fix** for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound and performance has improved. We’re continuing **monitor** the models to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have implemented a fix for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound however performance is still degraded. We’re continuing to investigate the models to determine the cause of the issues.
We have **implemented a fix** for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound and performance is improving. We’re now **monitoring** the models to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are currently investigating reports of an issue with our meta-llama/llama-4-scout-17b-16e-instruct & groq/compound models. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
November 2025(3 incidents)
Console, Website, & API Availability Issues: Global Cloudflare Outage
8 updates
We are resolving this incident as we have now observed ~30 minutes of normalized service availability and customer traffic. We will continue to closely monitor our systems for any signs of regression as [Cloudflare's statuspage incident remains active ](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7)and in monitoring.
We are continuing to monitor an issue with our services that is impacting customer experiences due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage"). Cloudflare has just shared that "A fix has been implemented and we believe the incident is now resolved. We are continuing to monitor for errors to ensure all services are back to normal.". We are monitoring to confirm the restoration of service availability and successful requests and are continuing to monitor for signs of persistent recovery as Cloudflare now claims to have implemented a fix.
We are continuing to monitor an issue with our services that is impacting customer experiences due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage"). Cloudflare previously shared that "We are continuing to work towards restoring other services" and "The issue has been identified and a fix is being implemented". We are continuing to observe higher-than-normal error rates and are continuing to monitor for signs of recovery as Cloudflare works to implement a fix.
We are continuing to monitor an issue with our services that is impacting customer experiences due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage"). Cloudflare recently shared that "We are continuing to work towards restoring other services" and "The issue has been identified and a fix is being implemented". We are continuing to observe higher-than-normal error rates and are monitoring for signs of recovery as Cloudflare works to implement a fix.
We are continuing to monitor an issue with our services that is impacting customer experiences due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage"). Cloudflare recently shared that "The are continuing to investigate the issue". We are continuing to observe higher-than-normal error rates despite Cloudflare's previous update stating that they were seeing services recover.
We are continuing to monitor an issue with our services that is impacting customer experiences. Our team is has identified that is due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage") and they've recently shared that "We are seeing services recover, but customers may continue to observe higher-than-normal error rates as we continue remediation efforts". This aligns with recent improvements observed in our request traffic but higher-than-normal error rates still persist.
We are currently monitoring an issue with our services that is impacting customer experiences. Our team is has identified that is due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage") and is working to remediate the problem and will follow up with more details as they become available from Cloudflare.
We are currently investigating an issue with our services that we believe is impacting customer experiences. Our team is has identified that is due to a [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage") and is working to remediate the problem and will follow up shortly with more details.
Cloud API Degradation
4 updates
The issues causing degraded performance have been **resolved**. All models are now operating normally. Upon investigation we learned the issue was scoped to an internal feature yet to be released that is still in development. At this time we don't believe there was any customer impact as a direct result of this issue. We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for the performance degradation of our Cloud API and performance is improving. We’re now **monitoring** the fix to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have **identified the problem** causing performance degradation of our Cloud API. A fix is in progress, users may still see elevated error rates and or slow responses until the fix completes. Our team is working on getting this fix rolled out as soon as possible.
We are currently investigating an issue resulting in a performance degradation of our Cloud API. Users may be experiencing **elevated error rates and or slow responses**. Our team is analyzing logs and infrastructure to identify the cause and implement a fix as soon as possible.
Model Performance Issue: llama-3.1-8b-instant
3 updates
The issues affecting llama-3.1-8b-instant have been **resolved**. The model is operating normally. Root cause: Two recent changes to the inference-engine-instances repository that were contributing to the elevated Orion LoRA latencies were reverted . We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for llama-3.1-8b-instant service and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We have **identified an issue** affecting the lama-3.1-8b-instant service. A fix is in progress. Users may still experience latency until the fix completes.
October 2025(5 incidents)
openai/gpt-oss-120b Degraded Performance
5 updates
The latency issues affecting openai/gpt-oss-120b have been resolved. The model is now operating normally and latency has returned to expected ranges. We apologize for the disruption and thank you for your patience.
We are continuing to investigate issues with our openai/gpt-oss-120b model. Users may be experiencing or slow responses. Other models remain operational. Our team is analyzing traces, logs and infrastructure to identify the cause.
We are continuing to investigate issues with our openai/gpt-oss-120b model. Users may be experiencing elevated error rates and or slow responses. Other models remain operational. Our team is analyzing traces, logs and infrastructure to identify the cause.
We are continuing to investigate issues with our openai/gpt-oss-120b model. Users may be experiencing elevated error rates and or slow responses. Other models remain operational. Our team is analyzing traces, logs and infrastructure to identify the cause.
We are currently investigating reports of an issue with ouropenai/gpt-oss-120b. Users may be experiencing elevated error rates and or slow responses. Other models remain operational. Our team is analyzing traces, logs and infrastructure to identify the cause.
openai/gpt-oss-120b Degraded Performance
4 updates
The issues affecting openai/gpt-oss-120b have been **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We are actively implementing a fix for openai/gpt-oss-120b and we expect performance to begin improving shortly. We’ll be monitoring the model to ensure stability as this fix rolls out. If all remains normal, we will resolve the incident in the next update.
We have **identified a problem** affecting the openai/gpt-oss-120b service. The issue was traced to . A fix is in progress, which may take some time to implement. Users may still see elevated error rates or delayed responses until the fix is in place.
We are currently investigating reports of an issue with our openai/gpt-oss-120b model. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
meta-llama/llama-4-scout-17b-16e-instruct & groq/compound Degraded Performance
5 updates
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct have been **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We have implemented a fix for the issue impacting our meta-llama/llama-4-scout-17b-16e-instruct model and performance has improved. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are still working to identify the root cause of the issue with our meta-llama/llama-4-scout-17b-16e-instruct model. As part of the investigation our team has implemented additional logging to pinpoint the issue. Users may still be experiencing elevated error rates and or slow responses. Other models remain operational.
We are continuing to investigate an issue with our meta-llama/llama-4-scout-17b-16e-instruct model. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational.
We are currently investigating reports of an issue with our meta-llama/llama-4-scout-17b-16e-instruct model. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing metrics and logs to identify the cause.
Llama-3.1-8b-instant model Degraded Performance
4 updates
The issues affecting llama-3.1-8b-instant have been resolved. The model is operating normally. Thank you for your patience.
We have fix and performance is improving. We’re now monitoring the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We have identified the problem affecting the llama-3.1-8b-instant service. A fix is in progress. Users may still see impact until the fix completes.
We are currently investigating reports of an issue with the llama-3.1-8b-instant model. Users may be experiencing repeated performance degradation, with significant spikes in response times and increased 503 errors across several regions. The team is investigating.
llama-3.3-70b-versatile and llama-3.1-8b-instant Degraded Performance
3 updates
The issues affecting llama-3.3-70b-versatile and llama-3.1-8b-instant have been **resolved**. The models are operating normally. We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for the llama-3.3-70b-versatile and llama-3.1-8b-instant models and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are currently investigating reports of an issue with our **[**llama-3.3-70b-versatile and llama-3.1-8b-instant models. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
September 2025(6 incidents)
Degraded login via Vercel
3 updates
Vercel‑initiated login is operating normally. Other login methods were unaffected. This incident is now resolved.
A small number of Vercel‑initiated logins may error. Direct login at the Groq console works normally. We are coordinating on a fix with our identity provider. We will update this page if impact or timeline changes.
Following the login system upgrade earlier today, we’ve identified an issue with Vercel-initiated logins. Users starting login from Vercel may see an error. Other login methods are unaffected. The team is working to identify the cause to remediate. **Workaround:** Log in directly at the Groq console using your usual method. We will provide an update when there is a material change.
Elevated Latency on Search Tooling
3 updates
This issue has been resolved.
A fix has been implemented and performance is improving. We’re now monitoring the models to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are investigating an issue causing elevated latency on search tooling affecting the **groq/compound, groq/compound-mini, openai/gpt-oss-120b and openai/gpt-oss-20b** models. The team is working to identify the cause and remediate.
llama-3.3-70b-versatile Performance Issue
3 updates
This incident has been resolved.
We have implemented a fix and are no longer seeing degraded output from llama-3.3-70b. The team is monitoring this model to ensure continued performance.
We have identified an issue where a small portion of requests against llama-3.3-70b-versatile are returning degraded output. A fix is in progress, Users may still see the issue until the fix is fully deployed.
moonshotai/kimi-k2-instruct Unavailable or Degraded Performance
4 updates
The issues affecting moonshotai/kimi-k2-instruct have been **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for moonshotai/kimi-k2-instruct and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We have **identified the problem** affecting the moonshotai/kimi-k2-instruct model. A fix is in progress however, users may still see **persistent errors** until the fix completes.
We are currently investigating reports of an issue with our **moonshotai/kimi-k2-instruct-0905 model**. Users may be experiencing **elevated error rates**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
me-central-1 Data Center Failure Impacting Capacity
5 updates
We have **fully resolved** the failure in our me-central-1 region. All services are operating at normal capacity. The incident was caused by a **network failure**, which has been corrected. We apologize for the disruption and **appreciate your patience**.
We have **resolved the failure** in our me-central-1 region. The team has isolated the cause (related to networking issues) and has implemented a fix. We expect requests to be served successfully as services resume normal operations.
We have **identified a failure** in our me-central-1 region. The team has isolated the cause (related to networking issues) and is working on a fix now. Users continue to see persistent failures until the fix is fully implemented.
We have **identified a failure** in our me-central-1 region. The team has isolated the cause (related to **networking issues**) and is **working on restoring capacity**. Users **continue to see persistent failures**.
We are currently investigating an issue in our me-central-1 region. Users **may experience persistent failures** due to reduced resources in the region. Our engineering team is working to identify the root cause and remediate the problem.
API Errors Affecting Multiple Regions
4 updates
The authentication issue impacting multiple regions has been fully resolved. Services are operating normally, and error rates remain stable.
A fix has been implemented for the authentication issue affecting multiple regions. Error rates have returned to normal, and we are closely monitoring to ensure service remains stable.
We have **identified an issue** in multiple region's infrastructure that is causing **elevated error rates and incorrect 401s errors against api requests**. The team has isolated the cause (related to authentication issues) and is **working on restoring service**.
We are currently investigating a potential issue at our me-central-1 region that is impacting system capacity. Users **may experience elevated error rates** due to reduced resources. Our engineering team is working to identify the root cause and remediate the problem.
📡 Tired of checking Groq status manually?
Better Stack monitors uptime every 30 seconds and alerts you instantly when Groq goes down.