Groq LPU inference engine, GroqCloud API, and Groq platform services.
Official status data last reflected here:
[ STATUS AT A GLANCE ]
Current status of individual Groq services
When Groq has issues, Bifrost automatically routes your requests to a healthy alternative provider. Zero code changes. 99.999% effective uptime.
What Groq does, where the data on this page comes from, and recent reliability
[ ABOUT GROQ ]
Groq provides GroqCloud API, LPU inference, and Open model serving. Groq is often chosen for speed-sensitive applications, so status changes can immediately impact latency-sensitive user experiences and throughput targets.
This page pulls data from Groq's official status page to show current service health, any active incidents, and a history of recent issues, all in one view.
[ DATA SOURCES ]
Groq publishes detailed component status, a full incident archive, and scheduled maintenance data through their official status page.
[ RELIABILITY ]
[ COMMON USE CASES ]
Groq is often chosen for speed-sensitive applications, so status changes can immediately impact latency-sensitive user experiences and throughput targets.
Active incidents, scheduled maintenance, and incident history for Groq
Invalid Date to Invalid Date
Jul 1, 2026 · Resolved Jul 2, 2026
The **capacity issue has been resolved**. All services are operating at normal capacity. The incident was caused by a **power loss issue that led to a cooling system failure at one of our US Central data centers**. We apologize for any disruption and **appreciate your patience**.
We have **redistributed and allocated more capacity to production models** in **the US Central region.** Users should **see performance return to normal**. We’re now **monitoring** the infrastructure to ensure stability.
A power loss issue at approximately 22:20pm UTC led to a subsequent **cooling system failure** at one of our **US Central** data centers that is causing **reduced capacity**, primarily for the `openai/gpt-oss-20b` model. The team is **working on restoring capacity**.
We have **identified a cooling system failure** at one of our **US Central** data centers that is causing **reduced capacity**. The team is **working on restoring capacity**. Users **continue to see elevated latencies for specific models**.
We are currently investigating a potential issue at one of our US Central data centers that is impacting system capacity. Users **may experience higher latencies** due to reduced resources. Our engineering team is working to identify the root cause and remediate the problem.
Mar 19, 2026 · Resolved Mar 19, 2026
The issues affecting openai/gpt-oss-120b have been **resolved**. The model is operating normally. Actions were taken to cancel billing plans and restrict verification status of organizations engaged in coordinated abuse. We apologize for the disruption and thank you for your patience.
We have implemented a fix for openai/gpt-oss-120b and performance is improving. We’re now monitoring the model to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have identified the problem affecting the openai/gpt-oss-120b model. The issue has been traced to malicious traffic patterns. A fix is in progress, we are working to block the source of these requests. Users may still see degraded performance from requests made to openai/gpt-oss-120b as we identify and block the orgs from which this traffic is originating.
Feb 7, 2026 · Resolved Feb 7, 2026
This incident has been resolved. Thank you for your patience.
We have implemented a new fix for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound and performance has improved. We’re continuing monitor the models to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have implemented a fix for meta-llama/llama-4-scout-17b-16e-instruct however performance is still degraded. We’re continuing to investigate the models to determine the cause of the issues.
We are continuing to investigating an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing elevated error rates or slow response time. Other models remain operational. Our team is continuing to analyze logs and infrastructure to identify the cause.
We are currently investigating reports of an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Feb 5, 2026 · Resolved Feb 5, 2026
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct have been **resolved**. The model is once again operating normally. We apologize for the disruption and thank you for your patience.
meta-llama/llama-4-scout-17b-16e-instruct and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are currently investigating reports of an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Feb 5, 2026 · Resolved Feb 5, 2026
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct have been fully **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We’re continuing to **monitor** the scout model to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have **implemented a fix** for meta-llama/llama-4-scout-17b-16e-instruct and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We have **identified the problem** affecting the meta-llama/llama-4-scout-17b-16e-instruct. A fix is in progress, users may still see delayed responses and errors from Scout until the fix completes.
We are currently investigating reports of an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Jan 26, 2026 · Resolved Jan 27, 2026
This incident has been resolved. Thank you for your patience.
We have implemented a new fix for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound and performance has improved. We’re continuing monitor the models to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have implemented a fix for meta-llama/llama-4-scout-17b-16e-instruct however performance is still degraded. We’re continuing to investigate the models to determine the cause of the issues.
We are currently investigating an issue with meta-llama/llama-4-scout-17b-16e-instruct. Users may be experiencing elevated error rates or slow response time. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Jan 24, 2026 · Resolved Jan 24, 2026
This incident is now resolved.
This issue is now resolved. All services are operating normally. We will continue to monitor for stability and resolve this incident shortly. Thank you for your patience.
The issue impacting our SYD data center has been fixed. Users should gradually see performance return to normal as we continue to recover all models. We will continue to monitor.
The team is still working to resolve the issue at our SYD Data Center. Customers in the area may still be experiencing some model latency. We will provide further updates at they are available.
We have **identified an issue in our Sydney, AUS data center** that may be causing some latency for customers in this region. The issue was traced to **a network issue**. Users may still see latency until the fix completes.
Dec 24, 2025 · Resolved Dec 25, 2025
The issues affecting **llama-3.3-70b-versatile** have been **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for **llama-3.3-70b-versatile** and performance is improving. The team is continuing to work the issue and we hope to have it resolved shortly.
We’ve identified an issue causing 503s for the `llama-3.3-70b-versatile` production model. We’ve started mitigation. Error rates are improving, but some requests may still fail while we continue tuning and monitoring.
Dec 17, 2025 · Resolved Dec 14, 2025
**The data center capacity issue has been resolved. All services are operating at normal capacity. The incident was caused by a failed ARP entry on the network gateway device caused the Salam network link to go down in the DMM1 data center, which has been addressed. We apologize for the disruption and appreciate your patience.**
We have restored capacity in DMM 1 and services are beginning to recover. Users should gradually see network performance return to normal. We’re now monitoring the infrastructure to ensure stability.
We have identified a failure in our DMM1 infrastructure that is causing reduced capacity and network latency. The team has isolated the cause to the Salam link confirmed down and the Mobily link operating at full capacity, resulting in significant packet loss and is working on restoring capacity. Users continue to see Network latency.
We are currently investigating a potential network issue at our DMM1 that is impacting system capacity. Users may experience slower response times or intermittent failures due to reduced resources. Our engineering team is working to identify the root cause and remediate the problem.
Dec 6, 2025 · Resolved Dec 6, 2025
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct , meta-llama/llama-4-maverick-17b-128e-instruct, & groq/compound have been **resolved**. The models are operating normally. We apologize for the disruption and thank you for your patience.
We have implemented a new fix for meta-llama/llama-4-scout-17b-16e-instruct and performance and have observed improvement over the last 15 minutes. We’re continuing monitor to ensure stability persists. If all remains normal, we will resolve the incident in the next update. In addition to meta-llama/llama-4-scout-17b-16e-instruct, we realized there also was some impact to groq/compound, which has also recovered.
We are no longer seeing issues on meta-llama/llama-4-maverick-17b-128e-instruct. However, the team is still troubleshooting errors and latency issues with meta-llama/llama-4-scout-17b-16e-instruct.
We are now seeing similar degradation on meta-llama/llama-4-maverick-17b-128e-instruct. The team is still investigating the issue and working to resolve.
We are currently investigating reports of an issue with our meta-llama/llama-4-scout-17b-16e-instruct & groq/compound models. Users may be experiencing elevated error rates and or slow responses. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Dec 2, 2025 · Resolved Dec 2, 2025
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct & groq/compound have been **resolved**. Both of the models are operating normally. We apologize for the disruption and thank you for your patience.
We have **implemented a new fix** for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound and performance has improved. We’re continuing **monitor** the models to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have implemented a fix for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound however performance is still degraded. We’re continuing to investigate the models to determine the cause of the issues.
We have **implemented a fix** for meta-llama/llama-4-scout-17b-16e-instruct & groq/compound and performance is improving. We’re now **monitoring** the models to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are currently investigating reports of an issue with our meta-llama/llama-4-scout-17b-16e-instruct & groq/compound models. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Nov 18, 2025 · Resolved Nov 18, 2025
We are resolving this incident as we have now observed ~30 minutes of normalized service availability and customer traffic. We will continue to closely monitor our systems for any signs of regression as [Cloudflare's statuspage incident remains active ](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7)and in monitoring.
We are continuing to monitor an issue with our services that is impacting customer experiences due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage"). Cloudflare has just shared that "A fix has been implemented and we believe the incident is now resolved. We are continuing to monitor for errors to ensure all services are back to normal.". We are monitoring to confirm the restoration of service availability and successful requests and are continuing to monitor for signs of persistent recovery as Cloudflare now claims to have implemented a fix.
We are continuing to monitor an issue with our services that is impacting customer experiences due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage"). Cloudflare previously shared that "We are continuing to work towards restoring other services" and "The issue has been identified and a fix is being implemented". We are continuing to observe higher-than-normal error rates and are continuing to monitor for signs of recovery as Cloudflare works to implement a fix.
We are continuing to monitor an issue with our services that is impacting customer experiences due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage"). Cloudflare recently shared that "We are continuing to work towards restoring other services" and "The issue has been identified and a fix is being implemented". We are continuing to observe higher-than-normal error rates and are monitoring for signs of recovery as Cloudflare works to implement a fix.
We are continuing to monitor an issue with our services that is impacting customer experiences due to an ongoing [global cloudflare outage](https://www.cloudflarestatus.com/incidents/8gmgl950y3h7 "global cloudflare outage"). Cloudflare recently shared that "The are continuing to investigate the issue". We are continuing to observe higher-than-normal error rates despite Cloudflare's previous update stating that they were seeing services recover.
+ 3 more updates
Nov 5, 2025 · Resolved Nov 5, 2025
The issues causing degraded performance have been **resolved**. All models are now operating normally. Upon investigation we learned the issue was scoped to an internal feature yet to be released that is still in development. At this time we don't believe there was any customer impact as a direct result of this issue. We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for the performance degradation of our Cloud API and performance is improving. We’re now **monitoring** the fix to ensure stability persists. If all remains normal, we will resolve the incident in the next update.
We have **identified the problem** causing performance degradation of our Cloud API. A fix is in progress, users may still see elevated error rates and or slow responses until the fix completes. Our team is working on getting this fix rolled out as soon as possible.
We are currently investigating an issue resulting in a performance degradation of our Cloud API. Users may be experiencing **elevated error rates and or slow responses**. Our team is analyzing logs and infrastructure to identify the cause and implement a fix as soon as possible.
Nov 4, 2025 · Resolved Nov 4, 2025
The issues affecting llama-3.1-8b-instant have been **resolved**. The model is operating normally. Root cause: Two recent changes to the inference-engine-instances repository that were contributing to the elevated Orion LoRA latencies were reverted . We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for llama-3.1-8b-instant service and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We have **identified an issue** affecting the lama-3.1-8b-instant service. A fix is in progress. Users may still experience latency until the fix completes.
Oct 22, 2025 · Resolved Oct 22, 2025
The latency issues affecting openai/gpt-oss-120b have been resolved. The model is now operating normally and latency has returned to expected ranges. We apologize for the disruption and thank you for your patience.
We are continuing to investigate issues with our openai/gpt-oss-120b model. Users may be experiencing or slow responses. Other models remain operational. Our team is analyzing traces, logs and infrastructure to identify the cause.
We are continuing to investigate issues with our openai/gpt-oss-120b model. Users may be experiencing elevated error rates and or slow responses. Other models remain operational. Our team is analyzing traces, logs and infrastructure to identify the cause.
We are continuing to investigate issues with our openai/gpt-oss-120b model. Users may be experiencing elevated error rates and or slow responses. Other models remain operational. Our team is analyzing traces, logs and infrastructure to identify the cause.
We are currently investigating reports of an issue with ouropenai/gpt-oss-120b. Users may be experiencing elevated error rates and or slow responses. Other models remain operational. Our team is analyzing traces, logs and infrastructure to identify the cause.
Oct 21, 2025 · Resolved Oct 21, 2025
The issues affecting openai/gpt-oss-120b have been **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We are actively implementing a fix for openai/gpt-oss-120b and we expect performance to begin improving shortly. We’ll be monitoring the model to ensure stability as this fix rolls out. If all remains normal, we will resolve the incident in the next update.
We have **identified a problem** affecting the openai/gpt-oss-120b service. The issue was traced to . A fix is in progress, which may take some time to implement. Users may still see elevated error rates or delayed responses until the fix is in place.
We are currently investigating reports of an issue with our openai/gpt-oss-120b model. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Oct 14, 2025 · Resolved Oct 14, 2025
The issues affecting meta-llama/llama-4-scout-17b-16e-instruct have been **resolved**. The model is operating normally. We apologize for the disruption and thank you for your patience.
We have implemented a fix for the issue impacting our meta-llama/llama-4-scout-17b-16e-instruct model and performance has improved. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are still working to identify the root cause of the issue with our meta-llama/llama-4-scout-17b-16e-instruct model. As part of the investigation our team has implemented additional logging to pinpoint the issue. Users may still be experiencing elevated error rates and or slow responses. Other models remain operational.
We are continuing to investigate an issue with our meta-llama/llama-4-scout-17b-16e-instruct model. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational.
We are currently investigating reports of an issue with our meta-llama/llama-4-scout-17b-16e-instruct model. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing metrics and logs to identify the cause.
Oct 6, 2025 · Resolved Oct 7, 2025
The issues affecting llama-3.1-8b-instant have been resolved. The model is operating normally. Thank you for your patience.
We have fix and performance is improving. We’re now monitoring the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We have identified the problem affecting the llama-3.1-8b-instant service. A fix is in progress. Users may still see impact until the fix completes.
We are currently investigating reports of an issue with the llama-3.1-8b-instant model. Users may be experiencing repeated performance degradation, with significant spikes in response times and increased 503 errors across several regions. The team is investigating.
Oct 6, 2025 · Resolved Oct 6, 2025
The issues affecting llama-3.3-70b-versatile and llama-3.1-8b-instant have been **resolved**. The models are operating normally. We apologize for the disruption and thank you for your patience.
We have **implemented a fix** for the llama-3.3-70b-versatile and llama-3.1-8b-instant models and performance is improving. We’re now **monitoring** the model to ensure stability. If all remains normal, we will resolve the incident in the next update.
We are currently investigating reports of an issue with our **[**llama-3.3-70b-versatile and llama-3.1-8b-instant models. Users may be experiencing **elevated error rates and or slow responses**. Other models remain operational. Our team is analyzing logs and infrastructure to identify the cause.
Sep 18, 2025 · Resolved Sep 18, 2025
Vercel‑initiated login is operating normally. Other login methods were unaffected. This incident is now resolved.
A small number of Vercel‑initiated logins may error. Direct login at the Groq console works normally. We are coordinating on a fix with our identity provider. We will update this page if impact or timeline changes.
Following the login system upgrade earlier today, we’ve identified an issue with Vercel-initiated logins. Users starting login from Vercel may see an error. Other login methods are unaffected. The team is working to identify the cause to remediate. **Workaround:** Log in directly at the Groq console using your usual method. We will provide an update when there is a material change.
Check the status indicator at the top of this page. It pulls directly from Groq's official status page. If Groq is experiencing any issues, you'll see it reflected here. This real-time monitoring helps teams quickly identify whether performance problems are caused by Groq infrastructure or their own systems.
This page tracks GroqCloud API, LPU inference, and Open model serving using data from Groq's official status page. You can see current component health, active incidents, and a history of past issues. This visibility is crucial for teams building resilient AI applications that need to route around provider outages.
We check Groq's status page every 60 seconds to ensure you get near real-time status updates. How quickly issues show up here depends on how fast Groq updates their own official status. For production systems that need instant failover, Bifrost can automatically detect and route around degraded providers.
Groq is often chosen for speed-sensitive applications, so status changes can immediately impact latency-sensitive user experiences and throughput targets. Real-time status monitoring enables proactive incident response and helps teams decide when to route traffic to alternative providers for maximum uptime.
When Groq experiences an outage, the best practice is automatic failover to alternative AI providers. Bifrost is an open-source AI gateway that automatically detects Groq degradation and routes LLM traffic to healthy alternatives like Together AI, Fireworks AI, OpenAI, keeping your application running with zero manual intervention. This intelligent routing ensures your users never experience downtime from a single provider's issues.
Production AI applications should never depend on a single provider. Bifrost AI Gateway provides automatic multi-provider failover, intelligent load balancing, and health-based routing. When Groq degrades, Bifrost instantly routes requests to operational alternatives while maintaining API compatibility. This architecture approach is used by teams running mission-critical AI features.
Common causes of Groq issues include infrastructure scaling challenges, regional cloud provider problems, API gateway overload, and deployment errors. Monitor this page to stay informed, and consider implementing automatic failover with Bifrost to maintain uptime during Groq incidents.
When Groq experiences issues, Bifrost AI Gateway can automatically route your LLM traffic to these alternatives
💡 Build resilient AI apps: Configure Bifrost to automatically detect Groq outages and route requests to healthy alternatives. This multi-provider approach ensures your application maintains uptime even when individual providers experience issues. Learn more about Bifrost AI Gateway
Groq offers 11 models across 1 mode (11 chat)