Replicate model hosting, inference API, and machine learning platform services.
Official status data last reflected here:
[ STATUS AT A GLANCE ]
Current status of individual Replicate services
When Replicate has issues, Bifrost automatically routes your requests to a healthy alternative provider. Zero code changes. 99.999% effective uptime.
What Replicate does, where the data on this page comes from, and recent reliability
[ ABOUT REPLICATE ]
Replicate provides Replicate API, Hosted open models, and Inference jobs. Replicate is used to access a wide range of models, so availability issues can impact multiple AI features at once across image, audio, and text workflows.
This page pulls data from Replicate's official status page to show current service health, any active incidents, and a history of recent issues, all in one view.
[ DATA SOURCES ]
Replicate publishes detailed component status, a full incident archive, and scheduled maintenance data through their official status page.
[ RELIABILITY ]
[ COMMON USE CASES ]
Replicate is used to access a wide range of models, so availability issues can impact multiple AI features at once across image, audio, and text workflows.
Active incidents, scheduled maintenance, and incident history for Replicate
Jul 22, 2026 · Resolved Jul 23, 2026
This issue is resolved
This issue is now resolved
GPU usage is back below capacity. We're continuing to monitor.
We are over provisioned on H100s currently which is causing long queue times for some models running on H100s
Jul 16, 2026 · Resolved Jul 16, 2026
We have not seen elevated rates of HuggingFace/model setup errors for a couple hours now, so we believe this incident is cleared
Still seeing signs of improvement, most models are able to spin up with a few still waiting their turn from HuggingFace. Models should spin up after enough tries if they are failing for HuggingFace related reasons.
We have seen most HuggingFace calls start succeeding, except for a small subset that are being downloaded through [cas-bridge.xethub.hf.co](http://cas-bridge.xethub.hf.co "cas-bridge.xethub.hf.co"). Most models are now succeeding to spin up
HuggingFace is reporting recovery, which we are also seeing on our end. However, we still see minor degradation, which we believe is due to thundering herd once the main issue was resolved. We are still keeping an eye out, but believe we are recovering
We are aware of 504s being returned by models that reach out to HuggingFace during setup. We believe this is likely related to the disruption visible at [https://status.huggingface.co/](https://status.huggingface.co/ "https://status.huggingface.co/"). We are monitoring and will update when we can identify the HuggingFace CDN is back up/
Jul 13, 2026 · Resolved Jul 14, 2026
H100 capacity has returned to normal levels
We pushed a fix and are monitoring
Long queue times, especially on BFL models
Jul 10, 2026 · Resolved Jul 10, 2026
We are back under max capacity for H100 hardware. Thank you for your patience!
We are seeing high contention on H100 hardware which is resulting in delays on predictions and scale-out.
Jun 30, 2026 · Resolved Jun 30, 2026
The H100 capacity is back in good health. Thanks for your patience!
We recently received a sharp increase in demand for H100 capacity which is resulting in delayed scale-out and queue backup.
Jun 3, 2026 · Resolved Jun 3, 2026
System is back to operating normally
System is back to operating normally
We're seeing long setup times and high contention for models on some L40S and H200 clusters.
May 28, 2026 · Resolved May 28, 2026
This issue has been resolved and queue times are back to normal
Long queue times for black-forest-labs/flux-2-klein-4b resulting in canceled predictions
May 21, 2026 · Resolved May 21, 2026
Message flows are healthy.
Our message queues for prediction and training status updates are hitting capacity limits which are causing connection failures for queue consumers. We are in the process of bringing additional capacity online.
May 21, 2026 · Resolved May 21, 2026
H100 hardware contention has resolved. Thank you for your patience!
We are seeing heightened demand for H100 hardware which is causing severe queue delays.
May 12, 2026 · Resolved May 12, 2026
There is no more contention for H100 hardware. Thank you for your patience!
Demand for constrained H100 hardware is causing scaling delays. This impacts queue size and inference speed for any models running on H100s.
Apr 19, 2026 · Resolved Apr 19, 2026
All A100 capacity is back. Thanks for your patience!
All predictions and trainings targeting A100 hardware are experiencing degraded performance while control plane nodes restart.
Apr 9, 2026 · Resolved Apr 9, 2026
The maintenance is complete and all systems are reporting healthy. Thank you for your patience!
The persistent storage for all A100 hardware is under maintenance and is expected to be degraded until completion.
Mar 23, 2026 · Resolved Mar 24, 2026
Black Forest Labs has resolved the issue.
Some Black Forest Labs models are failing due to downstream errors from BFL. We are monitoring the situation and working on work arounds. BFL Status page: [https://status.bfl.ml/](https://status.bfl.ml/ "https://status.bfl.ml/")
Mar 10, 2026 · Resolved Mar 10, 2026
Flux Schnell requests are being served normally.
A GPU provider has an outage. Traffic is being rerouted and we are processing new Flux Schnell requests.
We are investigating an outage that is only affecting Flux Schnell.
Feb 20, 2026 · Resolved Feb 20, 2026
Models are once again operational
We have identified the root cause and have made an update. We are continuing to monitor as we start to see things improve.
We are currently investigating why a large number of models are not currently processing requests, with predictions stalled with a "starting" status.
Jan 26, 2026 · Resolved Jan 26, 2026
Model setup failures have dropped to normal levels since 10:40 UTC. Thank you for your patience!
We are seeing download issues from HuggingFace for a handful of T4 models, stopping new instances from starting. Requests for those models may be delayed as new capacity fails to come online. We will continue to monitor the situation.
The problem seems to be isolated to models downloading weights from HuggingFace.
We are seeing increased setup failure rates on models using T4 GPUs.
Jan 20, 2026 · Resolved Jan 21, 2026
All infrastructure problems have been resolved and model predictions and training are available again for the relevant hardware types.
We are seeing indications that multiple models are not receiving prediction or training requests correctly.
Jan 20, 2026 · Resolved Jan 21, 2026
The `black-forest-labs/flux-schnell` model is fully operational now. Thank you for your patience!
The infrastructure failure preventing predictions from getting scheduled has been resolved for the past hour. We are monitoring performance and ensuring all adjacent systems are working well.
The black-forest-labs/flux-schnell model is currently unavailable. We are actively investigating.
Jan 15, 2026 · Resolved Jan 15, 2026
Incident is resolved.
All models have returned to normal processing timeframes. Thank you for your patience.
We have shifted traffic for the Flux Schnell model to a region with additional capacity. While we work through the backlog of requests, new requests should begin to be served within a normal timeframe. This continues to impact only the Flux Schnell official model.
We have identified and remediated a data store related to queues that was wedged, preventing traffic from reaching it within one of our regions. Most inference and API usage as returned to normal with the exception of the Flux Schnell model. We are working on restoring service to the flux schnell model. The impact is large delays on predictions and elevated errors when submitting preditions.
We are aware of elevated errors for predictions being sent to one of our regions. This is impacting an subset of models targeting H100 and L40S Hardware types. We are investigating and will provide an update as soon as information is available.
Dec 18, 2025 · Resolved Dec 20, 2025
We are all clear now. Thank you for your patience!
We are seeing high demand for the H100 hardware type, with increased delays for scale-out and cold boot events.
Check the status indicator at the top of this page. It pulls directly from Replicate's official status page. If Replicate is experiencing any issues, you'll see it reflected here. This real-time monitoring helps teams quickly identify whether performance problems are caused by Replicate infrastructure or their own systems.
This page tracks Replicate API, Hosted open models, and Inference jobs using data from Replicate's official status page. You can see current component health, active incidents, and a history of past issues. This visibility is crucial for teams building resilient AI applications that need to route around provider outages.
We check Replicate's status page every 60 seconds to ensure you get near real-time status updates. How quickly issues show up here depends on how fast Replicate updates their own official status. For production systems that need instant failover, Bifrost can automatically detect and route around degraded providers.
Replicate is used to access a wide range of models, so availability issues can impact multiple AI features at once across image, audio, and text workflows. Real-time status monitoring enables proactive incident response and helps teams decide when to route traffic to alternative providers for maximum uptime.
When Replicate experiences an outage, the best practice is automatic failover to alternative AI providers. Bifrost is an open-source AI gateway that automatically detects Replicate degradation and routes LLM traffic to healthy alternatives like Stability AI, Hugging Face, keeping your application running with zero manual intervention. This intelligent routing ensures your users never experience downtime from a single provider's issues.
Production AI applications should never depend on a single provider. Bifrost AI Gateway provides automatic multi-provider failover, intelligent load balancing, and health-based routing. When Replicate degrades, Bifrost instantly routes requests to operational alternatives while maintaining API compatibility. This architecture approach is used by teams running mission-critical AI features.
Common causes of Replicate issues include infrastructure scaling challenges, regional cloud provider problems, API gateway overload, and deployment errors. Monitor this page to stay informed, and consider implementing automatic failover with Bifrost to maintain uptime during Replicate incidents.
When Replicate experiences issues, Bifrost AI Gateway can automatically route your LLM traffic to these alternatives
💡 Build resilient AI apps: Configure Bifrost to automatically detect Replicate outages and route requests to healthy alternatives. This multi-provider approach ensures your application maintains uptime even when individual providers experience issues. Learn more about Bifrost AI Gateway
Replicate offers 40 models across 1 mode (40 chat)