Try Bifrost Enterprise free for 14 days.
Request access
Replicate
[ LIVE STATUS ]

Is Replicate Down?

Replicate model hosting, inference API, and machine learning platform services.

Official status data last reflected here:

All Systems Operational
Live · updated just now

[ STATUS AT A GLANCE ]

Operational
Current Status
All Systems Operational
13
Components
Service areas tracked on this page
10
90d Incidents
Incidents reported in the last 90 days

System Components

Current status of individual Replicate services

BillingAll Systems Operational
system-wide incident history
90 days agoToday
Support TicketsAll Systems Operational
system-wide incident history
90 days agoToday
Streaming APIAll Systems Operational
system-wide incident history
90 days agoToday
HTTP APIAll Systems Operational
system-wide incident history
90 days agoToday
Replicate Registry (r8.im)All Systems Operational
system-wide incident history
90 days agoToday
CPU HardwareAll Systems Operational
system-wide incident history
90 days agoToday
Home PageAll Systems Operational
system-wide incident history
90 days agoToday
PlaygroundAll Systems Operational
system-wide incident history
90 days agoToday
A100 HardwareAll Systems Operational
system-wide incident history
90 days agoToday
T4 HardwareAll Systems Operational
system-wide incident history
90 days agoToday
Official ModelsAll Systems Operational
system-wide incident history
90 days agoToday
H100 HardwareAll Systems Operational
system-wide incident history
90 days agoToday
L40S HardwareAll Systems Operational
system-wide incident history
90 days agoToday
[ AUTOMATIC FAILOVER ]

Replicate down? Route around it.

When Replicate has issues, Bifrost automatically routes your requests to a healthy alternative provider. Zero code changes. 99.999% effective uptime.

About Replicate

What Replicate does, where the data on this page comes from, and recent reliability

[ ABOUT REPLICATE ]

About Replicate

Replicate provides Replicate API, Hosted open models, and Inference jobs. Replicate is used to access a wide range of models, so availability issues can impact multiple AI features at once across image, audio, and text workflows.

This page pulls data from Replicate's official status page to show current service health, any active incidents, and a history of recent issues, all in one view.

Replicate APIHosted open modelsInference jobs

[ DATA SOURCES ]

Full incident history available

Replicate publishes detailed component status, a full incident archive, and scheduled maintenance data through their official status page.

  • Data pulled from Replicate's official status page (www.replicatestatus.com)
  • Refreshed every 60 seconds
  • Covers Replicate API, Hosted open models, and Inference jobs
  • Includes full incident archive and scheduled maintenance history

[ RELIABILITY ]

Recent reliability

  • 10 incidents reported over the last 90 days.
  • Last reported incident was 1 day ago.
  • All 13 monitored components are currently operational.

[ COMMON USE CASES ]

How teams use Replicate

Replicate is used to access a wide range of models, so availability issues can impact multiple AI features at once across image, audio, and text workflows.

Hosted model APIs
Batch generation
Product experimentation

Incidents & Maintenance

Active incidents, scheduled maintenance, and incident history for Replicate

Past Incidents

Hitting GPU Capacity for H100s creating large queue times for some models

Jul 22, 2026 · Resolved Jul 23, 2026

Resolvedminor
ResolvedJul 23, 2026, 2:19 PM UTC

This issue is resolved

MonitoringJul 23, 2026, 3:05 AM UTC

This issue is now resolved

MonitoringJul 22, 2026, 5:11 PM UTC

GPU usage is back below capacity. We're continuing to monitor.

IdentifiedJul 22, 2026, 11:51 AM UTC

We are over provisioned on H100s currently which is causing long queue times for some models running on H100s

HuggingFace download issues

Jul 16, 2026 · Resolved Jul 16, 2026

Resolvednone
ResolvedJul 16, 2026, 4:04 PM UTC

We have not seen elevated rates of HuggingFace/model setup errors for a couple hours now, so we believe this incident is cleared

MonitoringJul 16, 2026, 9:46 AM UTC

Still seeing signs of improvement, most models are able to spin up with a few still waiting their turn from HuggingFace. Models should spin up after enough tries if they are failing for HuggingFace related reasons.

MonitoringJul 16, 2026, 9:28 AM UTC

We have seen most HuggingFace calls start succeeding, except for a small subset that are being downloaded through [cas-bridge.xethub.hf.co](http://cas-bridge.xethub.hf.co "cas-bridge.xethub.hf.co"). Most models are now succeeding to spin up

MonitoringJul 16, 2026, 9:10 AM UTC

HuggingFace is reporting recovery, which we are also seeing on our end. However, we still see minor degradation, which we believe is due to thundering herd once the main issue was resolved. We are still keeping an eye out, but believe we are recovering

InvestigatingJul 16, 2026, 8:31 AM UTC

We are aware of 504s being returned by models that reach out to HuggingFace during setup. We believe this is likely related to the disruption visible at [https://status.huggingface.co/](https://status.huggingface.co/ "https://status.huggingface.co/"). We are monitoring and will update when we can identify the HuggingFace CDN is back up/

H100 GPU shortage resulting in high queue times

Jul 13, 2026 · Resolved Jul 14, 2026

Resolvedminor
ResolvedJul 14, 2026, 5:27 PM UTC

H100 capacity has returned to normal levels

MonitoringJul 13, 2026, 5:48 PM UTC

We pushed a fix and are monitoring

InvestigatingJul 13, 2026, 4:26 PM UTC

Long queue times, especially on BFL models

High contention on H100 hardware

Jul 10, 2026 · Resolved Jul 10, 2026

Resolvedminor
ResolvedJul 10, 2026, 11:04 PM UTC

We are back under max capacity for H100 hardware. Thank you for your patience!

MonitoringJul 10, 2026, 2:50 PM UTC

We are seeing high contention on H100 hardware which is resulting in delays on predictions and scale-out.

Limited H100 capacity

Jun 30, 2026 · Resolved Jun 30, 2026

Resolvedminor
ResolvedJun 30, 2026, 7:29 PM UTC

The H100 capacity is back in good health. Thanks for your patience!

IdentifiedJun 30, 2026, 1:55 PM UTC

We recently received a sharp increase in demand for H100 capacity which is resulting in delayed scale-out and queue backup.

We're seeing long setup times and high contention for models on some L40S and H200 clusters.

Jun 3, 2026 · Resolved Jun 3, 2026

Resolvedmajor
ResolvedJun 3, 2026, 7:06 PM UTC

System is back to operating normally

InvestigatingJun 3, 2026, 7:06 PM UTC

System is back to operating normally

InvestigatingJun 3, 2026, 5:59 PM UTC

We're seeing long setup times and high contention for models on some L40S and H200 clusters.

Degraded performance on flux-2-klein-4b

May 28, 2026 · Resolved May 28, 2026

Resolvedminor
ResolvedMay 28, 2026, 2:17 PM UTC

This issue has been resolved and queue times are back to normal

InvestigatingMay 28, 2026, 12:30 PM UTC

Long queue times for black-forest-labs/flux-2-klein-4b resulting in canceled predictions

Prediction and Training status updates delayed

May 21, 2026 · Resolved May 21, 2026

Resolvedminor
ResolvedMay 21, 2026, 11:28 PM UTC

Message flows are healthy.

IdentifiedMay 21, 2026, 9:38 PM UTC

Our message queues for prediction and training status updates are hitting capacity limits which are causing connection failures for queue consumers. We are in the process of bringing additional capacity online.

Constrained H100 capacity

May 21, 2026 · Resolved May 21, 2026

Resolvedmajor
ResolvedMay 21, 2026, 10:05 PM UTC

H100 hardware contention has resolved. Thank you for your patience!

IdentifiedMay 21, 2026, 3:09 PM UTC

We are seeing heightened demand for H100 hardware which is causing severe queue delays.

Constrained capacity for H100 hardware

May 12, 2026 · Resolved May 12, 2026

Resolvedminor
ResolvedMay 12, 2026, 7:43 PM UTC

There is no more contention for H100 hardware. Thank you for your patience!

IdentifiedMay 12, 2026, 3:27 PM UTC

Demand for constrained H100 hardware is causing scaling delays. This impacts queue size and inference speed for any models running on H100s.

Degraded A100 hardware

Apr 19, 2026 · Resolved Apr 19, 2026

Resolvedminor
ResolvedApr 19, 2026, 3:02 PM UTC

All A100 capacity is back. Thanks for your patience!

MonitoringApr 19, 2026, 2:36 PM UTC

All predictions and trainings targeting A100 hardware are experiencing degraded performance while control plane nodes restart.

A100 capacity unavailable during storage maintenance

Apr 9, 2026 · Resolved Apr 9, 2026

Resolvedmajor
ResolvedApr 9, 2026, 5:30 PM UTC

The maintenance is complete and all systems are reporting healthy. Thank you for your patience!

InvestigatingApr 9, 2026, 4:43 PM UTC

The persistent storage for all A100 hardware is under maintenance and is expected to be degraded until completion.

Downstream errors for Black Forest Labs models

Mar 23, 2026 · Resolved Mar 24, 2026

Resolvedminor
ResolvedMar 24, 2026, 1:34 AM UTC

Black Forest Labs has resolved the issue.

IdentifiedMar 23, 2026, 3:57 PM UTC

Some Black Forest Labs models are failing due to downstream errors from BFL. We are monitoring the situation and working on work arounds. BFL Status page: [https://status.bfl.ml/](https://status.bfl.ml/ "https://status.bfl.ml/")

Degraded performance on Flux Schnell

Mar 10, 2026 · Resolved Mar 10, 2026

Resolvednone
ResolvedMar 10, 2026, 6:17 PM UTC

Flux Schnell requests are being served normally.

MonitoringMar 10, 2026, 1:11 PM UTC

A GPU provider has an outage. Traffic is being rerouted and we are processing new Flux Schnell requests.

InvestigatingMar 10, 2026, 12:56 PM UTC

We are investigating an outage that is only affecting Flux Schnell.

Model Predictions Stuck at "Starting"

Feb 20, 2026 · Resolved Feb 20, 2026

Resolvedmajor
ResolvedFeb 20, 2026, 1:12 PM UTC

Models are once again operational

MonitoringFeb 20, 2026, 12:59 PM UTC

We have identified the root cause and have made an update. We are continuing to monitor as we start to see things improve.

InvestigatingFeb 20, 2026, 12:13 PM UTC

We are currently investigating why a large number of models are not currently processing requests, with predictions stalled with a "starting" status.

Increased setup failures for T4 models

Jan 26, 2026 · Resolved Jan 26, 2026

Resolvednone
ResolvedJan 26, 2026, 3:16 PM UTC

Model setup failures have dropped to normal levels since 10:40 UTC. Thank you for your patience!

MonitoringJan 26, 2026, 10:10 AM UTC

We are seeing download issues from HuggingFace for a handful of T4 models, stopping new instances from starting. Requests for those models may be delayed as new capacity fails to come online. We will continue to monitor the situation.

InvestigatingJan 26, 2026, 8:53 AM UTC

The problem seems to be isolated to models downloading weights from HuggingFace.

InvestigatingJan 26, 2026, 7:33 AM UTC

We are seeing increased setup failure rates on models using T4 GPUs.

Predictions and training unavailable for multiple models

Jan 20, 2026 · Resolved Jan 21, 2026

Resolvedminor
ResolvedJan 21, 2026, 12:22 AM UTC

All infrastructure problems have been resolved and model predictions and training are available again for the relevant hardware types.

InvestigatingJan 20, 2026, 10:53 PM UTC

We are seeing indications that multiple models are not receiving prediction or training requests correctly.

Flux Schnell unavailable

Jan 20, 2026 · Resolved Jan 21, 2026

Resolvedminor
ResolvedJan 21, 2026, 12:07 AM UTC

The `black-forest-labs/flux-schnell` model is fully operational now. Thank you for your patience!

MonitoringJan 20, 2026, 9:39 PM UTC

The infrastructure failure preventing predictions from getting scheduled has been resolved for the past hour. We are monitoring performance and ensuring all adjacent systems are working well.

InvestigatingJan 20, 2026, 8:58 PM UTC

The black-forest-labs/flux-schnell model is currently unavailable. We are actively investigating.

Prediction Errors

Jan 15, 2026 · Resolved Jan 15, 2026

Resolvedmajor
ResolvedJan 15, 2026, 10:34 AM UTC

Incident is resolved.

InvestigatingJan 15, 2026, 10:33 AM UTC

All models have returned to normal processing timeframes. Thank you for your patience.

InvestigatingJan 15, 2026, 9:32 AM UTC

We have shifted traffic for the Flux Schnell model to a region with additional capacity. While we work through the backlog of requests, new requests should begin to be served within a normal timeframe. This continues to impact only the Flux Schnell official model.

InvestigatingJan 15, 2026, 8:34 AM UTC

We have identified and remediated a data store related to queues that was wedged, preventing traffic from reaching it within one of our regions. Most inference and API usage as returned to normal with the exception of the Flux Schnell model. We are working on restoring service to the flux schnell model. The impact is large delays on predictions and elevated errors when submitting preditions.

InvestigatingJan 15, 2026, 7:38 AM UTC

We are aware of elevated errors for predictions being sent to one of our regions. This is impacting an subset of models targeting H100 and L40S Hardware types. We are investigating and will provide an update as soon as information is available.

High demand for H100 hardware type

Dec 18, 2025 · Resolved Dec 20, 2025

Resolvedminor
ResolvedDec 20, 2025, 5:10 AM UTC

We are all clear now. Thank you for your patience!

MonitoringDec 18, 2025, 7:06 PM UTC

We are seeing high demand for the H100 hardware type, with increased delays for scale-out and cold boot events.

Frequently Asked Questions

Is Replicate down right now?

Check the status indicator at the top of this page. It pulls directly from Replicate's official status page. If Replicate is experiencing any issues, you'll see it reflected here. This real-time monitoring helps teams quickly identify whether performance problems are caused by Replicate infrastructure or their own systems.

What does this Replicate status page track?

This page tracks Replicate API, Hosted open models, and Inference jobs using data from Replicate's official status page. You can see current component health, active incidents, and a history of past issues. This visibility is crucial for teams building resilient AI applications that need to route around provider outages.

How often is Replicate status updated here?

We check Replicate's status page every 60 seconds to ensure you get near real-time status updates. How quickly issues show up here depends on how fast Replicate updates their own official status. For production systems that need instant failover, Bifrost can automatically detect and route around degraded providers.

Why monitor Replicate status?

Replicate is used to access a wide range of models, so availability issues can impact multiple AI features at once across image, audio, and text workflows. Real-time status monitoring enables proactive incident response and helps teams decide when to route traffic to alternative providers for maximum uptime.

What should I do when Replicate goes down?

When Replicate experiences an outage, the best practice is automatic failover to alternative AI providers. Bifrost is an open-source AI gateway that automatically detects Replicate degradation and routes LLM traffic to healthy alternatives like Stability AI, Hugging Face, keeping your application running with zero manual intervention. This intelligent routing ensures your users never experience downtime from a single provider's issues.

How can I prevent Replicate downtime from affecting my application?

Production AI applications should never depend on a single provider. Bifrost AI Gateway provides automatic multi-provider failover, intelligent load balancing, and health-based routing. When Replicate degrades, Bifrost instantly routes requests to operational alternatives while maintaining API compatibility. This architecture approach is used by teams running mission-critical AI features.

What are common causes of Replicate outages and degraded performance?

Common causes of Replicate issues include infrastructure scaling challenges, regional cloud provider problems, API gateway overload, and deployment errors. Monitor this page to stay informed, and consider implementing automatic failover with Bifrost to maintain uptime during Replicate incidents.

Alternative Providers for Automatic Failover

When Replicate experiences issues, Bifrost AI Gateway can automatically route your LLM traffic to these alternatives

💡 Build resilient AI apps: Configure Bifrost to automatically detect Replicate outages and route requests to healthy alternatives. This multi-provider approach ensures your application maintains uptime even when individual providers experience issues. Learn more about Bifrost AI Gateway

Replicate Models & Pricing

Replicate offers 40 models across 1 mode (40 chat)

40
Total Models
1
Modes
$0.03
Cheapest / 1M input
$15.00
Most expensive / 1M input