Home/Errors
Errors and Rate Limits: A Practical Playbook for the Uncensored API
Production code spends more time on failure paths than on the happy path. This playbook lists each status code the Uncensored API returns, says which ones deserve a retry, and gives you client-side throttling for the 300 requests per minute limit in Python asyncio and Node, plus a calm way to surface a 402 to your own users.
Updated
Key points
- Retry only 429, 503 and transport failures; fix 400, 401, 403, 404 and 402 at the cause.
- Errors arrive as JSON with an error object holding code and message, so branch on status first, then code.
- A token bucket refilling 5 tokens per second keeps one key under 300 requests per minute without server pushback.
- A 402 means the prepaid balance is gone or the trial ended; treat it as a product state, not a crash.
The error envelope
Every failure returns a JSON body with a single top-level error object. Two fields matter: code, a short machine string, and message, text meant for logs and humans.
{
"error": {
"code": "no_credit",
"message": "Balance used up or trial expired."
}
}Decide in this order. Look at the HTTP status first, because it tells you the class of problem. Use code only to separate cases that share a status, and never parse message for logic, since wording can change. For the shape of successful responses, see the API reference.
Status codes, causes and actions
| Status | Code | Typical cause | Retry? | Action |
|---|---|---|---|---|
| 400 | bad request | Malformed JSON, or prompt tokens plus max_tokens above 100,000 | No | Fix the payload; trim history or lower max_tokens |
| 401 | invalid key | Missing header, typo, or a key that was regenerated | No | Reload the key from your secret store |
| 402 | no_credit | Balance used up or free trial expired | No | Top up the prepaid balance, then resume |
| 403 | content_blocked | Request touches sexual content involving minors | No | Do not resend; show your own policy message |
| 404 | not found | Wrong path, such as a missing /v1 prefix | No | Check the base URL |
| 429 | rate limit | More than 300 requests in a minute on one key | Yes | Back off with jitter; add client throttling |
| 503 | upstream_busy | Temporary capacity pressure | Yes | Retry after a few seconds |
The 403 deserves emphasis: it is a policy decision, not a transient fault. Retrying the identical request wastes quota and will return the same answer. The 400 case that surprises people most is the context sum. A 60,000-token prompt with max_tokens of 8,000 overflows by 4,000 and fails even though each number is valid alone.
A retry policy that behaves
Use a small, bounded policy and apply it only where retrying can help.
- Whitelist retryable failures: 429, 503, connection resets and timeouts. Everything else raises immediately.
- Exponential backoff with jitter: wait 1, 2, 4, 8 seconds plus a random fraction of a second, capped at 30. Jitter prevents many workers from retrying in lockstep.
- Cap attempts at five. Beyond that, an outage is more likely than bad luck, and queued work should fail visibly.
- Disable SDK auto-retries (
max_retries=0) when you run your own loop, otherwise attempts multiply and your throttle loses count. - Retry streams from the start. A stream that fails midway cannot be resumed; discard the partial text and re-issue the request, or keep it and ask for a continuation.
For a 503, the guidance is to wait a few seconds, so start your first delay near one or two seconds rather than hammering immediately. For a 429, the cure is slowing down, which the next section automates.
Staying under 300 requests per minute
The limit is 300 requests per minute per key, which is five per second on average. A token bucket enforces exactly that: it holds up to a small burst of tokens, refills at five per second, and makes each request take one token. If the bucket is empty the caller sleeps until a token exists.
The Python version uses asyncio. All twenty tasks share one bucket and a lock, so concurrent coroutines cannot overdraw it. Note that the bucket counts every attempt, retries included.
import asyncio, os, random, time
from openai import AsyncOpenAI, APIStatusError, APIConnectionError
class TokenBucket:
"""Allows `rate` acquisitions per `per` seconds, with a small burst."""
def __init__(self, rate=300, per=60.0, burst=10):
self.fill = rate / per # tokens added per second (5.0)
self.cap = burst
self.tokens = burst
self.stamp = time.monotonic()
self.lock = asyncio.Lock()
async def acquire(self):
async with self.lock:
while True:
now = time.monotonic()
self.tokens = min(self.cap, self.tokens + (now - self.stamp) * self.fill)
self.stamp = now
if self.tokens >= 1:
self.tokens -= 1
return
await asyncio.sleep((1 - self.tokens) / self.fill)
client = AsyncOpenAI(base_url="https://api.uncensoredapis.com/v1",
api_key=os.environ["API_KEY"], max_retries=0)
bucket = TokenBucket()
RETRYABLE = {429, 503}
async def ask(prompt, attempts=5):
for n in range(attempts):
await bucket.acquire()
try:
r = await client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": prompt}],
max_tokens=200,
)
return r.choices[0].message.content
except APIStatusError as e:
if e.status_code not in RETRYABLE or n == attempts - 1:
raise
except APIConnectionError:
if n == attempts - 1:
raise
await asyncio.sleep(min(30, 2 ** n) + random.uniform(0, 1))
async def main():
prompts = [f"Give one-line fact number {i} about tides." for i in range(1, 21)]
results = await asyncio.gather(*(ask(p) for p in prompts))
print(len(results), "answers; first:", results[0])
asyncio.run(main())Run it with API_KEY exported. Twenty requests fit inside the burst plus a few seconds of refill, so you will see it finish quickly; raise the count to 600 and the run stretches to roughly two minutes, which is the limit doing its job.
The Node version, using the official package in an ES module, chains acquisitions through a promise queue so callers are served in order.
import OpenAI from "openai";
class TokenBucket {
constructor(rate = 300, perMs = 60000, burst = 10) {
this.fillPerMs = rate / perMs;
this.cap = burst;
this.tokens = burst;
this.stamp = Date.now();
this.queue = Promise.resolve();
}
acquire() {
// Serialize callers so the bucket math never races.
this.queue = this.queue.then(async () => {
for (;;) {
const now = Date.now();
this.tokens = Math.min(this.cap, this.tokens + (now - this.stamp) * this.fillPerMs);
this.stamp = now;
if (this.tokens >= 1) { this.tokens -= 1; return; }
await new Promise((r) => setTimeout(r, Math.ceil((1 - this.tokens) / this.fillPerMs)));
}
});
return this.queue;
}
}
const client = new OpenAI({
baseURL: "https://api.uncensoredapis.com/v1",
apiKey: process.env.API_KEY,
maxRetries: 0,
});
const bucket = new TokenBucket();
async function ask(prompt, attempts = 5) {
for (let n = 0; n < attempts; n++) {
await bucket.acquire();
try {
const r = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: prompt }],
max_tokens: 200,
});
return r.choices[0].message.content;
} catch (err) {
const retryable = err.status === 429 || err.status === 503 || err.status === undefined;
if (!retryable || n === attempts - 1) throw err;
await new Promise((r) => setTimeout(r, Math.min(30000, 2 ** n * 1000) + Math.random() * 1000));
}
}
}
const prompts = Array.from({ length: 20 }, (_, i) => `Give one-line fact number ${i + 1} about tides.`);
const answers = await Promise.all(prompts.map((p) => ask(p)));
console.log(answers.length, "answers; first:", answers[0]);Size the burst conservatively. A burst of 10 allows a short spike while the sustained rate stays at five per second. If several processes share one key, divide the rate among them, for example 100 per minute each across three workers, or route all traffic through one queue.
Handling 402 gracefully in your product
A 402 with code no_credit means the prepaid balance reached zero or the free trial expired. It is the one error your end users should never see as a stack trace, and it usually signals something only you can fix: topping up.
Build three layers. First, detect it centrally in one wrapper so no call site has to remember. Second, translate it: the browser or mobile app receives a neutral state such as capacity_paused and shows a friendly screen, never the upstream message. Third, alert yourself, because every user is affected at once; send a notification to the account owner and flip a feature flag so the UI stops submitting requests.
# server side: translate the upstream 402 into a product state
from fastapi import FastAPI
from fastapi.responses import JSONResponse
from openai import AsyncOpenAI, APIStatusError
import os
app = FastAPI()
client = AsyncOpenAI(base_url="https://api.uncensoredapis.com/v1", api_key=os.environ["API_KEY"])
@app.post("/api/reply")
async def reply(body: dict):
try:
r = await client.chat.completions.create(
model="uncensored", messages=body["messages"], max_tokens=400)
return {"text": r.choices[0].message.content}
except APIStatusError as e:
if e.status_code == 402:
# Tell the client app to show a "temporarily unavailable" screen,
# and page yourself to top up the prepaid balance.
return JSONResponse({"state": "capacity_paused"}, status_code=503)
if e.status_code == 429 or e.status_code == 503:
return JSONResponse({"state": "busy_retry"}, status_code=503)
raiseAdd proactive protection as well. Sum usage.total_tokens per request into your own ledger and warn at a threshold you choose, so a top-up happens before a user is blocked. Remember that trial credit is $0.50 and lasts 7 days, so a prototype can hit 402 from expiry alone with money still unspent. Rates for planning are on the pricing page.
Once the balance is positive again, replay the work you queued while paused. Your wrapper should clear the paused flag on the first success.
What to log and alert on
Good records make incidents short. Log, per request: your own request id, the completion id when present, HTTP status, error code, attempt number, and token usage. Never log the Authorization header or full prompts that carry personal data.
Useful alerts, in priority order: any 402 (immediate), a sustained run of 401s (a regenerated key was not deployed everywhere), 429 rate above a few percent of calls (your throttle is mis-sized), and 503 clusters (back off wider, consider a queue). A dashboard of status counts per minute makes all four visible at a glance. Handled well, these codes become routine signals instead of outages.
Testing your failure paths before launch
Error handling that has never run is a guess. You can exercise most branches without waiting for a real outage, using cheap requests and your own test doubles.
- 401: call
GET /v1/modelswith a deliberately wrong key and confirm your app shows a configuration error to operators, not to customers. - 404: request
/v1/chat/completion(singular) once and confirm the log line names the path, so a typo in a base URL is obvious in minutes. - 400: send
max_tokensof 16,000 with a prompt padded to roughly 50,000 tokens. The sum passes 100,000, and you can watch the validation path fire. - 429 and 503: you cannot summon these on demand, so stub them. Wrap your HTTP client in a fake that raises the status for the first two calls, then succeeds, and assert that your loop sleeps, retries and returns the final answer.
- 402: stub it too, and confirm the UI shows the paused screen, the alert fires and nothing retries in a tight loop.
Write these as automated tests that run in continuous integration. A table-driven test with one row per status from the table above takes minutes to write and protects the policy from well-meaning refactors. Finally, run a short load test against your own throttle with the API stubbed out: fire 1,000 acquisitions and assert that no 60-second window contains more than 300 grants plus the burst. That proves the arithmetic before real traffic depends on it. When you later read the agent loop guide, reuse the same wrapper there, since agents multiply request counts and hit these limits first.
Questions and answers
Does a 429 count against my balance?
Check the usage figures in your own ledger rather than assuming. Either way, the fix is to slow down and retry with backoff.
Should I retry a 403 content_blocked response?
No. It is a policy block and an identical request returns the same result. Show your own message and move on.
My key suddenly returns 401. What changed?
Most often the key was regenerated, which invalidates the old one at once. Copy the new key from your account and redeploy it.
How many workers can share one key?
All of them share the 300 requests per minute limit. Split the rate across workers or funnel calls through one queue.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.