Home/Docs
Uncensored API Documentation: Quickstart Guide
Get started with the uncensored API by authenticating with your key and sending requests to the chat-completions endpoint. This guide covers the base URL, SDK integration, streaming, and limits for the uncensored model.
https://api.uncensoredapis.com/v1uncensored
Base URL and Authentication
Connect to the uncensored llm api using the official OpenAI-compatible base URL: https://api.uncensoredapis.com/v1. Authentication is handled via the Authorization header. You must include your API key as a bearer token. The key is generated immediately after you sign up on the Get API key page. Keep this key secure; it can be regenerated at any time, which revokes the old one. No card is needed to start, and trial credit is available for new accounts.
curl https://api.uncensoredapis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'Every request must target the /v1/chat/completions endpoint. This is a text-only API. There are no embeddings, image, audio, or video endpoints. The base URL remains constant regardless of your plan. Use standard HTTP headers for content type and authorization. This setup works with any OpenAI-compatible client by simply updating the base URL configuration.
Model ID and First Request
Specify the model ID as uncensored in your request body. This model is an open-weight large language model running on our own GPU servers, tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or any other vendor's model. The context window is 100,000 tokens, covering both prompt and completion. This large window supports high-volume nsfw llm api use cases requiring long context retention.
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredapis.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)The API accepts standard chat-completion parameters. Send your messages in a list format. The response returns generated text. There is no model routing; we serve a single dedicated model. This ensures consistent behavior. Check the pricing page for token costs. The model handles controversial or adult topics without standard refusals, except for the hard limit on sexual content involving minors.
Streaming Responses (SSE)
Enable streaming by setting stream: true in your request. The endpoint supports Server-Sent Events (SSE) for real-time token delivery. This is useful for chat interfaces where you want to display text as it generates. The uncensored API does not return streaming data by default in a synchronous call; you must explicitly request it. Streaming works with the same model ID and authentication. It is efficient for nsfw ai api applications needing immediate partial results. Monitor the stream for completion to know when the response ends.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Handle partial chunks in your client code. The final chunk usually contains the full message or an indicator of completion. This approach reduces perceived latency. Ensure your client can handle incremental text updates. The API key and base URL remain the same. Rate limits apply to the total request volume, not just the streamed response size.
Tool/Function Calling Support
The API supports tool/function calling for structured outputs. Define tools in your request body with their names, descriptions, and parameter schemas. The model returns structured JSON data alongside text when it decides to use a tool. This is distinct from raw text responses. Use the tools parameter in the chat-completions endpoint. Pass the tool outputs back to the model in subsequent turns. This enables complex workflows without external parsing logic. The uncensored model handles function definitions reliably.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uncensoredapis.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Structure your tool definitions according to the OpenAI schema. The API returns the tool call arguments in JSON format. Parse these arguments in your code to execute the desired function. Then, send the result back to the model. This cycle continues until the model generates a final text response. This feature is available on the same prepaid credit plan. No extra fees for tool usage.
Rate Limits and Body Size
Each API key is limited to 300 requests per minute. The maximum request body size is 8 MB. Exceeding these limits returns specific HTTP status codes. A 401 error indicates an invalid or missing API key. A 402 error means your prepaid credit balance is insufficient. A 429 error signals that you have exceeded the rate limit. Reduce your request frequency or wait for the window to reset. These limits apply per key, not per account. Regenerate your key to get a fresh start if needed.
The context window is 100,000 tokens. Ensure your combined prompt and completion fit within this limit. Large inputs may be truncated or rejected if they exceed the body size limit. Monitor your usage in the dashboard. The prepaid credit model means you pay only for what you use. No monthly fees or subscriptions. Credits never expire.
Pricing and Credits
The uncensored api uses a straightforward prepaid credit model. Input tokens cost $0.25 per 1M tokens. Output tokens cost $1.00 per 1M tokens. There are no monthly fees or subscriptions. Paid credits never expire. You can top up from $10 using crypto (USDT or USDC). Bonus credits are available: +5% for $50+ and +10% for $100+. New accounts receive $0.50 in trial credit for 7 days, no card needed. This model is ideal for high-volume nsfw ai api use cases.
Check your balance in the dashboard. Credits are deducted automatically with each request. The openrouter uncensored models alternative often uses different pricing; our direct model offers predictable costs. Use the API key to authenticate all requests. The base URL remains the same for all tiers. No hidden fees for streaming or tool calling. Track your usage to avoid unexpected charges.
What the API supports
A quick checklist for developers: format, limits, features, billing.
| Spec | Value |
|---|---|
| Protocol | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Model | uncensored |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.uncensoredapis.com/v1 |
| API key | Authorization: Bearer YOUR_KEY |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Streaming | Supported (stream: true), usage included at the end |
| Context window | 100,000 tokens (prompt + completion together) |
| Structured output | JSON object mode via response_format json_object |
| Other parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Requests per minute | 300 requests per minute per key |
| Parallel requests | up to 8 in parallel per key |
| Request size | 8 MB request body |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Subscription | no monthly fee; paid credit does not expire |
| Volume bonus | +5% from $50, +10% from $100 |
| Price | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Content | adult content allowed; sexual content involving minors is refused |
| Keys | one key per account, regenerate any time (the old one stops working) |
| Account | Google or e-mail and password |
When a request fails
Every error is JSON with a type you can switch on. You are never charged for an error.
| Code | Type | Meaning |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What is the context window size?
The context window is 100,000 tokens, covering both the prompt and the completion. This allows for long conversations or large document processing within a single request.
Does the uncensored API support streaming?
Yes, streaming is supported via Server-Sent Events (SSE). You must set the stream parameter to true in your request. This enables real-time token delivery for chat interfaces.
What happens if my credit runs out?
You will receive a 402 HTTP status code indicating no credit. Top up your account from the dashboard to resume usage. Credits never expire, so you can add funds whenever needed.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.