OpenAI-compatible uncensored LLM APIuncensoredapis.com

HomeDocs

Uncensored API Documentation: Quickstart Guide

Get started with the uncensored API by authenticating with your key and sending requests to the chat-completions endpoint. This guide covers the base URL, SDK integration, streaming, and limits for the uncensored model.

https://api.uncensoredapis.com/v1uncensored

Base URL and Authentication

Connect to the uncensored llm api using the official OpenAI-compatible base URL: https://api.uncensoredapis.com/v1. Authentication is handled via the Authorization header. You must include your API key as a bearer token. The key is generated immediately after you sign up on the Get API key page. Keep this key secure; it can be regenerated at any time, which revokes the old one. No card is needed to start, and trial credit is available for new accounts.

curl https://api.uncensoredapis.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Every request must target the /v1/chat/completions endpoint. This is a text-only API. There are no embeddings, image, audio, or video endpoints. The base URL remains constant regardless of your plan. Use standard HTTP headers for content type and authorization. This setup works with any OpenAI-compatible client by simply updating the base URL configuration.

Model ID and First Request

Specify the model ID as uncensored in your request body. This model is an open-weight large language model running on our own GPU servers, tuned to answer without content refusals for lawful adult use. It is not GPT, Claude, or any other vendor's model. The context window is 100,000 tokens, covering both prompt and completion. This large window supports high-volume nsfw llm api use cases requiring long context retention.

from openai import OpenAI

client = OpenAI(base_url="https://api.uncensoredapis.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

The API accepts standard chat-completion parameters. Send your messages in a list format. The response returns generated text. There is no model routing; we serve a single dedicated model. This ensures consistent behavior. Check the pricing page for token costs. The model handles controversial or adult topics without standard refusals, except for the hard limit on sexual content involving minors.

Streaming Responses (SSE)

Enable streaming by setting stream: true in your request. The endpoint supports Server-Sent Events (SSE) for real-time token delivery. This is useful for chat interfaces where you want to display text as it generates. The uncensored API does not return streaming data by default in a synchronous call; you must explicitly request it. Streaming works with the same model ID and authentication. It is efficient for nsfw ai api applications needing immediate partial results. Monitor the stream for completion to know when the response ends.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Handle partial chunks in your client code. The final chunk usually contains the full message or an indicator of completion. This approach reduces perceived latency. Ensure your client can handle incremental text updates. The API key and base URL remain the same. Rate limits apply to the total request volume, not just the streamed response size.

Tool/Function Calling Support

The API supports tool/function calling for structured outputs. Define tools in your request body with their names, descriptions, and parameter schemas. The model returns structured JSON data alongside text when it decides to use a tool. This is distinct from raw text responses. Use the tools parameter in the chat-completions endpoint. Pass the tool outputs back to the model in subsequent turns. This enables complex workflows without external parsing logic. The uncensored model handles function definitions reliably.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.uncensoredapis.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Structure your tool definitions according to the OpenAI schema. The API returns the tool call arguments in JSON format. Parse these arguments in your code to execute the desired function. Then, send the result back to the model. This cycle continues until the model generates a final text response. This feature is available on the same prepaid credit plan. No extra fees for tool usage.

Rate Limits and Body Size

Each API key is limited to 300 requests per minute. The maximum request body size is 8 MB. Exceeding these limits returns specific HTTP status codes. A 401 error indicates an invalid or missing API key. A 402 error means your prepaid credit balance is insufficient. A 429 error signals that you have exceeded the rate limit. Reduce your request frequency or wait for the window to reset. These limits apply per key, not per account. Regenerate your key to get a fresh start if needed.

The context window is 100,000 tokens. Ensure your combined prompt and completion fit within this limit. Large inputs may be truncated or rejected if they exceed the body size limit. Monitor your usage in the dashboard. The prepaid credit model means you pay only for what you use. No monthly fees or subscriptions. Credits never expire.

Pricing and Credits

The uncensored api uses a straightforward prepaid credit model. Input tokens cost $0.25 per 1M tokens. Output tokens cost $1.00 per 1M tokens. There are no monthly fees or subscriptions. Paid credits never expire. You can top up from $10 using crypto (USDT or USDC). Bonus credits are available: +5% for $50+ and +10% for $100+. New accounts receive $0.50 in trial credit for 7 days, no card needed. This model is ideal for high-volume nsfw ai api use cases.

Check your balance in the dashboard. Credits are deducted automatically with each request. The openrouter uncensored models alternative often uses different pricing; our direct model offers predictable costs. Use the API key to authenticate all requests. The base URL remains the same for all tiers. No hidden fees for streaming or tool calling. Track your usage to avoid unexpected charges.

What the API supports

A quick checklist for developers: format, limits, features, billing.

SpecValue
ProtocolOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
Modeluncensored
EndpointsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.uncensoredapis.com/v1
API keyAuthorization: Bearer YOUR_KEY
Function callingYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Completion length16,000 tokens max; 2,048 if max_tokens is not set
StreamingSupported (stream: true), usage included at the end
Context window100,000 tokens (prompt + completion together)
Structured outputJSON object mode via response_format json_object
Other parameterstemperature, top_p, stop, seed and the two penalties are passed through
Requests per minute300 requests per minute per key
Parallel requestsup to 8 in parallel per key
Request size8 MB request body
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Free trial$0.50 of credit valid 7 days, no card needed
Top-upUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
How you payprepaid credit, charged by real token usage; errors and refusals are free
Subscriptionno monthly fee; paid credit does not expire
Volume bonus+5% from $50, +10% from $100
Priceinput $0.25 / 1M tokens, output $1.00 / 1M tokens
Contentadult content allowed; sexual content involving minors is refused
Keysone key per account, regenerate any time (the old one stops working)
AccountGoogle or e-mail and password

When a request fails

Every error is JSON with a type you can switch on. You are never charged for an error.

CodeTypeMeaning
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busymodel busy — retry in a few seconds

Questions and answers

What is the context window size?

The context window is 100,000 tokens, covering both the prompt and the completion. This allows for long conversations or large document processing within a single request.

Does the uncensored API support streaming?

Yes, streaming is supported via Server-Sent Events (SSE). You must set the stream parameter to true in your request. This enables real-time token delivery for chat interfaces.

What happens if my credit runs out?

You will receive a 402 HTTP status code indicating no credit. Top up your account from the dashboard to resume usage. Credits never expire, so you can add funds whenever needed.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.