Home/API Reference
Uncensored API Reference, Explained Field by Field
This page documents every field you will send to or receive from the Uncensored API: the chat completions request body, the response object, finish_reason values, the usage block, streaming chunks and the models list. Each section pairs a short specification with an annotated example you can paste into a terminal and run.
Updated
Key points
- Two endpoints exist: POST /v1/chat/completions and GET /v1/models. Both authenticate with a Bearer key.
- The response follows the familiar chat completion shape: id, choices[].message, finish_reason and usage.
- Streaming uses server-sent events; a final chunk carries usage automatically and the stream ends with data: [DONE].
- Limits that shape requests: 100,000-token context, 2048 default and 16,000 maximum max_tokens, 8 MB body.
Conventions and transport
The base URL is https://api.uncensoredapis.com/v1. All bodies are JSON and UTF-8. Every request carries Authorization: Bearer <key>; a missing or wrong key returns 401. Only two routes exist, and any other path returns 404.
| Method | Path | Purpose |
|---|---|---|
| POST | /v1/chat/completions | Generate an assistant message, optionally streamed |
| GET | /v1/models | List the model id you can request |
The API is text only. There are no embeddings, image, audio or video routes and no fine-tuning, and there is exactly one model, addressed as uncensored. Because the wire format is compatible with the chat completions convention, the official Python and Node SDKs work when you override base_url. The docs cover first-run setup; this page is the field-level companion.
Request body fields
The table lists the fields the endpoint understands. Sampling fields are passed through to generation unchanged, so the values you pick behave as they would in any chat-completions-style API.
| Field | Type | Notes |
|---|---|---|
model | string | Always "uncensored". |
messages | array | Ordered objects with role (system, user, assistant, tool) and content. |
max_tokens | integer | Completion cap. Default 2048, maximum 16,000. |
temperature | number | Randomness. Lower is more deterministic. |
top_p | number | Nucleus sampling cutoff. |
stop | string or array | Sequences that end generation. |
stream | boolean | true switches the reply to server-sent events. |
tools | array | Function definitions in the standard JSON-schema format. |
Here is a complete annotated request. The system message sets behavior, the user message carries the task, and stop ends generation at the first blank line.
curl https://api.uncensoredapis.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [
{"role": "system", "content": "You are a terse release-notes writer."},
{"role": "user", "content": "Summarize: fixed login bug, added dark mode."}
],
"max_tokens": 120,
"temperature": 0.4,
"stop": ["\n\n"]
}'Two budget rules matter. First, the prompt plus the completion must fit in 100,000 tokens, so a request whose prompt tokens plus max_tokens exceed that is rejected with a 400. Second, the serialized body may not exceed 8 MB. If you send long histories, trim old turns on your side before they approach either limit.
The response object
A non-streaming call returns one JSON object. Read it top to bottom as follows.
{
"id": "chatcmpl-example123",
"object": "chat.completion",
"created": 1767225600,
"model": "uncensored",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "v1.4: Login bug fixed. Dark mode added."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 38,
"completion_tokens": 14,
"total_tokens": 52
}
}- id: a unique identifier for the completion. Log it next to your own request id when you need to correlate events.
- object: the constant
chat.completion. - created: Unix timestamp in seconds.
- model: echoes
uncensored. - choices: an array. You receive one entry, so
choices[0]is the answer.indexis its position. - choices[].message:
roleisassistant;contentholds the text, or isnullwhen the model returned tool calls instead. - choices[].finish_reason: why generation stopped; see the next section.
- usage: token accounting for billing, covered below.
finish_reason values
Branch on finish_reason before you trust the content. A truncated answer and a complete answer look identical until you check it.
| Value | Meaning | What your code should do |
|---|---|---|
stop | The model reached a natural end or hit one of your stop sequences. | Use the content as is. |
length | The max_tokens cap (or the context window) was reached mid-answer. | Raise max_tokens, shorten the prompt, or ask the model to continue. |
tool_calls | The model wants one or more functions run. | Execute them and send results back as tool messages. |
When the value is tool_calls, the message looks like this, with arguments delivered as a JSON string that you must parse yourself:
{
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [{
"id": "call_abc1",
"type": "function",
"function": {"name": "get_order", "arguments": "{\"order_id\": 8841}"}
}]
},
"finish_reason": "tool_calls"
}]
}Always wrap json.loads or JSON.parse on arguments in error handling, because generated JSON can occasionally be malformed.
The usage block and what it costs
usage reports prompt_tokens, completion_tokens and total_tokens. These are the numbers your balance is charged against: input at $0.25 per million tokens and output at $1.00 per million. Prepaid balance is deducted per request, with no subscription.
A worked example under stated assumptions: suppose a request uses 2,000 prompt tokens and 500 completion tokens. Input costs 2,000 / 1,000,000 x $0.25 = $0.0005. Output costs 500 / 1,000,000 x $1.00 = $0.0005. The request totals $0.0010. These token counts are assumptions for arithmetic, not a measurement of any real workload. The pricing page has the rates in one place.
Since output tokens cost four times as much as input tokens, a tight max_tokens is your most direct cost control.
Streaming chunk format
Set "stream": true and the response becomes text/event-stream. Each event is a line starting with data: followed by JSON, separated by a blank line. The object type changes to chat.completion.chunk and message is replaced by delta.
data: {"id":"chatcmpl-example123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-example123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"v1.4: "},"finish_reason":null}]}
data: {"id":"chatcmpl-example123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Login bug fixed."},"finish_reason":null}]}
data: {"id":"chatcmpl-example123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-example123","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":38,"completion_tokens":14,"total_tokens":52}}
data: [DONE]- The first chunk usually sets
delta.roletoassistantwith empty content. - Content chunks carry a fragment in
delta.content. Concatenate them in arrival order. - The closing content chunk has an empty
deltaand a non-nullfinish_reason. - A final chunk with an empty
choicesarray carriesusage. You do not request it; it is added automatically. - The literal
data: [DONE]terminates the stream.
Guard against that empty choices array: indexing choices[0] on the usage chunk raises an error in naive loops. The SDK example below handles it.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredapis.com/v1", api_key=os.environ["API_KEY"])
reason, usage = None, None
for chunk in client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Name three uses of a rubber duck in debugging."}],
stream=True,
):
if chunk.choices:
piece = chunk.choices[0].delta.content
if piece:
print(piece, end="", flush=True)
if chunk.choices[0].finish_reason:
reason = chunk.choices[0].finish_reason
if getattr(chunk, "usage", None):
usage = chunk.usage
print("\nfinish:", reason, "| tokens:", usage.total_tokens if usage else "n/a")For streamed tool calls, delta.tool_calls arrives in fragments; accumulate the arguments string per call index and parse once finish_reason is tool_calls.
The models endpoint
GET /v1/models needs only the Bearer header and returns a list object. Because there is one model, data holds one entry, uncensored. Clients that populate a model dropdown from this route will therefore show a single choice.
curl https://api.uncensoredapis.com/v1/models -H "Authorization: Bearer $API_KEY"
{
"object": "list",
"data": [
{"id": "uncensored", "object": "model"}
]
}It is also the cheapest way to verify a key: a 200 means the key is valid, a 401 means it is wrong or was regenerated. Calling this route does not generate tokens. For status handling and retry rules, see the errors and rate limits playbook.
Reference checklist for client authors
If you are writing a client library or wrapping the endpoint in a gateway, verify each of these behaviors against the specification above before shipping.
- Header hygiene. Send the key only in the Authorization header, never in the query string, and never log the header value.
- Parse defensively. Treat unknown extra fields in responses as ignorable. Pin your parser to the fields documented here rather than to the exact JSON key order.
- Check finish_reason on every reply. Surface a visible indicator when it equals
length, because users otherwise read cut-off text as the full answer. - Budget before sending. Estimate prompt tokens, add your
max_tokens, and compare the sum with 100,000. Rejecting locally is faster than waiting for a 400. - Handle the stream tail. Expect a chunk with an empty
choiceslist and ausageobject, then[DONE]. Close the connection after the sentinel. - Record usage. Store
total_tokensper request so you can reconcile your own metering against the prepaid balance later. - Respect the ceiling. The account allows 300 requests per minute per key, so a gateway fanning out many users needs its own queue.
None of these steps depend on a vendor SDK. A plain HTTP client that sets the Bearer header and decodes JSON is sufficient, and the SDK examples on this page are conveniences, not requirements. Pair this checklist with the error table in the playbook and you cover the whole request lifecycle: build, send, stream, parse, account and recover.
Questions and answers
Can I get more than one choice per request?
Expect a single entry in choices and read choices[0]. If you want alternatives, issue separate requests.
Why is message.content null?
The model returned tool_calls instead of text. Check finish_reason, run the functions and send the results back.
What happens if prompt plus max_tokens passes 100,000?
The request is rejected with a 400 bad request error. Trim the prompt or lower max_tokens.
Does streaming change the price?
No. Billing uses the same prompt and completion token counts, and the final usage chunk reports them.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.