Home/Agents
Agents on an Uncensored LLM API: Tool Calling, Frameworks and Guardrails
An agent is a loop: the model asks for a function, your code runs it, and the result goes back into the conversation. Because the Uncensored LLM API speaks tool calling in the standard chat format, you can build that loop in forty lines, then plug the same endpoint into LangChain or LlamaIndex. This guide covers both, plus the checks that keep tool execution safe.
Updated
Key points
- Pass tools in the request, watch for finish_reason equal to tool_calls, run the functions, append tool messages, and repeat until the model answers in text.
- LangChain needs ChatOpenAI with base_url; LlamaIndex needs OpenAILike with api_base and the function-calling flag set.
- The model proposes calls; your code decides whether they run. Allowlist, validate, time-limit and cap the loop.
- Agents multiply request counts, so budget for the 300 requests per minute limit and the 100,000-token context.
Anatomy of the tool-calling loop
Strip away frameworks and every agent does the same four things. You send the conversation plus a tools list. The model replies either with text, which means it is done, or with tool_calls, which means it wants something executed. You run each call and append one tool message per call, quoting the call's id. Then you ask the model again with the longer conversation.
The signal for the branch is finish_reason. When it equals tool_calls, the assistant message has content set to null and a tool_calls array, each entry holding an id, a function name and an arguments string of JSON. The reference page shows the exact shape. Two details trip people up: the arguments are a string that must be parsed, and the assistant message containing the calls must stay in the history before your tool results, in order.
A complete loop from scratch
The script below defines one function, describes it in JSON schema, and loops until the model produces plain text. It uses the official Python SDK pointed at the API. Every pass appends the assistant message, then a tool message per call, and a step counter stops runaway behavior.
import json, os
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredapis.com/v1", api_key=os.environ["API_KEY"])
TICKETS = {"T-101": "open", "T-102": "closed"}
def ticket_status(ticket_id: str) -> str:
return TICKETS.get(ticket_id, "unknown ticket")
TOOLS = [{
"type": "function",
"function": {
"name": "ticket_status",
"description": "Look up the status of a support ticket by id.",
"parameters": {
"type": "object",
"properties": {"ticket_id": {"type": "string"}},
"required": ["ticket_id"],
},
},
}]
REGISTRY = {"ticket_status": ticket_status}
messages = [{"role": "user", "content": "Are tickets T-101 and T-102 still open?"}]
for step in range(6): # hard cap on loop turns
r = client.chat.completions.create(
model="uncensored", messages=messages, tools=TOOLS, max_tokens=400)
msg = r.choices[0].message
messages.append(msg.model_dump(exclude_none=True))
if r.choices[0].finish_reason != "tool_calls":
print(msg.content)
break
for call in msg.tool_calls:
fn = REGISTRY.get(call.function.name)
try:
args = json.loads(call.function.arguments)
result = fn(**args) if fn else f"error: unknown tool {call.function.name}"
except Exception as exc:
result = f"error: {exc}"
messages.append({"role": "tool", "tool_call_id": call.id, "content": str(result)})
else:
print("stopped: step limit reached")Notice what the loop does on failure: errors become tool results as strings. That lets the model read the problem and try again, for example correcting a bad ticket id, instead of your program crashing. The for ... else clause fires only if the cap is hit without a final answer, which is how you detect a model stuck in circles.
Keep prompts focused. A short system message that states when to call a tool and when to answer directly reduces needless calls, and each skipped call saves a round trip, tokens and one slot of your rate limit.
Guarding tool execution
Model output is untrusted input. A tool name or argument chosen by the model may be wrong, odd, or influenced by text the model read from a web page or a user. Treat every call like a request arriving from the public internet.
- Allowlist by name. Look the function up in a registry you built. Never use the string to import or evaluate code.
- Separate read from write. Read-only tools run freely. Anything that sends mail, spends money, deletes data or touches files needs explicit human approval.
- Validate arguments against the schema you advertised: required keys, no extra keys, types and length limits.
- Bound time and output. Apply timeouts, and truncate results before feeding them back so a huge tool output cannot eat your 100,000-token context.
- Cap iterations and cost. Limit steps per task and tokens per session, using the
usageblock. - Log every call with name, arguments and outcome so you can audit behavior.
import json
from concurrent.futures import ThreadPoolExecutor, TimeoutError as FutTimeout
ALLOWED = {"ticket_status", "lookup_invoice"} # read-only tools only
NEEDS_APPROVAL = {"refund_order"} # side effects wait for a human
POOL = ThreadPoolExecutor(max_workers=4)
def validate(args: dict, schema: dict) -> None:
props = schema["properties"]
for key in schema.get("required", []):
if key not in args:
raise ValueError(f"missing argument: {key}")
for key, val in args.items():
if key not in props:
raise ValueError(f"unexpected argument: {key}")
if props[key]["type"] == "string" and not isinstance(val, str):
raise ValueError(f"{key} must be a string")
if isinstance(val, str) and len(val) > 200:
raise ValueError(f"{key} too long")
def run_tool(name, raw_args, registry, schemas, approve=lambda n, a: False):
if name in NEEDS_APPROVAL:
args = json.loads(raw_args)
if not approve(name, args):
return "denied: this action needs human approval"
elif name not in ALLOWED:
return f"error: tool {name} is not permitted"
args = json.loads(raw_args)
validate(args, schemas[name]["parameters"])
future = POOL.submit(registry[name], **args)
try:
return str(future.result(timeout=5))[:2000] # time and size limits
except FutTimeout:
return "error: tool timed out"Treat tool results as untrusted too. If a tool returns scraped text containing instructions, the model may follow them. Wrap such content in clear delimiters in the tool message and state in your system prompt that tool output is data, not commands.
Writing tool schemas the model can use
The schema is the only documentation the model sees, so it deserves the same care as a public API. Name functions with verbs and nouns that match the task, such as ticket_status or lookup_invoice, rather than generic labels like run or helper. Write the description as an instruction: what the function does, what it returns, and when not to use it.
- Keep parameters few and flat. Two or three simple fields are called correctly far more often than a nested object with ten.
- Use enums for closed sets, for instance
"enum": ["open", "closed"], so the model cannot invent a value. - Describe each property with a unit or example: "Ticket id such as T-101" beats a bare string.
- Return compact results. Strings or small JSON objects work best; long dumps burn context for little gain.
- Prefer several narrow tools over one omnibus tool with a mode flag.
When a call fails validation, return the validation message as the tool result rather than raising. The model usually repairs the call on the next turn, since it can read what was wrong. If it fails the same way twice, rewrite the description, not the loop.
LangChain configuration
LangChain's ChatOpenAI class accepts a base_url, so no custom provider package is needed. Set model to uncensored, pass the key from your environment, and bind tools with bind_tools. Install langchain-openai and langchain-core.
import os
from langchain_openai import ChatOpenAI
from langchain_core.tools import tool
@tool
def ticket_status(ticket_id: str) -> str:
"""Return the status of a support ticket."""
return {"T-101": "open", "T-102": "closed"}.get(ticket_id, "unknown ticket")
llm = ChatOpenAI(
model="uncensored",
base_url="https://api.uncensoredapis.com/v1",
api_key=os.environ["API_KEY"],
temperature=0.3,
max_tokens=400,
)
llm_with_tools = llm.bind_tools([ticket_status])
ai = llm_with_tools.invoke("What is the status of ticket T-101?")
for call in ai.tool_calls:
print(call["name"], call["args"], "->", ticket_status.invoke(call["args"]))The result of invoke is an AIMessage whose tool_calls list holds parsed names and arguments, so you skip the manual JSON step. To get a full autonomous loop, hand the same llm object to LangChain's agent constructors; the endpoint only has to honor tool calling, which it does. If an agent seems to ignore its tools, check that the docstring and type hints are present, since they become the schema the model sees.
LlamaIndex configuration
LlamaIndex ships OpenAILike for servers that follow the chat completions format. Three settings matter. api_base points at the API, is_chat_model=True selects the chat route, and is_function_calling_model=True tells the framework it may send tools. Set context_window to 100000 so its text splitters and memory buffers size themselves correctly. Install llama-index-llms-openai-like.
import os
from llama_index.llms.openai_like import OpenAILike
from llama_index.core.tools import FunctionTool
def ticket_status(ticket_id: str) -> str:
"""Return the status of a support ticket."""
return {"T-101": "open", "T-102": "closed"}.get(ticket_id, "unknown ticket")
llm = OpenAILike(
model="uncensored",
api_base="https://api.uncensoredapis.com/v1",
api_key=os.environ["API_KEY"],
is_chat_model=True,
is_function_calling_model=True,
context_window=100000,
max_tokens=400,
)
tool = FunctionTool.from_defaults(fn=ticket_status)
response = llm.predict_and_call([tool], "Check ticket T-102 for me.")
print(response)From here you can pass the same llm to an agent worker or a query engine. Remember that retrieval pipelines can stuff many chunks into one prompt; with a 100,000-token window shared between prompt and completion, cap the number of retrieved chunks and keep max_tokens modest.
Operating agents on a budget
An agent turns one user message into several requests, so plan capacity. A task with five tool rounds costs six model calls. Ten simultaneous tasks therefore produce sixty requests in a short span, and the 300 requests per minute ceiling arrives sooner than a chat app would reach it. Put the token-bucket throttle from the errors and rate limits playbook in front of your client.
Context grows with every round because each tool result is appended. Summarize or drop old tool outputs once they have been used, and keep an eye on usage.prompt_tokens; when it approaches 50,000, compress history. For cost, assume an illustrative task of 6 calls averaging 3,000 prompt tokens and 200 completion tokens: that is 18,000 input tokens (about $0.0045 at $0.25 per million) and 1,200 output tokens (about $0.0012 at $1.00 per million). Those token counts are assumptions for the sums, not measurements. If you are new to the endpoint, start with the docs and the free trial credit.
Testing and debugging an agent
Agents fail in ways unit tests on a single function miss, so test the conversation. Record the full message list for every run, including tool arguments and results, and replay failures from the log. Three habits pay off quickly.
- Fake the model. Stub the client so it returns scripted
tool_callsmessages. This tests your registry, validation and loop termination with no network and no spend. - Pin sampling. Use a low
temperature, such as 0.2, for tool-driven tasks. Lower randomness gives more repeatable argument JSON, and you can raise it later for writing-heavy steps. - Assert on trajectories. For a handful of fixed prompts, check which tools were called and in what order, not only the final text.
Watch finish_reason as well. A value of length in the middle of a tool call means the arguments were cut off; raise max_tokens (up to 16,000) and the JSON will complete. Streaming works with tools too, but accumulate argument fragments per call index and parse only after the stream closes. For a refresher on those chunks, see the streaming section of the reference.
Questions and answers
Does the API support parallel tool calls in one reply?
The message carries a tool_calls array, so handle every entry in it and send one tool message per call id.
Why does my agent loop forever?
Usually the tool errors are not informative or the system prompt never says when to stop. Add a step cap and return clear error strings.
Do I need LangChain or LlamaIndex?
No. A plain loop with the OpenAI SDK is enough, and frameworks only add orchestration on top of the same endpoint.
How do I stop the model running destructive tools?
Never register them without an approval step. Only the functions in your registry can run, so keep write actions behind human confirmation.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.