Integrate the Uncensored Local LLM API
Integrate the Uncensored Local LLM API by pointing your OpenAI-compatible client to our hosted endpoint. You get a 100k context window, streaming support, and raw model outputs without managing local GPU infrastructure.
- Base URL
- https://api.uncensoredlocalllm.com/v1
- Model
- uncensored
Base URL and Authentication
Our API is fully compatible with the OpenAI SDKs. You only need to update the base_url and provide your API key. Sign up on the Get API key page to receive your key immediately. No phone number or credit card is required to start with the trial credit. The base URL is https://api.uncensoredlocalllm.com/v1. Keep your key secure; you can regenerate it at any time, which instantly revokes the old one. This ensures you always have a single, stable endpoint for your applications without the complexity of local inference management.
First Request
Make your first API call using your preferred HTTP client. The endpoint supports standard chat-completions format. Below is a cURL example demonstrating a basic request. Replace YOUR_API_KEY with your actual key. This request sends a prompt and expects a text response. The model id is uncensored. This open-weight model is tuned to answer without content refusals for lawful adult use, providing direct outputs for your use case.
curl https://api.uncensoredlocalllm.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
The response will contain the generated text in the choices array. If you encounter a 401 error, your API key is invalid or missing. Ensure you are using the correct base URL and that your key has not been revoked after regeneration.
Python SDK Integration
Use the official OpenAI Python SDK to integrate quickly. Set the base_url and api_key in your client initialization. This approach allows you to use familiar methods like chat.completions.create(). The SDK handles JSON serialization and HTTP requests automatically. This is ideal for backend services or scripts that need to process large amounts of text. Remember that the model id remains uncensored regardless of the SDK version you use.
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredlocalllm.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
This method is efficient for batch processing or when you need to handle the response logic in Python. Ensure your environment has the openai package installed. The SDK supports both synchronous and asynchronous operations, fitting various architectural needs.
Node SDK Integration
For JavaScript or TypeScript environments, the OpenAI Node SDK provides a similar experience. Initialize the client with your base URL and API key. This is particularly useful for web applications or serverless functions. The SDK abstracts the HTTP details, allowing you to focus on application logic. You can call the completions endpoint with the same structure as other OpenAI-compatible services.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uncensoredlocalllm.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Use this integration for real-time web interfaces or Node.js microservices. The SDK supports promises and async/await patterns, making error handling straightforward. Ensure you handle potential network errors gracefully, especially in distributed systems.
Streaming Responses
Enable streaming by setting stream: true in your request. This allows you to receive partial responses as they are generated, improving perceived latency for users. The API supports Server-Sent Events (SSE) for streaming. This is crucial for chat interfaces where immediate feedback is expected. Streaming does not affect the model's output quality or context window.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Handle the stream events in your client code. Each chunk contains a portion of the final response. Aggregate these chunks to reconstruct the full message. This approach works seamlessly with the OpenAI SDKs in both Python and Node.js.
Rate Limits, Errors, and Context Window
The API enforces a limit of 300 requests per minute per key. The maximum request body size is 8 MB. You will receive a 429 error if you exceed the rate limit. A 402 error indicates insufficient prepaid credit. Ensure you monitor your usage and top up your account via the dashboard. Credit does not expire, and you can pay with crypto (USDT or USDC).
The context window is 100,000 tokens, covering both prompt and completion. This allows for extensive conversations or large document processing. If you exceed the limit, the API will reject the request. The model id is uncensored, and it supports tool/function calling for extended capabilities.
Capabilities and limits
If your tool speaks the OpenAI API, these are the details that matter.
| Parameter | Details |
|---|---|
| Compatibility | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model ID | uncensored |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| API key | Authorization: Bearer YOUR_KEY |
| Base URL | https://api.uncensoredlocalllm.com/v1 |
| Function calling | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Context window | 100,000 tokens, input and output combined |
| Max output | up to 16,000 tokens per request (default 2,048) |
| JSON mode | JSON object mode via response_format json_object |
| SSE streaming | Yes — server-sent events; the last chunk carries token usage |
| Request size | up to 8 MB per request |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Concurrency | up to 8 in parallel per key |
| Rate limit | 300/min per key |
| Bonus credit | +5% from $50, +10% from $100 |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Trial credit | $0.50 for 7 days, no card |
| How you pay | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Subscription | no monthly fee; paid credit does not expire |
| Account | sign in with Google or with e-mail + password |
| Keys | one key per account, regenerate any time (the old one stops working) |
| Content policy | adult content allowed; sexual content involving minors is refused |
Error reference
Every error is JSON with a type you can switch on. You are never charged for an error.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | temporary overload, retry shortly |
Questions and answers
What is the context window size?
The API supports a 100,000 token context window, which includes both the input prompt and the output completion. This allows for extensive conversations or processing large documents in a single request.
How do I handle streaming responses?
Set <code>stream: true</code> in your request to receive Server-Sent Events (SSE). The OpenAI SDKs in Python and Node.js handle streaming automatically when this flag is enabled. You can then process each chunk as it arrives.
What happens if I run out of credit?
You will receive a <code>402</code> error indicating no credit is available. You can top up your account with as little as $10 using crypto (USDT or USDC). Credits never expire, and bonuses are applied for larger top-ups.
uncensoredlocalllm.com
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.