Get API key

uncensoredlocalllm.comDocs

Integrate the Uncensored Local LLM API

Integrate the Uncensored Local LLM API by pointing your OpenAI-compatible client to our hosted endpoint. You get a 100k context window, streaming support, and raw model outputs without managing local GPU infrastructure.

Base URL
https://api.uncensoredlocalllm.com/v1
Model
uncensored

Base URL and Authentication

Our API is fully compatible with the OpenAI SDKs. You only need to update the base_url and provide your API key. Sign up on the Get API key page to receive your key immediately. No phone number or credit card is required to start with the trial credit. The base URL is https://api.uncensoredlocalllm.com/v1. Keep your key secure; you can regenerate it at any time, which instantly revokes the old one. This ensures you always have a single, stable endpoint for your applications without the complexity of local inference management.

First Request

Make your first API call using your preferred HTTP client. The endpoint supports standard chat-completions format. Below is a cURL example demonstrating a basic request. Replace YOUR_API_KEY with your actual key. This request sends a prompt and expects a text response. The model id is uncensored. This open-weight model is tuned to answer without content refusals for lawful adult use, providing direct outputs for your use case.

curl https://api.uncensoredlocalllm.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response will contain the generated text in the choices array. If you encounter a 401 error, your API key is invalid or missing. Ensure you are using the correct base URL and that your key has not been revoked after regeneration.

Python SDK Integration

Use the official OpenAI Python SDK to integrate quickly. Set the base_url and api_key in your client initialization. This approach allows you to use familiar methods like chat.completions.create(). The SDK handles JSON serialization and HTTP requests automatically. This is ideal for backend services or scripts that need to process large amounts of text. Remember that the model id remains uncensored regardless of the SDK version you use.

from openai import OpenAI

client = OpenAI(base_url="https://api.uncensoredlocalllm.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

This method is efficient for batch processing or when you need to handle the response logic in Python. Ensure your environment has the openai package installed. The SDK supports both synchronous and asynchronous operations, fitting various architectural needs.

Node SDK Integration

For JavaScript or TypeScript environments, the OpenAI Node SDK provides a similar experience. Initialize the client with your base URL and API key. This is particularly useful for web applications or serverless functions. The SDK abstracts the HTTP details, allowing you to focus on application logic. You can call the completions endpoint with the same structure as other OpenAI-compatible services.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.uncensoredlocalllm.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Use this integration for real-time web interfaces or Node.js microservices. The SDK supports promises and async/await patterns, making error handling straightforward. Ensure you handle potential network errors gracefully, especially in distributed systems.

Streaming Responses

Enable streaming by setting stream: true in your request. This allows you to receive partial responses as they are generated, improving perceived latency for users. The API supports Server-Sent Events (SSE) for streaming. This is crucial for chat interfaces where immediate feedback is expected. Streaming does not affect the model's output quality or context window.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Handle the stream events in your client code. Each chunk contains a portion of the final response. Aggregate these chunks to reconstruct the full message. This approach works seamlessly with the OpenAI SDKs in both Python and Node.js.

Rate Limits, Errors, and Context Window

The API enforces a limit of 300 requests per minute per key. The maximum request body size is 8 MB. You will receive a 429 error if you exceed the rate limit. A 402 error indicates insufficient prepaid credit. Ensure you monitor your usage and top up your account via the dashboard. Credit does not expire, and you can pay with crypto (USDT or USDC).

The context window is 100,000 tokens, covering both prompt and completion. This allows for extensive conversations or large document processing. If you exceed the limit, the API will reject the request. The model id is uncensored, and it supports tool/function calling for extended capabilities.

Capabilities and limits

If your tool speaks the OpenAI API, these are the details that matter.

ParameterDetails
CompatibilityOpenAI Chat Completions schema; official openai SDKs work unchanged
Model IDuncensored
EndpointsPOST /v1/chat/completions · GET /v1/models
API keyAuthorization: Bearer YOUR_KEY
Base URLhttps://api.uncensoredlocalllm.com/v1
Function callingSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Context window100,000 tokens, input and output combined
Max outputup to 16,000 tokens per request (default 2,048)
JSON modeJSON object mode via response_format json_object
SSE streamingYes — server-sent events; the last chunk carries token usage
Request sizeup to 8 MB per request
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Concurrencyup to 8 in parallel per key
Rate limit300/min per key
Bonus credit+5% from $50, +10% from $100
Top-upcrypto: USDT on TRON or USDC on Base, $10–$500, any whole sum
Token prices$0.25 per 1M input tokens · $1.00 per 1M output tokens
Trial credit$0.50 for 7 days, no card
How you paypay as you go from prepaid credit; nothing is charged for failed or refused requests
Subscriptionno monthly fee; paid credit does not expire
Accountsign in with Google or with e-mail + password
Keysone key per account, regenerate any time (the old one stops working)
Content policyadult content allowed; sexual content involving minors is refused

Error reference

Every error is JSON with a type you can switch on. You are never charged for an error.

StatusTypeReason
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedno key, wrong key, or a key replaced by a newer one
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busytemporary overload, retry shortly

Questions and answers

What is the context window size?

The API supports a 100,000 token context window, which includes both the input prompt and the output completion. This allows for extensive conversations or processing large documents in a single request.

How do I handle streaming responses?

Set <code>stream: true</code> in your request to receive Server-Sent Events (SSE). The OpenAI SDKs in Python and Node.js handle streaming automatically when this flag is enabled. You can then process each chunk as it arrives.

What happens if I run out of credit?

You will receive a <code>402</code> error indicating no credit is available. You can top up your account with as little as $10 using crypto (USDT or USDC). Credits never expire, and bonuses are applied for larger top-ups.

uncensoredlocalllm.com

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.