Get API key

How to Integrate an Unrestricted AI API for Uncensored Outputs

An unrestricted AI API allows developers to integrate large language models that do not apply standard content refusals, making them ideal for creative writing, roleplay, and adult-oriented applications. By using an OpenAI-compatible interface, you can swap in uncensored outputs without rewriting your existing infrastructure.

Updated

Key points

Why Use an Uncensored API?

Standard LLM providers often apply layered content filters that refuse benign but mature topics, leading to inconsistent behavior in creative or roleplay applications. An uncensored API removes these arbitrary refusals, allowing the model to generate full-length responses without early cutoffs or tone policing.

This approach is particularly valuable for developers building character-driven apps, adult fiction generators, or research tools that need raw model outputs. While some providers charge premium rates for high-quality reasoning models, uncensored APIs often prioritize accessibility and volume. However, you should note that 'uncensored' does not mean 'unlimited.' Providers still enforce hard limits on specific content types, such as sexual content involving minors, and may throttle abusive usage patterns.

When choosing an uncensored solution, consider whether you need raw text output or multimodal capabilities. Many uncensored services focus exclusively on text to keep latency low and costs predictable. If your application requires image generation or audio processing, you may need to integrate a separate service alongside your text API.

Understanding the OpenAI-Compatible Interface

Most modern LLM APIs follow a standardized REST structure. Using an OpenAI-compatible interface means your application can interact with the model using familiar endpoints like POST /v1/chat/completions. This compatibility reduces integration time significantly, as you likely already have code written for OpenAI's structure.

The interface typically supports streaming via Server-Sent Events (SSE), allowing you to display tokens as they are generated. This is critical for maintaining a responsive user experience, especially when the model is processing long contexts. You can also utilize standard parameters like temperature, top_p, and stop sequences to control generation behavior.

Unlike aggregators that route requests between multiple vendors, a direct uncensored API serves a single, tuned model. This eliminates the overhead of model selection and ensures consistent behavior. The model id is usually a simple string, such as "uncensored," making it easy to swap out in your configuration files.

Setting Up Your Unrestricted AI API Key

Accessing an uncensored API usually requires creating an account and generating an API key. The process is designed to be quick, often requiring only an email address or a Google login. Unlike traditional SaaS platforms, many uncensored services do not require a credit card for initial access, lowering the barrier to entry for experimentation.

Once you generate your key, it is typically displayed immediately. Store it securely, as it provides full access to your prepaid balance. Some providers limit you to one active key per account, replacing the previous key upon generation of a new one. This simplifies key rotation and security management.

You will need to configure your SDK or HTTP client to use the provider's base URL. For example, requests are sent to https://api.unrestrictedai.cc/v1. Ensure your client sends the API key in the Authorization header as a Bearer token. This setup works with the official OpenAI Python and Node.js SDKs, as well as generic HTTP clients like cURL.

Making Your First Request

A basic request involves sending a JSON payload with your system prompt, user message, and model identifier. The response will contain the generated text, token usage statistics, and completion details. If the request fails, the API returns a standard error code, allowing you to handle retries or inform the user appropriately.

Here is how you can make a simple request using cURL:

curl https://api.unrestrictedai.cc/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response will include the generated content in the choices[0].message.content field. You can also specify parameters like max_tokens to control the length of the output. The model supports a context window of 64,000 tokens, allowing for extensive conversations or long document processing.

Remember that the model is tuned for adult and unfiltered content, so it will respond to mature topics without unnecessary refusals. However, it does not offer embeddings, image generation, or fine-tuning capabilities. If your application requires these features, you will need to use a different service or integrate additional tools.

Handling Streaming and Tools

Streaming is essential for real-time applications. By setting stream: true, you receive a series of Server-Sent Events, each containing a fragment of the response. This allows you to display text to the user as it is generated, improving perceived performance.

Function calling, or tool use, is another powerful feature. You can define tools in your request, and the model will return structured JSON indicating which tool to call and with what arguments. This enables your application to perform actions like fetching weather data or updating a database based on the user's conversation.

Here is an example of streaming a response:

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Ensure your client handles the final chunk, which often contains the total token usage. This information is crucial for tracking costs and monitoring usage patterns. The API supports standard parameters like temperature and top_p to control the randomness and diversity of the output.

Pricing and Crypto Payments

Uncensored APIs often use a prepaid credit model, where you pay for tokens as you use them. This approach eliminates monthly subscription fees and allows you to control your spending precisely. Pricing is typically based on input and output tokens, with output tokens often costing more.

For example, the Unrestricted AI API charges $0.25 per 1 million input tokens and $1.00 per 1 million output tokens. Errors and refusals do not consume credit, ensuring you only pay for actual content generation. Credits never expire, so you can top up when needed without worrying about losing your balance.

Payments are accepted via cryptocurrency, specifically USDT on the TRC20 network or USDC on the Base network. You can top up with any whole amount between $10 and $500. Larger top-ups may include bonus credits, such as a 5% bonus for $50 or a 10% bonus for $100. This crypto-only model simplifies global payments and reduces transaction fees.

Rate Limits and Quotas

To ensure fair usage and maintain service stability, uncensored APIs implement rate limits. These limits prevent a single user from consuming excessive resources, which could impact other users on the same infrastructure.

Typical limits include requests per minute and concurrent requests per key. For instance, a single API key might be limited to 300 requests per minute and 8 concurrent requests. Request bodies are also capped, often at 8 MB, to prevent large payloads from causing delays.

If you exceed these limits, the API will return a 429 Too Many Requests error. You can handle this by implementing exponential backoff in your application. Monitoring your token usage and rate limit headers can help you optimize your request patterns and avoid disruptions.

Privacy and Data Usage

When using an uncensored API, consider how your data is handled. Many providers do not use your prompts for training their models, which is important for proprietary applications or sensitive content. However, data is still transmitted over the network and stored temporarily to generate the response.

Account creation usually requires minimal information, such as an email address. This reduces the amount of personal data you share with the provider. Some services offer trial credits, allowing you to test the API without immediate payment. These trials are often limited in duration and value, providing a good opportunity to evaluate the model's quality.

Refunds are typically not available for prepaid credits, as they are considered a service purchase rather than a physical good. If you encounter billing errors, such as a double charge, you can contact support for resolution. Always review the provider's privacy policy to understand how your data is stored and processed.

Questions and answers

What does 'uncensored' mean for an LLM?

It means the model is tuned to answer a wide range of topics, including adult or controversial subjects, without applying standard content refusals. However, it still enforces hard limits, such as blocking sexual content involving minors.

Does this API support image generation?

No, this API serves a text-only large language model. It does not offer image, audio, or video generation capabilities. You would need a separate service for multimodal tasks.

How do I pay for the API?

Payments are accepted via cryptocurrency only, specifically USDT (TRC20) or USDC (Base). You can top up with any whole amount from $10 to $500, and credits never expire.

Is there a free trial?

Yes, new accounts receive $0.50 of trial credit valid for 7 days. No credit card is required to start, making it easy to test the API before committing to a purchase.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.