Get API key
Get API key

Quickstart: Integrate the AI Girlfriend API

Integrate our uncensored LLM API in minutes using standard OpenAI-compatible endpoints. Get your API key, set your base URL, and start generating text responses for your roleplay application.

per 1M input tokens
$0.25
Output tokens / 1M
$1.00
token context
100,000
trial credit
$0.50
requests per minute
300

Base URL and Authentication

Our API follows the OpenAI chat-completions standard, making integration straightforward for existing clients. You must include your API key in the Authorization header of every request. The base URL for all endpoints is https://api.roleplayapi.com/v1. When using official SDKs, update the base_url configuration to point to this address and provide your key as the api_key. This setup ensures your requests reach our dedicated uncensored endpoint without modification to your standard logic.

Send a Completion

Initiate a conversation by sending a POST request to /v1/chat/completions. The model identifier is uncensored. You can pass any number of messages in the messages array to establish context or start a fresh dialogue. The response contains the model's text output, which you can stream or parse as a JSON object. This endpoint handles the core text-in, text-out interaction required for character-driven apps.

curl https://api.roleplayapi.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

Use the official openai Python library to interact with the API with minimal boilerplate. Configure the client with our base URL and your API key. Then, call chat.completions.create with the uncensored model ID. This approach leverages existing type hints and error handling mechanisms, allowing you to focus on your application logic rather than HTTP parsing. The response object provides access to the generated content directly.

from openai import OpenAI

client = OpenAI(base_url="https://api.roleplayapi.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node SDK Integration

For JavaScript and TypeScript projects, the OpenAI Node SDK provides a clean interface for API calls. Initialize the client with the custom base URL and your secret key. Invoke the chat completions endpoint using the uncensored model. This method is ideal for web servers or edge functions where you need to manage asynchronous requests efficiently. The SDK handles serialization and deserialization of the JSON payload automatically.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.roleplayapi.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Enable Streaming

Reduce perceived latency by enabling server-sent events (SSE). Set the stream parameter to true in your request. The API will return a stream of chunks instead of a single complete response. This is particularly useful for chat interfaces where users expect immediate character responses. You can process each chunk as it arrives, updating the UI incrementally. The stream ends when the model finishes generating the token sequence.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Limits, Errors, and Context

Your requests are subject to a rate limit of 300 requests per minute per key. If you exceed this, the API returns a 429 status code. Authentication failures result in a 401 error, while insufficient prepaid credit triggers a 402 error. The context window supports up to 100,000 tokens for the combined prompt and completion. If you need to check available models, send a GET request to /v1/models. Ensure your request body does not exceed 8 MB.

API facts in one table

If your tool speaks the OpenAI API, these are the details that matter.

ItemValue
ProtocolOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
Modeluncensored
Base URLhttps://api.roleplayapi.com/v1
AuthenticationBearer token in the Authorization header
EndpointsPOST /v1/chat/completions · GET /v1/models
SSE streamingYes — server-sent events; the last chunk carries token usage
Max output16,000 tokens max; 2,048 if max_tokens is not set
Function callingSupported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages
Max context100,000 tokens (prompt + completion together)
Sampling parameterstemperature, top_p, stop, seed and the two penalties are passed through
JSON modeJSON object mode via response_format json_object
Concurrency8 requests at the same time per key
Request sizeup to 8 MB per request
Requests per minute300/min per key
Response headersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Token prices$0.25 per 1M input tokens · $1.00 per 1M output tokens
Trial credit$0.50 for 7 days, no card
Top-upUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
Credit expiryno monthly fee; paid credit does not expire
Volume bonus+5% on $50+, +10% on $100+
Billingprepaid credit, charged by real token usage; errors and refusals are free
Keysone active key per account; a new key replaces the old one
Content policyadult content allowed; sexual content involving minors is refused
Sign-inGoogle or e-mail and password

Error codes

Errors come back as JSON with a stable type; failed and refused requests are not billed.

StatusTypeReason
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditout of credit; add credit and retry
403content_blockedrefused by the content policy
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largebody over 8 MB
429rate_limited · concurrencyover 300/min or 8 parallel — back off and retry
503upstream_busymodel busy — retry in a few seconds

Questions and answers

What is the model ID for the uncensored API?

The model ID is "uncensored". This is an open-weight model tuned for adult roleplay and is not GPT, Claude, or any other vendor's model. You must specify this ID in your requests.

How does the pricing work?

Pricing is pay-as-you-go with no monthly fees. Input tokens cost $0.25 per million, and output tokens cost $1.00 per million. Prepaid credit does not expire, and you can top up starting at $10.

Is the API truly uncensored?

Yes, the model does not refuse lawful adult, fictional, or controversial topics. However, sexual content involving minors is always blocked. The uncensored nature applies to the text generation without standard content filters.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.