Base URL & Authentication
Start by creating an account on the Get API key page. You only need an email and password; no phone number or credit card is required to activate the trial credit. Once registered, your API key is displayed immediately. Copy this key and store it securely. Use it in the Authorization header as a Bearer token for all requests. The base URL for all endpoints is https://api.gemma4api.com/v1. This URL works with any standard OpenAI-compatible client, allowing you to switch providers by simply changing the base URL and API key in your configuration.
First Request
Send your first prompt using curl to verify connectivity. Replace YOUR_API_KEY with your actual key. The model identifier is uncensored. This simple text-in, text-out request confirms your authentication is correct and the model is responding. If you receive a JSON response with a choices array, your setup is complete. You can now proceed to integrate the SDK for more complex workflows.
curl https://api.gemma4api.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Python SDK Integration
Install the official OpenAI Python package. Set the base URL and API key environment variables or pass them directly to the client constructor. Use the model ID uncensored to ensure you are calling our specific instance. This approach ensures your code remains compatible with standard OpenAI patterns while leveraging our uncensored model for content generation. The SDK handles retries and JSON parsing automatically.
from openai import OpenAI
client = OpenAI(base_url="https://api.gemma4api.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
Node.js SDK Integration
For JavaScript environments, install the OpenAI Node.js package. Configure the client with the base URL https://api.gemma4api.com/v1 and your API key. Use the model ID uncensored in your completion calls. This setup allows you to maintain standard OpenAI-compatible code structures while benefiting from the specific model behavior hosted on our servers. Ensure you handle asynchronous responses correctly within your Node.js event loop.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.gemma4api.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Streaming Responses
Enable streaming by setting the stream parameter to true in your API request. This returns a Server-Sent Events (SSE) stream, allowing you to display tokens as they are generated. This is ideal for chat interfaces where latency matters. Each chunk contains a partial delta of the response. Handle the stream end event to know when the generation is complete. This method reduces perceived latency significantly compared to waiting for the full response.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Limits, Errors & Context Window
Your requests support a context window of 100,000 tokens for both prompt and completion combined. The rate limit is 300 requests per minute per key, with a maximum request body size of 8 MB. Common errors include 401 for invalid keys, 402 if your prepaid credit is exhausted, and 429 if you exceed the per-minute limit. Regenerating your API key revokes the old one but does not reset the rate counter. Ensure your client handles these HTTP status codes appropriately to maintain smooth operation.
API specifications
Use this table to decide whether the API fits your project before you buy credit.
| Item | Value |
|---|---|
| Compatibility | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| API key | Bearer token in the Authorization header |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.gemma4api.com/v1 |
| Model | uncensored |
| Context window | 100,000 tokens (prompt + completion together) |
| SSE streaming | Supported (stream: true), usage included at the end |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Completion length | up to 16,000 tokens per request (default 2,048) |
| JSON mode | JSON object mode via response_format json_object |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Rate limit | 300/min per key |
| Max body | 8 MB request body |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Concurrency | 8 requests at the same time per key |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Subscription | no monthly fee; paid credit does not expire |
| Top-up | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Free trial | $0.50 for 7 days, no card |
| Bonus credit | +5% from $50, +10% from $100 |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Sign-in | Google or e-mail and password |
| Key management | one active key per account; a new key replaces the old one |
| Content policy | uncensored for adults; the only hard rule: no sexual content involving minors |
HTTP errors
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Does regenerating my API key reset my rate limit counter?
No. Regenerating your key revokes the previous one and issues a new credential, but the rate limit applies per account and persists regardless of key changes. You must wait for the minute window to reset if you hit the 300 requests/minute limit.
What model is running behind the 'uncensored' ID?
It is an open-weight large language model hosted on our own GPU servers. It is not GPT, Claude, Gemini, or any other vendor's model. It is tuned to answer without content refusals for lawful adult use.
Can I use this API for tool calling?
Yes. The <code>/v1/chat/completions</code> endpoint supports tool and function calling. You can pass tool definitions in the request, and the model will return structured JSON output compatible with standard OpenAI SDK implementations.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.