gemma4api.comGet API key

Gemma API: First request in five minutes

Get your first response in under five minutes using the official OpenAI SDKs with our uncensored model. This guide covers authentication, basic requests, streaming, and tool calling with zero configuration friction.

Base URL & Authentication

Start by creating an account on the Get API key page. You only need an email and password; no phone number or credit card is required to activate the trial credit. Once registered, your API key is displayed immediately. Copy this key and store it securely. Use it in the Authorization header as a Bearer token for all requests. The base URL for all endpoints is https://api.gemma4api.com/v1. This URL works with any standard OpenAI-compatible client, allowing you to switch providers by simply changing the base URL and API key in your configuration.

First Request

Send your first prompt using curl to verify connectivity. Replace YOUR_API_KEY with your actual key. The model identifier is uncensored. This simple text-in, text-out request confirms your authentication is correct and the model is responding. If you receive a JSON response with a choices array, your setup is complete. You can now proceed to integrate the SDK for more complex workflows.

curl https://api.gemma4api.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

Python SDK Integration

Install the official OpenAI Python package. Set the base URL and API key environment variables or pass them directly to the client constructor. Use the model ID uncensored to ensure you are calling our specific instance. This approach ensures your code remains compatible with standard OpenAI patterns while leveraging our uncensored model for content generation. The SDK handles retries and JSON parsing automatically.

from openai import OpenAI

client = OpenAI(base_url="https://api.gemma4api.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Node.js SDK Integration

For JavaScript environments, install the OpenAI Node.js package. Configure the client with the base URL https://api.gemma4api.com/v1 and your API key. Use the model ID uncensored in your completion calls. This setup allows you to maintain standard OpenAI-compatible code structures while benefiting from the specific model behavior hosted on our servers. Ensure you handle asynchronous responses correctly within your Node.js event loop.

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.gemma4api.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Streaming Responses

Enable streaming by setting the stream parameter to true in your API request. This returns a Server-Sent Events (SSE) stream, allowing you to display tokens as they are generated. This is ideal for chat interfaces where latency matters. Each chunk contains a partial delta of the response. Handle the stream end event to know when the generation is complete. This method reduces perceived latency significantly compared to waiting for the full response.

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Limits, Errors & Context Window

Your requests support a context window of 100,000 tokens for both prompt and completion combined. The rate limit is 300 requests per minute per key, with a maximum request body size of 8 MB. Common errors include 401 for invalid keys, 402 if your prepaid credit is exhausted, and 429 if you exceed the per-minute limit. Regenerating your API key revokes the old one but does not reset the rate counter. Ensure your client handles these HTTP status codes appropriately to maintain smooth operation.

API specifications

Use this table to decide whether the API fits your project before you buy credit.

ItemValue
CompatibilityOpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key
API keyBearer token in the Authorization header
EndpointsPOST /v1/chat/completions · GET /v1/models
Base URLhttps://api.gemma4api.com/v1
Modeluncensored
Context window100,000 tokens (prompt + completion together)
SSE streamingSupported (stream: true), usage included at the end
Tools / tool callsYes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool
Completion lengthup to 16,000 tokens per request (default 2,048)
JSON modeJSON object mode via response_format json_object
Other parameterstemperature, top_p, stop, seed, presence_penalty, frequency_penalty
Rate limit300/min per key
Max body8 MB request body
HeadersX-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency
Concurrency8 requests at the same time per key
Price$0.25 per 1M input tokens · $1.00 per 1M output tokens
Subscriptionno monthly fee; paid credit does not expire
Top-upUSDT (TRC20) or USDC (Base), any whole amount from $10 to $500
Free trial$0.50 for 7 days, no card
Bonus credit+5% from $50, +10% from $100
Billingpay as you go from prepaid credit; nothing is charged for failed or refused requests
Sign-inGoogle or e-mail and password
Key managementone active key per account; a new key replaces the old one
Content policyuncensored for adults; the only hard rule: no sexual content involving minors

HTTP errors

Errors come back as JSON with a stable type; failed and refused requests are not billed.

StatusTypeReason
400bad_requestinvalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend
401missing_key · invalid_key · key_revokedcheck the Authorization header or use your current key
402no_creditout of credit; add credit and retry
403content_blockedsexual content involving minors — refused, not billed
404not_foundonly /v1/chat/completions and /v1/models exist
413request_too_largerequest body larger than 8 MB
429rate_limited · concurrencyslow down: rate or parallel limit reached
503upstream_busymodel busy — retry in a few seconds

Questions and answers

Does regenerating my API key reset my rate limit counter?

No. Regenerating your key revokes the previous one and issues a new credential, but the rate limit applies per account and persists regardless of key changes. You must wait for the minute window to reset if you hit the 300 requests/minute limit.

What model is running behind the 'uncensored' ID?

It is an open-weight large language model hosted on our own GPU servers. It is not GPT, Claude, Gemini, or any other vendor's model. It is tuned to answer without content refusals for lawful adult use.

Can I use this API for tool calling?

Yes. The <code>/v1/chat/completions</code> endpoint supports tool and function calling. You can pass tool definitions in the request, and the model will return structured JSON output compatible with standard OpenAI SDK implementations.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key