Key points
- The API serves a single uncensored model accessible via standard OpenAI SDKs with a 100,000-token context window.
- Pricing is transparent at $0.25 per 1M input tokens and $1.00 per 1M output tokens with no monthly fees.
- The service supports streaming and tool calling but does not offer embeddings, images, or fine-tuning.
- Every new account receives $0.50 in trial credit valid for seven days without requiring a credit card.
What Does Uncensored Mean?
In the context of large language models, "uncensored" typically refers to a model that has been fine-tuned or prompted to reduce or eliminate the refusal mechanisms found in commercial models like GPT-4 or Claude. When you ask an uncensored model a question about a controversial topic, a niche hobby, or a mature theme, it is less likely to respond with a generic "I can't discuss that" or a lengthy safety disclaimer.
It is important to clarify what this does not mean. The model is not completely lawless. It still follows instructions and maintains logical coherence. The primary difference is that it prioritizes answering your specific prompt over adhering to a predefined set of content policies. For example, if you ask for a detailed description of a historical event that some providers might deem sensitive, an uncensored model will likely provide the full historical account without hedging.
However, there are boundaries. Our API enforces a hard content limit: requests containing sexual content involving minors are blocked. This is a standard safety baseline that applies to all inputs. For lawful adult content, political commentary, or creative fiction, the model aims to provide direct, unrestricted answers. This makes it ideal for use cases where you want the model's raw output without intermediate filtering layers.
Why Use an Uncensored Model?
Developers choose uncensored models for several practical reasons. First, it reduces the cognitive load of prompt engineering. You do not need to craft complex prompts to bypass refusal triggers. If you want a model to write a story with mature themes or analyze a controversial political stance, you can ask directly. This is particularly valuable for creative writing, role-playing applications, and research where nuance is often lost in safety filters.
- Creative Freedom: Generate stories, scripts, or dialogue without worrying about the model refusing to continue due to a specific keyword or theme.
- Research and Analysis: Get unbiased summaries of controversial topics without the model injecting its own "safe" narrative.
- Cost Efficiency: Uncensored models are often smaller or more efficiently tuned for specific tasks, which can lead to lower costs compared to large, general-purpose commercial models.
However, there are trade-offs. Uncensored models may occasionally produce output that is less factually rigorous or more prone to hallucination because they are not penalized as heavily for deviating from established facts. They also lack the nuanced conversational style that major providers polish into their products. If you need a model that politely apologizes for minor errors or strictly adheres to brand voice guidelines, an uncensored model might feel too raw. But if you need direct, unfiltered data or content, this approach is highly effective.
Gemma vs. GPT and Claude
When comparing an uncensored Gemma-based API to services like GPT-4 or Claude, the differences lie in architecture, policy, and pricing. Commercial models are often larger, more expensive, and heavily filtered. They are designed for general-purpose enterprise use, where safety and brand alignment are paramount. In contrast, our uncensored Gemma API is built for developers who want control over the output without paying for a broad suite of features they may not use.
| Feature | Uncensored Gemma API | Commercial Models (GPT/Claude) |
|---|---|---|
| Content Filtering | Minimal; hard limit on minor content | Strict; frequent refusals |
| Pricing Model | Pay-as-you-go, prepaid credit | Per-token, often subscription-based |
| Context Window | 100,000 tokens | Varies (e.g., 128k+ for some commercial models) |
| Open Weight | Yes | No (proprietary) |
Commercial models offer robust ecosystems with embeddings, vision, and audio capabilities. Our API focuses solely on text completion. If you need multimodal features, you will need a different provider. However, for pure text generation, our API offers a dedicated, uncensored endpoint that works with the same SDKs. The key advantage is transparency: you know exactly what you are getting—a single, powerful model that doesn't talk back to you for asking a mature question.
How to Connect the API
Connecting to our API is straightforward because it is fully OpenAI-compatible. You can use the official OpenAI SDKs for Python, Node.js, or any other language that supports the OpenAI protocol. The only changes required are the base URL and the API key.
First, sign up for an account at gemma4api.com to get your API key. Then, configure your client to point to our base URL: https://api.gemma4api.com/v1. Use the model identifier uncensored in your requests. This ensures that the model serving your request is our uncensored Gemma variant.
- Base URL:
https://api.gemma4api.com/v1 - Model ID:
uncensored - Endpoint:
POST /v1/chat/completions
This setup allows you to swap in our API into existing projects with minimal code changes. The API supports standard parameters like temperature, max_tokens, and stop. It also supports function calling, allowing you to define tools that the model can invoke. This makes it suitable for building agents or complex workflows that require structured output.
Streaming and Tool Calling
Our API supports both streaming and tool calling, providing flexibility for different use cases. Streaming is enabled by setting the stream parameter to true in your request. This returns data in chunks as they are generated, which is essential for real-time user interfaces where you want to display text as it is being written.
Tool calling allows you to define a schema for functions the model can call. When the model decides a function is needed, it returns a structured object with the function name and arguments. You then execute the function and pass the result back to the model for the next turn. This is powerful for building applications that interact with external systems, such as databases, calendars, or custom APIs.
Both features work seamlessly with the standard OpenAI-compatible format. You do not need to learn a new protocol or write custom parsers. The API handles the JSON formatting for tool calls and the SSE (Server-Sent Events) for streaming. This reduces development time and ensures compatibility with existing tooling and libraries.
Pricing and Limits
Our pricing is transparent and straightforward. There are no monthly subscriptions or hidden fees. You pay only for the tokens you use, with prepaid credit that never expires.
- Input Tokens: $0.25 per 1 million tokens
- Output Tokens: $1.00 per 1 million tokens
You can top up your account starting from $10. We offer bonuses for larger deposits: +5% bonus credit for deposits of $50 or more, and +10% bonus credit for deposits of $100 or more. This means you get more value for your money as you scale.
There are specific limits to keep in mind. You are allowed 300 requests per minute per API key. The request body size is limited to 8 MB. Each account is limited to one API key, which can be regenerated at any time if needed. This ensures fair usage and prevents abuse. The trial credit of $0.50 is valid for 7 days and does not require a credit card, making it easy to test the API before committing.
Privacy and Data Usage
Privacy is a key consideration for many developers. Our API requires only an email and password to create an account. We do not collect phone numbers or require payment information for the trial. Your API key is the sole identifier for your usage.
Regarding your data, we do not use your prompts for training. This is a significant difference from some commercial providers who may use input data to improve their models. Your conversations remain private to your account. We do not share your data with third parties unless required by law. This makes our API suitable for sensitive applications where data privacy is critical.
Additionally, our uncensored model does not infer content preferences based on your usage. Each request is processed independently, ensuring consistent behavior. The hard content limit for minor sexual content is applied uniformly, but otherwise, your data is yours. We do not retain logs of your conversations beyond what is necessary for billing and debugging, and you can request deletion of your data at any time.
Common Use Cases
Uncensored APIs are ideal for a variety of applications. Here are some common use cases where our API excels:
- Creative Writing: Generate novels, screenplays, or short stories with mature themes without content filters interrupting the flow.
- Role-Playing Games: Create dynamic NPCs that respond naturally to any topic, including violence, romance, or political debate.
- Data Extraction: Extract information from text without the model refusing to process certain keywords or entities.
- Research: Analyze controversial topics or historical events with minimal bias or filtering.
These use cases benefit from the model's ability to provide direct, unfiltered answers. For example, in a role-playing game, the NPC can react to a player's violent action without saying, "In this fictional context, it is acceptable." The model simply responds as the character would. Similarly, in data extraction, the model can process any text, regardless of its content, making it versatile for diverse datasets.
Getting Started Today
Getting started with the uncensored Gemma API is simple. Visit gemma4api.com and sign up with your email and password. You will receive an API key immediately, along with $0.50 in trial credit valid for 7 days. No credit card is required for the trial.
Use the API key to configure your OpenAI-compatible client. Point it to https://api.gemma4api.com/v1 and use the model ID uncensored. Make your first request and see the difference. If you need more capacity, top up your account with crypto (USDT or USDC). Bonuses are available for larger deposits.
Our API is designed for developers who want control, transparency, and uncensored output. Whether you are building a creative writing tool, a research assistant, or a role-playing game, our API provides the foundation you need. Sign up today and start exploring the possibilities of uncensored AI.
Questions and answers
Is the model open-weight?
Yes, the model is open-weight, meaning its architecture and weights are accessible. However, it is hosted on our servers, so you access it via the API rather than running it locally unless you download the weights yourself.
Can I use multiple API keys?
No, each account is limited to one API key. You can regenerate the key at any time, which will revoke the old one. This simplifies management and security.
Are there any hidden fees?
No, there are no monthly subscriptions or hidden fees. You pay only for the tokens you use, with prepaid credit that never expires. Bonuses are applied automatically for larger deposits.
What is the context window size?
The context window is 100,000 tokens, which includes both the prompt and the completion. This allows for long conversations or processing large documents.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.