SoreQen Platform
API access to all three models. OpenAI-compatible, so most existing code works after changing two lines.
Pricing
| Model | Input / 1M | Output / 1M |
|---|---|---|
| SoreQen S1 Mini | $0.05 | $0.15 |
| SoreQen S1 | $0.10 | $0.30 |
| SoreQen S1 Mega | $0.25 | $0.75 |
Pay for what you use. No monthly minimum and no subscription. Reasoning tokens are billed as output: asking for more thinking is asking the GPU for more work.
Every key, every capability
Nothing here is a tier above yours. The API is separate from the chat plans (a subscription buys chat, and credit buys calls), so a key reaches all three models with the whole feature set from the first request.
- Vision, with images in the request
- Tool calling, in the OpenAI shape
- Structured output against a JSON Schema
- Reasoning, at every effort setting
- 256k context window
- 32k tokens per reply
What a model is serving at a given moment depends on how its workers were launched; GET /v1/models reports the context window and capabilities in force today.
Compatible by design
The API speaks the OpenAI chat-completions protocol. If you already call an OpenAI-compatible endpoint, point the base URL here and change the model name.
from openai import OpenAI
client = OpenAI(
api_key="soreqen-live-...",
base_url="https://soreqen.com/v1",
)
r = client.chat.completions.create(
model="soreqen-s1",
messages=[{"role": "user", "content": "yaar ek regex samjha do"}],
)Keys
Keys are shown once, at creation, and stored only as a hash. There is no reveal-key button and there cannot be one. Each key carries its own spend cap, so a leaked key cannot run up an unbounded bill, and revocation takes effect within about ten seconds because the gateway caches a resolved key for that long, which is what keeps a burst of requests from becoming a burst of database reads.