Quickstart
ProfusAPI gives you one OpenAI-compatible endpoint for every model. If you already use an OpenAI SDK, change the base URL and the API key — nothing else.
- Create an account.
- Add credits on the Credits page.
- Create a key on the API Keys page.
- Pick a model from the catalog and send a request:
from openai import OpenAI
client = OpenAI(
base_url="https://profusapi.com/api/v1",
api_key="<PROFUSAPI_API_KEY>",
)
completion = client.chat.completions.create(
model="openai/gpt-5.5",
messages=[{"role": "user", "content": "What is the meaning of life?"}],
)
print(completion.choices[0].message.content)Authentication
Send your key as a Bearer token on every request. Keys start with sk-pf-v1-. Never expose them in client-side code.
Authorization: Bearer sk-pf-v1-...API Reference
Base URL:
https://profusapi.com/api/v1POST /chat/completions
Creates a model response for the given conversation. Accepts the standard OpenAI chat completion parameters (temperature, max_tokens, top_p, tools, response_format, …) which are forwarded to the model.
| model | string, required | Model ID, e.g. anthropic/claude-sonnet-5.5 |
| messages | array, required | List of {role, content} messages |
| stream | boolean | Stream tokens as server-sent events |
| temperature | number | Sampling temperature, 0–2 |
| max_tokens | integer | Maximum tokens to generate |
The response is the standard chat completion object. usage.cost contains the amount (USD) billed for the request.
{
"id": "gen_...",
"model": "openai/gpt-5.5",
"choices": [{ "index": 0, "message": { "role": "assistant", "content": "..." }, "finish_reason": "stop" }],
"usage": { "prompt_tokens": 14, "completion_tokens": 120, "total_tokens": 134, "cost": 0.0012175 }
}Streaming
Set "stream": true to receive server-sent events. The stream ends with data: [DONE]. Add "stream_options": {"include_usage": true} to get token usage in the final chunk.
curl https://profusapi.com/api/v1/chat/completions \
-H "Authorization: Bearer $PROFUSAPI_API_KEY" \
-H "Content-Type: application/json" \
-N -d '{"model":"openai/gpt-5.5","stream":true,"messages":[{"role":"user","content":"Hello"}]}'Conversation memory
ProfusAPI does not store conversation content: every request is processed on its own and your messages don't stay on our servers afterwards. One customer's conversations can never reach another customer.
To let a model "remember" earlier conversations, keep the history on your side and send it as the messages array with each request. Tools like Claude Code, Cursor and Cline already do this on your own machine. A minimal example for your own app:
import json, os, requests
HISTORY = "history.json" # stays on your own machine
messages = json.load(open(HISTORY)) if os.path.exists(HISTORY) else []
messages.append({"role": "user", "content": "What did we work on last time?"})
r = requests.post(
"https://profusapi.com/api/v1/chat/completions",
headers={"Authorization": f"Bearer {os.environ['PROFUSAPI_API_KEY']}"},
json={"model": "gemini-3-flash", "messages": messages},
)
reply = r.json()["choices"][0]["message"]
messages.append({"role": "assistant", "content": reply["content"]})
json.dump(messages, open(HISTORY, "w"))
print(reply["content"])Longer history means more tokens (and cost) per request. For long projects, replace old messages with a short summary.
Listing models
GET /models returns every available model with pricing (USD per token) and context length. No key required.
curl https://profusapi.com/api/v1/modelsKey & credit info
GET /key returns the usage and limit of the key making the request. GET /credits returns your account balance.
curl https://profusapi.com/api/v1/key -H "Authorization: Bearer $PROFUSAPI_API_KEY"
curl https://profusapi.com/api/v1/credits -H "Authorization: Bearer $PROFUSAPI_API_KEY"Errors
Errors return a JSON body: {"error": {"code": 402, "message": "..."}}
| 400 | Invalid request body or parameters |
| 401 | Missing or invalid API key |
| 402 | Insufficient credits, or the key reached its credit limit |
| 403 | The API key is disabled |
| 404 | Unknown model |
| 429 | Too many requests for this key — slow down and retry |
| 502 | The upstream model provider failed or is unreachable |
| 503 | Inference is temporarily unavailable |