Documentation
How to use the ChatGPT API
Last updated:
Sending the first request is easy. The hard part comes after: keeping the model aware of the conversation, dealing with a 429, knowing which parameters move the bill. This is everything you need once the first request works.
What a request is made of
A request is an array of messages. Each one carries a role, and the model reads them in order. Roles are how it tells an instruction apart from a line of conversation.
system- The instruction that governs the whole conversation: tone, output format, limits. It goes first and stays the same between turns.
user- The person's message. What the model is answering right now.
assistant- The model's own earlier replies. They are what let it remember what it already said.
curl https://api.llm-gate.tech/v1/chat/completions \
-H "Authorization: Bearer $CHATGPT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6.1-sol",
"messages": [
{ "role": "system", "content": "Отвечай кратко, по-русски." },
{ "role": "user", "content": "Что такое токен?" }
]
}'Running a conversation
The API is stateless. The model remembers nothing from the previous call, so you resend the conversation every time: the system instruction, every earlier turn, and the new question. You append the model's reply to the array yourself, otherwise it won't see it on the next step.
messages = [
{"role": "system", "content": "Отвечай кратко, по-русски."},
{"role": "user", "content": "Что такое токен?"},
# ответ модели возвращаем обратно в массив
{"role": "assistant", "content": "Токен это кусок текста примерно в 4 символа."},
{"role": "user", "content": "А сколько их в слове «программирование»?"},
]
response = client.chat.completions.create(
model="gpt-6.1-sol",
messages=messages,
max_completion_tokens=300,
)History grows and the bill grows with it: you pay for every input token on every request. On long conversations, trim old turns or fold them into a short summary.
Streaming: output as it is generated
By default the answer arrives in one piece once the model has finished. On long answers that reads as a freeze lasting tens of seconds. With streaming the text arrives in chunks right away, the way the ChatGPT web interface behaves.
stream = client.chat.completions.create(
model="gpt-6.1-sol",
messages=messages,
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content
if delta:
print(delta, end="", flush=True)- Chat interfaces where someone is watching the screen and waiting
- Long outputs: prose, code, detailed breakdowns
- Not needed for background work: classification, field extraction, scheduled jobs
The parameters that actually matter
There are more, but in practice you will touch these four. The rest matter only in narrow cases.
| Parameter | What it does | When to change it |
|---|---|---|
model | Picks the model | The main lever on cost and quality. Start cheap, move up when it fails |
max_completion_tokens | Caps the answer length | Always set it: it protects you from an unexpectedly long, expensive reply |
temperature | Spread of wording | Lower for facts and code, higher for prose and ideas |
stream | Deliver the answer in chunks | Turn it on wherever a person is waiting |
Errors and what they mean
The status code tells you whose problem it is and whether retrying helps. Half of these clear on a retry; half never will.
| Code | Cause and what to do |
|---|---|
| 401 | The key is wrong or missing. Check the Authorization header and the key itself. Retrying won't help |
| 404 | Model not found. Check the spelling against our model list. Retrying won't help |
| 429 | Rate limit hit. Retry after a pause, growing it with each attempt |
| 5xx | A fault on the service side. Retry after a pause; it usually clears |
| Timeout | No answer in time. Raise the client timeout: long reasoning answers take longer |
Retries with growing delay
Retrying immediately is pointless: on a rate limit you get the same error and burn quota. Grow the pause with each attempt and cap the number of tries. The official SDKs do this out of the box; if you call the API directly, the shape is this.
import time
for attempt in range(5):
try:
response = client.chat.completions.create(
model="gpt-6.1-sol",
messages=messages,
max_completion_tokens=300,
)
break
except Exception:
if attempt == 4:
raise
time.sleep(2 ** attempt) # 1, 2, 4, 8 секундKeeping the bill down
The bill tracks tokens, not requests. Four things account for most of the savings.
Cap the answer
Without max_completion_tokens the model may run to its maximum length and you pay for it. One parameter removes the worst cases.
Don't resend history you don't need
Every request pays for the whole message array. On long conversations, trim old turns or replace them with a summary.
Keep the stable part of the prompt first
A repeated prefix lands in the cache and costs a fraction of normal input. Put the variable data last.
Match the model to the job
Classification and field extraction have no business running on the flagship: the price gap reaches a hundredfold for the same result.
FAQ
Why doesn't the model remember previous messages?+
Because the API is stateless. Every request is processed from scratch, and everything the model knows about the conversation is what you pass in the messages array. Append its earlier answers with the assistant role.
What do I do about a 429?+
That is a rate limit. Retry after a pause and grow it each time: one second, two, four. If 429s are constant, lower your request rate or spread the load out over time.
Do I need to rewrite my code for your API?+
No. The request and response format matches OpenAI, so the official SDKs work after you swap the base URL and key. Every example here works against both the official API and ours.
How do I see how many tokens a request used?+
The response carries a usage field with input and output token counts. Log it: that is the only way to see where the money goes before the invoice arrives.
Does streaming cost more?+
No, the per-token price is identical. Only the delivery changes: chunks as they are generated instead of one package at the end.