Quickstart
Your first chat completion with the official OpenAI SDK. Change the base URL and the key. Nothing else.
Last reviewed 2026-08-25
Set your API key
Keys look like ms_live_ followed by 40 characters. Put the key in your environment. Never commit it.
export MSERVE_API_KEY="ms_live_your_key"Check the key
GET /mserve/v1/key returns the key's label, limits, usage today and this month, and the account balance. See Authentication.
Create a chat completion
Point the SDK at https://api.mserve.ai/openai/v1 and request the qwen3.8-27b model.
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.mserve.ai/openai/v1", api_key=os.environ["MSERVE_API_KEY"])
response = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Write one sentence about GPUs."}],
)
print(response.choices[0].message.content)Read the response
The response is the OpenAI chat completion object. model echoes the id you sent. usage counts prompt and completion tokens, and those are what you are billed for.
{
"id": "chatcmpl-8f1c…",
"object": "chat.completion",
"model": "qwen3.8-27b",
"choices": [
{ "index": 0, "message": { "role": "assistant", "content": "GPUs run thousands of small calculations at once." }, "finish_reason": "stop" }
],
"usage": { "prompt_tokens": 14, "completion_tokens": 11, "total_tokens": 25 }
}Next steps
Migrate an existing app
Copy the prompt and let your coding agent change the base URLs and model ids.
Speak the answer
Send the text to /openai/v1/audio/speech and stream the audio.
Stream tokens
Server-sent events with usage on the final chunk.
Compatibility
Which parameters reach the model and which return 400.