Chat Completions
POST /v1/chat/completions is the most widely compatible chat endpoint for existing OpenAI SDKs and clients.
curl
bash
# Send a non-streaming POST request to the complete Chat endpoint.
curl https://sprelaytoken.com/v1/chat/completions \
-H "Authorization: Bearer $SPRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"messages": [
{"role": "system", "content": "Answer concisely."},
{"role": "user", "content": "What is an API gateway?"}
]
}'-H adds headers and -d supplies the JSON body. The system message sets behavior; the following user message asks the question. JSON cannot contain comments, so preserve its punctuation when editing.
Python SDK
python
# Read the secret from the environment and import the compatible SDK.
import os
from openai import OpenAI
# The SDK appends /chat/completions to this base URL.
client = OpenAI(
api_key=os.environ["SPRELAY_API_KEY"],
base_url="https://sprelaytoken.com/v1",
timeout=60.0,
)
# create() waits for one complete non-streaming response.
response = client.chat.completions.create(
model="gpt-5.6-sol",
messages=[
{"role": "system", "content": "Answer concisely."},
{"role": "user", "content": "What is an API gateway?"},
],
)
# Print only the text from the first candidate.
print(response.choices[0].message.content)JavaScript SDK
javascript
// Import the SDK after installing the openai package.
import OpenAI from 'openai'
// Read the key from the Node.js process environment.
const client = new OpenAI({
apiKey: process.env.SPRELAY_API_KEY,
baseURL: 'https://sprelaytoken.com/v1',
timeout: 60_000,
})
// await pauses until the non-streaming response is complete.
const response = await client.chat.completions.create({
model: 'gpt-5.6-sol',
messages: [
{ role: 'system', content: 'Answer concisely.' },
{ role: 'user', content: 'What is an API gateway?' },
],
})
// Print only the first assistant message.
console.log(response.choices[0].message.content)model and messages are required. Add optional fields such as stream, temperature, max_tokens, or tools only after confirming that the selected model supports them. For multi-turn chat, send the relevant history again and control its size to limit latency and cost.
