Skip to Content
API referenceOpenAI CompatibleChat Completions

Chat Completions

Create chat completion responses. Supports text generation, multimodal input, Function Calling, streaming, and more.

For new projects, we recommend the Responses API. The Responses API separates instructions from input, so system instructions benefit from Prompt Caching automatically, with a higher cache hit rate and noticeably lower cost and latency. Chat Completions remains supported long term; existing integrations do not need to migrate.

Endpoint

POST https://api.ofox.io/v1/chat/completions

Request Parameters

ParameterTypeRequiredDescription
modelstring✅Model identifier, e.g. openai/gpt-6.1-sol
messagesarray✅Message array
temperaturenumber—Sampling temperature 0-2, default 1
max_tokensnumber—Maximum tokens to generate
streamboolean—Enable streaming response
top_pnumber—Nucleus sampling parameter
frequency_penaltynumber—Frequency penalty -2 to 2
presence_penaltynumber—Presence penalty -2 to 2
toolsarray—Tool definitions (Function Calling)
tool_choicestring/object—Tool selection strategy
response_formatobject—Response format (JSON Mode)
providerobject—OfoxAI extension: routing and fallback config

Message Format

interface Message { role: 'system' | 'user' | 'assistant' | 'tool' content: string | ContentPart[] // Text or multimodal content name?: string tool_calls?: ToolCall[] // Tool calls in assistant messages tool_call_id?: string // Call ID in tool messages } // Multimodal content type ContentPart = | { type: 'text'; text: string } | { type: 'image_url'; image_url: { url: string; detail?: 'auto' | 'low' | 'high' } }

Request Examples

Terminal
curl https://api.ofox.io/v1/chat/completions \ -H "Authorization: Bearer $OFOX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-6.1-sol", "messages": [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Explain what an API Gateway is"} ], "temperature": 0.7 }'

Response Format

{ "id": "chatcmpl-abc123", "object": "chat.completion", "created": 1703123456, "model": "openai/gpt-6.1-sol", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "An API Gateway is a..." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 25, "completion_tokens": 150, "total_tokens": 175 } }

Streaming

Set stream: true to enable SSE streaming responses:

stream.py
stream = client.chat.completions.create( model="openai/gpt-6.1-sol", messages=[{"role": "user", "content": "Tell me a story"}], stream=True ) for chunk in stream: content = chunk.choices[0].delta.content if content: print(content, end="", flush=True)

Streaming Response Format

Each chunk is sent via SSE:

data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]} data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" there"},"finish_reason":null}]} data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]} data: [DONE]

Multimodal Input (Vision)

Send images for model analysis:

response = client.chat.completions.create( model="openai/gpt-6.1-sol", messages=[{ "role": "user", "content": [ {"type": "text", "text": "What's in this image?"}, {"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}} ] }] )

Models with vision capabilities include openai/gpt-6.1-sol, anthropic/claude-sonnet-5.5, google/gemini-3.8-flash, and more. See the Vision guide for details.

Function Calling

See the Function Calling guide for details.

Last updated on