20% off all GPT ๐ŸŽ‰ 30% off DeepSeekLearn more

Z.ai: GLM-5.3-FlashX

Chat-59%
z-ai/glm-5.3-flashx

GLM-5.3-FlashX is z.ai's lightweight, high-speed model on the international site (api.z.ai), built for efficient coding and long-horizon Agent tasks, with inference speeds of up to 200 tokens/s for a faster, smoother experience. It supports native multimodal input (images, video), 1M context, always-on thinking (which cannot be disabled, with low, high, and max reasoning effort levels), prompt caching, tool calling, and web search, and connects through both the OpenAI-compatible and Anthropic protocols (the Responses endpoint is not supported).

Context Window
1M
Max Output Tokens
131K
Released
2026-08-26
Capabilities
VisionFunction CallingReasoningPrompt CachingWeb SearchVideo Input
Available Providers
Z.ai
Supported Protocols
openaianthropic

Providers

Z.aiprovider.type: "zai"
-59%
Input Tokens
$0.15/M
$0.37/M
Output Tokens
$0.5/M
$1.25/M
Cache Read
$0.03/M
$0.075/M
Web Search
$0.01/R
Protocols
openai/v1/chat/completions
anthropic

Code Examples

from openai import OpenAI
client = OpenAI(
base_url="https://api.ofox.io/v1",
api_key="YOUR_OFOX_API_KEY",
)
response = client.chat.completions.create(
model="z-ai/glm-5.3-flashx",
messages=[
{"role": "user", "content": "Hello!"}
],
)
print(response.choices[0].message.content)

Frequently Asked Questions

Z.ai: GLM-5.3-FlashX on Ofox.ai costs $0.37/M per million input tokens and $1.25/M per million output tokens. Pay-as-you-go, no monthly fees.