Models/glm-5.3-flashx

glm-5.3-flashxActive · Cloud

OpenRouter · Released Sep 2026 · Proprietary

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Context
1.0M
Max output
131K
Input
$0.37
Output
$1.25
Cached in
$0.07
Latency
-

Capabilities

Core
Streaming
Function calling
JSON mode
Structured outputs
Multimodal
Vision
Image input
Audio
Image generation
Video
Advanced
Reasoning
Tool calling
MCP
Prompt caching
Batch API
Embeddings

Details

ProviderOpenRouter
ReleasedSep 2026
LicenseProprietary
InputText
OutputText
Throughput-
Availability100%

Related models