GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window.
Modalities
Context
1.0M
Released
Sep 18, 2026
Token volume and request traffic to this model over time.
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window.
GLM 5.3 FlashX has a 1,048,576 token context window.
GLM 5.3 FlashX accepts text, images, and video as input and returns text.
GLM 5.3 Flash, GLM 5.3, GLM 5.2 and 11 more are other text models from Z.ai.
GLM 5.3 FlashX was released on September 18, 2026.