GLM-5.3-Flash benchmarks

Posts tagged with GLM-5.3-Flash benchmarks

GLM-5.3-Flash: 320B Params, 18B Active, 10× Cheaper Than GLM-5.3 — and It Was "Ox Alpha" All Along (2026)
Z.ai's GLM-5.3-Flash is the first natively multimodal GLM-5 model: 320B total / 18B active, hybrid sparse + linear attention, $0.15 per million input tokens, and MIT weights on day one. It beats GLM-5.2 on every published benchmark at a tenth of the price — and it's the anonymous "Ox Alpha" that took over OpenRouter.