Z.ai’s GLM 5.3 is now generally available on Amazon Bedrock for complex coding and long-running agentic workloads. The open-weight mixture-of-experts model has 753 billion total parameters, with roughly 40 billion active per token.
It supports a one-million-token context window, up to 128K output tokens and selectable reasoning effort levels for balancing performance, latency and token usage. Bedrock also provides explicit prompt caching to reduce latency and input costs when applications repeatedly reuse context.
Eligible enterprise customers can access GLM 5.3 through US and Global cross-Region inference profiles and supported Amazon Bedrock APIs.




