Modal has announced day-one support for Moonshot AI's Kimi K3, making the open-weight model available through its Shared API and dedicated Auto Endpoints.
The platform includes a custom DFlash speculative decoding model optimized for Kimi K3, enabling inference speeds of up to 460 tokens per second while improving throughput for long-running agentic workloads.
Developers can access Kimi K3 through OpenAI-compatible APIs with token-based pricing or deploy dedicated infrastructure that automatically scales with demand. The release highlights growing ecosystem support for frontier open-weight models and faster production deployment across enterprise AI applications.





