OpenAI’s GPT-6 Astra now supports UltraFast mode on Amazon Bedrock, providing a higher-speed option for latency-sensitive AI workloads. According to OpenAI, UltraFast delivers up to 6x faster inference through the API and generates up to 300 tokens per second.
The speed tier targets applications where response time directly affects the user experience, including real-time coding assistants, interactive AI agents and customer-facing applications.
Teams can access GPT-6 Astra UltraFast through the Amazon Bedrock console or supported APIs while using existing AWS controls for workload security, access governance and model invocation auditing across production AI deployments.




