Generative AI Inference Recommendation for Amazon SageMaker now available in the SageMaker AI Studio
This reduces the time required to find validated production settings from weeks to hours. Teams can deploy generative AI models faster by avoiding manual trial-and-error benchmarking.
- Provides a visual, low-code and no-code workflow extending the April 2026 API-based launch
- Benchmarks configurations on real GPU infrastructure using NVIDIA AIPerf
- Applies optimization techniques such as speculative decoding and kernel tuning
- Ranks recommendations by TTFT, inter-token latency, throughput, and cost
- Available at no additional cost beyond standard compute across multiple AWS regions in the US, Europe, and Asia Pacific