
vLLM Multi-LoRA Boosts MoE Serving on AWS
AWS implemented multi-LoRA inference for Mixture of Experts (MoE) models in vLLM, including kernel-level optimizations. This enables efficient serving of dozens of fine-tuned models on Amazon SageMaker AI and Amazon Bedrock. GPT-OSS 20B serves as the primary example.




