SheepNav
新上线6个月前0 投票

Efficiently serve dozens of fine-tuned models with vLLM on Amazon SageMaker AI and Amazon Bedrock

In this post, we explain how we implemented multi-LoRA inference for Mixture of Experts (MoE) models in vLLM, describe the kernel-level optimizations we performed, and show you how you can benefit from this work. We use GPT-OSS 20B as our primary example throughout this post.

延伸阅读

  1. 通用编码计算迎来学习理论新基础:应对分布式系统中的慢节点问题
  2. 等变层神经网:在图上学习几何传输的新方法
  3. 无监督潜空间对齐:超球面测地线匹配新方法HGA
查看原文