Lab 11: KServe Inference with HAMi DRA GPU Sharing
Deploy a KServe Standard vLLM service and run two Predictor replicas on one NVIDIA GPU through native HAMi DRA claims.
Deploy a KServe Standard vLLM service and run two Predictor replicas on one NVIDIA GPU through native HAMi DRA claims.
Install HAMi on a GPU cluster and schedule SGLang inference services with GPU partitioning.
Install HAMi v2.10.0 and verify per-Pod MIG placement, mixed profiles, selective reclamation, restart recovery, and multi-GPU spillover.
Simulate 8 A100 GPUs with HAMi scheduling features, no real GPU needed.