Running Multiple AI Workloads on One GPU with HAMi: Architecture and Gotchas
- Time
- 2026-08-09 11:20 ~ 11:50
- Speaker
- Reza jelveh
- Room
- TR214
- Co-write
Abstract
GPUs are expensive. Kubernetes doesn't share them well yet, DRA is still work in progress.
HAMi: Heterogeneous GPU sharing for Kubernetes.
What I'll Tell You:
How it hijacks CUDA calls without touching your app code Why memory isolation matters Some real production use cases Real Results: Teams cut GPU costs 40-60%
Who Should Show Up: K8s operators, platform engineers, anyone watching GPUs sit idle
github.com/Project-HAMi/HAMi
Speaker
Reza jelveh
Reza Jelveh is a solutions engineer at HAMi (project-hami.io), focusing on the mechanics, constraints, and resource allocation of AI workloads on Kubernetes.
His background spans three decades—from reverse-engineering VCRs to building systems for semiconductor testing, seismic mitigation, healthcare, and data center construction. He has served as CTO for startups and a €bn-revenue government institution, navigating everything from bare-metal hardware tradeoffs to the operational friction of legacy middleware.