SemiAnalysis 列出多租户 AI 云基础设施的常见错误设计,包括仅靠容器或 VM 隔离租户、缺少 VXLAN 与按租户 VPC、共享 Kubernetes 控制面(kube-apiserver、scheduler、etcd)、无硬件固件 OS 自动化配置导致机器在租户间回收复用。
Here are some common examples of bad designs:
🟠 Containers on shared hardware as the only level of isolation between tenants
🟠 VMs on shared hardware as the only level of isolation between tenants (less serious than containers, but still risky)
🟠 Lack of VXLANs or improper configuration, and no concept of a per-tenant VPC on the frontend network, relying on firewalls
🟠 Multi-tenant Kubernetes control planes, where components such as kube-apiserver, the scheduler, helm charts controlling cluster-wide services such as a shared GPUOperator or NetworkOperator, and etcd are shared across tenants
🟠 No hardware, firmware, and OS provisioning automation, leading to a practice where machines are recycled between tenants instead of provisioning things from scratch
🟠 Backend storage on a shared network with no isolation (the VPC comment again)
🟠 Giving any tenant access to the BMC network (IPMI, Redfish) or to any management port on Bluefield DPUs or any SmartNICs
🟠 Leaving Bluefield DPUs in their default host-trusted mode, where anyone with root on the host can reach the DPU’s Arm cores over RShim
🟠 Giving any tenant access to login to backend or frontend switches on a shared network
🟠 Incorrect configuration of InfiniBand security keys such as PKey, MKey and SAKey
🟠 Multi-tenant dashboards where infrastructure is shared between internal-facing logs and customer-facing logs
🟠 Controls where provider employees can receive privileged access to tenant logs and monitoring dashboards without tenant permission. (4/5)
来源:@SemiAnalysis_ · x.com