Indosat Business

Private GPU infrastructure, designed on NVIDIA reference architecture

From facilities readiness and cluster design through Kubernetes, GPU Operator, vLLM, monitoring, and handover — sized on NVIDIA best practice from a single Hopper or Blackwell node to a multi-rack SuperPOD, with a clear expansion path toward Rubin.

Rack, fabric, and platform — one design from node to SuperPOD

The solution is layered so a single validated rack and a multi-rack SuperPOD share the same facility, fabric, and platform design — only the scale changes.

LayerWhat it covers
Facility layerRack space, power distribution, direct liquid cooling, cabling, and physical security sized for high-density Hopper, Blackwell, or Rubin racks.
Compute layerGPU nodes or NVL72 rack systems with GPU, CPU, and memory sized per workload, forming one NVLink domain per rack (8 to 72 GPUs).
Fabric layerNon-blocking InfiniBand or Spectrum-X Ethernet, BlueField DPUs, and SuperNICs linking racks into one low-latency, multi-rack GPU domain.
Platform layerKubernetes or RKE2, NVIDIA GPU Operator, MIG, and topology-aware scheduling so jobs place along the fastest NVLink and fabric paths.
Serving layervLLM and model-serving endpoints for training, fine-tuning, and inference workloads, exposed through governed APIs and ingress.
Management & security layerPrometheus, Grafana, DCGM metrics, IAM/RBAC, backup, audit logging, and the runbooks used at operational handover.

From assessment to operational handover

Assessment & discovery

Capture business goals, workloads, data sensitivity, current platforms, operating model, and deployment constraints.

Workload sizing

Estimate GPU, CPU, memory, storage, network, and availability requirements for training, inference, VDI, HPC, or mixed workloads.

GPU architecture design

Design the GPU cluster topology, node profiles, isolation model, capacity plan, and expansion path.

Procurement support

Support bill of materials, vendor coordination, delivery planning, and commercial alignment for the selected architecture.

Sized on NVIDIA generations, from Hopper to Rubin

Architecture decisions follow NVIDIA reference designs for each GPU generation, so today’s node design has a validated path to rack-scale and multi-rack SuperPOD deployment instead of a rebuild.

GenerationRepresentative systemsScale-out unitBest-fit deployment
Hopper (2022)H100 / H200 SXM, 4th-gen NVLinkSingle node, 8-GPU NVLink domainDepartmental training, fine-tuning, and high-throughput inference
Blackwell (2024)B200, GB200 / GB300 NVL72, 5th-gen NVLinkNVL72 rack, 72-GPU NVLink domainRack-scale generative AI and frontier-scale training
Rubin (2026)Vera Rubin NVL72, 6th-gen NVLink, ConnectX-9, BlueField-4NVL72 rack, scale-out via Spectrum-X / InfiniBandAgentic AI, reasoning, and long-context inference at SuperPOD scale

Data center, network, and storage readiness

Data center readiness

Review rack space, power, cooling, cabling, physical security, and operational constraints for high-density GPU systems.

Network design

Plan management, storage, cluster, inference, ingress, and east-west traffic patterns with adequate throughput and segmentation.

Storage design

Define object, file, block, local, and high-throughput storage patterns for datasets, checkpoints, models, vectors, and backups.

A phased path from one rack to a SuperPOD

High-end Hopper, Blackwell, and Rubin deployments are validated in phases against NVIDIA reference architecture, so scale-out stays predictable as GPU count grows from one rack to many.

Rack-scale power and liquid cooling

Plan direct liquid cooling, coolant distribution, redundant power feeds, and rack density for high-power NVLink rack systems such as NVL72, ahead of delivery.

NVLink domain sizing

Size the GPU-to-GPU NVLink domain per rack, then design the scale-out fabric that links multiple racks into one addressable SuperPOD.

Non-blocking fabric design

Design leaf-spine InfiniBand or Spectrum-X Ethernet fabrics with BlueField DPUs and SuperNICs for non-blocking, low-latency, multi-rack GPU communication.

Topology-aware scheduling

Configure Kubernetes, GPU Operator, MIG, and NCCL topology awareness so training and inference jobs place workloads along the fastest NVLink and fabric paths.

Reference-architecture validation

Validate delivered racks against NVIDIA DGX SuperPOD / DGX BasePOD reference designs with burn-in, NCCL bandwidth tests, and acceptance benchmarks before handover.

Phased capacity buildout

Start with a single validated rack or proof-of-concept cluster, then expand rack by rack toward a multi-rack SuperPOD as workload demand and budget allow.

GPU software stack and operations

Kubernetes / RKE2 deployment

Deploy a production-oriented Kubernetes foundation with namespaces, RBAC, ingress, storage classes, and platform policies.

NVIDIA GPU Operator

Install and validate GPU drivers, device plugin, DCGM integration, MIG support where applicable, and GPU scheduling.

vLLM / LLM deployment

Deploy inference endpoints for Qwen, Llama, or selected models using GPU-backed serving patterns.

Monitoring stack

Implement Prometheus, Grafana, DCGM metrics, alerting, dashboards, and capacity indicators.

Handover and operations guide

Provide runbooks, operational guide, knowledge transfer, and platform validation evidence for customer teams.

Frequently asked questions

When is on-premises GPU infrastructure appropriate?

On-premises infrastructure is appropriate for persistent utilization, dedicated control, private data paths, facility-specific requirements, or internal policies that require owned environments.

What is assessed before deployment?

Assessment covers workloads, facilities, rack density, power, cooling, cabling, storage, networking, security, Kubernetes, operations, support model, and expansion path.

Which NVIDIA GPU generation should we plan for — Hopper, Blackwell, or Rubin?

Hopper (H100/H200) fits departmental training and inference on a single node. Blackwell (GB200/GB300 NVL72) is built for rack-scale generative AI with a 72-GPU NVLink domain per rack. Rubin (Vera Rubin NVL72) extends that design with 6th-generation NVLink and faster scale-out networking for agentic AI and long-context inference. The right starting generation depends on workload, budget, and facility readiness — architecture design should still leave a validated path to the next generation.

Can a GPU On-Prem deployment scale into a SuperPOD?

Yes. Deployments are designed against NVIDIA reference architecture from day one — NVLink domain sizing, non-blocking InfiniBand or Spectrum-X fabric, liquid cooling, and topology-aware Kubernetes scheduling — so capacity can expand rack by rack from a single validated cluster to a multi-rack SuperPOD without re-architecting.

Can Indosat Business help with handover?

Yes. Handover can include validation evidence, architecture records, runbooks, monitoring dashboards, access model, and knowledge transfer.

Contact Indosat Business