Enterprise Orchestration
Kubernetes & Slurm Clusters
Orchestrating multi-GPU inference pods and HPC batch adaptation jobs on private sovereign infrastructure.
Kubernetes Deployment Manifest
apiVersion: apps/v1
kind: Deployment
metadata:
name: moro-inference
spec:
replicas: 2
template:
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
args: ["--model", "/models/moro-finetune", "--gpu-memory-utilization", "0.9"]
resources:
limits:
nvidia.com/gpu: 1