Enterprise Orchestration

Kubernetes & Slurm Clusters

Orchestrating multi-GPU inference pods and HPC batch adaptation jobs on private sovereign infrastructure.

Kubernetes Deployment Manifest

apiVersion: apps/v1
kind: Deployment
metadata:
  name: moro-inference
spec:
  replicas: 2
  template:
    spec:
      containers:
      - name: vllm
        image: vllm/vllm-openai:latest
        args: ["--model", "/models/moro-finetune", "--gpu-memory-utilization", "0.9"]
        resources:
          limits:
            nvidia.com/gpu: 1