HPC Consultant – Kubernetes / GPU / AWS
Umanist StaffingHPC Consultant – Kubernetes / GPU / AWS
Job Type: Full-Time
Location: Remote
Work Model: Remote-first; occasional on-site visits to customer offices
Job Summary
We are looking for an experienced HPC Consultant / Cloud Infrastructure Engineer with strong hands-on expertise in Kubernetes, HPC/GPU infrastructure, Terraform, and AWS.
This is a Kubernetes-heavy infrastructure role supporting large-scale, multi-cloud GPU platforms. The ideal candidate will have experience operating production Kubernetes clusters at meaningful scale, troubleshooting complex cluster issues, automating infrastructure through Infrastructure-as-Code, and supporting high-performance computing workloads.
Key Responsibilities
-
Operate production Kubernetes platforms including EKS, CKS, and GKE.
-
Manage Kubernetes cluster lifecycle, node pools, upgrades, networking policies, and platform stability.
-
Troubleshoot complex Kubernetes issues including:
-
Scheduler problems
-
CNI/networking issues
-
Node failures
-
Resource allocation
-
Cluster upgrades
-
-
Provision and manage HPC/GPU infrastructure using CI/CD and Terraform.
-
Support infrastructure across AWS, GCP, OCI, CoreWeave, and other cloud providers.
-
Manage GPU compute resources for AI/ML training and inference workloads.
-
Monitor cluster and infrastructure health and maintain SLIs/SLOs.
-
Build and maintain monitoring, alerting, and operational dashboards.
-
Participate in production incident response and severity escalations.
-
Conduct post-incident reviews and implement corrective actions.
-
Work closely with Networking, Storage, Security, and AI/ML Platform teams.
-
Develop Python-based tooling and automation for infrastructure operations.
Required SkillsKubernetes – MUST HAVE
-
4+ years of infrastructure engineering, cloud platform, or HPC experience.
-
Strong hands-on production Kubernetes experience.
-
Experience operating large-scale Kubernetes clusters.
-
Experience with:
-
EKS / GKE / CKS
-
Node pool management
-
Kubernetes scheduling
-
CNI/networking troubleshooting
-
Rolling upgrades
-
Cluster lifecycle management
-
Production troubleshooting
-