Skip to main content
Umanist Staffing logo

HPC Consultant – Kubernetes / GPU / AWS

Umanist Staffing
1 day ago
Full-time
On-site
Canada

HPC Consultant – Kubernetes / GPU / AWS

Job Type: Full-Time
Location: Remote
Work Model: Remote-first; occasional on-site visits to customer offices

Job Summary

We are looking for an experienced HPC Consultant / Cloud Infrastructure Engineer with strong hands-on expertise in Kubernetes, HPC/GPU infrastructure, Terraform, and AWS.

This is a Kubernetes-heavy infrastructure role supporting large-scale, multi-cloud GPU platforms. The ideal candidate will have experience operating production Kubernetes clusters at meaningful scale, troubleshooting complex cluster issues, automating infrastructure through Infrastructure-as-Code, and supporting high-performance computing workloads.

Key Responsibilities

  • Operate production Kubernetes platforms including EKS, CKS, and GKE.

  • Manage Kubernetes cluster lifecycle, node pools, upgrades, networking policies, and platform stability.

  • Troubleshoot complex Kubernetes issues including:

    • Scheduler problems

    • CNI/networking issues

    • Node failures

    • Resource allocation

    • Cluster upgrades

  • Provision and manage HPC/GPU infrastructure using CI/CD and Terraform.

  • Support infrastructure across AWS, GCP, OCI, CoreWeave, and other cloud providers.

  • Manage GPU compute resources for AI/ML training and inference workloads.

  • Monitor cluster and infrastructure health and maintain SLIs/SLOs.

  • Build and maintain monitoring, alerting, and operational dashboards.

  • Participate in production incident response and severity escalations.

  • Conduct post-incident reviews and implement corrective actions.

  • Work closely with Networking, Storage, Security, and AI/ML Platform teams.

  • Develop Python-based tooling and automation for infrastructure operations.

Required SkillsKubernetes – MUST HAVE

  • 4+ years of infrastructure engineering, cloud platform, or HPC experience.

  • Strong hands-on production Kubernetes experience.

  • Experience operating large-scale Kubernetes clusters.

  • Experience with:

    • EKS / GKE / CKS

    • Node pool management

    • Kubernetes scheduling

    • CNI/networking troubleshooting

    • Rolling upgrades

    • Cluster lifecycle management

    • Production troubleshooting