Share this job
Head of GPU Cloud
Apply for this job

Head of GPU Cloud

Remote - Nationwide


Our client is building the software foundation for a next generation accelerated computing platform designed to make large-scale AI infrastructure easier to consume, operate, and scale. They are seeking a Head of GPU Cloud to lead the engineering organization responsible for transforming significant GPU capacity into high-performance cloud services for AI workloads. This is an opportunity to shape platform architecture, production inference, developer experiences, and engineering strategy at a stage where technical decisions will directly influence customer adoption, infrastructure economics, and long-term growth.


This Role Offers

  • Opportunity to define the architecture and operating model behind large scale AI inference services.
  • Direct influence over GPU utilization, platform economics, customer experience, and technical strategy.
  • Close collaboration with leaders across infrastructure, networking, product, operations, and commercial functions.
  • A highly technical environment where software engineering intersects with accelerated computing, distributed systems, and AI infrastructure.


Focus

  • Build and scale the engineering organization behind a nationwide GPU cloud platform, with ownership across inference services, orchestration, APIs, platform reliability, and technical execution.
  • Set the architecture for moving AI workloads efficiently from customer request to accelerator, including routing, placement, model lifecycle management, caching, and memory aware scheduling.
  • Lead production model serving and optimization across technologies such as vLLM, TensorRT LLM, and TGI, improving throughput, latency, availability, accelerator utilization, and cost.
  • Own Kubernetes based GPU infrastructure strategy across workload scheduling, elasticity, observability, deployment automation, multi-tenant isolation, and operational resilience.
  • Develop platform capabilities for hosted models, private inference environments, customized deployments, usage measurement, customer controls, and developer facing services.
  • Design orchestration policies that match workloads to accelerators based on memory requirements, performance objectives, capacity, cluster topology, and infrastructure economics.
  • Partner with networking, systems, data center, product, and commercial leaders to align software decisions with accelerator architecture, fabric performance, storage, and customer requirements.
  • Establish engineering standards, service objectives, capacity planning, incident readiness, and team accountability while recruiting and developing senior technical talent.


Skill Set

  • 12 or more years of progressive software engineering experience, including substantial leadership responsibility across cloud platforms, distributed infrastructure, HPC, or similarly complex production systems.
  • 5 or more years leading engineering teams responsible for business critical infrastructure, platform services, or other mission critical technology products.
  • Demonstrated production experience operating large scale model inference using vLLM, TensorRT LLM, TGI, or equivalent serving stacks.
  • Strong expertise in model serving optimization, including dynamic batching, decoding acceleration, reduced precision execution, compilation, memory reuse, and request scheduling.
  • Advanced knowledge of Kubernetes and containerized infrastructure, including scheduling, elasticity, telemetry, deployment practices, and production reliability.
  • Practical accelerator infrastructure knowledge covering GPU memory behavior, high bandwidth networking, InfiniBand, RoCE, storage performance, and cluster topology.
  • Experience architecting highly available distributed platforms with automated resource allocation, programmatic interfaces, multi customer support, and detailed consumption measurement.
  • Strong technical and leadership judgment with the ability to balance performance, reliability, security, customer experience, and infrastructure economics across multidisciplinary teams.


Additional Experience That Stands Out

  • Leadership experience in GPU cloud, AI infrastructure, hosted model platforms, or accelerated computing environments.
  • Experience delivering elastic inference services, dedicated AI capacity, model customization workflows, or managed AI products.
  • Familiarity with open model ecosystems and the operational differences among model families, serving configurations, and hardware profiles.
  • Experience creating consistent developer interfaces across multiple model backends.
  • Track record improving accelerator utilization, workload density, capacity forecasting, and compute economics across multiple GPU generations.
  • Experience scaling engineering organizations in fast moving environments where software requirements and infrastructure capacity evolve together.


About Blue Signal:  

Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS 


Apply for this job