AI Infra/HPC Engineer
Location: On Site in the Bay Area, CA
Are you passionate about building the infrastructure powering the next generation of artificial intelligence? Our client is an emerging, venture backed technology company developing advanced GPU cloud infrastructure designed for large scale AI training and inference workloads. As one of the organization's early infrastructure engineers, you will play a foundational role in architecting and operating the systems that enable cutting edge AI applications. This is a unique opportunity to help shape a rapidly growing platform, influence core technical decisions, and work alongside an experienced founding team tackling some of the most complex infrastructure challenges in AI.
This Role Offers
- Competitive base salary, annual bonus, comprehensive benefits, and meaningful equity participation.
- Opportunity to join an early stage, well-funded company at a pivotal stage of growth.
- Direct ownership of mission critical AI infrastructure supporting enterprise customers.
- Highly collaborative engineering culture with significant technical autonomy.
- Opportunity to influence architecture, tooling, and operational standards from the ground up.
- Exposure to some of the industry's most advanced GPU computing technologies.
Focus
- Design, deploy, and operate large scale Kubernetes and or Slurm environments supporting production AI workloads.
- Build and enhance GPU orchestration capabilities, including scheduling optimization, topology aware placement, and GPU lifecycle automation.
- Architect and manage distributed storage platforms utilizing object storage, NVMe clusters, and high performance networking.
- Develop reliable infrastructure supporting large scale distributed training and inference environments.
- Partner directly with customers and internal engineering teams to design infrastructure solutions that meet demanding AI workload requirements.
- Own production reliability, incident response, performance optimization, and infrastructure observability.
- Create operational documentation, automation, and best practices that improve platform scalability and maintainability.
- Continuously evaluate emerging infrastructure technologies that improve efficiency, performance, and operational resilience.
Required Qualifications
- 3 to 6 years of experience building and operating large scale Kubernetes and or Slurm clusters.
- Experience designing, implementing, or managing GPU orchestration platforms, including custom scheduling, topology aware placement, or GPU lifecycle automation.
- Deep understanding of distributed object storage, NVMe storage clusters, storage architectures, and high bandwidth networking for AI infrastructure.
- Experience operating production GPU infrastructure at scale.
- Strong systems engineering background supporting distributed AI workloads.
- Comfortable owning production reliability, troubleshooting, and infrastructure operations.
- Strong Linux systems administration and infrastructure automation experience.
- Excellent troubleshooting, communication, and collaboration skills.
Preferred Qualifications
- Experience supporting large scale AI inference platforms.
- Familiarity with GPU performance optimization and workload scheduling.
- Experience with infrastructure as code and production automation.
- Exposure to power aware scheduling, GPU power management, or energy optimized infrastructure is a plus.
About Blue Signal:
Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS