Share this job
ML Architect
United States
Apply for this job

ML Architect

Location: Remote, Nationwide


Our client is building advanced AI systems designed to move sophisticated machine learning capabilities from experimentation into dependable, real-world production. They are seeking an ML Architect to take hands on technical ownership of large-scale model architecture, distributed computing, GPU performance, and the systems required to operate modern AI reliably at scale. This is an opportunity for a deeply technical ML leader to solve difficult zero to one challenges, influence engineering standards, and help shape how next generation intelligent products are built.


This Role Offers

  • High impact architectural ownership across the complete machine learning lifecycle, from data and training through production inference.
  • A highly technical environment where individual contribution, sound judgment, and rapid experimentation are valued.
  • Competitive cash and equity compensation with health, dental, life insurance, and 401(k) benefits.


Focus

  • Define and evolve production ML architecture spanning data preparation, model development, evaluation, inference, deployment, and operational feedback.
  • Engineer scalable distributed training and inference environments across GPU infrastructure using modern parallelization and systems approaches.
  • Lead GPU and serving optimization efforts focused on memory utilization, quantization, mixed precision, throughput, latency, cost, and production capacity.
  • Build dependable ML platforms and services with strong standards for maintainability, observability, reliability, and production readiness.
  • Own ambiguous, zero to one ML initiatives and make pragmatic architectural tradeoffs involving model quality, performance, cost, reliability, and safety.
  • Partner with research and application engineering teams to translate emerging model capabilities into stable, scalable production experiences.
  • Develop robust data and evaluation workflows that support model quality measurement, responsible deployment, experimentation, and continuous improvement.
  • Set technical direction across complex ML systems, resolve difficult architecture and performance challenges, and establish scalable engineering practices for other technical contributors.


Skill Set

  • Advanced knowledge of AI model development, including practical experience building, optimizing, and implementing complex neural network solutions in real-world environments.
  • Strong background leveraging scalable machine learning platforms and parallel computing tools to support complex model development and high-volume AI workloads.
  • Demonstrated GPU performance expertise across memory efficiency, quantization, mixed precision, latency reduction, throughput improvement, and production scaling.
  • Strong software engineering fundamentals with a track record of building robust, maintainable, production grade machine learning systems.
  • Proven ability to own architecture end to end across data, training, evaluation, inference, deployment, and ongoing production operations.
  • Principal level technical influence, including establishing engineering standards, guiding technical contributors, and solving complex architectural and performance problems.
  • Strong command of PyTorch, JAX, or another modern ML framework, with the ability to work effectively across evolving machine learning technology stacks.
  • Experience with technologies such as high-performance LLM serving frameworks, GPU kernels or compilers, reinforcement learning workflows, multimodal models, or large scale data processing is beneficial.


About Blue Signal:

Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS


Apply for this job