ML Architect
Location: Remote, Nationwide
Our client is building advanced AI systems designed to move sophisticated machine learning capabilities from experimentation into dependable, real-world production. They are seeking an ML Architect to take hands on technical ownership of large-scale model architecture, distributed computing, GPU performance, and the systems required to operate modern AI reliably at scale. This is an opportunity for a deeply technical ML leader to solve difficult zero to one challenges, influence engineering standards, and help shape how next generation intelligent products are built.
This Role Offers
- High impact architectural ownership across the complete machine learning lifecycle, from data and training through production inference.
- A highly technical environment where individual contribution, sound judgment, and rapid experimentation are valued.
- Competitive cash and equity compensation with health, dental, life insurance, and 401(k) benefits.
Focus
- Define and evolve production ML architecture spanning data preparation, model development, evaluation, inference, deployment, and operational feedback.
- Engineer scalable distributed training and inference environments across GPU infrastructure using modern parallelization and systems approaches.
- Lead GPU and serving optimization efforts focused on memory utilization, quantization, mixed precision, throughput, latency, cost, and production capacity.
- Build dependable ML platforms and services with strong standards for maintainability, observability, reliability, and production readiness.
- Own ambiguous, zero to one ML initiatives and make pragmatic architectural tradeoffs involving model quality, performance, cost, reliability, and safety.
- Partner with research and application engineering teams to translate emerging model capabilities into stable, scalable production experiences.
- Develop robust data and evaluation workflows that support model quality measurement, responsible deployment, experimentation, and continuous improvement.
- Set technical direction across complex ML systems, resolve difficult architecture and performance challenges, and establish scalable engineering practices for other technical contributors.
Skill Set
- Advanced knowledge of AI model development, including practical experience building, optimizing, and implementing complex neural network solutions in real-world environments.
- Strong background leveraging scalable machine learning platforms and parallel computing tools to support complex model development and high-volume AI workloads.
- Demonstrated GPU performance expertise across memory efficiency, quantization, mixed precision, latency reduction, throughput improvement, and production scaling.
- Strong software engineering fundamentals with a track record of building robust, maintainable, production grade machine learning systems.
- Proven ability to own architecture end to end across data, training, evaluation, inference, deployment, and ongoing production operations.
- Principal level technical influence, including establishing engineering standards, guiding technical contributors, and solving complex architectural and performance problems.
- Strong command of PyTorch, JAX, or another modern ML framework, with the ability to work effectively across evolving machine learning technology stacks.
- Experience with technologies such as high-performance LLM serving frameworks, GPU kernels or compilers, reinforcement learning workflows, multimodal models, or large scale data processing is beneficial.
About Blue Signal:
Blue Signal is an award-winning, executive search firm specializing in various specialties. Our recruiters have a proven track record of placing top-tier talent across industry verticals, with deep expertise in numerous professional services. Learn more at bit.ly/46Gs4yS