Jobgether logo

Senior Principal AI Engineer

Jobgether US


No Relocation

Posted: August 7, 2026

Additional Content

Job Description
  • This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Principal AI Engineer based in United States. This role offers the opportunity to build the foundation behind the next generation of large-scale artificial intelligence systems. You will design, optimize, and operate distributed training infrastructure for advanced neural networks across large GPU clusters. The position combines deep systems engineering, machine learning infrastructure, and performance optimization to enable faster and more reliable AI development. You will work closely with research and applied machine learning teams to productionize large-model training pipelines. Your expertise will directly impact the ability to train larger models, accelerate experimentation, and improve AI scalability. This is a high-impact technical leadership role for an engineer passionate about solving complex distributed computing challenges.
  • Accountabilities: As a Senior Principal AI Engineer, you will lead the design and optimization of large-scale distributed AI training systems. You will focus on improving performance, reliability, and scalability across GPU infrastructure while collaborating with technical teams to advance AI capabilities. Design, build, and operate distributed training systems for large neural networks, including autoregressive models, diffusion models, and state space models. Optimize multi-node and multi-GPU execution environments to maximize GPU utilization, throughput, and training efficiency. Diagnose and resolve complex performance bottlenecks involving compute, memory, storage, and networking layers. Improve training reliability through fault-tolerant architectures, failure recovery strategies, and operational best practices. Partner with research and applied machine learning teams to productionize scalable model training pipelines. Build and optimize GPU cluster orchestration solutions using technologies such as Slurm, Kubernetes, Ray, and RunAI. Ensure efficient scheduling, workload isolation, and resource allocation across large-scale training environments. Optimize distributed communication systems using technologies including NCCL, RDMA, InfiniBand, and NVLink. Scale AI training workloads using frameworks such as PyTorch Distributed, Megatron-LM, and DeepSpeed. Own multi-node training configurations, performance tuning, deployment processes, and recovery mechanisms. Apply advanced memory optimization techniques, including activation checkpointing, ZeRO optimization, and offload strategies. Balance compute, memory, and communication requirements to enable larger models and higher batch sizes. Establish best practices that improve training stability, scalability, and operational efficiency. Drive technical improvements that allow AI models to be trained faster, at greater scale, and with increased reliability. Requirements: The ideal candidate is a highly experienced AI infrastructure or distributed systems engineer with deep expertise in scaling machine learning workloads. You have a strong systems background and enjoy solving complex performance and reliability challenges. Extensive hands-on experience building and operating distributed systems or machine learning infrastructure. Proven experience running large-scale workloads on GPU clusters in production environments. Strong experience with distributed training using PyTorch Distributed. Deep understanding of parallelism strategies, including data parallelism, tensor parallelism, and pipeline parallelism. Strong knowledge of GPU communication architectures and networking technologies. Experience with GPU orchestration platforms such as Slurm, Kubernetes, Ray, or RunAI. Expertise with communication libraries and infrastructure including NCCL, RDMA, InfiniBand, and NVLink. Experience with large-scale training frameworks such as Megatron-LM and DeepSpeed. Knowledge of memory optimization approaches including activation checkpointing and ZeRO offload techniques. Ability to troubleshoot complex issues involving hardware, networking, compute resources, and distributed workloads. Strong understanding of AI infrastructure challenges, scalability limitations, and production ML environments. Experience working with large language models, foundation models, or generative AI systems is highly preferred. Backgrounds such as ML Systems Engineer, Distributed Systems Engineer, AI Infrastructure Engineer, or HPC Engineer transitioning into machine learning are highly relevant. Strong problem-solving skills, technical ownership, and ability to collaborate across research and engineering teams. Benefits: Competitive compensation package based on experience and qualifications. Fully remote work opportunity within the United States. Opportunity to work on cutting-edge AI infrastructure and large-scale machine learning systems. High-impact role contributing directly to the future of AI development and innovation. Collaborative environment with experienced engineers, researchers, and technical leaders. Opportunity to solve challenging distributed computing and scalability problems. Access to modern AI technologies and advanced infrastructure platforms.
  • How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
  • We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
  • apply for this job