
Data Engineering Specialist – AI
Jobgether • US
No Relocation
Posted: July 21, 2026
Additional Content
Job Description
- This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Data Engineering Specialist – AI based in United States. This role offers the opportunity to design and operate advanced data infrastructure powering next-generation artificial intelligence systems. You will build scalable pipelines that support AI training, evaluation, and continuous model improvement across diverse data types. The position combines deep data engineering expertise with an understanding of machine learning workflows and large-scale distributed systems. You will work closely with AI researchers and engineering teams to improve data quality, efficiency, and reliability. The ideal candidate will help shape the infrastructure behind high-performance AI applications while solving complex technical challenges. This is a fully remote opportunity for an experienced engineer passionate about building the future of AI through robust data platforms.
- Accountabilities: The Data Engineering Specialist – AI will be responsible for designing, developing, and maintaining large-scale data systems that enable AI development and operational excellence. This role focuses on building reliable data pipelines, improving data quality, and ensuring efficient delivery of datasets for advanced machine learning workflows. Design and operate large-scale data pipelines supporting AI training, evaluation, and continuous improvement processes. Build ingestion systems capable of handling diverse data modalities, including text, images, audio, video, and structured datasets. Develop data cleaning, deduplication, filtering, and quality assurance processes at significant scale. Implement dataset versioning, lineage tracking, and provenance systems to ensure reproducible AI workflows. Build high-throughput data loading solutions that optimize accelerator and GPU utilization during training. Develop labeling workflows, active learning systems, and human-in-the-loop processes to improve dataset quality. Design storage architectures that balance cost, scalability, throughput, and performance requirements. Create evaluation dataset pipelines with strong integrity controls and contamination prevention measures. Implement privacy, security, redaction, and consent mechanisms throughout data workflows. Collaborate with machine learning researchers and engineers to align infrastructure capabilities with AI development goals. Monitor and improve data quality, pipeline reliability, and system performance through observability practices. Optimize infrastructure costs through compression strategies, data formats, and caching solutions. Maintain documentation for data systems, schemas, and operational procedures. Stay current with emerging AI data infrastructure technologies, research, and open-source solutions. Requirements: The ideal candidate will have extensive experience building data engineering platforms and supporting AI or machine learning workloads. They should combine strong software engineering skills with expertise in distributed data systems, large-scale processing, and modern AI infrastructure practices. Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field. 6+ years of professional experience in data engineering, including significant experience supporting AI or machine learning environments. Strong programming skills in Python and experience with at least one JVM-based or systems programming language. Hands-on experience with large-scale data processing frameworks such as Spark, Ray, or Beam. Experience operating large-scale storage systems and high-volume data pipelines. Strong understanding of distributed systems, data modeling, and modern storage formats. Experience implementing dataset versioning, lineage, and reproducibility solutions for machine learning workflows. Familiarity with high-performance data loading architectures for accelerator-based model training. Strong software engineering practices, including testing, CI/CD workflows, and code review processes. Excellent problem-solving, communication, and cross-functional collaboration skills. Experience working with multimodal datasets at scale is preferred. Knowledge of data quality tooling, dataset evaluation methods, and AI data governance practices is a plus. Exposure to privacy-preserving data systems and regulated data environments is preferred. Experience contributing to open-source data infrastructure projects or supporting advanced AI model development is beneficial. Benefits: Competitive annual salary range of approximately $100,000 - $150,000. Fully remote work opportunity across the United States. Opportunity to work on cutting-edge AI and data engineering initiatives. Exposure to large-scale machine learning systems and advanced technology projects. Collaborative environment with experienced engineers and technical professionals. Career growth opportunities within a technology-focused organization. Opportunity to contribute to impactful AI solutions shaping future applications.
- How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best! Why Apply Through Jobgether? Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time. #LI-CL1
- We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
- apply for this job