Data Engineer
paretocaptiveservicesllc • Remote
Posted: August 18, 2026
Job Description
Position Summary:
The Data Engineer will design, build, and maintain serverless data pipelines and data models on AWS that make high‑quality, analytics‑ready data available to our Analytics and AI teams. This role will own ingestion, transformation, and storage patterns using services such as Lambda, Glue, Athena, and S3, ensuring data is reliable, well‑documented, and aligned with business and model‑training needs.
Key Responsibilities:
- Design, implement, and maintain scalable, fully serverless data pipelines on AWS using Lambda, Glue, Athena, Step Functions, and S3 to support reporting, analytics, and AI use cases.
- Build and evolve data models and schemas that enable performant querying and downstream consumption by Analytics, Data Science, and AI engineering teams.
- Develop ETL/ELT workflows to ingest, cleanse, transform, and load data from internal applications, third‑party sources, and event streams into our data lake and analytical layers.
- Partner closely with Product, Underwriting, Analytics, and AI stakeholders to understand data requirements and translate them into robust data structures, contracts, and SLAs.
- Implement data quality controls, monitoring, and alerting to ensure accuracy, completeness, timeliness, and lineage of critical datasets and features used by models.
- Optimize serverless workloads for cost, performance, and scalability, including query tuning in Athena and efficient storage formats/partitioning in S3.
- Contribute to and enforce data engineering best practices, including version control, code review, CI/CD for data pipelines, and Infrastructure as Code with AWS CDK.
- Collaborate with Analytics and AI teams to design and maintain feature stores and other reusable data assets that accelerate experimentation and model deployment.
- Troubleshoot pipeline issues, resolve data‑related incidents, and provide ongoing support for production data workflows and model‑driven applications.
- Document data models, pipelines, and data contracts, and help evangelize data literacy and self‑service analytics across the organization.
Required Qualifications / Skills:
- Strong experience with AWS data and serverless services (Lambda, Glue, Athena, S3, Step Functions or similar orchestration tools).
- Proficiency in Python and TypeScript or a similar language for data engineering, including building ETL/ELT jobs and reusable libraries.
- Solid understanding of data modeling principles (dimensional, normalized, wide‑table, and event‑driven designs) for analytics and AI workloads.
- Experience designing and operating data pipelines at scale, including batch and near‑real‑time ingestion.
- Familiarity with SQL and query optimization in columnar data stores and engines such as Athena.
- Knowledge of data quality, governance, and security practices, including handling sensitive healthcare and financial data.
- Hands‑on experience with Infrastructure as Code, preferably AWS CDK, for provisioning and managing data infrastructure.
- Ability to collaborate with Analytics and AI teams, understand model data needs, and design data solutions that enable experimentation and productionization.
- Strong communication skills and the ability to translate complex data concepts into clear, actionable language for business stakeholders.
Minimum Requirements:
- Bachelor’s degree in Computer Science, Engineering, Mathematics, or related field, or equivalent practical experience.
- 3+ years of experience in data engineering or software engineering roles focused on data pipelines and analytics platforms.
- 2+ years of hands‑on experience with AWS data services in a production environment (e.g., Lambda, Glue, Athena, S3).
- Experience building and maintaining data solutions that support analytics, BI, and/or AI/ML initiatives.
Perks & Benefits:
- Fully paid medical, dental, and vision benefits.
- Flexible PTO
- 401k company contribution
- Tuition reimbursement
- Professional development allowance
- Transportation allowance and daily parking reimbursement
- Engaging hybrid work environment
Additional Content
Position Summary:
The Data Engineer will design, build, and maintain serverless data pipelines and data models on AWS that make high‑quality, analytics‑ready data available to our Analytics and AI teams. This role will own ingestion, transformation, and storage patterns using services such as Lambda, Glue, Athena, and S3, ensuring data is reliable, well‑documented, and aligned with business and model‑training needs.
Key Responsibilities:
- Design, implement, and maintain scalable, fully serverless data pipelines on AWS using Lambda, Glue, Athena, Step Functions, and S3 to support reporting, analytics, and AI use cases.
- Build and evolve data models and schemas that enable performant querying and downstream consumption by Analytics, Data Science, and AI engineering teams.
- Develop ETL/ELT workflows to ingest, cleanse, transform, and load data from internal applications, third‑party sources, and event streams into our data lake and analytical layers.
- Partner closely with Product, Underwriting, Analytics, and AI stakeholders to understand data requirements and translate them into robust data structures, contracts, and SLAs.
- Implement data quality controls, monitoring, and alerting to ensure accuracy, completeness, timeliness, and lineage of critical datasets and features used by models.
- Optimize serverless workloads for cost, performance, and scalability, including query tuning in Athena and efficient storage formats/partitioning in S3.
- Contribute to and enforce data engineering best practices, including version control, code review, CI/CD for data pipelines, and Infrastructure as Code with AWS CDK.
- Collaborate with Analytics and AI teams to design and maintain feature stores and other reusable data assets that accelerate experimentation and model deployment.
- Troubleshoot pipeline issues, resolve data‑related incidents, and provide ongoing support for production data workflows and model‑driven applications.
- Document data models, pipelines, and data contracts, and help evangelize data literacy and self‑service analytics across the organization.
Required Qualifications / Skills:
- Strong experience with AWS data and serverless services (Lambda, Glue, Athena, S3, Step Functions or similar orchestration tools).
- Proficiency in Python and TypeScript or a similar language for data engineering, including building ETL/ELT jobs and reusable libraries.
- Solid understanding of data modeling principles (dimensional, normalized, wide‑table, and event‑driven designs) for analytics and AI workloads.
- Experience designing and operating data pipelines at scale, including batch and near‑real‑time ingestion.
- Familiarity with SQL and query optimization in columnar data stores and engines such as Athena.
- Knowledge of data quality, governance, and security practices, including handling sensitive healthcare and financial data.
- Hands‑on experience with Infrastructure as Code, preferably AWS CDK, for provisioning and managing data infrastructure.
- Ability to collaborate with Analytics and AI teams, understand model data needs, and design data solutions that enable experimentation and productionization.
- Strong communication skills and the ability to translate complex data concepts into clear, actionable language for business stakeholders.
Minimum Requirements:
- Bachelor’s degree in Computer Science, Engineering, Mathematics, or related field, or equivalent practical experience.
- 3+ years of experience in data engineering or software engineering roles focused on data pipelines and analytics platforms.
- 2+ years of hands‑on experience with AWS data services in a production environment (e.g., Lambda, Glue, Athena, S3).
- Experience building and maintaining data solutions that support analytics, BI, and/or AI/ML initiatives.
Perks & Benefits:
- Fully paid medical, dental, and vision benefits.
- Flexible PTO
- 401k company contribution
- Tuition reimbursement
- Professional development allowance
- Transportation allowance and daily parking reimbursement
- Engaging hybrid work environment