L

AI Data Engineer (ML Data Pipelines)

LAK Technology Inc

IT Services and IT Consulting · 11-50 employees

9 h ago
Remote data-engineer Mid (2-5 yrs) Full-time
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design and build scalable data pipelines to support machine learning workflows, including feature engineering and real-time data ingestion. Collaborate with cross-functional teams to ensure data quality, monitoring, and efficient model deployment.

What they look for

Python SQL Spark Databricks Airflow Feature Engineering Data Pipelines Data Quality AWS Azure GCP Kafka Machine Learning MLOps Distributed Systems Cloud Computing

Requirements

Requires 4+ years of experience in data engineering with strong proficiency in Python, SQL, and distributed processing frameworks like Spark. Candidates should have hands-on experience with cloud environments and orchestration tools such as Airflow.

Full description

This is a remote position.

We are seeking an AI Data Engineer to design and build production-grade data pipelines that power machine learning systems. This role focuses on creating scalable ingestion, transformation, and feature engineering workflows that support model training, evaluation, and real-time inference.

You will work closely with Data Scientists, Machine Learning Engineers, and Platform teams to ensure high-quality, reliable, and efficient data flows across cloud environments. The ideal candidate understands both traditional data engineering and the unique data needs of ML systems.

Key Responsibilities:

  • Design and build scalable data pipelines for ML workflows
  • Develop feature engineering and data preparation processes
  • Implement batch and real-time data ingestion systems
  • Ensure data quality, validation, and monitoring
  • Collaborate with ML engineers to support model training and deployment
  • Integrate pipelines with orchestration tools (Airflow or similar)
  • Optimize pipeline performance and cloud cost efficiency
  • Maintain documentation and version control of data workflows

Requirements

Requirements

  • 4+ years of experience in Data Engineering
  • Strong Python and SQL skills
  • Experience building data pipelines for ML or analytics systems
  • Hands-on experience with Spark, Databricks, or similar distributed processing frameworks
  • Experience with orchestration tools (Airflow or similar)
  • Experience in AWS, Azure, or GCP environments
  • Familiarity with data quality validation and monitoring frameworks
  • Understanding of feature engineering and model data lifecycle

Preferred Qualifications:

  • Experience with streaming systems (Kafka, Kinesis, Pub/Sub)
  • Experience supporting model deployment and MLOps workflows
  • Experience with feature stores or vector databases
  • Familiarity with ML frameworks (TensorFlow, PyTorch)

Similar roles