About the role
Build and maintain scalable ETL pipelines using Python and PySpark on AWS platforms while orchestrating complex workflows. Collaborate with cross-functional teams to design storage solutions and implement data quality monitoring and performance optimization.
What they look for
Requirements
Requires 4-8 years of software development experience with strong proficiency in Python, PySpark, and SQL. Candidates must have hands-on experience with AWS services, data warehousing architectures, and CI/CD deployment practices.
Full description
Build and maintain ETL pipelines using Python and PySpark on AWS Glue and related platforms.
Orchestrate workflows using AWS Step Functions and Lambda. - Implement messaging and event-driven integrations using SNS and SQS. Design and optimize storage and querying solutions in Amazon Redshift, RDS, Oracle and S3-based architectures.
Write efficient SQL for transformations, validation, and reporting. - Integrate data from APIs and process structured and semi-structured JSON data. - Implement data quality checks, monitoring, and operational support processes.
Participate in CI/CD and version control practices for deployment and release management.
Collaborate with cross-functional teams to translate business requirements into technical solutions.
8 years of software development experience across the appropriate platform.
Strong hands-on experience with Python, PySpark, API’s and SQL.
Experience with ETL/data pipeline development and Orchestration using Step functions / AirFlow .
Working knowledge of AWS services including Glue, Lambda, Step Functions, Redshift, S3, SNS, and SQS.
Experience with Athena, EMR, Kinesis, DynamoDB, or RDS. - Good Knowledge on CloudWatch, logging, and production support.
Understanding of data warehousing, data lakes, Lake House and query optimization.
Experience with GitLab/Terraform or similar and CI/CD workflows.
Good understanding of using AI tools like Github Copilot or similar for code productivity .
Exposure to enterprise data lake or cloud migration initiatives.
Have an eye to solving complex problems, great communication with stakeholders .
Have a good understanding of performance engineering of code pipelines and near real time systems - Good understanding on Agents and MCP".
A Bachelor’s Degree in Information Technology, Computer Science, Business Administration, or a related field is required.
Similar roles
-
Senior Data Engineer, Vice President
State Street Boston, Massachusetts, United States · $110K–$208K/yr
-
AI Data Engineer II (Business Data Analyst II)
UKG Bengaluru, Karnataka, India
-
Databricks Data Engineer
Elevate Government Solutions Washington, District of Columbia, United States
-
Senior Data Engineer
Iris Software Noida, Uttar Pradesh, India
-
Data Engineer II, GenAI
Travelers Hartford, Connecticut, United States · $126K–$209K/yr
-
Data Engineer - BODS
Weekday AI Bengaluru, Karnataka, India