AI Data Engineer - Senior
Cummins Pune, Maharashtra, India
Motor Vehicle Manufacturing · 10,001+ employees
About the role
Leads the design, development, and maintenance of data and analytics platforms to process and store data for business stakeholders. Implements automated data pipelines, data governance processes, and scalable storage solutions using modern cloud and distributed technologies.
What they look for
Requirements
Requires 5 to 8 years of experience in data engineering with strong expertise in ETL/ELT processes and building scalable data pipelines. Proficiency in Python, SQL, Spark, and experience with Big Data platforms and cloud-based implementations is essential.
Full description
Job Summary:
Leads projects for design, development and maintenance of a data and analytics platform. Effectively and efficiently process, store and make data available to analysts and other consumers. Works with key business stakeholders, IT experts and subject-matter experts to plan, design and deliver optimal analytics and data science solutions. Works on one or many product teams at a time.
Key Responsibilities:
Designs and automates deployment of our distributed system for ingesting and transforming data from various types of sources (relational, event-based, unstructured). Designs and implements framework to continuously monitor and troubleshoot data quality and data integrity issues. Implements data governance processes and methods for managing metadata, access, retention to data for internal and external users. Designs and provide guidance on building reliable, efficient, scalable and quality data pipelines with monitoring and alert mechanisms that combine a variety of sources using ETL/ELT tools or scripting languages. Designs and implements physical data models to define the database structure. Optimizing database performance through efficient indexing and table relationships. Participates in optimizing, testing, and troubleshooting of data pipelines. Designs, develops and operates large scale data storage and processing solutions using different distributed and cloud based platforms for storing data (e.g. Data Lakes, Hadoop, Hbase, Cassandra, MongoDB, Accumulo, DynamoDB, others). Uses innovative and modern tools, techniques and architectures to partially or completely automate the most-common, repeatable and tedious data preparation and integration tasks in order to minimize manual and error-prone processes and improve productivity. Assists with renovating the data management infrastructure to drive automation in data integration and management. Ensures the timeliness and success of critical analytics initiatives by using agile development technologies such as DevOps, Scrum, Kanban Coaches and develops less experienced team members.
Responsibilities
Competencies: Security & Compliance Principles - Applies standards, tools, and best practices to embed security, privacy, and compliance into the design/build/test/operate lifecycle for products, services, apps, systems, software, and configurations—balancing protection, efficiency, and cost. Programming Principles - Applies programming languages, frameworks, and patterns to design, write, configure, test, and maintain software/solutions/systems that are efficient, secure, scalable, and reliable. Data Principles - Governs, models, secures, implements, and observes data flows to ensure integrity, quality, and compliance—enabling trusted, scalable, cost-conscious data use. Modern Development Practices - Applies modern engineering practices and tools—such as Agile/DevSecOps, CI/CD, automated testing, and infrastructure as code—to accelerate delivery, improve quality, and reduce risk across the SDLC. Solution Design - Translate business requirements into integrated designs, architectures, patterns, and system interactions that deliver customer value and align with enterprise standards and subject-matter platforms. Demonstrating Mastery - Maintains essential knowledge and proficiency in relevant domains, tools, technologies, methodologies, or frameworks through targeted credentials and rigorous proficiency, future-proofing organizational skills against strategic needs. Strategic and Innovative Thinking - Evaluates business and technology trends, anticipates future needs, develops creative approaches, and frames innovations to shape strategy and create durable value with cost-aware innovation. Technical Passion & Drive - Models curiosity and excitement for technology by self-initiating continuous development, experimenting with emerging technologies, and identifying insertion opportunities that accelerate business performance. Driving Effective Outcomes - Takes ownership, acts with urgency, and initiates action to turn goals into clear plans, decisions, guardrails, and cadences while navigating ambiguity and change to drive momentum and deliver consistent results. Engaging with Impact - Communicates with clarity and purpose to align stakeholders, foster collaboration, build trust, and influence coordinated action across teams and functions to accelerate outcomes. Values Differences - Recognizing the value that different perspectives and cultures bring to an organization. Ensuring Customer Success - Embraces a customer-first mindset to deliver outcomes by linking customer needs and business priorities to aligned solutions, delivery, adoption, satisfaction, and realized value through sustained engagement that builds partnership and trust.
Education, Licenses, Certifications: College, university, or equivalent degree in relevant technical discipline, or relevant equivalent experience required. This position may require licensing for compliance with export controls or sanctions regulations.
Experience: Intermediate experience in a relevant discipline area is required. Knowledge of the latest technologies and trends in data engineering are highly preferred and includes: - Familiarity analyzing complex business systems, industry requirements, and/or data regulations - Background in processing and managing large data sets - Design and development for a Big Data platform using open source and third-party tools - SPARK, Scala/Java, Map-Reduce, Hive, Hbase, and Kafka or equivalent college coursework - SQL query language - Clustered compute cloud-based implementation experience - Experience developing applications requiring large file movement for a Cloud-based environment and other data extraction tools and methods from a variety of sources - Experience in building analytical solutions Intermediate experiences in the following are preferred: - Experience with IoT technology - Experience in Agile software development - Experience with continuous improvement across cost optimization, performance tuning and scalability of Data Engineering pipelines. - Experience with enabling self service data engineering pipeline implementation capabilities for end users preferred. - Experience with using Co-pilot/AI capabilities to improve the productivity of Data Engineering pipeline development/testing activities.
Qualifications
Experience:
5 to 8 years of experience in data engineering, with strong expertise in building and optimizing scalable data pipelines, ETL/ELT processes, and data integration solutions. Skilled in designing robust architectures that support advanced analytics, reporting, and data-driven applications.
Technical Skills:
Required:
- Knowledge of the latest technologies and trends in data science is highly preferred.
- Hands on experiences in the following are preferred:
- Exposure to Big Data open source - Clustered compute cloud-based implementation experience
- Familiarity analyzing complex business systems, industry requirements, and/or data regulations
- Understanding of AI/ML concepts and tools
- Experience in ETL/ELT Data Engineering Technologies
- Background in processing and managing large data sets
- Design and development for a Big Data platform using open source and third-party tools
- Proficiency in Python, SQL, and Spark (PySpark preferred).
- Hands-on experience integrating with platforms like Palantir, Snowflake, Neo4j, etc.
- Solid knowledge of machine learning workflows, model deployment, and advanced analytics (regression, clustering, time-series analysis).
- Understanding of data governance, data cataloging tools (e.g., Azure Purview, Alation), and metadata management.
- SQL query language
- Clustered compute cloud-based implementation experience
- Experience developing applications requiring large file movement for a Cloud-based environment and other data extraction tools and methods from a variety of sources
- Take full ownership of the developed data pipelines, providing ongoing support for enhancements and performance optimization
Nice to have:
- Experience with graph data modeling and graph databases (e.g., Neo4j, TigerGraph) and familiarity with Palantir Ontology is a strong plus
- Understanding of data governance, data cataloging tools (e.g., Azure Purview, Alation), and metadata management.
Additional Key Responsibilities:
- Stay current with AI trends and suggest improvements to existing systems and workflows
- Excellent verbal and written communication skills
- Demonstrated self-starter with a proactive, problem-solving mindset
Candidate need to work from Cummins Pune IOC-B office for 3 days a week. There will be an overlap of few hours with US EST time zone.
Cummins is an equal opportunity employer. Our policy is to provide equal employment opportunities to all qualified persons without regard to race, sex, color, disability, national origin, age, religion, union affiliation, sexual orientation, veteran status, citizenship, gender identity, or other status protected by law.
Similar roles
-
Data Engineer II
Interstates Sioux Center, Iowa, United States
-
#EG AI Data Engineer
NCS Singapore, Singapore
-
Lead Data Engineer – PySpark/Palantir Foundry
Logic20/20 Inc. Seattle, Washington, United States · $156K–$175K/yr
-
Lead Geospatial Data Engineer
Logic20/20 Inc. Seattle, Washington, United States · $156K–$175K/yr
-
Data Engineer / Data Scientist
Jobgether India
- Data Engineer - Energy Data