Senior AI Data Engineer
UKG Bengaluru, Karnataka, India
Software Development · 10,001+ employees
About the role
Lead the design, development, and operation of scalable data platforms, pipelines, and lakehouse solutions across cloud environments. Translate business and AI requirements into technical designs while mentoring engineers and upholding high standards for data quality and performance.
What they look for
Requirements
Requires 6+ years of experience in data engineering with strong proficiency in Python, SQL, and distributed processing frameworks. Candidates must have deep knowledge of cloud data architecture and experience leading technical delivery in production environments.
Benefits
Full description
Why UKG:
At UKG, the work you do matters. The code you ship, the decisions you make, and the care you show a customer all add up to real impact. Today, tens of millions of workers start and end their days with our workforce operating platform. Helping people get paid, grow in their careers, and shape the future of their industries. That’s what we do.
We never stop learning. We never stop challenging the norm. We push for better, and we celebrate the wins along the way. Here, you’ll get flexibility that’s real, benefits you can count on, and a team that succeeds together. Because at UKG, your work matters—and so do you.
Role Overview We are looking for a Lead Data Engineer to lead the design, development, and operation of reliable, scalable data platforms and products. This role combines hands-on engineering with technical leadership, guiding a team to deliver secure, well-governed data pipelines and curated datasets for analytics, reporting, and AI use cases. The ideal candidate brings deep experience with cloud data platforms, distributed processing, and modern data architecture. They can translate business and analytical needs into clear technical designs, set engineering standards, mentor engineers, and work across teams to deliver production-ready data solutions. Experience with GCP, Azure, Databricks, Python, PySpark, SQL, and Generative AI data patterns is valuable.
Key Responsibilities
- Lead the architecture and delivery of scalable batch and streaming data pipelines, data products, and lakehouse solutions across cloud environments.
- Translate business, reporting, analytics, and AI requirements into technical designs, data models, interfaces, and delivery plans.
- Design ingestion and transformation patterns for relational databases, APIs, event streams, files, and cloud storage, including structured and semi-structured data.
- Build and review production-grade data solutions using Python, PySpark, SQL, Spark, and platforms such as GCP BigQuery, Azure Databricks, and Azure Data Lake.
- Establish reusable engineering patterns for orchestration, modular transformations, metadata, schema evolution, incremental processing, and backfills.
- Guide data modeling across raw, refined, and curated layers, ensuring datasets are understandable, reusable, performant, and aligned with domain needs.
- Define and uphold standards for coding, peer review, testing, version control, CI/CD, release management, and operational readiness.
- Own reliability and performance outcomes for critical pipelines, including monitoring, alerting, recovery, capacity planning, and cost optimization.
- Implement data quality controls, reconciliation, lineage, and observability so that completeness, accuracy, freshness, and consistency can be measured.
- Partner with security, governance, architecture, platform, analytics, and business teams to address access controls, privacy, retention, and compliance requirements.
- Lead migration and modernization work from legacy data platforms to cloud environments, including dependency analysis, parity validation, cutover planning, and decommissioning support.
- Mentor and coach data engineers; provide technical direction, unblock delivery, and help the team grow its engineering and platform skills.
- Break down complex initiatives into milestones, estimate effort, surface risks and dependencies early, and communicate delivery progress to stakeholders.
- Investigate and resolve complex production issues, conduct root-cause analysis, and ensure corrective actions prevent recurrence.
- Prepare governed, well-documented datasets for AI and Generative AI use cases, including retrieval, feature, and semantic search workflows where appropriate.
- Evaluate tools and design options pragmatically, documenting trade-offs and recommending approaches that meet scale, security, cost, and maintainability needs.
Required Skills
- 6+ years of experience in data engineering, data platform engineering, or a closely related discipline, including ownership of production data systems.
- Demonstrated experience leading technical delivery, mentoring engineers, and coordinating work across multiple teams or domains.
- Strong programming skills in Python and advanced SQL; hands-on experience with PySpark or Spark for distributed data processing.
- Deep understanding of cloud data architecture and lakehouse concepts, including storage, compute, partitioning, file formats, and workload design.
- Hands-on experience with one or more major cloud data platforms; GCP and Azure experience is preferred, including BigQuery, Databricks, and Azure Data Lake.
- Experience designing and operating reliable ETL/ELT pipelines, workflow orchestration, data transformations, and reusable ingestion frameworks.
- Strong data modeling skills, including dimensional modeling and the design of curated datasets for analytics and downstream applications.
- Experience with data quality, reconciliation, observability, lineage, metadata, and production monitoring practices.
- Working knowledge of security and governance practices such as role-based access, encryption, sensitive data handling, and auditability.
- Experience with Git-based development, automated testing, CI/CD, infrastructure or configuration management, and controlled deployments.
- Ability to troubleshoot performance and reliability issues across distributed data systems and explain technical findings clearly.
- Strong communication, planning, prioritization, and stakeholder management skills; able to make technical topics accessible to non-engineering partners.
Candidates should be willing to learn and work with:
- Generative AI and Large Language Models (LLMs) in enterprise data environments
- Preparing governed structured and unstructured data for AI applications
- Embeddings, vector search, semantic search, and retrieval-augmented generation (RAG)
- AI-assisted engineering, data discovery, and workflow automation
- New cloud data services and evolving data governance and observability practices
Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field; an advanced degree is a plus.
- 6+ years of relevant experience in data engineering or data platform roles, with evidence of technical leadership and successful production delivery.
- Experience owning solutions through design, implementation, deployment, and ongoing operations.
- Ability to balance hands-on technical contribution with team guidance and cross-functional collaboration.
Company Overview:
UKG is the Workforce Operating Platform that puts workforce understanding to work. With the world's largest collection of workforce insights, and people-first AI, our ability to reveal unseen ways to build trust, amplify productivity, and empower talent, is unmatched. It's this expertise that equips our customers with the intelligence to solve any challenge in any industry — because great organizations know their workforce is their competitive edge. Learn more at ukg.com.
UKG is proud to be an equal opportunity employer and is committed to promoting diversity and inclusion in the workplace, including the recruitment process.
Disability Accommodation in the Application and Interview Process
For individuals with disabilities that need additional assistance at any point in the application and interview process, please email UKGCareers@ukg.com
Similar roles
-
Data Engineer II
Interstates Sioux Center, Iowa, United States
-
#EG AI Data Engineer
NCS Singapore, Singapore
-
Lead Data Engineer – PySpark/Palantir Foundry
Logic20/20 Inc. Seattle, Washington, United States · $156K–$175K/yr
-
Lead Geospatial Data Engineer
Logic20/20 Inc. Seattle, Washington, United States · $156K–$175K/yr
-
Data Engineer / Data Scientist
Jobgether India
- Data Engineer - Energy Data