Data Engineer
CommIT Warsaw, Masovian Voivodeship, Poland
Software Development · 501-1,000 employees
About the role
You will own the lakehouse architecture and design reliable data pipelines to ensure accurate, consistent, and traceable data for internal teams. Additionally, you will manage data ingestion, storage, retention, and governance while monitoring system health and performance.
What they look for
Requirements
Candidates must have 3+ years of experience in a production lakehouse environment with expertise in open table formats and cloud data warehousing. Strong proficiency in SQL, data modeling, streaming ingestion, and cloud object storage is essential for this role.
Full description
We’re looking for a Mid Data Engineer to take ownership of our data lake — the system of record for millions of financial events every day, including bets, wallet transactions, and live odds, serving 12M+ active users. You’ll be responsible for designing and maintaining reliable data pipelines and defining how data is ingested, stored, retained, reconciled, and governed across AWS S3 and Snowflake/Databricks. Your work will ensure that Analytics, Finance, and Regulatory teams have access to accurate, consistent, and fully traceable data — with zero drift from source systems.
Location: Kraków, Poland. Hybrid — 2 days per week from the office.
What you will do:
- Own the lakehouse architecture: bronze/silver/gold layers, Iceberg/Delta tables, schema evolution.
- Land operational data via CDC streaming (Kafka, Debezium), handling late and duplicate events.
- Design data layout for speed and cost: partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
- Own retention and archival: storage tiering, regulatory retention, immutability, GDPR deletion.
- Guarantee correctness: freshness SLAs, drift detection, reconciliation against the source wallet and ledger systems.
- Own governance: catalog and lineage, row/column access control, PII masking, encryption, audit trails.
- Monitor ingestion health, data anomalies, and cloud storage/compute spend.
Requirements
Must-have:
- 3+ years hands-on in a production lakehouse environment.
- Lakehouse architecture — bronze/silver/gold layering, an open table format (Iceberg, Delta, or Hudi), schema evolution.
- Data layout & query optimization at TB+ scale — partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
- Cloud lakehouse/DWH in production — Snowflake, Databricks, or BigQuery.
- CDC & streaming ingestion — Kafka + Debezium or equivalent; late, duplicate and out-of-order events.
- Strong SQL and data modeling — enough relational grounding to reason about the OLTP systems you capture from. Critical for financial ledgers.
- Correctness — freshness SLAs, drift detection, reconciliation against source wallet/ledger systems.
- Governance — catalogs, lineage, row/column access control, PII masking, retention, GDPR deletion.
- Cloud object storage — S3 or GCS, plus storage tiering and archival.
- Python and an orchestrator — Airflow or Dagster, as tools.
Nice to have:
- Fintech, iGaming, or another regulated, audit-heavy environment.
- Cost monitoring / FinOps for storage and compute spend.
- Hudi specifically; Dagster specifically.
- Immutability / WORM regulatory retention.
Similar roles
-
Data Engineer II
Interstates Sioux Center, Iowa, United States
-
#EG AI Data Engineer
NCS Singapore, Singapore
-
Lead Data Engineer – PySpark/Palantir Foundry
Logic20/20 Inc. Seattle, Washington, United States · $156K–$175K/yr
-
Lead Geospatial Data Engineer
Logic20/20 Inc. Seattle, Washington, United States · $156K–$175K/yr
-
Data Engineer / Data Scientist
Jobgether India
- Data Engineer - Energy Data