Ellison Institute of Technology

Data Engineer (Autonomous Systems)

Ellison Institute of Technology Oxford, England, United Kingdom

Research Services · 201-500 employees

3 h ago
Remote data-engineer Senior (5-10 yrs) Full-time Visa sponsorship United Kingdom
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build and maintain high-throughput data pipelines for autonomous laboratory systems, integrating edge-to-cloud telemetry and hardware data. You will collaborate with researchers and engineers to create model-ready datasets and knowledge graphs that support automated scientific discovery.

What they look for

Python Data engineering Kafka MQTT Streaming APIs System design Cloud infrastructure GPU infrastructure Data modeling Robotics Telemetry Schema design Version control FastAPI gRPC GraphQL

Requirements

The role requires strong Python programming experience and a deep understanding of real-time data streaming and system architecture. Candidates should have experience with high-frequency data capture and a commitment to robust, reproducible engineering practices.

Benefits

Travel allowance Bonus Enhanced holiday Pension Life assurance Income protection Private medical insurance Hospital cash plan Employee discounts Electric car scheme Nursery salary sacrifice scheme Cycle to work scheme Family planning Neurodiversity support Coaching and therapy services

Full description

Join us at EIT:

At the Ellison Institute of Technology (EIT), we’re on a mission to translate scientific discovery into real world impact. We bring together visionary scientists, technologists, engineers, researchers, educators and innovators to tackle humanity’s greatest challenges in four transformative areas:

  • Health, Medical Science & Generative Biology
  • Food Security & Sustainable Agriculture
  • Climate Change & Managing CO₂
  • Artificial Intelligence & Robotics

This is ambitious work - work that demands curiosity, courage, and a relentless drive to make a difference. At EIT, you’ll join a community built on excellence, innovation, tenacity, trust, and collaboration, where bold ideas become real-world breakthroughs. Together, we push boundaries, embrace complexity, and create solutions to scale ideas from lab to society. Explore more at www.eit.org.

The Scientific Compute and Data Team

The Scientific Compute and Data team builds the compute, data and tooling foundation behind EIT's science. We run the cloud and GPU infrastructure that the models train on, the data infrastructure that instruments, robots and researchers use, and the data products and shared data models that let one team's results be built on by the next. We provide hands-on expertise in platform engineering, data engineering, and data management, and create common tooling for every lab to accelerate their experiments across gene editing, battery chemistry, plant biology and more. 

Your Role:

At EIT we are seeking an experienced Data Engineer to join our Scientific Compute and Data team. Our Data Engineers in Autonomous Systems work within a wider project, AutoLabs. This is where EIT automates the scientific discovery loop: experimental design, execution, analysis, and the learning that feeds back into the next experiment and the lab itself. The AutoLabs data platform captures, stores and serves high-frequency multimodal data. It spans an edge-to-cloud telemetry stack (MQTT, Kafka, media streams such as HLS), streaming APIs, and a growing set of hardware integrations across multiple sites including robotic systems, sensor rigs and lab instruments. 

Our labs produce raw data for a fast-growing set of models: vision-language-action models, orchestration and monitoring models, and world models that learn the dynamics of a laboratory to simulate experiments before they are run.  

As a Data Engineer, you’ll work directly alongside AI researchers, software engineers, robotics engineers, and domain scientists to build the data layer that makes our autonomous operations possible. You will contribute to an engineering culture that values robust system design, testing, and deep collaboration, but allows flexibility for rapid prototyping and responsiveness to changing landscapes. 

  

Day-to-Day, You Might: 

  • Build high-throughput, low-latency ingestion for live video, sensor telemetry, robot joint state and instrument control signals, integrating validation at the edge before data reaches the cloud. 
  • Work across the edge-to-cloud stack, from edge devices and MQTT brokers through to Kafka topics, cloud storage, and the APIs that serve scientists and ML pipelines. 
  • Turn streams of robot state, instrument output and video into structured database transactions: which arm did what to which sample, when, and what came out. 
  • Build the knowledge graph on top of that, linking inputs to results, so that a fact can always be traced back to the video segment and sensor readings behind it. 
  • Make the output model-ready: the formats, sampling, labelling and provenance that training and evaluation need, working directly with the researchers who consume it. 

What Makes You a Great Fit 

Nobody checks every box - if you’re not sure if you’re qualified, we still encourage you to apply.  

  • You have strong programming experience in Python, and value code quality, reliability, and performance while adapting to an evolving industry.  
  • You think in systems and own them end-to-end - from device output to APIs - and embrace long-term engineering rather than one-off scripts. 
  • Experience with real-time data streaming - Kafka or Pub/Sub (or equivalent), and the challenges that come with high-frequency, low-latency data capture. 
  • You treat reproducibility, version control, schema design and lineage as the things that make good science possible. 
  • A clear, respectful communicator who can collaborate with scientists, roboticists and researchers. 

  

It would be great if you also have: 

  • Automated lab or life sciences environments - instrument APIs, lab protocols, and what scientists expect of data quality. 
  • Data pipelines for physical systems: robotics (e.g. MQTT), autonomous vehicles, lab or clinical instruments, industrial control, or similar environments. 
  • Knowledge, entity, search or feed systems where correctness under continuous update was the hard part. 
  • Stateful stream processing (Flink, Kafka Streams, Beam or similar): event time, watermarks, exactly-once. 
  • APIs for data ingestion and consumption with FastAPI, gRPC or GraphQL. 
  • Graph stores and entity resolution or curation at scale, where the interesting decisions are about what to keep, merge or drop. 
  • Experience working close to model development, such as curating training or evaluation data, or building the data systems that feed models. You don’t need to be a machine learning engineer, but you should know what makes a dataset good. 

We offer the following benefits:

  • Salary dependent on experience + travel allowance + bonus 
  • Enhanced holiday. Our annual leave allowance is 25 days plus 8 bank holidays and an additional 3 days between Christmas and New Year. You will also have the opportunity to purchase an additional 5 days annual leave in January and July.
  • Pension - Employer contribution 7.5%, minimum employee contribution 5%
  • Life Assurance.
  • Income Protection
  • Private Medical Insurance as standard for you, your partner and any dependents. Including hospital Cash Plan
  • Employee discounts
  • Electric car scheme
  • Nursery Salary Sacrifice scheme
  • Cycle to Work Scheme
  • Family Planning
  • Neurodiversity support including advise and assessments
  • Coaching & Therapy services

 

Working Together – What It Involves:

  • 2 days per week in Oxford office, 1 day per week in employee’s choice of Oxford or London office, 2 days per week work from any UK location. 
  • You must have the right to work permanently in the UK with a willingness to travel as necessary. In certain cases, we can consider sponsorship, and this will be assessed on a case-by-case basis.

Similar roles