Clera

Research Engineer, Synthetic Data

Clera Singapore, Singapore · $100K–$170K/yr

Technology, Information and Internet · 11-50 employees

Yesterday
Mid (2-5 yrs) Full-time Visa sponsorship Singapore
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build and maintain pipelines to generate, validate, and improve synthetic training tasks for AI agents. Additionally, you will collaborate with subject-matter experts and analyze agent performance to refine model capabilities.

What they look for

Python Linux Docker Synthetic data Machine learning AI research Data pipelines Evaluation frameworks LLM Reinforcement learning Agentic AI Data validation Software engineering Automated systems

Requirements

Candidates should have two to four years of experience in software or machine learning engineering with hands-on expertise in synthetic data generation. Proficiency in Python, Linux, and containerization tools like Docker is required.

Benefits

Visa sponsorship

Full description

About the Role

Join an engineering team at an early-stage AI company building infrastructure for training and evaluating AI agents. You will develop synthetic data pipelines and methods that turn real-world professional workflows into useful training tasks, helping improve AI capabilities across technical and professional domains.

What You'll Do

  • Build pipelines that generate realistic, structured, and challenging synthetic training tasks from domain-specific workflows.
  • Collaborate with subject-matter experts to create tasks across professional and technical domains.
  • Design methods and tools to generate, mutate, validate, and improve synthetic tasks.
  • Analyze agent performance to understand what tasks teach and where models fail.
  • Develop metrics for task diversity, realism, learnability, and overall quality.

What We're Looking For

  • Two to four years of relevant experience in software engineering, machine learning engineering, or AI research.
  • Hands-on experience applying synthetic data methods to build end-to-end data generation pipelines for AI or machine learning applications.
  • Proficiency in Python, Linux, and containerization tools such as Docker.
  • Experience with synthetic data quality criteria, evaluation metrics, and the limitations of synthetic data.
  • Experience designing or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
  • A track record of independently delivering technical projects and building automated systems to generate, validate, or process structured data at scale.
  • Strong attention to detail, first-principles reasoning, and clear communication for collaboration across time zones.
  • Familiarity with reinforcement learning, agentic AI workflows, or LLM post-training is useful.

Compensation & Benefits

Compensation is USD 100,000 to USD 170,000 annually. Visa sponsorship is available.

Location

On-site in Singapore, Singapore.