Snorkel AI FDE

Senior/Staff FDE - Synthetic Data Generation

  • Location New York City, NY (Hybrid); San Francisco, CA (Hybrid)
  • Posted 2026-08-27

Original posting ↗ You apply on the company site — we never collect applications.

Extracted automatically from public job postings. Always verify details on the original posting before applying.

Hard requirements to check first

  • Clearance:not mentioned in the posting
  • Work auth:not mentioned in the posting

Skills

PythonEvalsAgents

Excerpt from the original posting

About Snorkel 

Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine technology with research-driven AI data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes. 

Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!

About the Role

Snorkel AI is hiring a Forward Deployed Engineer focused on Synthetic Data Generation to partner with leading AI labs and enterprises on their most critical AI initiatives.

In this role, you will lead the technical execution of complex customer engagements where synthetic data is used to improve model training, evaluation, and performance. You will translate ambiguous model and data challenges into effective data strategies, build scalable generation and evaluation pipelines, and use experimentation to continuously improve data quality and downstream model outcomes.

You will work across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements.

Main Responsibilities

Synthetic Data Generation & Evaluation

- Design and build scalable synthetic data generation, transformation, filtering, and evaluation pipelines for complex AI use cases

- Translate model objectives, failure modes, and data gaps into synthetic data strategies, experiments, and technical specifications

- Develop LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets across targeted behaviors, domains, and edge cases

- Build automated evaluators, quality checks, and measurement frameworks to assess correctness, relevance, diversity, coverage, and adherence to customer requirements

- Design and run experiments to measure the impact of synthetic data on downstream model performance and iteratively improve generation approaches

- Package and deliver production-grade datasets with standardized formats, quality assurance, and clear documentation

Forward Deployed Engineering & Customer Partnership

- Lead technical workstreams from initial solution design through production delivery, navig…

→ Snorkel AI · Greenhouse

Verify before you apply

  • The title, level and compensation match the original posting.
  • The role is still open — postings often close without notice, we drop them within 10 days of disappearing.
  • What "remote" means here: sometimes it is remote within one country only. The location field shows the employer’s own wording.
  • Work authorization, visa sponsorship and security clearance — these are the most common reasons an application goes nowhere.