Snorkel AI Applied AI
Research Scientist - Frontier Benchmarks
- Location New York City, NY (Hybrid); San Francisco, CA (Hybrid); United States (Remote)
- Seniority senior
- Posted 2026-05-29
Original posting ↗ You apply on the company site — we never collect applications.
Extracted automatically from public job postings. Always verify details on the original posting before applying.
role details
Design datasets and benchmarks for frontier model evaluation, and translate insights into clear narratives for customer-facing presentations.
Summary generated by AI from the original posting.
Hard requirements to check first
- Clearance:not mentioned in the posting
- Work auth:not mentioned in the posting
- Travel:customer-facing; travel not stated
Skills
AI/ML evaluationNLPExperimental designRigorous researchData operationsProductEngineeringStrategy
Excerpt from the original posting
About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! ABOUT THE ROLE We're looking for a Research Scientist to lead the design of next-generation benchmarks and datasets that push the boundaries of frontier model evaluation. You'll define what "good" looks like across a range of hard tasks, drawing on conversations with customers and academic partners to ground your datasets in real performance gaps. You'll build both the benchmarks that measure those gaps and the training data that helps close them, then partner with delivery, product, and go-to-market to bring what you build into production. This role is ideal for someone who wants to shape how the field measures progress in frontier AI, and who's energized by doing that work inside a fast-moving, cross-functional startup. MAIN RESPONSIBILITIES - Design state of the art datasets that drive frontier model training and evaluation based on current model performance and academic partnerships - Translate benchmark insights into clear, compelling narratives that articulate the ROI of expert-curated data for customer-facing presentations, technical reports, and go-to-market materials. - Work cross-functionally with data operations, product, engineering, and strategy to surface research findings that inform the company roadmap. - Stay at the frontier of LLM evaluation research and bring best practices into Snorkel's workflows - Represent Snorkel's research externally through publications, blog posts, conference talks, and customer engagements that advance the conversation around data-centric AI PREFERRED QUALIFICATIONS - Strong research background in AI/ML evaluation, NLP, or related fields, with a track record of…