Research Scientist, Data
Menlo Park, Canada · Jornada completa
Sé el primero en postularte
- Experiencia
- Cualquier
- Salario
- USD 250,000 – USD 350,000 / year
- Vacantes
- 1
- Al corriente
- Hace 4 horas
- Modo de trabajo
- En la oficina
- Educación
- Bachelor's degree or equivalent experience
- Reanudar
- Se requiere solicitud
Dónde trabajarás
Descripción del trabajo
About Periodic Labs
Periodic Labs is an innovative AI and physical sciences company dedicated to accelerating breakthroughs in materials, energy, and related scientific fields by developing cutting-edge models. Supported by top-tier investors, the company is expanding rapidly and operates at the forefront of scientific innovation, with a team committed to deep expertise, strong ownership, and a relentless pursuit of scientific advancement.
Role Overview
This role focuses on a critical component of Scientific AI development: the design and management of evaluations and datasets. Responsibilities include creating advanced evaluations for scientific use cases, acquiring external datasets, integrating experimental data into training processes, and constructing training environments for reinforcement learning. The goal is to ensure that the AI teams have optimal assets for model assessment and improvement.
Key Responsibilities
- Lead the evaluation and data strategy within the training pipeline, identifying gaps in capabilities and collaborating with scientific and AI research leaders to develop a strategic roadmap.
- Translate complex scientific workflows into precise benchmarks, evaluations, and reinforcement learning environments by working closely with domain experts.
- Source, assess, and obtain external datasets spanning chemistry, physics, materials science, mathematics, simulation data, and laboratory instrumentation outputs.
- Develop and maintain stable data pipelines that facilitate ingestion, cleaning, and transformation of diverse large-scale datasets suitable for training.
- Create tools and analytical workflows that enable researchers to examine data quality, identify model failure patterns, and prioritize the creation of new evaluations and datasets.
Qualifications and Skills
- Experience designing evaluations, benchmarks, or RL environments for language models, agents, or scientific AI applications.
- Proficiency in building and managing large-scale data pipelines for various phases of large language model (LLM) training and evaluation.
- Strong discernment regarding dataset and evaluation quality, evaluating criteria such as scientific relevance, comprehensiveness, data origin, licensing constraints, and contamination risks.
- Advanced software engineering and data handling skills, including data processing at scale, dataset versioning, and lineage tracking.
- A research-oriented approach involving hypothesis formation on data, conducting controlled experiments, quantitative evaluation of model performance, and iterative refinement.
Additional Information
- Minimum qualification: Bachelor's degree or equivalent experience.
- Primary location: Menlo Park, California, with additional locations in Montreal, Canada, and soon San Francisco.
- Compensation ranges from $250,000 to $350,000 annually, plus equity options.
- Visa sponsorship is available for eligible candidates.