LLM Engineer (Data Generation)
remotemid
via Ashby
About this role
ABOUT THE TEAM & MISSION
LLM Engineer (Data Generation)는 Data-Centric AI 관점에서 모델 성능 향상에 필요한 학습 데이터를 설계·생성·평가·개선하는 역할을 수행합니다.
모델의 성능 병목과 Failure case를 분석하여 데이터 요구사항을 정의하고, Instruction Data, Preference Data, Reasoning Data, Domain-specific Data 등 목적에 맞는 학습 데이터를 구축하며, Data Generation Pipeline, 학습 결과 기반 Data Evaluation, Data Curation을 통해 대규모 학습 데이터의 품질을 체계적으로 개선하고, 차세대 Generative AI 모델의 성능 향상에 기여합니다.
RESPONSIBILITIES
- 모델 성능 개선을 위한 데이터 설계 및 생성
- Research 및 Model Training 팀과 협업하여 모델의 성능 병목, Failure Case, 학습 목표를 분석하고 데이터 요구사항을 정의합니다.
- Instruction Data, Preference Data, Reasoning Data, Domain-specific Data 등 목적에 맞는 학습 데이터를 설계·생성·정제합니다.
- 생성된 데이터가 모델 성능에 미치는 영향을 실험적으로 분석하고, 결과를 바탕으로 데이터 생성 전략을 반복적으로 개선합니다.
- Data Generation Pipeline 구축…
What we'd score you on
reqspace match rubricFive dimensions, recruiter-grade. Upload your resume and we'll generate a written explanation of where you fit and where the gaps are.
1
Skills match
For this role: python, spark, airflow, openai, ray
2
Level fit
This role is mid-level. We check your trajectory against it.
3
Domain experience
Your work in the role's domain matters more than your years total. We weight recent and direct experience.
4
Recency
A skill you used last quarter weighs more than one from five years ago. We grade on recency, not lifetime.
5
Location fit
This role is remote-eligible — we factor in your stated location and time-zone overlap.
Score yourself on this role.
Free · no card · written explanation included
Skills in this role
Pulled from the job description. These are the keywords we'll weight when scoring your fit.
pythonsparkairflowopenairay
