Site Reliability Engineer (Junior)
junior
via Ashby
About this role
ABOUT THE ROLE
VESSL AI의 Senior Site Reliability Engineer는 VESSL GPU 클라우드 플랫폼의 가용성과 성능을 책임집니다. 메트릭·로그·트레이스 기반의 Observability 체계를 구축해 이상 징후를 조기에 탐지하고, 장애 발생 시 On-call 대응과 Root Cause Analysis를 통해 신속하게 복구하며, 반복적인 운영 작업은 자동화로 대체해 시스템 전반의 안정성과 운영 효율을 함께 끌어올립니다.
GPU 클라우드 플랫폼은 대규모 GPU 클러스터, InfiniBand·RoCE 기반 고성능 네트워크, Kubernetes·Slurm 기반 스케줄링 시스템 등 여러 레이어로 구성됩니다. GPU 워크로드는 몇 시간에서 몇 주까지 이어지는 경우가 많고, 노드 하나의 장애나 네트워크 지연만으로도 진행 중이던 학습·추론 작업 전체가 중단될 수 있어 일반적인 웹 서비스보다 훨씬 높은 수준의 안정성이 요구됩니다. 이를 위해서는 GPU, NIC, 드라이버, 커널에서부터 스케줄러, 네트워크 패브릭에 이르기까지 시스템 전 레이어의 상태를 지속적으로 모니터링하고, 장애가 발생했을 때 원인을 빠르게 좁혀 복구하며, 반복되는 운영 작업을 자동화로 줄여나가는 기능이 필요합니다.
WHAT YOU WILL DO
- Reliability & Observability: 메트릭·로그·트레이스 등 Observability 스택을 구축하고, SLI/SLO를 정의해 플랫폼의 안정성을 정량적으로 추적…
What we'd score you on
reqspace match rubricFive dimensions, recruiter-grade. Upload your resume and we'll generate a written explanation of where you fit and where the gaps are.
1
Skills match
For this role: python, go, kubernetes, docker, terraform…
2
Level fit
This role is junior-level. We check your trajectory against it.
3
Domain experience
Your work in the role's domain matters more than your years total. We weight recent and direct experience.
4
Recency
A skill you used last quarter weighs more than one from five years ago. We grade on recency, not lifetime.
5
Location fit
This role is based in a specific location. We weight your proximity and willingness to relocate.
Score yourself on this role.
Free · no card · written explanation included
Skills in this role
Pulled from the job description. These are the keywords we'll weight when scoring your fit.
pythongokubernetesdockerterraformansiblegrafanaprometheusopentelemetrygit
