Senior Deep Learning Architect, LLM Inference
USonsitesenior
Posted today · via Workday
About this role
We are now looking for a Senior Deep Learning Architect, LLM Inference! NVIDIA is at the forefront of the generative AI revolution. The Inference Benchmarking (IB) team specifically focuses on inference server performance optimization for Large Language Models (LLMs). If you're passionate about pushing the boundaries of GPU hardware and software performance and understand terms like disaggregated serving, data parallel attention, MoE, Qwen3.5, DeepSeek, GPT-OSS, then this is a great role for you! What you'll be doing: You will do workload characterization of the latest LLMs and inference servers like vLLM, SGLang and TRT-LLM to ensure NVIDIA maintains its leadership position.…
What we'd score you on
reqspace match rubricFive dimensions, recruiter-grade. Upload your resume and we'll generate a written explanation of where you fit and where the gaps are.
1
Skills match
For this role: pytorch, openai, claude, teams
2
Level fit
This role is senior-level. We check your trajectory against it.
3
Domain experience
Your work in the role's domain matters more than your years total. We weight recent and direct experience.
4
Recency
A skill you used last quarter weighs more than one from five years ago. We grade on recency, not lifetime.
5
Location fit
This role is based in US. We weight your proximity and willingness to relocate.
Score yourself on this role.
Free · no card · written explanation included
Skills in this role
Pulled from the job description. These are the keywords we'll weight when scoring your fit.
pytorchopenaiclaudeteams
More at Nvidia
- View →Senior Software Engineer, PlatformsUS, CA, Santa Clara
- View →Software and System ArchitectUS, CA, Santa Clara
- View →PCB Design Layout EngineerUS, CA, Santa Clara
- View →Software and System ArchitectIsrael, Raanana
- View →Senior System Software Engineer - Data Platform ObservabilityUS, CA, Santa Clara
- View →Senior Chip ArchitectIsrael, Yokneam
