Snowflake
Senior Software Engineer
Overview
Design and build distributed systems and infrastructure for the Cortex LLM post-training platform, turning scarce GPU capacity into a simple, composable service.
About Snowflake
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work.
Requirements & Eligibility
- 3 + years (Intermediate) | 6+ years (Senior) building and shipping production ML systems
- Strong distributed systems and infrastructure foundation - designing scalable, fault-tolerant services and operating them on Kubernetes in production
- Familiarity with GPU and LLM infrastructure - e.g., PyTorch, DeepSpeed/FSDP, Ray, CUDA/NCCL, vLLM
- Demonstrated ability to harden complex systems for reliability, throughput, and cost efficiency
- BS in Computer Science or a related field (MS/PhD a plus)
Key Responsibilities
- Design and build across the full stack - from the public training APIs and SDK through the control plane to the GPU data plane
- Scale the distributed systems that make GPU compute serverless - multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools
- Drive end-to-end performance at scale - keep the training, inference, and RL loops fast and the data plane responsive
- Productionize research building blocks - partner with Snowflake Research to turn state-of-the-art training and inference techniques into reliable, composable components
Disclaimer: Trace Hiring is an independent job board. We are not directly affiliated with Snowflake. Please verify all details on the official company application portal.