Role Overview
NVIDIA is hiring engineers to build and scale the infrastructure that supports our Electronic Design Automation (EDA) workloads. We are looking for engineers with strong programming skills, a deep understanding of distributed systems, experience operating large-scale production infrastructure, and excellent communication and planning abilities.
About Nvidia
NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry.
Key Responsibilities
- Design and build platforms that automate the provisioning, configuration, operation, and lifecycle management of large-scale GPU and CPU compute infrastructure.
- Develop monitoring, health-management, and remediation systems that improve the reliability, availability, and utilization of EDA compute environments.
- Automate hardware deployment, operating-system configuration, firmware and software updates, cluster enrollment, and recovery workflows.
- Build reliable services and workflows that integrate with workload schedulers, infrastructure management systems, and observability platforms.
- Use hardware diagnostics, operating-system signals, scheduler data, and network and storage telemetry to identify failures and return unhealthy systems to service.
- Work with EDA, infrastructure, networking, storage, and hardware engineering teams to deliver scalable solutions for critical chip-design workloads.
- Participate in incident response, root-cause analysis, capacity planning, and the continuous improvement of production services.
Requirements & Eligibility
- 5+ years of software engineering or infrastructure engineering experience supporting large-scale production systems.
- A BS in Computer Science, Engineering, Physics, Mathematics, or a related field, or equivalent experience.
- Strong programming experience in Go or Python, including a solid understanding of data structures, algorithms, testing, and software design.
- Experience designing automation for distributed systems and large fleets of Linux-based compute nodes.
- Understanding of performance, security, reliability, fault tolerance, state management, and data consistency in complex systems.
- Experience with infrastructure automation, software deployment, observability, and operational recovery.
- Strong communication skills and the ability to work effectively across teams, organizations, and geographic regions.
- A systematic approach to problem solving, with a strong sense of ownership and an emphasis on reducing operational toil.
Required Skills & Tech
GoPythonLinuxDistributed SystemsSlurmLSFKubernetesBright Cluster ManagerInfrastructure Automation