Replit
Staff Site Reliability Engineer
Overview
As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability.
About Replit
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.
Requirements & Eligibility
- 8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering).
- Strong programming skills in languages like Python or Go.
- Deep understanding of distributed systems.
- Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies.
- Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions.
- Strong incident management skills with extensive experience leading incident response for complex systems.
- Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools.
- Excellent written and verbal communication skills.
Key Responsibilities
- Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions.
- Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
- Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution.
- Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work.
- Optimize Performance on Kubernetes: Collaborate with core infrastructure and product teams to performance-tune and optimize our large-scale cloud deployments.
- Debug and Harden Distributed Systems: Dive deep into debugging extremely difficult technical problems across the stack.
- Provide Staff-Level Guidance: Review feature and system designs from across the company.
- Educate and Mentor: Educate, mentor, and hold accountable the broader engineering team to improve the reliability of our systems.
Disclaimer: Trace Hiring is an independent job board. We are not directly affiliated with Replit. Please verify all details on the official company application portal.