Autodesk
Site Reliability Engineer
Overview
An exciting opportunity is available for a Site Reliability Engineer to join Autodesk's Product Design and Manufacturing Solutions (PDMS) Platform Site Reliability Engineering (SRE) team. In this role, you will wear multiple hats, including first responder, performance analyst, system architect, capacity planner, and monitoring expert.
About Autodesk
Welcome to Autodesk! Amazing things are created every day with our software – from the greenest buildings and cleanest cars to the smartest factories and biggest hit movies. We help innovators turn their ideas into reality, transforming not only how things are made, but what can be made.
Requirements & Eligibility
- 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related role supporting cloud-based applications
- Bachelor's degree in Computer Science or a related technical field
- Advanced hands-on experience with Linux administration, including monitoring, troubleshooting, reliability, performance, and security
- Experience managing large-scale cloud infrastructure, preferably on Amazon Web Services (AWS)
- Strong scripting skills using languages such as Bash, Python, or Perl
- Expert-level knowledge of AWS services, including Amazon Elastic Compute Cloud (EC2), Elastic Container Service (ECS), Elastic Kubernetes Service (EKS), AWS Lambda, Elastic Load Balancing (ELB), Amazon Simple Storage Service (S3), Identity and Access
- Hands-on experience with Docker, Kubernetes, and container technologies
- Proficiency with Infrastructure as Code (IaC) tools such as Terraform and AWS CloudFormation
Key Responsibilities
- Architect and implement hosting solutions for highly dynamic Software as a Service (SaaS) web applications, ensuring reliability, scalability, and performance
- Design, implement, and maintain Infrastructure as Code (IaC) solutions to support scalable, reliable, and secure global environments
- Develop and maintain well-documented engineering standards, processes, and best practices
- Implement infrastructure and application security best practices, including system hardening and the principle of least privilege
- Use modern infrastructure management tools such as Docker, Terraform, Amazon Web Services (AWS) CloudFormation, and AWS Cloud Development Kit (CDK) to manage and deploy containers and virtual machines
- Collaborate with Development, Quality Assurance, and Documentation teams throughout the product development lifecycle to ensure quality and reliability
- Automate operational processes and integrate new technologies to improve efficiency, reliability, and scalability
- Define and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) and manage error budgets to ensure reliability goals are achieved
Disclaimer: Trace Hiring is an independent job board. We are not directly affiliated with Autodesk. Please verify all details on the official company application portal.