Remote within the United States
We are currently seeking a Senior Director, Site Reliability Engineering to build our Site Reliability Engineering Product Reliability Teams as we support and manage our global cloud platform infrastructure. These teams will be multi-skilled SRE, embedded in our product engineering teams and driving enduring availability and scalable reliability. Zscaler’s cloud platform is one of the world’s largest private clouds delivering Security-as-a-Service to the world’s leading enterprise companies.
You will have the opportunity to learn and challenge yourself technically working in a complex technical environment. Job Description
You will lead a number of geographically dispersed embedded SRE teams who will work hand in hand with product engineering teams. Those teams will ensure products are built to provide enduring availability and scalable reliability as well as:
True observability within the product
High Fidelity and low frequency alerting associated with availability issues
Low touch operational automation services across the Zscaler stack which provides predictable and scalable results and removes humans from many interactions with the cloud
Standardized low touch patching, capacity implementation, and operationalization to allow for seamless scalability and world class Reliability
AI/ML capabilities to improve speed of implementation and adoption
You will be responsible for building embedded SRE capabilities across the Zscaler product lines
Deep experience in driving SRE principles in product design and build
Deep expertise in building Infrastructure as a Service to deliver repeatable outcomes with rapid deployment of services
Be a forcing factor within the organization as we transform to a world leading SRE organization
Take the lead in driving organizational excellence by building scalable process frameworks.
Provide vision, leadership and direction to the team to ensure high performance and fast delivery of team objectives
Bring your passion for, and experience in, SRE to collaborate with the wider SRE organization and engineering teams leadership to ensure SRE is a driving force in allowing the company to scale.
Working with product teams, ensure new services and iterations of existing services are built for reliability, scalability and ease of operations.
Build technical training capabilities to ensure the SRE organization is a destination for people in their career while providing a continuous learning culture.
Support and troubleshoot multiple large-scale distributed software applications and networks
Champion best practices for reliability within Engineering Department
Qualifications
Highly experienced in the culture of SRE in comparable technology environments.
Demonstrable experience in building large scale operational tooling for Cloud Operations or SRE organizations
Demonstrable experience in building world class SRE teams.
Strong communication and interpersonal skills
Decisiveness and ability to make sound judgments under pressure
Problem-solving and analytical skills
Execution focus and proven track record of timely and impactful deliveries
Demonstrable experience in building organizational strategies and ensuring it aligns with company strategies.
Demonstrable experience in driving automation first environments with the ability to build tooling and automations at scale.
Demonstrable experience in leading high performing and geographically dispersed teams
Bachelor’s degree in Computer Science, a related technical field involving computer systems engineering, or equivalent practical experience.
LI-remote LI-AM12
$55,000 — $95,000/year
To apply for this job, please visit the application page

