Staff Site Reliability Developer, Protected Data SRE – Google, Waterloo, ON
Location: Waterloo, ON | Company: Google
Google is hiring a Staff Site Reliability Developer to join the Protected Data SRE team in Waterloo, Ontario. This senior technical role focuses on building and operating large-scale, distributed and fault-tolerant systems while improving the reliability, performance and safety of Google’s critical infrastructure.
As a technical anchor for the Protected Data SRE team, you will provide leadership across complex systems, guide cross-functional initiatives and help manage systemic production risks. The position combines software engineering, systems engineering, automation and large-scale infrastructure expertise.
About the Role
Google Site Reliability Engineering combines software and systems engineering to build and operate massively distributed systems. SRE teams work to maintain appropriate levels of reliability and uptime while continuously improving system capacity, performance and operational efficiency.
In this position, you will help reduce infrastructure complexity, develop reusable solutions, improve production safety and provide technical direction for developers working across Google’s infrastructure stack.
Key Areas of Responsibility
The Staff Site Reliability Developer will provide technical leadership for the Protected Data SRE team while developing infrastructure capabilities that improve reliability, observability and operational safety across complex systems.
Reliability Strategy
Drive strategies that reduce complexity across the ecosystem and encourage solution and component reuse to prevent new production risks.
Infrastructure Engineering
Apply software and systems engineering principles to improve the reliability, scalability, performance and maintainability of large distributed systems.
Production Safety
Design capabilities for change safety, distributed observability, large-scale data repair and control plane safety.
Cross-Functional Leadership
Partner with executive development stakeholders and cross-functional programs to balance product reliability requirements with regulatory deadlines.
Technical Direction
Provide technical direction and mentorship to developers in Waterloo while encouraging collaboration across different infrastructure stacks.
Automation & Scale
Help optimize existing systems, build infrastructure and eliminate repetitive operational work through software and automation.
Technical Environment
This role operates within Google’s Technical Infrastructure organization, which develops and maintains the architecture behind Google’s products and services. The team works across infrastructure technologies ranging from large-scale data systems such as Spanner to Google Front End (GFE).
The position requires strong knowledge of Linux and Unix systems, networking, software development and large-scale system design, along with the ability to solve reliability challenges across complex production environments.
Who We’re Looking For
Candidates should have a strong technical background in computer science, software development, systems engineering and infrastructure reliability, together with experience working with Linux or Unix systems and computer networking.
Bachelor’s degree in Computer Science, a related technical field or equivalent practical experience. A Master’s degree in Computer Science or a related technical field is preferred.
At least three years of experience working with Unix/Linux operating system internals and administration.
Experience with computer networking technologies including DNS, load balancing, TCP/IP and routing.
Experience programming in at least one of C, C++, Java, Python or Go.
Five years of experience including product demand and supply planning, production and inventory management, as specified in the official posting.
Availability
This opportunity is based in Waterloo, Ontario, Canada. Candidates should be prepared to collaborate with developers, technical stakeholders and cross-functional programs while providing senior-level technical leadership for the Protected Data SRE team.
Physical Requirements
The official job posting does not specify any particular physical requirements for this position.
Why You’ll Love Working Here
This role provides the opportunity to solve reliability and infrastructure challenges at Google scale while working on systems that support critical internal and externally visible services. You will contribute to company-wide capabilities designed to improve change safety, observability and production reliability.
Google lists a base salary range of $216,000 to $221,000 CAD per year, plus a 20% bonus target, equity and benefits. Individual compensation is determined by factors including skills, experience and relevant education or training.
How to Apply
Interested candidates can submit their application through the official Google Careers website. Review the qualifications and technical requirements carefully before applying and provide an updated resume outlining relevant software development, systems engineering, Linux/Unix, networking and programming experience.
