Site Reliability Manager, Data Center Networking, SRE – Google – Waterloo, ON
Google’s Site Reliability Engineering (SRE) team in Waterloo, Ontario is looking for a Site Reliability Manager, Data Center Networking to lead a high-performing team at the intersection of software engineering and large-scale infrastructure. This is a senior leadership role within Google’s broader SRE organization, where your work directly shapes the reliability and performance of systems that serve millions of GCP customers worldwide.
In this position, you’ll be responsible for building a mission-first culture, scaling your leadership through trusted Tech Leads and domain experts, and driving meaningful improvements to incident detection and mitigation across Software-Defined Networking (SDN) infrastructure. It’s a role that demands both deep technical expertise and the people leadership skills to move a complex, distributed organization forward.
About the Role: Site Reliability Manager, Data Center Networking
As the Site Reliability Manager for Data Center Networking, you’ll serve as the ultimate execution owner for your team’s strategic initiatives. Your primary focus will be on dramatically improving Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM) for incidents — leveraging advanced signaling, tooling, and the integration of new signals into auto-mitigation systems. You’ll also set the standard for developer excellence across the SDN ecosystem, influencing the design and rollout of new network products to ensure they’re introduced safely and deliver high reliability.
Collaboration is central to this role. You’ll partner closely with the PLANET team and sibling SRE shards to define and monitor Network Service Level Objectives (SLOs), co-own blameless postmortems, and establish end-to-end repair coverage for network infrastructure. Google’s SRE culture is grounded in intellectual curiosity, psychological safety, and a commitment to eliminating toil through automation and systems thinking.
Benefits and Salary
This role offers a salary range of CAD $216,000 – $221,000 per year, plus a 20% bonus target, equity, and a comprehensive benefits package. Individual compensation is determined by job-related skills, experience, and relevant education or training. For full details on Google’s benefits, visit their official careers page.
Job Details
🏢 Company: Google
📍 Location: Waterloo, ON, Canada
🆔 Requisition ID: 92399333910946502
💰 Pay: CAD $216,000 – $221,000 per year + 20% bonus target + equity + benefits
Responsibilities
This role spans strategic leadership, technical influence, and operational ownership. You’ll be expected to drive reliability improvements at scale while mentoring and empowering the people around you — all within a culture that values blameless learning and continuous improvement.
- Build and sustain a cohesive, mission-first culture across multiple locations by scaling leadership through trusted Tech Leads and domain experts
- Actively prioritize the team’s workload to ensure sustained high performance and healthy on-call rotations
- Own execution of the team’s strategic efforts, with a focus on drastically improving MTTD and MTTM for incidents through advanced signaling and tooling
- Integrate new signals into auto-mitigation systems to reduce incident impact and manual intervention
- Set the bar for developer excellence across the SDN ecosystem, influencing the design and safe rollout of new network products (NPIs)
- Partner with PLANET and sibling SRE shards to define and monitor Network SLOs and establish end-to-end repair coverage for network infrastructure
- Co-own blameless postmortems to drive systemic improvements following incidents
Requirements / Skills
Google is looking for a leader who brings both deep software engineering expertise and proven experience managing technical teams in complex, distributed environments. The ideal candidate thrives in ambiguity, leads with curiosity, and has a track record of improving system reliability at scale.
- Bachelor’s degree in Computer Science or a related technical field, or equivalent practical experience
- 8 years of experience in software development, including work with data structures and algorithms
- 3 years of experience managing people or teams, leading projects, and designing, analyzing, and troubleshooting distributed systems
- Master’s degree or PhD in Computer Science, Engineering, or a related field is preferred
How to Apply
To apply, visit the official Google Careers posting using the link below. Make sure your resume is up to date and reflects your experience with distributed systems and technical leadership before submitting.
Share This Opportunity
Know someone who might be interested? Share this job posting and help them join Google in Waterloo.
Job Summary & Tips for Applying
Quick Summary & What to Highlight: This Site Reliability Manager, Data Center Networking role at Google in Waterloo is perfect for candidates who excel in distributed systems engineering, technical team leadership, and incident management. On your resume, emphasize any experience with SRE practices, SDN infrastructure, or large-scale system reliability, your approach to on-call health, and your ability to lead in a fast-paced, high-stakes environment. If you’ve previously worked in site reliability engineering, network infrastructure, or engineering management, make sure to highlight specific achievements and responsibilities that align with this position.
Resume & Application Tips: Before applying, tailor your resume to match the job description. Include keywords like Site Reliability Engineering, Mean Time to Detect/Mitigate, and Service Level Objectives (SLOs) that appear in the posting. Quantify your achievements where possible (e.g., “reduced MTTM by 40% through auto-mitigation tooling” or “managed a team of 12 engineers across 3 time zones”). Write a brief cover letter expressing your genuine interest in Google and why you’re excited about this opportunity in Waterloo. Double-check your application for spelling errors and ensure your contact information is current.
Interview Preparation: If selected for an interview, research Google‘s SRE philosophy, published books on Site Reliability Engineering, and the company’s approach to blameless postmortems beforehand. Prepare specific examples using the STAR method (Situation, Task, Action, Result) to demonstrate your technical leadership, incident response, and systems design experience. Common questions may include scenarios about managing on-call escalations, driving reliability improvements under pressure, and influencing cross-functional partners. Dress appropriately for a technology environment, arrive 10–15 minutes early (or log in early for virtual rounds), and bring copies of your resume. Prepare thoughtful questions about the SRE team structure, tooling roadmap, and growth opportunities. After the interview, send a thank-you email within 24 hours reiterating your interest in the position.
Recommended Job Offers
More Google openings near Waterloo, ON