Software Engineering Manager II, Site Reliability Engineering, Infra Bigtable SRE
- Location
- New York
- Workplace
- On-site
- Compensation
- $207k – $300k
About this role
Minimum qualifications:
- Bachelor’s degree in Computer Science, a related field, or equivalent practical experience.
- 8 years of experience with software development in one or more programming languages.
- 3 years of experience managing people or teams.
- 3 years of experience leading projects.
- 3 years of experience designing, analyzing, and troubleshooting distributed systems.
Preferred qualifications:
- Master's degree in Computer Science or Engineering.
- Experience managing multiple user-facing services and the teams which developed and operated them.
- Experience with code, networking, operating systems, and storage, and experience working with executive teams on strategy.
- Experience partnering and building trust with technical teams across the company.
- Experience recruiting and managing a team of engineers on large-scale projects.
About the job:
Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance.
To learn more: check out our books on Site Reliability Engineering or read a career profile about why a Software Engineer chose to join SRE.
Infra Bigtable SRE is responsible for running Bigtable Service, Google's largest structured storage system, and delivering durability, security, reliability, and efficiency.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities:
- Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.
- Own end-to-end availability and performance of key services and build automation to prevent problem recurrence. Automate response to all non-exceptional service conditions.
- Lead by example, mentor the team and establish credibility through quality technical execution.
- Manage on-call rotations across continents, using a follow-the-sun model.
- Design, write and deliver software to improve the availability, scalability, latency and efficiency of Google's services.
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?