Senior Site Reliability Engineer
ASAPPGenerative AI company
New York, United StatesSenior
Fidelity Investments
Dragoneer Investment Group
Emergence Capital
March Capital
John Doerr
Euclidean Capital
Software Engineering
About the role
TL;DR
Ensure the performance and reliability of ASAPP's infrastructure and products.
- •ASAPP seeks a Senior Site Reliability Engineer to ensure the performance and reliability of our infrastructure and products.
- •Key Responsibilities Work with product engineering teams on service architecture and implementation Deliver Infrastructure configuration as code and automate everything Direct and implement monitoring and alerting systems to support rapid problem diagnosis Perform Root Cause Analysis and design and deliver resolutions Work on our Kubernetes / AWS infrastructure to support our product engineers Design secure and performant networking solutions in our production systems Requirements +4 years of relevant experience bringing software to production at high scale Participation in on-call rotation, triaging and addressing production issues Obsession with automation and instrumentation Understanding of complex systems and failure scenarios Excellent communication skills Knowledge of AWS services, containers and container management frameworks Familiarity with Message Bus based systems and distributed architectures Proficiency in Terraform, Python and/or Go
Required skills
AWSTerraformPythonKubernetesCI/CDPrometheusGrafanaDatadog
Nice-to-have skills
Go
Benefits & perks
Competitive compensation with stock options, Comprehensive medical, vision, and dental insurance, 401k matching, Fitness and wellness stipend, Mobile phone reimbursement, Mental well-being benefits, Professional learning and development stipend, Parental leave, including adoptive and foster parents, 3 weeks paid time off (increases with tenure) and unlimited sick leave
Tech stack
AWSTerraformPythonKubernetesCI/CDPrometheusGrafanaDatadog