Lead SRE to operate and automate YugabyteDB DBaaS, ensuring availability and reliability.
•Operate and automate the life cycle of YugabyteDB DBaaS, ensuring database availability and reliability.
•Key Responsibilities Define and drive the technical vision for YugabyteDB’s DBaaS.
•Lead, design, develop, and maintain components of the DBaaS cloud infrastructure.
•Establish processes for handling and leading response to incidents on databases or infrastructure.
•Automate and manage regular maintenance operations such as upgrades.
•Utilize SRE golden signals to analyze and optimize the DBaaS system’s performance and reliability strategies.
•Requirements Strong software design and implementation skills in building infrastructure frameworks.
•Experience in building and managing large-scale distributed systems.
•Experience building and operating data systems for production applications, including fault tolerant designs, software lifecycles, and automation of critical operations.
•Experience with: Relational Database systems (PostgreSQL preferred), Public cloud infrastructure (AWS, GCP, and/or Azure), Containerization tooling, theory and design (Docker, Kubernetes), Infrastructure as Code (Terraform preferred), Configuration Management Tooling (Ansible preferred), Automation Scripting (Python and Bash preferred), Monitoring systems (Prometheus preferred), Version control systems (git preferred), CI/CD systems (GitHub Actions preferred).
•Solid understanding of Linux systems operations and troubleshooting.
•Willingness and ability to learn new languages and concepts.