Site Reliability Engineer
Responsibilities
Serve as the technical Subject Matter Expert (SME) for implementing, operating, and optimizing microservices on Kubernetes-based cloud platforms.
Collaborate with Cloud Engineering, Software Development, and DevOps teams to deploy and support applications across multi-cloud environments.
Perform load, stress, and resilience testing (including chaos engineering where applicable) to ensure the scalability, availability, and reliability of microservices.
Build observability for microservices and cloud platforms such as AWS, OCI, Azure, and GCP.
Develop, maintain, and validate disaster recovery plans in collaboration with Development and DevOps teams.
Analyze and resolve production performance, scalability, and reliability issues, including CPU, memory, Kubernetes scheduling, autoscaling (HPA), JVM tuning, and resource optimization.
Develop and maintain automation scripts and tools using Python, Go, Bash, or similar scripting languages.
Define, implement, and continuously improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs) with development teams to improve service reliability and business outcomes.
Create and maintain technical documentation, including architecture diagrams, design documents, operational runbooks, and standard operating procedures.
Ensure adherence to security and compliance standards, including ISO 27001, SOC 2, GDPR, and internal security policies where applicable.
Lead production incident response, troubleshooting, and service restoration activities to minimize customer impact.
Conduct post-incident reviews and root cause analyses, driving corrective and preventive actions to improve system reliability.
Evaluate new technologies and participate in proof-of-concept (POC) initiatives to support platform evolution.
Adapt to evolving technologies, processes, and tools while driving continuous improvement.
Mentor and provide technical guidance to junior engineers and team members.
Participate in an on-call rotation and provide production support outside business hours, including weekends, when required.
Perform other duties as assigned.
Requirements
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
5+ years of hands-on experience as a Site Reliability Engineer (SRE).
Proficiency in programming and scripting languages like Java, Python, Bash, or PowerShell.
Hands-on experience in SRE, DevOps, cloud operations, and cloud security best practices.
Strong knowledge of security technologies, including Identity and access management, Network security, Application security, and Data protection.
Strong problem-solving and analytical skills, with the ability to work independently and as part of a team.
Experience in developing and maintaining technical documentation and implementing compliance requirements.
Additional Skills (Preferred)
Expert-level cloud certifications include AWS Solutions Architect, Professional, Azure Solutions
Architect Expert, and GCP Professional Cloud Architect.
Experience with container orchestration technologies (e.g., Kubernetes).
Employer questions
- Which of the following statements best describes your right to work in Singapore?
- What's your expected monthly basic salary?
- Which of the following types of qualifications do you have?
- How many years' experience do you have as a Site Reliability Engineer?
- How many years' experience do you have in an Engineering Role?
- Which of the following programming languages are you experienced in?
- How many years' experience do you have in a DevOps role?