Overview
Site Reliability Engineer Jobs in Sandton, Gauteng, South Africa at Sun International
Title: Site Reliability Engineer
Company: Sun International
Location: Sandton, Gauteng, South Africa
The Site Reliability Engineer (SRE) is a technical contributor responsible for improving the reliability, resilience and operational performance of Sun International’s technology services, platforms and infrastructure. The role focuses on observability, monitoring, alerting, event management, reliability analytics, automation and auto-remediation, proactively identifying potential issues before they impact the business. Working across software engineering, infrastructure, enterprise applications and service management teams, the SRE implements monitoring and automation solutions that strengthen service reliability and reduce operational.
Core behavioural & Technical / proficiency competencies:
- Cloud platform management across Azure, AWS and/or GCP.
- Containerisation technologies, including Docker and Kubernetes.
- Monitoring, observability and alerting tools such as Prometheus, Grafana and ELK.
- Scripting and operational automation using Python, Bash and/or PowerShell.
- Development of automation, auto-remediation capabilities and operational runbooks.
- Incident management and proactive identification of service reliability risks.
- Root Cause Analysis (RCA) and analysis of recurring operational failures.
- Linux and Windows system administration.
- Networking fundamentals.
- Git version control.
- Troubleshooting and problem-solving across technology services and infrastructure.
- Analysis of operational telemetry, event data and service performance trends.
- Application of reliability standards, security requirements, operational controls and governance practices.
- Cross-functional collaboration with software engineering, infrastructure, enterprise applications and service management teams.
- Analytical thinking and evidence-based decision-making.
- Continuous improvement and operational excellence.
- Collaboration and knowledge sharing.
- Operational excellence and accountability.
Qualifications:
- Degree in Computer Science, Engineering, Information Technology or a related discipline (required)
- Cloud platform certification, e.g. Azure Administrator Associate or AWS Certified SysOps Administrator (preferred)
- ITIL Foundation Certification (preferred)
Experience:
- 2-5 years’ experience in infrastructure operations, technology operations, monitoring platforms, cloud operations, Site Reliability Engineering, observability tooling or a related technology discipline.
- Practical experience with monitoring and observability, cloud platforms, automation/scripting, incident management and troubleshooting aligned to an SRE or technology operations environment
#ZT