What will the Engineer Do
- Infrastructure
as Code:
Architect, develop, and optimize scalable Azure-based infrastructure using
Terraform to ensure consistent and reproducible environments.
- Systems
Reliability: Serve
as a core member of the SRE team, maintaining 24/7 mission-critical
availability for a global digital ecosystem.
- Kubernetes
Orchestration: Deploy,
scale, and manage containerized applications using Kubernetes,
applying industry best practices for cluster health and security.
- Observability
& Monitoring: Implement
and refine advanced monitoring frameworks; collaborate with application
teams to embed deep-link metrics and actionable alerts into services.
- CI/CD
& Automation: Streamline
software delivery pipelines using tools like GitLab, Helm,
and Ansible to accelerate release cycles.
- Incident
Management: Participate
in a "Follow-the-Sun" on-call rotation, leading rapid incident
response and conducting thorough Root Cause Analysis (RCA) to
prevent recurrence. Hands-on experience with observability and monitoring
practices and best standards.
Required Qualifications:
- Education: Bachelor's or Master's
degree in Computer Science, Software Engineering, or a related discipline.
- Experience: 5+ years of professional experience as a
DevOps/Site Reliability Engineer role.
- Cloud
& Containers: Deep technical expertise in Azure (preferred) or
AWS, with hands-on
mastery of Kubernetes administration.
- Automation: Proficient in scripting
using languages such
as Python, Bash, or PowerShell
- Infrastructure
as Code: Demonstrated
experience with Infrastructure as Code (IaC) tools such as Terraform, Helm, and
Ansible.
Preferred Qualifications:
- Experience
with observability tools such as Datadog, Grafana, Prometheus, ELK Stack.
- Experience
using on-call
management platforms such as PagerDuty, Zenduty