DevOps and Platform Engineer
We are seeking a skilled DevOps Specialist with 3 + years of hands-on experience in managing production systems, automation, and cloud-native infrastructure. The ideal candidate will have a strong foundation in operations, along with deep expertise in Kubernetes and container orchestration, to ensure scalability, reliability, and performance of system
Key Responsibilities
Must Have:
- Kubernetes – strong production experience
- Linux administration
- Terraform
- Azure / Azure DevOps – hands-on experience required
- CI/CD pipelines
- Bash / Python / Shell scripting
- Prometheus & Grafana
- Docker
- Git
- Networking fundamentals – DNS, Load Balancing, TCP/IP
- Incident management & Root Cause Analysis
What we're looking for:
- Experience running Kubernetes in a real production environment
- Hands-on incident troubleshooting
- Strong RCA and problem-solving capabilities
- Experience managing cloud/platform infrastructure
Good to Have: AWS/GCP, Ansible, ARM/Bicep, Jenkins, GitHub Actions, GitLab CI, ELK, Azure Monitor, Istio/Linkerd, CKA/CKAD or Azure certifications.
Experience and Qualifications
- 3–5 years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or System Engineering roles.
- Hands-on experience managing production environments with a focus on availability, scalability, and operational excellence.
- Experience deploying and managing Kubernetes-based container platforms in cloud or on-premises environments.
- Experience working with at least one major cloud platform (Azure, AWS, or GCP).
- Hands-on experience implementing Infrastructure as Code (IaC) using tools such as Terraform, ARM, CloudFormation, or Ansible.
- Experience designing and maintaining CI/CD pipelines using tools such as Azure DevOps, Jenkins, GitHub Actions, or GitLab CI.
- Experience with Linux system administration, containerization technologies (Docker), and scripting using Bash, Python, or Shell.
- Experience working in Agile/DevOps teams and collaborating with development, infrastructure, and security teams.
- Experience supporting production incidents, troubleshooting complex technical issues, and performing root cause analysis.