System Engineer (CaaS)– Role Overview
We are looking for a System Engineer specializing in Container as a Service (CaaS) to design, implement, and manage enterprise-grade container platforms. The role focuses on building and operating scalable, secure, and highly available container orchestration environments (primarily Kubernetes) to support modern application workloads, CI/CD pipelines, and cloud-native services.
Key Responsibilities
1. CaaS Platform Engineering
- Design, deploy, and maintain Container as a Service (CaaS) platform across cloud and on-prem environments.
- Build and manage Kubernetes clusters (AKS, EKS, GKE, OpenShift, Rancher).
- Ensure high availability, scalability, and fault tolerance of container platforms.
- Manage container runtime environments (Docker, containerd, CRI-O).
2. Infrastructure & Automation
- Implement Infrastructure as Code (IaC) using tools like Terraform, ARM, Bicep, or CloudFormation.
- Automate provisioning, scaling, and lifecycle management of container infrastructure.
- Configure networking components:
- Ingress controllers (NGINX, Traefik)
- Service meshes (Istio, Linkerd)
- Integrate storage solutions (persistent volumes, CSI drivers).
3. Platform Operations & Support
- Monitor system performance, capacity, and availability.
- Perform troubleshooting and root cause analysis for platform issues.
- Manage upgrades, patches, and cluster lifecycle activities.
- Maintain SLA, SLO, and SLIs for platform reliability.
4. Security & Compliance
- Implement container security best practices:
- Image scanning
- Runtime protection
- Pod security policies / admission controllers
- Integrate with DevSecOps pipelines for automated security checks.
- Manage identity and access control (RBAC, IAM integration).
- Ensure compliance with enterprise security standards and policies.
5. CI/CD & Integration
- Integrate CaaS platform with CI/CD pipelines (Azure DevOps, Jenkins, GitHub Actions).
- Enable automated deployment strategies:
- Rolling updates
- Blue-green and canary deployments
- Support developers in onboarding applications to Kubernetes.
6. Observability & Monitoring
- Implement monitoring and logging:
- Prometheus, Grafana
- ELK/EFK stack
- Cloud-native monitoring tools
- Track metrics for:
- Cluster health
- Application performance
- Setup alerting and incident response mechanisms.
7. Cost Optimization & Capacity Planning
- Optimize compute, storage, and networking utilization.
- Implement auto-scaling (HPA, VPA, Cluster Autoscaler).
- Manage cost efficiency across cloud environments.
8. Collaboration & Enablement
- Work with DevOps, application, and security teams.
- Provide platform guidance, documentation, and best practices.
- Conduct knowledge sharing and support onboarding.
Qualifications & Experience
Experience:
- 5–8 years of experience in Platform Engineering, DevOps, or Cloud Infrastructure.
- Minimum 3 years of hands-on Kubernetes/CaaS platform administration.
- Experience with Azure, AWS, or GCP cloud platforms.
- Experience with CI/CD, Infrastructure as Code, and Linux administration.
- Experience working in Agile development and operations teams.
- Experience supporting enterprise-scale production platforms.
Technical Skills
- Strong experience in Kubernetes administration
- Hands-on experience with:
- Docker / container runtimes
- Kubernetes (AKS/EKS/GKE/OpenShift)
- Experience with cloud platforms (Azure, AWS, or GCP)
- Experience in Infrastructure as Code (Terraform, ARM, Bicep)
- Knowledge of Linux system administration
DevOps & Automation
- CI/CD pipeline integration experience
- Scripting skills (Bash, PowerShell, Python)
- Experience with Git-based version control
Security & Networking
- Understanding of:
- Kubernetes RBAC
- Network policies
- TLS, certificates
- Experience with container security tools (Trivy, Aqua, Prisma)