System Engineer (CaaS)– Role Overview

We are looking for a System Engineer specializing in Container as a Service (CaaS) to design, implement, and manage enterprise-grade container platforms. The role focuses on building and operating scalable, secure, and highly available container orchestration environments (primarily Kubernetes) to support modern application workloads, CI/CD pipelines, and cloud-native services.

Key Responsibilities

1. CaaS Platform Engineering

  • Design, deploy, and maintain Container as a Service (CaaS) platform across cloud and on-prem environments.
  • Build and manage Kubernetes clusters (AKS, EKS, GKE, OpenShift, Rancher).
  • Ensure high availability, scalability, and fault tolerance of container platforms.
  • Manage container runtime environments (Docker, containerd, CRI-O).

2. Infrastructure & Automation

  • Implement Infrastructure as Code (IaC) using tools like Terraform, ARM, Bicep, or CloudFormation.
  • Automate provisioning, scaling, and lifecycle management of container infrastructure.
  • Configure networking components:
    • Ingress controllers (NGINX, Traefik)
    • Service meshes (Istio, Linkerd)
  • Integrate storage solutions (persistent volumes, CSI drivers).

3. Platform Operations & Support

  • Monitor system performance, capacity, and availability.
  • Perform troubleshooting and root cause analysis for platform issues.
  • Manage upgrades, patches, and cluster lifecycle activities.
  • Maintain SLA, SLO, and SLIs for platform reliability.

4. Security & Compliance

  • Implement container security best practices:
    • Image scanning
    • Runtime protection
    • Pod security policies / admission controllers
  • Integrate with DevSecOps pipelines for automated security checks.
  • Manage identity and access control (RBAC, IAM integration).
  • Ensure compliance with enterprise security standards and policies.

5. CI/CD & Integration

  • Integrate CaaS platform with CI/CD pipelines (Azure DevOps, Jenkins, GitHub Actions).
  • Enable automated deployment strategies:
    • Rolling updates
    • Blue-green and canary deployments
  • Support developers in onboarding applications to Kubernetes.

6. Observability & Monitoring

  • Implement monitoring and logging:
    • Prometheus, Grafana
    • ELK/EFK stack
    • Cloud-native monitoring tools
  • Track metrics for:
    • Cluster health
    • Application performance
  • Setup alerting and incident response mechanisms.

7. Cost Optimization & Capacity Planning

  • Optimize compute, storage, and networking utilization.
  • Implement auto-scaling (HPA, VPA, Cluster Autoscaler).
  • Manage cost efficiency across cloud environments.

8. Collaboration & Enablement

  • Work with DevOps, application, and security teams.
  • Provide platform guidance, documentation, and best practices.
  • Conduct knowledge sharing and support onboarding.

Qualifications & Experience

Experience:

  • 5–8 years of experience in Platform Engineering, DevOps, or Cloud Infrastructure.
  • Minimum 3 years of hands-on Kubernetes/CaaS platform administration.
  • Experience with Azure, AWS, or GCP cloud platforms.
  • Experience with CI/CD, Infrastructure as Code, and Linux administration.
  • Experience working in Agile development and operations teams.
  • Experience supporting enterprise-scale production platforms.

Technical Skills

  • Strong experience in Kubernetes administration
  • Hands-on experience with:
    • Docker / container runtimes
    • Kubernetes (AKS/EKS/GKE/OpenShift)
  • Experience with cloud platforms (Azure, AWS, or GCP)
  • Experience in Infrastructure as Code (Terraform, ARM, Bicep)
  • Knowledge of Linux system administration

DevOps & Automation

  • CI/CD pipeline integration experience
  • Scripting skills (Bash, PowerShell, Python)
  • Experience with Git-based version control

Security & Networking

  • Understanding of:
    • Kubernetes RBAC
    • Network policies
    • TLS, certificates
  • Experience with container security tools (Trivy, Aqua, Prisma)