About the role
This role owns the platform, from cluster architecture and delivery pipelines to the production systems that run beneath them. It sets the engineering standard across teams while remaining deeply hands-on. This is a leadership role by ownership and technical influence rather than team size.
You'll define how the platform is built, deployed, operated, and improved, helping engineering teams ship reliable software while maintaining the security, stability, and governance expected by enterprise and government clients.
What you'll do
- Design and run production Kubernetes clusters: networking, ingress, autoscaling, security.
- Own GitOps delivery with Argo CD: app-of-apps, sync policies, drift detection, clean rollbacks.
- Manage the container supply chain: image design, registries, scanning, signing.
- Write infrastructure as code (Terraform) that others can safely change.
- Run zero-downtime migrations and cutovers.
- Build observability that matters: metrics, logs, traces, SLOs, actionable alerts.
- Harden the platform for regulated GCC clients: data residency, access, audit trails.
- Carry the on-call rotation and lead incident reviews.
Must have
- Around six years in infrastructure or platform engineering, including at least three years operating Kubernetes in production, and experience being the engineer the team relies on before making production changes.
- Deep Kubernetes: scheduling, networking, RBAC, and how clusters actually fail.
- Strong Docker skills: image design, layering, registry hygiene.
- Argo CD in production.
- GitOps as your working model, not a slide.
- Terraform, or equivalent infrastructure as code.
- CI/CD ownership: pipelines you designed.
- Cloud depth (Azure or similar) with solid networking and IAM.
- On-call temperament: calm and methodical under pressure.
Nice to have
- You've built or migrated a platform for a government body, a bank, a telco, or another highly regulated organization where change control and formal approvals are simply part of the job.
- Experience working in audited environments with release windows, separation of duties, and evidence-based delivery.
- Experience delivering through a systems integrator (SI) or alongside multiple vendors where your platform formed one part of a larger programme.
- Service mesh, progressive delivery, or policy-as-code (Istio, Argo Rollouts, OPA/Kyverno).
- Platform work behind AI or GPU workloads.
- FinOps instinct: you know what the cluster costs.
- Supply-chain security (SBOMs, image signing).
- GCC regulatory experience (DIFC, ADGM).
How you work
We value these qualities as highly as your technical skills.
- You write clearly. Whether it's a design proposal, an incident review, or an email to a client's security team, people understand it the first time.
- You take ownership of your work and raise concerns early when something needs attention.
- You enjoy sharing knowledge and helping other engineers grow.
- You can work comfortably with project managers, architects, security teams, and compliance officers without hiding behind technical jargon.
- You stay calm under pressure, and your incident reviews focus on learning and prevention.
- You're honest about risk. If something isn't ready to ship, you'll say so before it reaches production.
- You treat client systems, access, and data with the care they deserve.
- You work at a sustainable pace. We believe in shared ownership, proper handovers, and solving problems through engineering rather than heroics.
- You know when the elegant solution is the right one, and when the practical solution is the better choice.
- Good working English. Arabic is an advantage for client-facing engagements.
Stack
- Kubernetes
- Docker
- Argo CD
- Terraform
- GitHub Actions
- AWS
- Prometheus
- Grafana
- Linux
