On-Prem AI Infra Engineer | Kubernetes & GitOps
MBR Partners
Our client is a young high-tech company incorporated in the heart of one of the world's fastest-growing
tech hubs—Dubai, UAE.
As the exclusive software partner to one of the world's largest ODMs in the networking equipment space, they develop the Network Operating Systems that power critical data centre and telecom routing & switching infrastructure. Building on this foundation, they have recently launched an AI division focused on designing our own chips to accelerate inference and
training workloads.
What sets them apart is their unique position at the centre of a historic development: our ODM partner
is establishing the first networking equipment factory of its kind in the GCC region, and they are the
software engine driving this groundbreaking initiative.
They are not just building technology—they are building a true networking vendor that serves regional interests while meeting the growing demandfor networking equipment across the MENA region and further.
Their long-term vision extends beyond products to people: creating a thriving ecosystem forembedded systems and ASIC design talent that will produce generations of world-classprofessionals, establishing our region as a global centre of excellence for Enterprise Computeinnovation.
As a rapidly growing company at the forefront of AI hardware innovation, they are constantly seeking talented and motivated individuals to join their team. We offer a dynamic and challenging work environment, with opportunities to make a significant impact on the future of AI technology.
Your Mission
Own the end-to-end design and operation of our on-premise infrastructure for AI and enterpriseworkloads—built as code, automated, observable, and secure. You will architect and runKubernetes clusters for training/inference, manage servers, networks, and core services, andenable developers with reliable CI/CD and platform tooling. This is where minutes, time-to-
recovery and cost-per-job directly impact AI velocity at scale.
Responsibilities
Design and operate on-prem infrastructure as code: author reusable Terraform/Ansible/Helm
modules; build GitOps workflows (e.g., Argo CD) for repeatable, audited changes acrossenvironments.
Build and run Kubernetes for AI: configure multi-tenant GPU clusters (MIG/GPUDirect RDMA,
NVIDIA device plugins/DCGM), scheduling/quotas, HPA/Cluster Autoscaler (where applicable),
and workload isolation.
Administer servers, networks, and core services: OS lifecycle (Linux), identity/SSO
(Keycloak/LDAP), secrets (Vault), DNS/DHCP/NTP, artifact registries, and internal package
mirrors.
Provide storage for AI pipelines: integrate and operate high-bandwidth/low-latency storage,
tune for dataset staging and checkpointing patterns.
Enable CI/CD: partner with developers to design fast, reproducible pipelines (GitLab CI/GitHub
Actions), caching and runners on GPU/CPU nodes, artifact provenance (SBOM, SLSA).
You'll Collaborate WithPlatform and ML engineers running training/inference at scale, silicon and systems teams
integrating hardware in the lab, security engineers safeguarding credentials and supply chain,
application developers delivering services via CI/CD, and site ops supporting data centre
deployments — together, we turn infrastructure into a product that accelerates the business.
Minimum Qualifications
5+ years in DevOps/SRE/Platform Engineering with hands-on ownership of on-prem
environments.
Proven experience operating Kubernetes in production (multi-tenant RBAC, networking/CNI,
storage, ingress, monitoring).
Proficiency with IaC and automation (Terraform, Ansible, Helm; GitOps with Argo CD/Flux).
Strong Linux administration, scripting (Bash/Python), and troubleshooting across the stack
(compute, network, storage).
CI/CD expertise (GitLab CI/GitHub Actions), container build security (SBOM, image signing), and
artifact management.
Solid networking fundamentals (L2/L3, routing, BGP, VLANs, EVPN/VXLAN, load balancing,
TLS/mTLS).
Experience implementing observability (Prometheus/Grafana, logs, tracing) and running incident
response.
Preferred (Nice-to-Haves)
GPU cluster operations for AI (NVIDIA drivers/operator, DCGM, MIG, GPUDirect RDMA, Slurm
integration).
Storage for data-intensive workloads (Ceph, parallel filesystems, NVMe-oF) and performance
tuning.
Secrets/identity platforms (Vault, Keycloak/LDAP/SSO), policy-as-code (OPA/Gatekeeper,
Kyverno).
Security/compliance practices (CIS benchmarks, SLSA, supply-chain scanning) and zero-trust
networking.
Data centre experience (rack/stack, power/cooling basics) and remote site rollout automation.
Familiarity with configuration management for network devices and API-driven switches/routers.
Reproducible environments by default: any engineer can spin up an identical dev/test stack
(K8s namespace, storage, secrets, runners) from Git in ≤30 minutes, with audit trails for every
change.
Solid CI/CD for AI workflows: model/build/test pipelines are deterministic and cache-efficient;
median pipeline time down 30–50%, with artifact provenance (SBOM, signatures) and traceable
datasets/checkpoints.
Predictable GPU orchestration: fair-share scheduling, quotas, and isolation (MIG/namespace
policies) keep queues short; cluster utilization increases >20% without starving latency-
sensitive jobs.
Lab-to-cluster continuity: hardware bring-up images, drivers, and firmware are versioned and
promoted through the same pipelines; new boards/nodes join clusters with push-button
automation.
Actionable observability: dashboards and alerts reflect SLOs meaningful to researchers
(throughput, time-to-first-token, I/O wait, GPU mem pressure); MTTR <30 minutes for priority
services.
Cost & toil reduction: infra tasks automated to eliminate recurring manual work; fewer “custom
one-offs,” more reusable modules; quarterly infra spend per GPU hour trends down.
Clear docs & self-service: engineers rely on concise runbooks and service catalogs; >80% of
routine requests resolved via self-service workflows rather than ad-hoc ops support.
Please note that the client can obtain work visas for Dubai
Please ignore the salary level - there is flexibility depending on the person's profile
#J-18808-Ljbffr- ...OpenShift/Kubernetes Design & Architecture · Extensive experience designing and architecting OpenShift and Kubernetes platforms to support... ...Tekton, Argo CD, Jenkins, and GitLab CI. · Experienced in GitOps methodologies for automating application deployments and infrastructure...
- ...a strong sense of ownership. If this feels like your kind of crew, you’ll probably fit right in. About the Role: As an AI Engineer at spiderSilk, you will design, develop, and deploy artificial intelligence and machine learning models that enhance our cybersecurity...
- ...Navix AI is seeking a talented AI Engineer to join its onsite team in Dubai, United Arab Emirates. This is an exciting opportunity for professionals who are passionate about developing production-ready Artificial Intelligence solutions and building innovative products...
- ...devoted to innovation, excellence, and empowering our clients to achieve their financial goals. Role Overview We are seeking an AI Engineer to design, build, and deploy intelligent AI agents and automated workflows that address real business problems. This is a hands-...
- ...classical ML pipelines (and increasingly, AI pipelines built around LLM calls),... .... You'll manage a team of roughly 5–7 engineers and QAs, growing them technically and professionally... ...and orchestration — Docker and Kubernetes Ownership of deployment and delivery practices...
- ...enterprise-grade applications and platforms with AI built into the core design rather than... ...will directly influence how quickly engineering teams can build, govern, and scale AI-... ...architecture, including containers, Kubernetes, infrastructure automation, observability...
AED6000 - 8000 per month
...We are looking for an AI Engineer to design, build, and own production-grade systems powered by modern AI development tools, automation platforms, databases, and intelligent agents. This role is suited to an experienced, hands-on developer who thinks in databases,...- ...Role Overview We are seeking a Senior AI Architect to design, build, and scale... ...code, build data pipelines, and set the engineering standards the broader AI team will follow... ...Gateway, Istio (service mesh, mTLS); Docker, Kubernetes (Helm, Kustomize); C4 Model (Structurizr...
- ...Implement real-time features using WebSockets and SSE, including AI-driven chat and streaming UX Optimize for performance (... ..., testing, and prototyping Contribute to our AI-assisted engineering culture , exploring and integrating new tools to increase velocity...
- ...Job Summary We are seeking an AI Engineer to lead the introduction and adoption of agentic AI technologies across the group. You will design, develop, and implement enterprise AI solutions that improve productivity, automation, decision-making, and operational efficiency...
- ...integrations such as Rithmic and CQG, and build configurable rule engines that enable risk and operations teams to manage futures trading... ..., you will lead squad delivery, stakeholder collaboration, and AI-native engineering practices to ensure scalable, high-performance...
- ...Future Edge Group is building iPulse, an AI market surveillance hub for investors. Website: We are looking for an AI Engineer to help advance the multi-agent market intelligence platform behind iPulse. The role focuses on applied AI systems for investment research...
- ...enterprise-grade applications and platforms with AI built into the core design rather than... ...will directly influence how quickly engineering teams can build, govern, and scale AI-... ...architecture, including containers, Kubernetes, infrastructure automation, observability...
- ...succeed. You will work with customers, identify high‑impact AI use cases in their SAP landscape, and build end‑to‑end solutions... ...close collaboration with multiple roles at SAP within product & engineering and customer services & delivery. You Will Work directly...
- ### Senior DevOps Engineer (APAC, Remote) The Senior DevOps Engineer... ...responsible for executing Mercans’ AI-native infrastructure strategy... .... The position focuses on Kubernetes-based infrastructure, GitLab... ...environments. * Familiarity with GitOps workflows, deployment...
- ...Ziina is looking for a Senior Data Platform Engineer to join our team. This role is an... ...decision-making, analytics, and future ML/AI capabilities across the company. We're at... ...for hosting our cloud infrastructure and Kubernetes for orchestrating our workloads Terraform...
- ...Role Overview We are seeking a Senior Machine Learning Engineer to join our AI team as a technical owner of ML products and infrastructure... ...and custom pipelines MLOps and Production : Docker, Kubernetes, MLflow, Weights and Biases, Airflow, Dagster, Prefect, GitHub...
- ...Industry to build the next generation of intelligent, cloud-first data platforms. We're looking for a seasoned Enterprise Data & AI Architect who can drive enterprise-wide data strategy, AI innovation, and cloud modernization. Location: UAE (Dubai / Abu Dhabi)...
- ...Senior Software Engineer — Backend Engineering · Dubai, UAE · Full-time Every line... ...the UAE and KSA, and we’re building an AI-first engineering organization. We’re looking... ...and container experience (AWS, Docker, Kubernetes) Comfort with Git-based workflows and...
- ...access global markets at scale. Our Platform Engineering team owns the infrastructure,... ...Hands-on experience operating services on Kubernetes, EKS preferred. Familiar with Infrastructure... ...constraints clearly. X-Factor: AI-Native Engineering You actively use modern...
- ...The Role Ziina is looking for an iOS Engineer to join our team. We have laid a strong foundation upon which we’ve shipped our product... ...searching. AWS for hosting our cloud infrastructure and Kubernetes for orchestrating our workloads. Terraform for IaC Github...
- ...revenue company with 7000+ employees on all continents. Our leading AI technology is the backbone of our award-winning enterprise... ...Architect, you are part of the industry’s most formidable solution engineering force. You work across the IFS AI portfolio, applying deep...
- ...As an AI Product Manager, you will own the discovery, design, and delivery of AI-powered products and workflows that make freight... ...where AI can solve real operational problems, partner closely with engineering and operations to build solutions that hold up in production,...
- ...We are looking for a talented Graphic Designer with strong presentation design skills and good knowledge of AI creative tools. The ideal candidate should be able to create premium, client-ready presentations, realistic mockups and high-quality visual content. Key...
- ...Position: Cloud Platform Engineer Date Posted: June 29, 2026 Industry: Information Technology / Cloud... .... • Experience with CI/CD pipelines, GitOps, monitoring, logging, and observability tools. • Kubernetes and container platform experience is highly preferred...
- ...to reach Net Zero by 2040. What you'll be doing Layla AI is Chalhoub Group's AI-powered beauty assistant, live on the FACES... ...-functional program, supporting coordination across Product, Engineering, Design, CRM, Commerce, Marketing, and Retail as Layla's...
- ...range of tasks await your commitment: As a Cloud Platform Engineer (m/f/d), you design, operate, and continuously enhance our cloud... ...on-premises and Microsoft Azure environments using Docker and Kubernetes (K8s). You configure and troubleshoot Kubernetes and Azure...
- ...Our Mission We're growing our Platform Engineering team and looking for a DevSecOps... ...facing SaaS products running on AWS and Kubernetes, fronted by Cloudflare at the edge. This... ...code, automates relentlessly, and uses AI tooling to move faster without cutting corners...
- ...Overview The Product Manager owns the vision, strategy, and roadmap for spiderSilk’s AI-driven security products. This role sits at the heart of the company — partnering with Engineering, Design, Threat Intelligence, and Go-to-Market teams to turn complex security...
- ...Position: AI Trainer Date Posted: July 12, 2026 Industry: Artificial Intelligence | Technology | Professional Training Employment Type: Full Time Experience: Experienced Professional with Proven AI Training Experience Qualification: Bachelor’s Degree...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to On-Prem AI Infra Engineer | Kubernetes & GitOps. Be the first to apply!
