Platform / DevOps / MLOps engineer (SDE II). I write the Python services, pipelines and Kubernetes clusters behind a production computer-vision platform: 50+ GPU nodes and 2,000+ cameras for a Saudi municipality, plus air-gapped government clusters with no internet egress. Some recent work:
- Cut a 55h model-training run to 5h by moving training to PyTorch DDP on on-demand k8s Jobs (pause/resume, MLflow, Kafka progress events).
- Found that every meeting-recording bot pod was pinned to a single node by a ReadWriteOnce volume plus a required podAffinity. Bots now record locally and the API streams recordings out as a tar over the k8s exec channel, so they spread across all workers.
- Caught a checked-in manifest that would have deleted 22 live production secrets, shipped a strategic-merge patch instead and reconciled the repo to zero drift.
- Took full platform setup from 3 days to 20 minutes, and built Vault-backed, approval-gated CI/CD that saves a client SAR 300K/yr.
Open source: kuiqctl, a kubeadm cluster that stays Ready when your laptop changes Wi-Fi (<a href="https://github.com/Zafeeruddin/kuiqctl" rel="nofollow">https://github.com/Zafeeruddin/kuiqctl), and causeway, which lets you view and record cameras behind VPNs and jump hosts with per-customer network namespaces and 298 tests (<a href="https://github.com/Zafeeruddin/causeway" rel="nofollow">https://github.com/Zafeeruddin/causeway).
Looking for: platform, infrastructure, SRE or MLOps roles, especially where GPUs, on-prem or messy networks are involved.
itszafeer · · focus · HN ↗
Remote: Yes
Willing to relocate: Yes, open to discussing (UAE / KSA or elsewhere)
Technologies: Python, Bash, TypeScript, Kubernetes (kubeadm, bare-metal HA), Docker, Argo CD, Jenkins, GitHub Actions, Terraform, Ansible, Vault, Harbor, NGINX/HAProxy/Keepalived, PyTorch DDP, MLflow, Kafka, ClickHouse, Postgres, Redis, Prometheus/Grafana, AWS (EKS), Alibaba Cloud, Cloudflare
Résumé/CV: <a href="https://zafeer.dev/resume" rel="nofollow">https://zafeer.dev/resume
Website: <a href="https://zafeer.dev" rel="nofollow">https://zafeer.dev
Email: mohammed.xafeer@gmail.com
---
Platform / DevOps / MLOps engineer (SDE II). I write the Python services, pipelines and Kubernetes clusters behind a production computer-vision platform: 50+ GPU nodes and 2,000+ cameras for a Saudi municipality, plus air-gapped government clusters with no internet egress. Some recent work:
- Cut a 55h model-training run to 5h by moving training to PyTorch DDP on on-demand k8s Jobs (pause/resume, MLflow, Kafka progress events).
- Found that every meeting-recording bot pod was pinned to a single node by a ReadWriteOnce volume plus a required podAffinity. Bots now record locally and the API streams recordings out as a tar over the k8s exec channel, so they spread across all workers.
- Caught a checked-in manifest that would have deleted 22 live production secrets, shipped a strategic-merge patch instead and reconciled the repo to zero drift.
- Took full platform setup from 3 days to 20 minutes, and built Vault-backed, approval-gated CI/CD that saves a client SAR 300K/yr.
Open source: kuiqctl, a kubeadm cluster that stays Ready when your laptop changes Wi-Fi (<a href="https://github.com/Zafeeruddin/kuiqctl" rel="nofollow">https://github.com/Zafeeruddin/kuiqctl), and causeway, which lets you view and record cameras behind VPNs and jump hosts with per-customer network namespaces and 298 tests (<a href="https://github.com/Zafeeruddin/causeway" rel="nofollow">https://github.com/Zafeeruddin/causeway).
Looking for: platform, infrastructure, SRE or MLOps roles, especially where GPUs, on-prem or messy networks are involved.