Distributed Systems Engineer III @Mozn
Software Development
Salary unspecified
Remote Location
Employment Type full-time
Posted Today

[Hiring] Distributed Systems Engineer III @Mozn

Today - Mozn is hiring a remote Distributed Systems Engineer III. πŸ’Έ Salary: unspecified πŸ“Location: Egypt

Role Description

We are looking for a highly motivated Distributed Systems Engineer III to join our Cloud Platform Engineering team. This role focuses on building and operating reliable, scalable, and resilient cloud platforms for distributed and data-intensive workloads. You will work across Kubernetes, cloud infrastructure, messaging systems, databases, automation, and platform reliability.

The role requires strong hands-on experience with Kafka, Kubernetes, and at least one relational database such as MySQL or PostgreSQL, along with a solid understanding of distributed-systems fundamentals.

As our platform evolves, you will also contribute to AI and data infrastructure, helping build the underlying platform capabilities required to run data-intensive and AI.

What you'll do

  • Cloud Platform & Distributed Systems
    • Build, operate, and continuously improve production cloud-native platforms running distributed workloads.
    • Work hands-on with Kubernetes, including upgrades, node pools, workload lifecycle, troubleshooting, and platform operations.
    • Deploy and manage workloads using ArgoCD, Helm, GitOps, Terraform, and automation.
    • Design and operate systems with a focus on scalability, availability, resilience, performance, and operational simplicity.
    • Troubleshoot complex issues across Kubernetes, cloud infrastructure, networking, storage, applications, and distributed services.
  • Kafka & Data Infrastructure
    • Operate and troubleshoot Apache Kafka in production across high-throughput and distributed workloads.
    • Work with topics, partitions, replication, consumer groups, retention, throughput, latency, and failure recovery.
    • Integrate Kafka with databases and applications using technologies such as Kafka Connect, Debezium, or similar CDC/event-streaming platforms.
    • Operate and troubleshoot MySQL and/or PostgreSQL, including replication, high availability, backup, recovery, performance, and migrations.
    • Support data-intensive workloads and analytical platforms such as StarRocks, ClickHouse, Apache Doris, or similar technologies.
  • Reliability, DR & Multi-Tenant Platforms
    • Design and operate platforms that remain resilient across node, service, zone, and infrastructure failures.
    • Implement and validate backup, recovery, disaster recovery, failover, and business-continuity capabilities.
    • Understand the fundamentals of multi-zone, multi-region, and active-active architectures and apply them where appropriate.
    • Build platforms that support multiple tenants and workloads, with appropriate isolation, scalability, resource management, and reliability.
    • Participate in DR exercises, failure simulations, migrations, and other resilience initiatives.
    • Understand distributed-system trade-offs involving replication, consistency, availability, partitioning, fault tolerance, latency, and throughput.
  • Automation, Observability & Operations
    • Automate infrastructure and platform lifecycle operations using Terraform, Python, Bash, Go, or similar technologies.
    • Build reliable deployment and GitOps workflows and reduce manual operational effort.
    • Implement effective monitoring, logging, alerting, and observability for distributed workloads.
    • Participate in production incident response, root-cause analysis, and long-term reliability improvements.
  • AI & Emerging Platform Infrastructure
    • Contribute to the infrastructure needed to support AI, machine-learning, and data-intensive workloads.
    • Help evolve cloud and Kubernetes platforms to support AI workloads, data pipelines, model-serving infrastructure, and associated platform services.
    • Understand the infrastructure requirements around compute, GPUs, networking, storage, data movement, observability, and workload isolation for AI platforms.
    • Work with engineering teams to build reusable platform capabilities that enable AI and data workloads to run reliably at scale.
    • Stay current with emerging infrastructure patterns across AI platforms, distributed data systems, and cloud-native technologies.

Qualifications

  • 4–7 years of experience in Platform Engineering, Infrastructure Engineering, Distributed Systems, SRE, Backend Engineering, Data Infrastructure, or a related field.
  • Strong production experience with Apache Kafka β€” mandatory.
  • Hands-on production experience with Kubernetes β€” mandatory.
  • Strong experience with at least one of MySQL or PostgreSQL β€” mandatory.
  • Solid understanding of distributed-systems fundamentals including replication, partitioning, consistency, availability, fault tolerance, scalability, and failure recovery.
  • Experience with ArgoCD/GitOps and infrastructure-as-code such as Terraform.
  • Experience operating workloads on a public cloud such as GCP, OCI, AWS, or Azure.
  • Strong production troubleshooting and incident-resolution skills.
  • Experience with automation or scripting using Python, Bash, Go, Java, or similar languages.
  • Understanding of high-availability, disaster-recovery, and multi-tenant architecture fundamentals.
  • Experience with observability and operational tooling such as Prometheus, Grafana, OpenSearch/ELK, LGTM, or equivalent.
  • Strong understanding of infrastructure and networking fundamentals in cloud-native environments.

Good to Have

  • Experience with Kafka Connect, Debezium, Kafka Streams, or CDC platforms.
  • Experience with distributed analytical databases such as StarRocks, ClickHouse, Apache Doris, or similar.
  • Experience with Flink, Spark, or other distributed data-processing systems.
  • Experience executing large-scale data, database, application, or infrastructure migrations.
  • Experience with active-active, multi-zone, or multi-region systems.
  • Experience operating stateful workloads on Kubernetes.
  • Experience supporting AI/ML infrastructure or GPU-based workloads.
  • Experience with cloud networking, service mesh, ingress, load balancing, or storage platforms.
  • Contributions to Kubernetes, Kafka, distributed-systems, or other open-source infrastructure projects.

Benefits

  • You will be at the forefront of an exciting time for the Middle East, joining a high-growth rocket-ship in an exciting space.
  • You will be given a lot of responsibility and trust. We believe that the best results come when the people responsible for a function are given the freedom to do what they think is best.
  • The fundamentals will be taken care of: competitive compensation, top-tier health insurance, and an enabling culture so that you can focus on what you do best.
  • You will enjoy a fun and dynamic workplace working alongside some of the greatest minds in AI.
  • We believe strength lies in difference, embracing all for who they are and empowered to be the best version of themselves.
Before You Apply
️
remote Be aware of the location restriction for this remote position: Egypt
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Distributed Systems Engineer III @Mozn
Software Development
Salary unspecified
Remote Location
Employment Type full-time
Posted Today
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
️
remote Be aware of the location restriction for this remote position: Egypt
β€Ό Beware of scams! When applying for jobs, you should NEVER have to pay anything. Learn more.
Apply for this position
Did not apply βœ“
Applied βœ“
Sent Follow-Up βœ“
Interview Scheduled βœ“
Interview Completed βœ“
Offer Accepted βœ“
Offer Declined βœ“
Application Denied βœ“
Unlock 125,000+ Remote Jobs
Γ—

Apply to the best remote jobs
before everyone else

Access 125,000+ vetted remote jobs and get daily alerts.

4.9 β˜…β˜…β˜…β˜…β˜… from 500+ reviews

⚑ 127,694+ remote jobs, refreshed hourly

πŸ”” Real-time alerts: Apply first

πŸ›‘οΈ Vetted companies, no scams, true remote only

Unlock All Jobs Now

Maybe later