Role Description
This is a hands-on senior engineering role focused on improving production resilience, strengthening security, driving operational excellence, and enhancing the developer experience across the organization.
In this role, you will:
-
Design, build, and evolve the foundational systems, tooling, and operational practices that enable engineering teams to ship secure, reliable, and scalable software with confidence.
-
Help establish reliability standards, define service level objectives (SLOs), improve observability, automate operational processes, and drive incident management and post-incident learning practices that strengthen platform stability over time.
-
Partner closely with Engineering, Security, Platform, and Product teams to architect scalable distributed systems, optimize Kubernetes and AWS-based infrastructure, and build automated delivery pipelines that support rapid and safe software releases.
-
Play a key role in reducing operational toil, improving system performance, increasing platform reliability, and ensuring that our infrastructure can support continued business growth.
Qualifications
-
8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or related cloud-native engineering roles.
-
Deep expertise in AWS services, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security best practices.
-
Advanced experience operating and scaling production Kubernetes environments.
-
Strong hands-on experience with Istio service mesh, including traffic management, security, observability, and resiliency.
-
Proven expertise with Infrastructure as Code (IaC), preferably using AWS CDK.
-
Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms.
-
Strong troubleshooting, performance optimization, and incident management experience in distributed systems.
-
Excellent communication, collaboration, and technical leadership skills.
Requirements
-
Experience designing and operating monitoring, logging, tracing, and alerting solutions for cloud-native platforms.
-
Strong knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling.
-
Experience defining and operationalizing SLIs, SLOs, alerting strategies, runbooks, and reliability metrics.
-
Proven ability to leverage observability data to improve service reliability, reduce incident impact, and optimize operational performance.
-
Strong proficiency in TypeScript and Node.js for platform engineering, automation, and operational tooling.
-
Experience building and maintaining scalable backend services, APIs, and event-driven systems.
-
Deep understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking.
-
Experience implementing zero-trust architectures, mTLS, and service-to-service security controls.
-
Commitment to high-quality engineering practices, including automated testing, code reviews, and observability-driven development.
-
Strong understanding of resilience engineering, including autoscaling, disruption management, failure testing, and safe deployment strategies.
Benefits
-
Innovation is at our core. We work with cutting-edge technology in accounting and financial reporting, constantly pushing the boundaries to create impactful software solutions.
-
We are committed to a collaborative culture, where your ideas are valued, and knowledge sharing is encouraged within a supportive, inclusive team.
-
Work-life balance is important to us. We offer flexible work options, remote opportunities, and generous time-off policies to ensure a healthy work-life balance.
-
We offer competitive compensation, including a competitive salary and comprehensive benefits such as health insurance and retirement plans.
-
We are driven by impactful work. Your contributions directly affect how our clients manage financial processes and drive their success.
-
Recognition and rewards matter to us. We celebrate hard work through recognition programs, performance bonuses, and opportunities for career growth.
-
We embrace global opportunities. Work on international projects and collaborate with a diverse, global team.