Role Description
Are you ready to unlock intelligence? If you donβt think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - weβre looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.
Senior SRE, Managed Gateways
The Mission: Kong's Managed Gateways is the fastest-growing product in the Kong portfolio, a SaaS offering with ARR growing multifold. As a Senior Site Reliability Engineer focused on Managed Gateways, you'll be instrumental in architecting and maintaining the resilient, scalable infrastructure that powers Kong's mission-critical managed services β and you'll be the technical face of that product for the enterprise customers who depend on it, directly ensuring the reliability and performance that lets them build the next generation of connected applications at global scale.
What Youβll Do
-
Own production reliability for the platform and act as the senior technical authority for enterprise customers.
-
Lead, mentor, and inspire a high-performing team of Site Reliability Engineers dedicated to Kong's Managed Gateway offerings.
-
Architect and implement robust, scalable, and fault-tolerant cloud-native systems using technologies like Kubernetes, Golang, and major cloud providers.
-
Own the end-to-end operational lifecycle, from proactive monitoring and alerting to incident response and blameless post-mortems, ensuring continuous service improvement.
-
Drive a culture of developer delight by implementing automation, self-service tooling, and streamlined workflows for deploying and managing API gateways.
-
Define, track, and report on key SLOs and SLIs to ensure optimal performance and reliability of Managed Gateways.
-
Champion technical debt prevention and advocate for architectural best practices that enhance system resilience and reduce operational toil.
-
Collaborate cross-functionally with Product, engineering, and Customer Success to influence roadmap decisions and ensure operational readiness for new features.
-
Partner directly with enterprise customers to drive end-to-end onboarding and implementation of Cloud Gateways.
-
Bring deep, cross-cloud breadth (AWS, GCP, Azure) to handle unique customer topologies and turn complex setups into successful, production-ready deployments.
-
Serve as the technical owner of the customer relationship through implementation, primarily supporting our North America customer base.
-
Feed real-world implementation patterns and customer constraints back to Product to further contribute to the roadmap.
Qualifications
-
Extensive experience as a Site Reliability Engineer, focusing on highly available and distributed systems.
-
Deep expertise with Kubernetes and cloud-native architectures, preferably across multiple public cloud providers (AWS, GCP, Azure).
-
Strong proficiency in Golang or similar modern programming languages for automation and tool development.
-
Proven track record in building and maintaining CI/CD pipelines and infrastructure as code (Terraform, Ansible).
-
In-depth knowledge of monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK stack, Datadog).
-
Experience with managed services, API gateways, or similar network infrastructure is highly desirable.
Requirements
-
You take immense ownership of your systems, treating reliability as a first-class feature.
-
You operate with a sense of urgency, especially in critical situations, and drive quick, effective resolutions.
-
You thrive in a collaborative environment, actively sharing knowledge and elevating the entire team.
-
Kong moves fast, and our team's spread across continents and time zones β plans shift mid-flight, and things don't always line up neatly.
-
You bring your own calm to the noise, figure things out as you go, and help the people around you do the same.
Bonus Points
-
Experience with Service Mesh technologies (e.g., Istio, Linkerd).
-
Familiarity with database administration for high-throughput systems (PostgreSQL, Cassandra).
-
Contributions to open-source SRE tools or projects.
-
Relevant cloud certifications (e.g., AWS Certified DevOps Engineer, CKA).