Role Description
The Senior Site Reliability Engineer (SRE) will implement, secure, and operate the cloud infrastructure that supports CenCore Groupβs proprietary enterprise SaaS platform. This role is responsible for maintaining a scalable, highly available, secure, and reliable cloud environment as the platform grows and supports enterprise customers.
Key Responsibilities
-
Manage, maintain, and improve AWS-based cloud infrastructure supporting enterprise SaaS operations.
-
Operate and support Kubernetes environments, including Amazon EKS.
-
Own platform reliability, scalability, availability, disaster recovery readiness, and operational resilience.
-
Design and support cloud networking, load balancing, routing, traffic management, and related infrastructure components.
-
Implement and maintain monitoring, alerting, logging, and observability solutions to support proactive issue detection and response.
-
Establish and document operational standards, Service Level Objectives (SLOs), incident response processes, and reliability best practices.
-
Partner with software engineering and product teams to improve application performance, platform stability, and deployment reliability.
-
Apply security best practices across IAM, secrets management, encryption, vulnerability remediation, access controls, and production operations.
-
Support production operations, troubleshoot critical issues, and participate in incident resolution as needed.
Qualifications
-
Active Top Secret clearance with SCI eligibility.
-
Professional experience supporting cloud infrastructure, site reliability, DevOps, platform engineering, or systems engineering functions.
-
Hands-on experience with AWS cloud services and production cloud operations.
-
Experience administering or operating Kubernetes environments.
-
Working knowledge of infrastructure reliability, availability, scalability, incident response, and operational support practices.
-
Experience implementing monitoring, logging, alerting, or observability tools.
-
Ability to troubleshoot complex production issues and coordinate resolution across technical teams.
-
Strong understanding of cloud security fundamentals, including identity and access management, encryption, secrets management, and vulnerability remediation.
-
Ability to document technical processes, standards, and operational procedures.
Preferred Qualifications
-
Experience with AWS services such as EKS, ALB, VPC, CloudFront, Route 53, RDS/Aurora, S3, and IAM.
-
Experience with Terraform or other Infrastructure as Code tools.
-
Experience with monitoring platforms such as Datadog, CloudWatch, Grafana, Prometheus, or similar tools.
-
PostgreSQL administration, performance tuning, or database operations experience.
-
Experience supporting enterprise SaaS, cloud-native applications, or customer-facing production platforms.
-
Experience developing disaster recovery, operational readiness, or production support documentation.
Skills / Competencies
-
Cloud infrastructure operations and automation.
-
Platform reliability, scalability, and performance optimization.
-
Kubernetes administration and containerized application support.
-
Monitoring, observability, and incident response.
-
Cloud security and operational risk awareness.
-
Technical troubleshooting and root cause analysis.
-
Cross-functional collaboration with engineering, product, and operations teams.
-
Clear technical documentation and process improvement.
Work Environment and Physical Requirements
This role is primarily performed in a professional office or remote technology environment, depending on business needs and position requirements. Work involves regular use of a computer, collaboration tools, and cloud-based systems. The position may require participation in production support, incident response, or after-hours troubleshooting as needed. Physical requirements are generally sedentary and include prolonged periods of sitting, computer use, and communicating with internal teams.