Role Description
We are seeking a Site Reliability Engineer to architect, scale, and maintain DeviantArtβs high-throughput AWS infrastructure, supporting over 1.5 billion monthly page views. In this role, you will drive infrastructure automation, optimize high-scale database performance, harden system security, and ensure maximum uptime through robust deployment pipelines and proactive site reliability management.
-
High-Scale AWS Infrastructure:
Architect and maintain auto-scaling AWS systems designed to reliably serve DeviantArt's 1.5B+ monthly page views.
-
High Availability & Incident Response:
Maximize platform uptime and rapidly resolve performance degradation or service outages.
-
Database & Search Optimization:
Scale, maintain, and troubleshoot sharded MySQL databases and search infrastructure (e.g., Vespa.ai) for low-latency queries.
-
CI/CD & Dev Parity:
Build zero-downtime pipelines (Terraform, Kubernetes, GitHub Actions, CodeBuild) and maintain production-parity dev environments.
-
Infrastructure Automation:
Automate resource provisioning and configuration management, backed by clear documentation and automated testing.
-
Security & DDoS Mitigation:
Enforce robust security protocols, defend against targeted DDoS attacks, and execute routine patch management.
-
Cloud Cost Optimization:
Monitor and tune AWS resource utilization to minimize infrastructure spend without compromising site performance.
-
24/7 Operations & On-Call:
Ensure continuous reliability as part of a 3-engineer 24/7 on-call rotation.
Qualifications
-
Minimum of 10 years of experience working with systems at scale in either a Dev Ops, Platform Engineer, or Site Reliability Engineer type role.
-
Excellent analytical skills with the ability to troubleshoot complex problems, analyze system bottlenecks, and implement effective solutions, from frontend through backend systems, sometimes during production degradation or outage.
-
Exceptional command line Linux skills.
-
In-depth knowledge of AWS services, infrastructure as code using Terraform, and container orchestration with Kubernetes.
-
Experience building Docker images, and composing Helm charts.
-
Proficiency in Python and Bash for scripting and automation.
-
Experience with sharded MySQL databases.
-
Proven track record in security compliance standards on large-scale web infrastructures and in-depth knowledge of DDOS mitigation strategies.
-
A proactive mindset in identifying potential issues and taking pre-emptive actions to prevent downtime or performance degradation.
-
Excellent communication skills, open minded, and capable of effectively collaborating with cross-functional teams and articulating technical concepts to non-technical stakeholders.
-
Bonus points: if you have experience with GitOps tools and methodologies.
Benefits
-
Comprehensive health & life benefits including fully covered medical, dental and vision.
-
Generous and flexible time off for personal recharge and special life moments.
-
Emotional and mental health support programs.
-
Allowances supporting nutrition, fitness and productivity.
-
401K matching, with no vesting period.
-
Equity & ESPP - We grant RSUs (restricted share units) to all of our employees.
-
All employees are also eligible for the ESPP (Employee Stock Purchase Plan) allowing employees a unique opportunity to benefit from Wixβs success by purchasing Wix shares at a discounted rate.