Role Description
We're looking for a Staff Site Reliability Engineer to help design, implement, and operate the infrastructure that powers Bluesky and atproto. This is a hands-on role for someone who has operated high-scale production systems, understands how distributed systems fail, and wants to build the operational foundation for an open social network.
-
Work across bare-metal systems, cloud services, data infrastructure, observability, incident response, capacity planning, and reliability engineering for systems serving millions of users.
-
Own reliability, availability, and operational excellence for our production systems, including observability, incident response, deployment, and rollback systems.
-
Improve production readiness for services, migrations, and infrastructure changes.
-
Develop software that pushes the state of the art in performance, automation, observability, and other areas.
-
Scale systems running on dense, latest-generation, bare-metal servers in our own colocation facilities.
-
Reduce toil through automation, tooling, and thoughtful engineering practices.
-
Partner with engineers across all our teams to help design services with strong operational characteristics.
-
Lead incident reviews and turn contributing factors into concrete engineering improvements as we practice continuous improvement.
-
Perform capacity planning and cost management across compute, storage, database, and networking workloads.
-
Manage various vendor relationships to ensure we can provide high quality services at a reasonable TCO.
-
Mentor engineers on reliability, operability, debugging, and distributed systems practices and help define a culture of operational excellence across the org.
Qualifications
-
+10 years experience operating high-scale production systems, including bare metal.
-
Strong fundamentals in Linux, networking, storage, databases, and distributed systems.
-
Built and operated high-scale systems where correctness, latency, throughput, and availability were critical.
-
Can write production-quality software in Go.
-
Comfortable debugging across application code, operating systems, databases, networks, and hardware.
-
Experience with observability systems, alert design, incident response, capacity planning, Kubernetes, and production automation.
-
Like working on very small, fast-moving teams at a startup.
-
Have read the AT Protocol docs, feel aligned with the mission, and want to contribute!
Requirements
-
Fully remote team, but an overlap of working hours with PST is required.
-
Willingness to travel to team meetups once every 3-4 months.
-
On-site component may be required for interviews.
Benefits
-
Health, dental, and vision insurance.
Additional Notes
-
Equity will be considered in the total compensation package.
-
The final base salary for this role will be based on the individual's geographic location, as well as experience level, skill set, training, licenses, and certifications.