Role Description
Weβre looking for a Cloud & Infrastructure Engineer to build, secure, and maintain scalable platform infrastructure and delivery pipelines. The ideal candidate combines hands-on experience with managed PaaS platforms and core AWS services (RDS, S3, Lambda) with strong capabilities in Infrastructure as Code, automated CI/CD, and transactional email deliverability. This role will collaborate closely with project leads and software engineers, establish comprehensive observability and security controls, and deliver reliable, automated, and well-documented cloud environments.
Qualifications
-
3+ years of engineering experience
-
PaaS - production experience running applications on a managed platform, and clear judgment about where the platform's abstractions stop being enough.
-
AWS RDS - provisioning, tuning, backup and recovery, and diagnosing the database problems that surface as application slowness.
-
AWS S3 - bucket policies, lifecycle management, encryption, and secure handling of user-uploaded content.
-
AWS Lambda - building, deploying, and operating serverless functions, including their less obvious failure modes.
-
SMTP services - configuring transactional email, domain authentication (SPF, DKIM, DMARC), and deliverability management.
-
CI/CD - you have built pipelines a team actually relied on, and have made deploying boring for people who were previously nervous about it.
Requirements
-
Assess the current infrastructure β what is running, how it is configured, what is manual, and where the risk sits β and produce a written assessment with prioritized findings.
-
Define and implement the target infrastructure across PaaS platform, e.g., Heroku, Render, Elastic Beanstalk, App Runner and supporting AWS services, with environment separation for dev / staging / production.
-
Define infrastructure as code so environments are reproducible rather than hand-configured. Terraform / CloudFormation / CDK / Pulumi, as agreed.
-
Implement IAM roles and least-privilege access, secrets management, and network configuration (VPC, security groups, TLS).
-
Provision and tune the database β instance sizing, parameter groups, connection pooling, read replicas if warranted β and configure automated backups, point-in-time recovery, and encryption at rest.
-
Run and document a restore test. A backup that has never been restored is not a backup.
-
Configure buckets for uploads / assets / exports / logs with correct access policies, lifecycle rules, versioning, and encryption. No public buckets that should not be public.
-
Set up CDN delivery for static and user-uploaded assets where appropriate.
-
Build and deploy the functions in scope β e.g., scheduled jobs, event-driven processing, webhook handlers, image or file processing β with appropriate triggers, concurrency limits, and timeouts.
-
Handle failure properly: retries, dead-letter queues, idempotency, and alerting on repeated failure.
-
Bring Lambda deployments into the same CI/CD pipeline and IaC definitions as the rest of the stack, not a separate manual path.
-
Configure transactional email through SES / SendGrid / Postmark / other for e.g., account, notification, and system emails.
-
Set up domain authentication β SPF, DKIM, DMARC β and manage sending reputation, warmup, and dedicated IP if warranted.
-
Configure bounce, complaint, and suppression handling, plus deliverability monitoring so a drop in inbox placement is visible before users report it.
-
Build pipelines in GitHub Actions / GitLab CI / CircleCI / other that run tests, lint, build, and deploy on merge, with environment promotion from staging to production.
-
Implement deployment strategy and rollback: blue-green / rolling / canary, as agreed, with a rollback that is a single documented action rather than an improvisation.
-
Automate database migrations as part of deploy, safely and reversibly.
-
Reduce pipeline runtime so the team is not waiting on it.
-
Set up centralized logging, metrics, and alerting through CloudWatch / Datadog / other, with alerts that page on real problems and stay quiet otherwise.
-
Define and monitor the health indicators that matter for the team: uptime, error rate, latency, queue depth, job failure.
-
Review the stack against HIPAA / PCI / SOC 2 / other requirements where applicable, and document what is in place and what is not.
-
Establish cost visibility and tagging; identify and remove waste.
Deliverables
-
Infrastructure assessment with prioritized findings, reviewed with project lead.
-
Architecture diagram of the target infrastructure, with data flow and trust boundaries.
-
Infrastructure as code covering all environments, in repository.
-
Configured and hardened RDS, S3, Lambda, and PaaS environments, with a documented and tested restore procedure.
-
Transactional email configured and authenticated, with deliverability monitoring in place.
-
CI/CD pipelines deploying dev / staging / production with automated tests, migrations, and one-step rollback.
-
Monitoring, logging, and alerting, with an on-call runbook covering the top failure modes and what to do about each.
-
Cost baseline and tagging scheme, with identified savings.
-
Handoff package: infrastructure documentation, access inventory, known limitations, recommended next steps, and a walkthrough with the team.
Nice to Haves
-
Infrastructure as code at depth in Terraform / CloudFormation / CDK.
-
Containers and orchestration where the stack calls for it.
-
Experience with HIPAA / PCI / SOC 2 compliance in a cloud environment.
-
Domain experience in medical software and/or e-commerce.
-
Prior contract work where you owned infrastructure end to end and handed it off cleanly.