Cloud DevOps Engineer
We are looking for a Cloud & DevOps Automation Engineer to own the platform layer beneath our automation practice - one of the core service categories on the REWORK platform. Client automations, agents, and pipelines have to run somewhere: you will provision that somewhere with infrastructure-as-code, wire the CI/CD that ships changes safely, containerize and deploy workloads across serverless and VPS targets (including self-hosted n8n and agent runtimes), and build the observability that catches failures before clients do. This is an SRE-minded role: least-privilege IAM and secrets management, backups and restore drills, cost budgets with alerting, and runbooks another engineer can execute at 2 a.m.
Core Responsibilities
Infrastructure & Deployment
- Provision client and internal environments with infrastructure-as-code (Terraform) - reproducible, reviewable, destroyable
- Build CI/CD pipelines (GitHub Actions) with environment promotion, rollbacks, and secret hygiene
- Containerize and deploy workloads across serverless (Cloud Run, Lambda) and VPS targets, including self-hosted n8n and agent runtimes
Reliability & Observability
- Instrument logging, metrics, uptime checks, and alerting so failures page before clients notice
- Design backup/restore and disaster-recovery paths - and actually run the restore drills
Security & Cost
- Enforce least-privilege IAM, rotate credentials, and manage secrets across environments
- Track and optimize cloud spend with budgets, alerts, and right-sizing - automation margins live and die on infra cost
Required Qualifications
Required Experience
- Production infrastructure you provisioned, deployed, and kept running for real workloads
- 1+ years hands-on with at least one major cloud (GCP or AWS), Docker, and a CI/CD system, plus solid Linux fundamentals
System Design & Problem-Solving
- Thinks in failure modes: blast radius, rollback paths, idempotent deploys, and graceful degradation
- Comfortable debugging across the stack - DNS and TLS, container networking, IAM denials, and noisy-neighbor performance
Documentation & Nice-to-Haves
- Writes runbooks, architecture diagrams, and incident notes another engineer can execute without you
- Nice to have: Kubernetes, Fly.io / Railway, Firebase / Cloud Functions ecosystems, SOC 2-style hardening experience
Key Deliverables
Environments provisioned as code and reproducible from scratchDeploys shipping through CI/CD with rollback paths that have been exercisedMonitoring and alerting that catches failures before clients report themDocumented runbooks, backups verified by restore drills, and infra cost kept in budget
Tech Stack & Skills
Terraform / IaCGitHub Actions CI/CDDockerGCP (Cloud Run / Functions)AWS (Lambda / ECS)Linux & VPS opsn8n self-hostingNginx / CaddyGrafana / CloudWatchSecrets managementIAM & least privilegeBash / Python
Expectations
- Work asynchronously in a remote-first environment (Slack, email, documented reports)
- Operate highly independently - evaluated on proof-of-work and reliability of deliverables
- Collaborate during core hours, 10:00 AM – 4:00 PM Eastern Time
- Keep infrastructure, pipelines, and runbooks clear, well-documented, and easy to navigate