DevOps Software Engineer
The job description
Tech stack. CI/CD design, Kubernetes, Terraform, observability (Prometheus, Grafana), incident management, deployment strategies, scripting (Python/Bash), cloud platforms, GitOps
About the role
You will own the systems and practices that let engineers ship software safely and frequently, working at the intersection of development and operations. DevOps engineers here build the deployment pipelines, observability platforms, and reliability practices that the entire engineering organization depends on every day. You will automate away toil relentlessly, make deployments boring and routine, and lead incident response with calm precision when things go wrong. Your customers are your fellow engineers, and their productivity plus their confidence in the platform is your scorecard. Great DevOps work is invisible when it succeeds and invaluable when it prevents the outage nobody ever hears about.
What you will achieve
- Build CI/CD pipelines that take code from commit to production safely: comprehensive automated tests, progressive rollouts, and one-click rollbacks
- Deliver observability platforms with meaningful dashboards, actionable alerts, and distributed tracing that make debugging production issues fast
- Reduce deployment lead time and change failure rate measurably through pipeline improvements tracked with DORA-style metrics over time
- Lead incident response with clear coordination, genuinely blameless postmortems, and systemic fixes that prevent entire classes of future failures
- Cut operational toil through aggressive automation: toil budgets, self-service tooling, and elimination of manual repetitive operational work
What you will bring
Must-haves
- 2 to 5 years in DevOps, SRE, or platform engineering roles with direct production systems responsibility
- Strong CI/CD skills: designing pipelines in GitHub Actions, GitLab CI, Jenkins, or similar with proper testing and deployment stages
- Experience with Kubernetes in production: deployments, networking, persistent storage, troubleshooting, and cluster operations
- Familiarity with infrastructure as code using Terraform: module design, state management, and safe apply workflows with reviews
- Knowledge of observability practices: Prometheus metrics, Grafana dashboards, log aggregation, and alert design that avoids fatigue
- Scripting proficiency in Python or Bash for automation, operational tooling, and glue code between systems
- BS in Computer Science or equivalent experience
Nice-to-haves
- Experience with progressive delivery techniques: canary deployments, feature flags, and automated rollback strategies
- Familiarity with service mesh technologies such as Istio or Linkerd for traffic management and security
- Knowledge of chaos engineering practices and resilience validation performed safely in production
- Experience with on-call program design: humane rotations, escalation policies, and burnout prevention
Google
Meta
Apple
Microsoft
Amazon
Oracle
Netflix
NVIDIA