Infrastructure Software Engineer
The job description
Tech stack. Go or Python, Kubernetes operators, Terraform, distributed systems fundamentals, cloud APIs, CI/CD for infrastructure, observability, capacity planning, internal platforms, SLO design, error budgets
About the role
You will build the foundational software that the entire engineering organization runs on: internal platforms, developer tooling, and infrastructure services that make everyone faster. Infrastructure engineers write real production software, not just configuration: controllers, operators, CLIs, and APIs that abstract cloud complexity into paved roads engineers want to use. You will treat internal teams as customers, measuring success in their velocity, satisfaction, and the reliability of what you provide. The work combines deep systems thinking with genuine product thinking: understanding what engineers need, then building platforms they choose because they are truly better. You will also define the SLO framework that measures platform reliability from the engineer's perspective, since internal adoption depends on trust earned through transparent performance data.
What you will achieve
- Build internal platforms for deployment, environment provisioning, and developer self-service that measurably reduce time from idea to production
- Deliver Kubernetes operators or controllers that automate complex operational workflows reliably, observably, and safely
- Reduce infrastructure costs through smarter abstractions: shared services, autoscaling policies, and efficient resource utilization across all teams
- Improve platform reliability with SLIs, error budgets, and incident practices applied to internal systems with the same rigor as products
- Drive platform adoption through excellent documentation, hands-on migration support, and feedback loops that make the paved road genuinely attractive
What you will bring
Must-haves
- 2 to 5 years in infrastructure, platform, or backend engineering with production systems experience and operational ownership
- Strong programming skills in Go or Python for building tools, controllers, and automation held to production quality standards
- Experience with Kubernetes: custom resources, operator patterns, or deep operational knowledge of cluster behavior under stress
- Familiarity with infrastructure as code using Terraform and programmatic use of cloud provider APIs
- Understanding of distributed systems concepts: consensus basics, leader election, retry strategies, and graceful degradation
- Customer empathy for internal developers: gathering requirements honestly, designing usable abstractions, and supporting adoption patiently
- BS in Computer Science or equivalent experience
Nice-to-haves
- Experience with platform engineering metrics: DORA, SPACE frameworks, or developer satisfaction measurement programs
- Familiarity with service mesh, API gateways, or internal developer portal frameworks
- Knowledge of multi-cluster or multi-cloud architecture patterns and their operational trade-offs
- Experience with chaos engineering or resilience validation applied to platform infrastructure
Google
Meta
Apple
Microsoft
Amazon
Oracle
Netflix
NVIDIA