Distributed Systems Engineer
The job description
Tech stack. Go/Java/Rust, consensus protocols, distributed databases, message streaming, service mesh, chaos engineering, distributed tracing, capacity planning, fault injection, replication design, partition tolerance, leader election, gossip protocols
About the role
You will design and build the distributed systems that keep working when individual machines fail, networks partition, and traffic spikes without warning. Distributed systems engineers think in terms of failure modes first: every serious design starts with what happens when things break, because they will. You will work on consensus, replication, sharding, and the coordination problems that make distributed computing genuinely difficult and deeply interesting. This is the specialty for engineers who find single-machine programming insufficiently challenging and want their systems to survive contact with the real world. You will also run the chaos engineering program that injects failures into production on schedule, proving the system's resilience claims instead of merely asserting them.
What you will achieve
- Design distributed services with correct handling of network partitions, node failures, and clock skew, validated through systematic fault-injection testing
- Deliver measurable availability improvements: higher uptime percentages, faster automated failover, and graceful degradation under partial failures
- Build data systems with well-reasoned consistency guarantees: choosing the right model for each workload and documenting the trade-offs explicitly
- Implement effective load distribution: sharding strategies, online rebalancing, and hot-spot mitigation backed by production metrics
- Create runbooks and automation for distributed failure scenarios so on-call engineers resolve incidents quickly, correctly, and confidently
What you will bring
Must-haves
- 2 to 5 years building distributed systems in production: multi-node services with genuine consistency and availability requirements
- Deep understanding of distributed fundamentals: CAP theorem implications, consensus protocols like Raft and Paxos, vector clocks, and quorum design
- Strong programming skills in Go, Java, or Rust with attention to concurrency correctness, networking behavior, and performance under load
- Experience operating distributed data systems: Cassandra, CockroachDB, Kafka, or similar, including their operational realities and failure modes
- Familiarity with failure testing practices: chaos engineering, fault injection frameworks, and regular game-day exercises
- Knowledge of distributed tracing and observability techniques for understanding emergent behavior across service boundaries
- BS in Computer Science or equivalent experience
Nice-to-haves
- Experience implementing consensus protocols or contributing to distributed systems open-source projects
- Familiarity with CRDTs and their application to collaborative editing or offline-first systems
- Knowledge of formal verification methods such as TLA+ for validating distributed protocol designs
- Experience with edge computing architectures or geo-distributed system design challenges
Google
Meta
Apple
Microsoft
Amazon
Oracle
Netflix
NVIDIA