Datacenter Operations Technician
The job description
Tech stack. Monitoring platforms (DCIM, BMS), incident response procedures, ticketing systems, Linux basics, power and cooling systems awareness, shift handover discipline, runbook execution, alerting tools
About the role
You will monitor and respond as an operations technician at a cloud infrastructure company running multiple availability zones serving millions of end users. You are the eyes on the fleet around the clock: watching dashboards, triaging alerts by severity and business impact, executing runbooks precisely, and escalating with complete context when incidents exceed your scope. Your calm, disciplined response during incidents is what keeps small anomalies from becoming customer-visible outages, and your handover quality determines whether the next shift starts strong or starts blind. You will develop deep familiarity with the fleet's failure signatures and seasonal patterns, becoming the operator who senses trouble before the dashboards show it. Your incident notes will be the primary source for post-mortems, so write them like the engineers investigating will read every word.
What you will achieve
- Maintain alert triage discipline with acknowledgment times inside SLA on every shift, ensuring no critical alert goes unnoticed or unowned regardless of hour or workload.
- Execute incident runbooks accurately under pressure, restoring service or escalating with complete diagnostic context within target timeframes measured in minutes.
- Reduce false-positive alert noise meaningfully by identifying misconfigured thresholds and stale monitors, working with engineering teams to tune them.
- Deliver clean shift handovers with documented system state, ongoing issues, and watch items, so the incoming shift starts fully informed and nothing falls through cracks.
- Contribute to post-incident reviews with accurate timelines, factual observations, and improvement suggestions that make future response faster and calmer.
What you will bring
Must-haves
- 2 to 5 years in datacenter operations, NOC, or similar 24/7 infrastructure monitoring roles.
- Experience with monitoring platforms, alerting workflows, and incident management processes.
- Ability to follow runbooks precisely while thinking critically about whether the runbook actually fits the situation in front of you.
- Basic understanding of datacenter power and cooling systems, their alarms, and their failure implications.
- Composure under pressure and clear, concise communication during active incidents.
- Willingness to work rotating shifts including nights, weekends, and holidays.
Nice-to-haves
- ITIL Foundation or similar incident management training.
- Linux or networking fundamentals certification (Linux+, Network+).
- Experience with on-call rotation and paging tools such as PagerDuty or Opsgenie.
- Familiarity with status page communication during customer-impacting events.
Google
Meta
Microsoft
Amazon
Equinix
Digital Realty