Data Software Engineer
The job description
Tech stack. Python, SQL, Airflow or similar orchestration, Spark, data warehousing (Snowflake, BigQuery), dbt, data quality frameworks, stream processing, data modeling
About the role
You will build the data pipelines and platforms that turn raw events into trustworthy datasets the entire company makes decisions on. Data engineers here own the complete data lifecycle: ingestion, transformation, quality validation, and serving, with the reliability engineering that keeps pipelines running around the clock. You will work with analysts, data scientists, and product teams who depend on your data being correct, fresh, and well-documented. When the numbers are wrong, the business makes wrong decisions, so your standards for correctness effectively become the company's standards for truth.
What you will achieve
- Build reliable data pipelines with defined SLAs for freshness and completeness, monitored continuously, alerted on failure, and documented for stakeholders
- Deliver measurable data quality improvements through validation frameworks, anomaly detection, and contract testing on pipeline inputs and outputs
- Reduce pipeline operating costs through optimization: smarter partitioning, incremental processing patterns, and right-sized compute backed by usage analysis
- Enable self-service analytics with well-modeled, thoroughly documented datasets that analysts can trust without needing to ask you questions
- Cut incident resolution time with genuine pipeline observability: lineage tracking, data quality dashboards, and clear, tested runbooks
What you will bring
Must-haves
- 2 to 5 years building data pipelines in production with real SLAs and stakeholders depending on the output daily
- Strong SQL skills: complex analytical queries, window functions, query optimization, and data modeling for analytical workloads
- Experience with orchestration tools such as Airflow, Dagster, or Prefect: DAG design, retry policies, alerting, and historical backfills
- Familiarity with data warehousing concepts: dimensional modeling, slowly changing dimensions, and effective partition strategies
- Proficiency in Python for data processing with pandas, PySpark, or similar, written with attention to correctness and performance
- Understanding of data quality practices: validation checks, schema enforcement, and anomaly detection built into pipelines
- BS in Computer Science or equivalent experience
Nice-to-haves
- Experience with stream processing: Kafka, Flink, or Spark Streaming powering real-time data products
- Familiarity with dbt for transformation modeling, testing, and documentation as code
- Knowledge of data governance: cataloging, access controls, and privacy compliance with GDPR or CCPA
- Experience with lakehouse architectures: Delta Lake, Iceberg, or Hudi in production
Google
Meta
Apple
Microsoft
Amazon
Oracle
Netflix
NVIDIA