Machine Learning Software Engineer
The job description
Tech stack. Python, PyTorch or TensorFlow, scikit-learn, MLflow, model serving (TorchServe, Triton), data pipelines, experiment tracking, GPU training, evaluation methodology, feature stores
About the role
You will build machine learning systems that work reliably in production, not just in notebooks, bridging the gap between model development and solid software engineering. ML software engineers here own the full lifecycle: data pipelines, training infrastructure, evaluation frameworks, model serving, and monitoring for drift and degradation. You will collaborate with data scientists on modeling choices while bringing software engineering rigor to everything around the model: versioning, testing, reproducibility, and safe rollback. The models are only as valuable as the systems that serve them, and you own those systems.
What you will achieve
- Ship ML features to production with complete pipelines: data ingestion, training, evaluation, deployment, and monitoring for performance drift over time
- Deliver measurable model quality improvements through better training data, thoughtful feature engineering, and rigorous offline evaluation before deployment
- Build training infrastructure that scales: distributed training jobs, experiment tracking, and reproducible runs any team member can audit and reproduce
- Reduce time from experiment to production through standardized serving patterns, model registries, and automated deployment pipelines with validation gates
- Establish evaluation discipline: held-out test sets, online business metrics, and shadow deployments that catch regressions before users ever see them
What you will bring
Must-haves
- 2 to 5 years in machine learning engineering or data science with production model deployments, not just research prototypes or Kaggle competitions
- Strong Python skills and deep familiarity with PyTorch or TensorFlow for model development, debugging, and performance tuning
- Understanding of ML fundamentals: train/validation/test splits, overfitting diagnosis, regularization, and choosing appropriate metrics for the problem
- Experience with data pipelines: cleaning messy data, validation checks, feature engineering, and handling the reality of production data drift
- Familiarity with model serving trade-offs: batch versus real-time inference, latency requirements, and scaling inference workloads cost-effectively
- Knowledge of experiment tracking and reproducibility: versioning data, code, hyperparameters, and results together as a unit
- BS in Computer Science, Statistics, or related field; ML coursework or equivalent applied experience
Nice-to-haves
- Experience with LLMs: fine-tuning, prompt engineering, RAG systems, or evaluation methodologies for generative models
- Familiarity with MLOps platforms such as Kubeflow, SageMaker, or Vertex AI for managed training and serving
- Knowledge of model optimization: quantization, pruning, distillation, or ONNX conversion for efficient inference
- Experience with feature stores and real-time feature computation architectures
Google
Meta
Apple
Microsoft
Amazon
Oracle
Netflix
NVIDIA