AI/ML CPU Architect
The job description
Tech stack. ML workload analysis, vector and matrix extensions, gem5, PPA analysis, Python, PyTorch and ONNX profiling, Arm SVE and RISC-V V, quantization-aware design
About the role
You will architect CPUs optimized for AI and machine learning workloads at a semiconductor company, defining how general-purpose processors handle the inference and training tasks that increasingly dominate compute demand. AI/ML CPU architects bridge the architecture and ML communities: they understand model architectures, quantization, and serving constraints deeply enough to shape hardware around them, while keeping the CPU general enough for everything else it must run. You will define math acceleration, memory system behavior for ML access patterns, and the software interfaces that make it all usable.
What you will achieve
- Define CPU architecture features for ML acceleration including vector and matrix datapaths, numeric format support, and memory hierarchy tuning for model serving access patterns.
- Profile production ML models across vision, language, recommendation, and emerging workloads, translating their computational signatures into concrete architectural requirements.
- Model inference latency and throughput on candidate architectures, optimizing for the serving constraints that matter: tail latency, queries per watt, and total cost per inference.
- Partner with ML framework and library teams to ensure architectural features are exploited by PyTorch, ONNX Runtime, and vendor kernels, closing the gap between hardware capability and realized performance.
- Specify the programming model and software interfaces for ML acceleration features, balancing ease of use for framework developers against hardware implementation cost.
- Validate shipped ML performance against targets using industry benchmarks and customer models, demonstrating competitive positioning with honest, reproducible measurements.
What you will bring
Must-haves
- 5 to 8 years of experience in CPU or accelerator architecture with significant ML workload focus and shipped products.
- Deep understanding of ML model architectures, quantization techniques, and inference serving characteristics.
- Knowledge of dense math datapath design including precisions, sparsity handling, and dataflow implications.
- Familiarity with ML frameworks and the path from model to optimized kernels to hardware instructions.
- Performance modeling skills with the discipline to measure realized performance rather than peak theoretical throughput.
- Ability to communicate across the hardware-software boundary with both architecture and ML engineering teams.
- Strong quantitative judgment about which ML features earn their silicon cost across the product's lifetime.
Nice-to-haves
- Experience with MLPerf submissions or competitive ML benchmarking.
- Background in NPU or GPU ML architecture for cross-domain insight.
- Contributions to open-source ML kernels, compilers, or runtimes.
Apple
NVIDIA
Qualcomm
Intel
AMD
Arm