Vector/Tensor CPU Architect
1
Which company are you aiming at?
To generate precise rejection-prevention feedback and application strategy, tell us where you're targeting.
2
Send your application
You need an account so we know where to send the feedback.
The job description
Tech stack. Vector and matrix ISA extensions, SIMD microarchitecture, gem5, SystemVerilog, PPA analysis, Python, ML workload characterization, Arm SVE and RISC-V V extensions
About the role
You will architect vector and tensor processing capabilities for CPUs at a semiconductor company, designing the SIMD units, matrix engines, and ISA extensions that accelerate AI, HPC, and media workloads on general-purpose processors. As AI inference increasingly runs on CPUs in datacenters and at the edge, the architects who define how CPUs handle dense math determine real product differentiation. You will design the datapath, define the programming model with software teams, model performance on ML workloads, and specify the implementation.
What you will achieve
- Define the vector and matrix execution architecture for a CPU program including datapath width, register file organization, lane interconnect, and the supported precisions from FP64 down to INT4.
- Design or extend the ISA for dense math acceleration, balancing the performance opportunity against decoder complexity, verification cost, and software enablement effort, with each decision backed by workload data.
- Model ML inference and training kernels on candidate architectures, demonstrating measurable speedups on representative models while keeping area and power within program budgets.
- Partner with compiler, library, and framework teams to ensure the architecture is genuinely programmable, with intrinsics, auto-vectorization support, and kernel libraries ready when silicon arrives.
- Specify the microarchitecture of vector units including chaining, predication, gather-scatter, and exception handling, with the precision implementation teams require.
- Validate the shipped implementation against performance targets using MLPerf-style benchmarks and customer models, and document lessons for the next generation.
What you will bring
Must-haves
- 5 to 8 years of experience in vector, SIMD, GPU, or accelerator architecture with shipped products.
- Deep understanding of SIMD microarchitecture including lane organization, shuffles, predication, and memory access patterns.
- Knowledge of ML workload characteristics including GEMM shapes, quantization, sparsity, and the roofline implications for CPU design.
- Familiarity with vector ISA extensions such as Arm SVE, RISC-V V, or x86 AVX, including their software ecosystem implications.
- Performance modeling skills applied to dense math workloads with honest accounting of utilization versus peak claims.
- Ability to work effectively with software teams on programming models, intrinsics, and library enablement.
- Strong quantitative decision making that weighs performance gains against the full lifetime cost of ISA and hardware complexity.
Nice-to-haves
- Experience with sparse acceleration techniques or mixed-precision datapaths.
- Background in GPU tensor core or NPU architecture for cross-pollination of ideas.
- Contributions to ML compiler or kernel library projects.
Apple
NVIDIA
Qualcomm
Intel
AMD
Arm