AI/ML ASIC Verification Engineer
The job description
Tech stack. SystemVerilog, UVM, tensor datapath verification, MAC/systolic arrays, numerical accuracy checking (FP8, BF16, INT8), Python, NumPy, Synopsys VCS
About the role
You will verify AI accelerator silicon at a fabless semiconductor company building chips for training and inference workloads. AI/ML verification centers on massive parallel datapaths: systolic MAC arrays, tensor cores, on-chip SRAM hierarchies, and the numerical precision behavior of reduced-precision arithmetic. Your testbenches prove that matrix multiplications land bit-accurately, that data movement keeps the compute units fed, and that the accelerator delivers its promised TOPS on real model shapes. You will build the bit-accurate reference models that define correctness, run representative neural-network workloads through the design, and hunt the numerical bugs that only appear at the intersection of arithmetic and architecture. The industry's AI ambitions run on silicon like this, and your verification is what makes it trustworthy.
What you will achieve
- Verify tensor datapaths and MAC arrays with bit-accurate reference models, closing numerical correctness on FP8, BF16, and INT8 precisions.
- Build workload-driven testbenches that run representative neural-network layer shapes (GEMM, convolution, attention) through the accelerator.
- Prove on-chip memory hierarchy behavior: SRAM banking, double-buffering, and data reuse patterns that sustain compute utilization above 80 percent.
- Catch numerical accuracy bugs (rounding, saturation, overflow handling) that functional testing of control logic would never expose.
- Drive performance verification confirming the chip hits its TOPS and TOPS-per-watt targets on benchmark model configurations, and track utilization metrics that guide architecture tuning on future revisions.
What you will bring
Must-haves
- 2 to 5 years of verification experience on datapath-heavy designs such as DSPs, GPUs, NPUs, or AI accelerators.
- Strong SystemVerilog and UVM skills with experience verifying wide parallel datapaths.
- Understanding of reduced-precision arithmetic (FP8, BF16, INT8/INT4) including rounding modes, saturation, and overflow behavior.
- Ability to build bit-accurate reference models, in SystemVerilog or via DPI-C with Python/NumPy golden models.
- Knowledge of neural-network compute patterns: matrix multiplication, convolution, and attention dataflows.
- Familiarity with on-chip memory architectures: SRAM banks, scratchpads, and DMA-driven data movement.
- Scripting skills for generating workload stimulus and analyzing numerical results.
Nice-to-haves
- Experience with ML frameworks (PyTorch, TensorFlow) for generating reference workloads.
- Knowledge of sparsity, quantization, or other inference optimization techniques.
- Exposure to interconnect verification for scale-out AI systems (NVLink, Ethernet fabrics).
- Understanding of distributed training communication patterns such as all-reduce.
Apple
NVIDIA
Qualcomm
AMD
Broadcom
Marvell