GPU/CPU Integration Architect
The job description
Tech stack. Heterogeneous system architecture, cache coherency, shared virtual memory, Arm and RISC-V ISAs, SystemVerilog, performance modeling, interconnect protocols, unified memory
About the role
You will architect the integration between CPU and GPU at a semiconductor company building heterogeneous processors for AI, graphics, and high-performance computing. The boundary between CPU and GPU is where some of the hardest system problems live: memory coherence across different consistency models, unified virtual memory, fine-grained task dispatch, and quality of service between latency-sensitive CPU threads and throughput-hungry GPU kernels. You will define how these processors work as one system rather than two chips sharing a package.
What you will achieve
- Define the coherency architecture between CPU and GPU domains, selecting protocols and consistency guarantees that enable efficient sharing without imposing impossible verification burdens.
- Architect unified memory and shared virtual memory schemes including page migration policies, fault handling, and TLB coherence, validated through modeling and workload analysis.
- Design the dispatch and synchronization mechanisms for fine-grained CPU-GPU cooperation, minimizing launch latency and enabling programming models that treat the system as unified.
- Model heterogeneous workload performance across candidate integration options, quantifying the benefits of tighter coupling against its costs in complexity, power, and area.
- Specify the hardware-software interfaces for heterogeneous compute including driver-visible registers, interrupt schemes, power management coordination, and debug support.
- Partner with GPU architecture, CPU architecture, software, and system teams to resolve cross-domain tradeoffs with decisions grounded in application-level performance data.
What you will bring
Must-haves
- 5 to 8 years of experience in CPU or GPU architecture, SoC architecture, or heterogeneous system design with shipped products.
- Strong understanding of both CPU memory models and GPU execution and memory models, and the challenges of bridging them.
- Experience with cache coherency protocols and their extension across heterogeneous agents.
- Knowledge of GPU programming models such as CUDA, ROCm, or Vulkan compute and their hardware implications.
- Performance modeling skills applied to heterogeneous workloads with the discipline to validate against real applications.
- Ability to write specifications spanning multiple architecture teams with different cultures and assumptions.
- Systems-level thinking and the communication skills to align CPU, GPU, and software organizations behind shared decisions.
Nice-to-haves
- Experience with chiplet-based heterogeneous integration or advanced packaging.
- Familiarity with AI accelerator architectures and their memory system requirements.
- Background in interconnect or fabric architecture for heterogeneous SoCs.
Apple
NVIDIA
Qualcomm
Intel
AMD
Arm