CENG 5410 Advanced Computer Architecture

Version:     2026 Fall
Lecture:     M15:30-18:15      Venue: ELB 401

Course Instructor:      Prof. Zhengrong Wang      zhengrongwang@cuhk.edu.hk

Course Tutors:
    Xingyu Jiang      xyjiang26@cse.cuhk.edu.hk
    Yifan Guo      yfguo26@cse.cuhk.edu.hk

Description

Computer architecture is the glue that binds the worlds of hardware and software. Developing new architectures is the art and science of understanding the intrinsic properties of applications, developing and interconnecting hardware components to meet cost and performance goals, and creating abstractions that enable efficient use of hardware without sacrificing generality.

Several important specific goals are:

We use piazza for questions.

Course Requirements

References

There is no required textbook. You could refer to Synthesis Lectures and a lot of papers we are going to review in this course. If you do want a textbook, Computer Architecture: A Quantitative Approach is an excellent choice.

Lectures

Subjected to changes.

Week Date Topic Reading Review Reference Material
1 Sep. 07 L01 Intro Cramming More Components Onto Integrated Circuits – Moore 65
Power: A First-Class Architectural Design Constraint – Mudge 01
2 Sep. 14 L02 Tech Compiler-friendly ISAs – Wulf 81
3 Sep. 21 L03 ISA Synth Lec: Processor Microarchitecture (chap 1-3) Power Struggles – Blem 13 Branch Prediction Taxonomy – Yeh 92
4 Sep. 28 L04 Pipeline Synth Lec: Processor Microarchitecture (chap 4,5) IA-64 – Huck 00 RUU – Sohi 90
5 Oct. 05 L05 Wide issue Synth Lec: Processor Microarchitecture (chap 6,7,8) Memory Dependence Prediction – Moshovos 97 Implementing Precise Interrupts – Smith 90
6 Oct. 12 Understanding Scheduling Replay – Kim 04 Continual Flow Pipelines – Srinivasan 04
7 Oct. 19 Haswell IEEE Micro (2014) Alternate Path Fetch
8 Oct. 26 Synth Lec: Multi-Core Cache Hierarchies (Chap 1) Victim Cache & Stream Buffers – Jouppi 90
9 Nov. 02 Synth Lec: Primer on Prefetching (Chap 1, 2.1-2.2, 3.1-3.2) ZCache – Sanchez 10 Criticality-based Prefetching – Nori 18
10 Nov. 09 Virtual memory in contemporary microprocessors Hawkeye – Jain 16 Virtual-real Cache – Wang 89
11 Nov. 16 Synth Lec: Multithreading Architecture (read ch1-5, skim 6-7)
12 Nov. 23 Synth Lec: Primer on Consistency and Cache Coherance (read ch1-3, skim 4-8) Niagara – Kongetira 05 Victim replication – Zhang 05
13 Nov. 30 Thread-level Speculation (TLS) – Mowry 98
14 Dec. 07 NVIDIA Tesla – Lindholm 08 Hopper GPU
15 Dec. 14 MIT RAW – Waingold 97 Wavescalar – Swanson 03 Inefficiency in general-purpose chips – Hameed 10

Review / Project Report

Please submit your final project through blackboard. If the submission is late but within 5 calendar days after the deadline, there will be a deduction of 10% per day from the marks awarded for the submitted piece of work.

Important Dates

Course Projects

The final project is an opportunity to gain hands-on experience with architecture simulation and evaluation. You may work on one of the topics below, propose a variation, or develop an open-ended idea of your own. A good project should make one focused change or investigation, evaluate it on meaningful workloads, and explain the resulting performance tradeoffs.

Simulation, Instrumentation, and Visualization

  1. Visualizing the FlashGPU-Sim instruction pipeline: Build a timeline or interactive visualization for instruction issue, stalls, dependencies, and memory events. A useful starting point is Konata for gem5; Google pprof provides an example of a mature profile visualization workflow.
  2. FlashGPU-Sim and gem5 co-simulation: Connect GPU and CPU simulation and study synchronization, data movement, and end-to-end performance. Start with FlashGPU-Sim and gem5.
  3. Integrated CPU-GPU simulation: Extend the previous idea to model a heterogeneous system such as a DGX Spark. Compare application behavior under different CPU/GPU balance, interconnect, and memory assumptions. DGX Spark technical information and NVIDIA’s Grace Blackwell platform overview provide system context.

AI Workloads and Software Stacks

  1. Characterizing emerging AI workloads: Characterize the execution and memory behavior of a recent mixture-of-experts, long-context, or reasoning workload. Possible starting points include DeepSeek-V4, Kimi K3, and Loop Transformer. Focus on a question that can be answered with available traces or a reduced representative workload rather than attempting to reproduce an entire production model.
  2. FlashGPU-Sim integration with a high-level framework: Integrate the simulator with vLLM or SGLang and quantify the effects of batching, scheduling, KV-cache management, or speculative decoding on a simulated accelerator.
  3. Architecture-aware LLM serving: Use a simulator or analytical model to study one serving policy, such as continuous batching, paged attention, or disaggregated prefill and decode. PagedAttention and Sarathi-Serve are useful references.

GPU Memory and Multi-Die Architecture

  1. Coherence between private L2 caches on two GPU dies: Design and evaluate a coherence or communication protocol for a multi-die GPU. Compare directory, broadcast, and software-assisted alternatives using locality, traffic, and synchronization metrics. You can target AMD’s chiplet and NVIDIA’s Blackwell architecture for contemporary multi-die designs.
  2. Spatial GPU placement and simulation: Model the impact of placing compute units, caches, memory controllers, or chiplets at different physical locations. Explore how placement changes latency, bandwidth, contention, or thermal/power proxies.
  3. CXL-tiered memory for LLMs: Model fast and slow memory tiers and implement a page or tensor migration policy. Evaluate capacity, latency, bandwidth, migration overhead, and tail latency.
  4. HBF-based memory-system simulation: Explore a high-bandwidth flash or hybrid-buffer memory design, including its controller, parallelism, scheduling, and failure or endurance tradeoffs. HBFSim can be a starting point.

Open-Ended Projects

  1. A project inspired by the CS251a project list: The CS251a course project page contains additional examples covering gem5, value prediction, prefetching, cache policies, CXL memory, processing-in-memory, and performance counters.
  2. Open topic: Propose a research question related to computer architecture, simulation, systems for machine learning, or hardware-software co-design. Projects connected to ongoing research are welcome, provided that the work and evaluation performed for this course are clearly scoped.

Project Proposal

Please submit a short proposal describing:

  1. Question: What problem are you studying, and why is it interesting?
  2. Method: What simulator, implementation, workloads, baselines, and metrics will you use?
  3. Plan: What are the implementation and evaluation milestones? What is the smallest complete version if the ambitious version cannot be finished?

The proposal should identify a tractable scope. Reproducing or extending an existing result is a valid project, as long as the report explains the original design and provides a careful evaluation.