Version: 2026 Fall
Lecture:
M15:30-18:15 Venue: ELB 401
Course Instructor: Prof. Zhengrong Wang zhengrongwang@cuhk.edu.hk
Course Tutors:
Xingyu Jiang xyjiang26@cse.cuhk.edu.hk
Yifan Guo yfguo26@cse.cuhk.edu.hk
Description
Computer architecture is the glue that binds the worlds of hardware and software. Developing new architectures is the art and science of understanding the intrinsic properties of applications, developing and interconnecting hardware components to meet cost and performance goals, and creating abstractions that enable efficient use of hardware without sacrificing generality.
Several important specific goals are:
- Be able to reason about the relationship between application properties and how they are exploited by architecture/microarchitecture mechanisms.
- Gain intuition and reasoning skills regarding fundamental architecture tradeoffs of hardware design choices (performance/area/power/complexity/generality).
- Understand microarchitecture techniques behind extracting instruction level parallelism and mechanisms for exploiting locality.
- Gain appreciation for state-of-the-art microprocessors.
- Learn about evaluation methods, including simulation, analytical modeling, and mechanistic models. Specific focus on gem5/flashgpu-sim.
We use piazza for questions.
Course Requirements
- Reading Reviews (30%), Midterm (30%), Final Project (40%).
- A student must gain at least 50% of the full marks in order to pass the course.
References
There is no required textbook. You could refer to Synthesis Lectures and a lot of papers we are going to review in this course. If you do want a textbook, Computer Architecture: A Quantitative Approach is an excellent choice.
Lectures
Subjected to changes.
Review / Project Report
Please submit your final project through blackboard. If the submission is late but within 5 calendar days after the deadline, there will be a deduction of 10% per day from the marks awarded for the submitted piece of work.
Important Dates
- Nov. 2: In-class midterm exam.
- Dec. 13: Final project due on 11:59 pm.
Course Projects
The final project is an opportunity to gain hands-on experience with architecture simulation and evaluation. You may work on one of the topics below, propose a variation, or develop an open-ended idea of your own. A good project should make one focused change or investigation, evaluate it on meaningful workloads, and explain the resulting performance tradeoffs.
Simulation, Instrumentation, and Visualization
- Visualizing the FlashGPU-Sim instruction pipeline: Build a timeline or interactive visualization for instruction issue, stalls, dependencies, and memory events. A useful starting point is Konata for gem5; Google pprof provides an example of a mature profile visualization workflow.
- FlashGPU-Sim and gem5 co-simulation: Connect GPU and CPU simulation and study synchronization, data movement, and end-to-end performance. Start with FlashGPU-Sim and gem5.
- Integrated CPU-GPU simulation: Extend the previous idea to model a heterogeneous system such as a DGX Spark. Compare application behavior under different CPU/GPU balance, interconnect, and memory assumptions. DGX Spark technical information and NVIDIA’s Grace Blackwell platform overview provide system context.
AI Workloads and Software Stacks
- Characterizing emerging AI workloads: Characterize the execution and memory behavior of a recent mixture-of-experts, long-context, or reasoning workload. Possible starting points include DeepSeek-V4, Kimi K3, and Loop Transformer. Focus on a question that can be answered with available traces or a reduced representative workload rather than attempting to reproduce an entire production model.
- FlashGPU-Sim integration with a high-level framework: Integrate the simulator with vLLM or SGLang and quantify the effects of batching, scheduling, KV-cache management, or speculative decoding on a simulated accelerator.
- Architecture-aware LLM serving: Use a simulator or analytical model to study one serving policy, such as continuous batching, paged attention, or disaggregated prefill and decode. PagedAttention and Sarathi-Serve are useful references.
GPU Memory and Multi-Die Architecture
- Coherence between private L2 caches on two GPU dies: Design and evaluate a coherence or communication protocol for a multi-die GPU. Compare directory, broadcast, and software-assisted alternatives using locality, traffic, and synchronization metrics. You can target AMD’s chiplet and NVIDIA’s Blackwell architecture for contemporary multi-die designs.
- Spatial GPU placement and simulation: Model the impact of placing compute units, caches, memory controllers, or chiplets at different physical locations. Explore how placement changes latency, bandwidth, contention, or thermal/power proxies.
- CXL-tiered memory for LLMs: Model fast and slow memory tiers and implement a page or tensor migration policy. Evaluate capacity, latency, bandwidth, migration overhead, and tail latency.
- HBF-based memory-system simulation: Explore a high-bandwidth flash or hybrid-buffer memory design, including its controller, parallelism, scheduling, and failure or endurance tradeoffs. HBFSim can be a starting point.
Open-Ended Projects
- A project inspired by the CS251a project list: The CS251a course project page contains additional examples covering gem5, value prediction, prefetching, cache policies, CXL memory, processing-in-memory, and performance counters.
- Open topic: Propose a research question related to computer architecture, simulation, systems for machine learning, or hardware-software co-design. Projects connected to ongoing research are welcome, provided that the work and evaluation performed for this course are clearly scoped.
Project Proposal
Please submit a short proposal describing:
- Question: What problem are you studying, and why is it interesting?
- Method: What simulator, implementation, workloads, baselines, and metrics will you use?
- Plan: What are the implementation and evaluation milestones? What is the smallest complete version if the ambitious version cannot be finished?
The proposal should identify a tractable scope. Reproducing or extending an existing result is a valid project, as long as the report explains the original design and provides a careful evaluation.