Class hours: MF 8-8:50 AM Tu 9-9:50 AM in KD 101
| Name | |
| Swarnendu Biswas | swarnendu@cse.iitk.ac.in |
| Name | |
| Srinjoy Sarkar | srinjoys23@cse.iitk.ac.in |
| Soham Rajesh Kelaskar | sohamrk25@cse.iitk.ac.in |
| Varad Prabhakar Shinde | vpshinde25@cse.iitk.ac.in |
| Vishnu Vardhan Reddy Tavanam | tvreddy25@cse.iitk.ac.in |
To achieve good performance, one needs to write correct yet scalable parallel programs using programming language abstractions such as threads. In addition, the developer needs to be aware of and utilize many architecture-specific features, such as vectorization, to realize the full performance potential. This course will discuss programming language abstractions with architecture-aware development to learn to write scalable parallel programs. This is not a "programming tips and tricks" course.
We will have 3-5 assignments to use the concepts learned in class and appreciate the challenges in extracting performance.
| Prerequisites |
|
The course will focus on a subset of the following topics.
The following is a tentative allocation and might change slightly depending on the strength of the class. Grading is relative.
| Assignments | 35% |
| Midsem | 30% |
| Endsem | 35% |
I am open to feedback about the course content and presentation. Feel free to provide suggestions for improvements.
| Date | Topic | Resources | Recommended Reading |
|---|---|---|---|
| First course handout | FCH | ||
| 03/08, 04/08 | Compiler Challenges for Parallel Architectures | Slides | AK 1.1-1.6 | 04/08, 07/08 | Cache Memory | Slides |
HP APP B.1-B.4, 2.1--2.3 CSAPP 6.2-6.4 |
| 10/08, 11/08, 11/08 | Write Cache-Friendly Code | Slides |
CSAPP 6.5-6.6 DRAG 11.1-11.2 |
| 14/08, 17/08, 18/08 | Cache Coherence | Slides | MCM Chapters 2, 6, 8 (IITK has subscribed to the ebook) |
| 21/08, 24/08 | Dependence Testing | Slides |
AK Chap 2, 3
DRAG 11.6 |
| Parallelizing Loops | Slides |
AK 5.2-5.4, 5.7.2, 5.9, 6.2.1, 6.2.2, 6.2.5,
6.3.1-6.3.4 AP 4.1, 4.2, 4.5, 5.1-5.6 HP 4.1, 4.2, 4.5 Compiler Transformations for High-Performance Computing Program Optimization Through Loop Vectorization Topics in Loop Vectorization |
|
| Shared-Memory Synchronization | Slides |
MP 2.3, 2.4, 2.6, 7.1-7.5, 8.3
SMS 4.1, 4.2, 4.3.1, 6.1 |
|
| OpenMP | Slides |
PP Chapter 5 (IITK has subscribed to the ebook) LLNL OpenMP Tutorial Tim Mattson: Introduction to OpenMP Nitya Hariharan: Core, Advanced OpenMP Application Programming Interface v6.0 OpenMP Application Programming Interface Examples v6.0 |
| [CSAPP] | Computer Systems: A Programmer's Perspective, 3rd edition - R. Bryant and D. O'Hallaron |
| [DRAG] | Compilers: Principles, Techniques, and Tools - A. Aho, M. Lam, R. Sethi, and J. Ullman |
| [HP] | Computer Architecture: A Quantitative Approach, 6th edition - J. Hennessy and D. Patterson |
| [AK] | Optimizing Compilers for Modern Architectures - R. Allen and K. Kennedy |
| [SMS] | Shared-Memory Synchronization, 2nd edition - M. Scott and T. Brown. |
| [AP] | Automatic Parallelization: An Overview of Fundamental Compiler Techniques - Samuel P. Midkiff |
| [PP] | An Introduction to Parallel Programming - Peter S. Pacheco |
| [KH] | Programming Massively Parallel Processors: A Hands-on Approach, 3rd edition - David B. Kirk and Wen-mei W. Hwu |
| [MCM] | A Primer on Memory Consistency and Cache Coherence, 2nd edition - Vijay Nagarajan, Daniel J. Sorin, Mark D. Hill and David A. Wood |
| [MP] | The Art of Multiprocessor Programming, 1st edition - Maurice Herlihy and Nir Shavit |