COE 592Lecture 01Reference
Reference sheet
Why efficient deep learning compressed onto one page: the definitions, formulas and numbers to have in your head before a quiz or exam.
The course question and the three pillars
Keep accuracy as high as possible under the budget the target device imposes. Progress comes from algorithm, hardware and data improving together; efficiency work reshapes the first two. Part 01: Accuracy at a cost
| Pillar | AlexNet 2012 | Efficiency lever |
|---|---|---|
| Algorithm | Deep CNN, ReLU, dropout, SGD | Pruning, quantization, distillation, NAS |
| Hardware | Two GTX 580 GPUs, 3 GB each, about six days | Accelerators: cloud GPU, phone NPU, Jetson, microcontroller |
| Data | ImageNet, about 1.2 million images, 1000 classes | Data that stays on the device and drives local learning |
- Privacy: raw data stays where it was collected.
- Cost: fewer GPU hours, smaller cloud bills, less carbon (OFA: one specialized net per target emits the CO2 of five cars).
- Connectivity: the model works where the network does not.
- Energy: milliwatt devices need models sized for milliwatts.
ImageNet metrics and the winners
| Property | Top-5 error | Top-1 accuracy |
|---|---|---|
| Counts as correct when | Any of five returned labels matches | The single top label matches |
| Strictness | Lenient, always the better score | Strict, always the worse score |
| Used on slide | 4 | 5, 6 |
| Typical pairing | about 5% error | 75 to 80% accuracy |
| Year | Winner | Top-5 error (%) |
|---|---|---|
| 2010 | NEC | 28.2 |
| 2011 | XRCE | 25.8 |
| 2012 | SuperVision (AlexNet), first deep CNN winner | 16.4 |
| 2013 | Clarifai | 11.7 |
| 2014 | GoogLeNet, 22 layers | 6.7 |
| 2015 | ResNet, 152 layers, first below human | 3.57 |
| 2016 | Trimps-Soushen | 2.99 |
| 2017 | SENet | 2.25 |
| Human | One trained annotator | 5.1 |
MACs, parameters and diminishing returns
A MAC is one multiply plus one accumulate, counted per input, independent of hardware. Parameters count storage. One MAC is two FLOPs when a source counts them separately; the ResNet paper labels multiply-adds as FLOPs, so its 3.8 x 10^9 equals Sze's 3.9 GMACs. Part 01: Accuracy at a cost
| Model | MACs (B) | Top-1 (%) | Parameters |
|---|---|---|---|
| MobileNetV1 | 0.57 | 70.6 | 4.2 M |
| ResNet-50 | 3.9 | 76.0 | 25.5 M |
| InceptionV3 | 5.7 | 78.8 | about 24 M |
| ResNet-101 | 7.6 | 76.4 | 44.5 M |
| Xception | 8.4 | 79.0 | about 23 M |
| Once-for-All (searched) | 0.595 | 80.0 | about 7 M |
Points of top-1 per extra billion MACs
- MobileNetV1 to ResNet-50
- (76.0 - 70.6) / (3.9 - 0.57) = 1.6 points per billion
- ResNet-50 to Xception
- (79.0 - 76.0) / (8.4 - 3.9) = 0.67 points per billion
- ResNet-50 to ResNet-101
- (76.4 - 76.0) / (7.6 - 3.9) = 0.11 points per billion
Efficient models and the Pareto frontier
Reduction ratio is reference MACs divided by new MACs in one unit. Once-for-All: 80.0% at 595 M MACs, inside the 600 M mobile setting. Part 01: Accuracy at a cost
| Reference | MACs | Top-1 (%) | Reduction |
|---|---|---|---|
| Xception | 8400 M | 79.0 | 8400 / 595 = 14.1x |
| ResNet-50 | 3900 M | 76.0 | 3900 / 595 = 6.6x |
| MobileNetV1 | 569 M | 70.6 | 0.96x, 9.4 points less accurate |
TinyML and on-device learning
MCUNet runs mask and person detection on an OpenMV Cam (Cortex-M7). Reference MCU STM32F746: 320 kB SRAM, 1 MB flash, no OS, no DRAM. TinyNAS searches under the memory limits; TinyEngine runs without interpreter overhead. Part 01: Accuracy at a cost
| Tier | Memory | Storage | Gap from the tier above |
|---|---|---|---|
| Cloud AI (V100) | 16 GB | TB to PB | reference |
| Mobile AI (iPhone 11) | 4 GB | more than 64 GB | 4x memory, 1000x storage |
| Tiny AI (STM32F746) | 320 kB | 1 MB | 3100x memory, 64000x storage |
Why nothing off the shelf fits
- ResNet-50 FP32 weights
- 25.5 M x 4 B = 102 MB, 102x over 1 MB flash
- ResNet-50 INT8 weights
- 25.5 MB, still 25x over
- MobileNetV2 peak activation
- 6.8 MB FP32 (22x), 1.7 MB INT8 (5.3x) over 320 kB SRAM
- MCUNet result
- 3.5x less SRAM, 5.7x less flash; 70.7% top-1 on STM32H743
| Resource | Inference | Training |
|---|---|---|
| Compute per example | One forward pass | forward + backward, about 3x |
| Activations | Freed layer by layer | All kept until the backward pass returns |
| Extra state | Weights only | Weights, gradients, optimizer moments |
| Memory scaling | Largest single layer | Grows with depth and batch size |
Naive training state on a microcontroller
- MobileNetV1 weights, FP32
- 4.2 M x 4 B = 16.8 MB
- Gradients and momentum
- 2 more copies = 33.6 MB
- Total before activations
- 50.4 MB, about 157x over 320 kB SRAM
- Lin et al. 2022
- training under 256 kB
SAM and EfficientViT-SAM
SAM is an image encoder (ViT-H, runs once per image, dominates cost), a prompt encoder and a lightweight mask decoder (together about 50 ms on a browser CPU). EfficientViT-SAM replaces only the encoder. Zero-shot means COCO was never in training; prompts are ViTDet boxes. Part 02: Generative and large models
| Model | Encoder params | Encoder MACs | img/s | mAP | Speedup |
|---|---|---|---|---|---|
| SAM-ViT-H | 641.1 M | 2973 G | 11 | 46.5 | 1x |
| EfficientViT-SAM-XL1 | 203.3 M | 322 G | 182 | 47.8 | 16.5x |
| EfficientViT-SAM-L2 | 61.3 M | 69 G | 538 | 46.6 | 48.9x |
| EfficientSAM | 25.3 M | 247 G | 183 | 44.4 | 16.6x |
| MobileSAM | 9.8 M | 39 G | 278 | 38.7 | 25.3x |
Generative compute: tokens times steps
| Model | Parameters | Workload | MACs |
|---|---|---|---|
| ViT-H | 0.6 B | 518 x 518 image, about 1,370 tokens | 1.02 T |
| Llama 3 8B | 8 B | 1024 context + 512 output, about 1,500 tokens | 11.2 T |
| FLUX.1 (one step) | 12 B | 1K x 1K latent, about 4,000 tokens | 37.2 T |
| CogVideoX (one step) | 5.6 B | 49 frames, 13 latent frames, about 17,500 tokens | 332 T |
Stable Diffusion training run, 2022
- GPUs
- 256 x A100
- GPU-hours
- about 150,000
- Wall clock
- 150,000 / 256 = 586 h, about 24 days
- Cost
- about $600,000, so $4 per GPU-hour
- MosaicML 2023 rerun
- 21,000 A100-hours, under $50,000
| Side | Pixels | MACs per step | MACs ratio |
|---|---|---|---|
| 1024 | 1x | about 30 T | 1x |
| 2048 | 4x | about 200 T | 6.7x |
| 3072 | 9x | about 720 T | 24x |
| 4096 | 16x | about 1,950 T | 65x |
DistriFusion
Each GPU denoises one patch but attends to its neighbours' activations from the previous step (displaced patch parallelism); the first step is synchronous, later communication is asynchronous and hidden. Part 02: Generative and large models
| Setup | MACs per device | Work per device | Latency | Speedup | Quality |
|---|---|---|---|---|---|
| Original, 1 GPU | 907 T | 1x | 12.3 s | 1x | reference |
| Naive patches, 4 GPUs | 190 T | 4.8x less | 3.14 s | 3.9x | duplicated subjects |
| DistriFusion, 4 GPUs | 227 T | 4.0x less | 4.16 s | 3.0x | no artifacts |
| DistriFusion, 8 GPUs | 113 T | 8.0x less | 2.74 s | 4.5x | no artifacts |
Fast-LiDARNet and the frame budget
| Configuration | Rate | Frame time | Lever | Against 33.3 ms |
|---|---|---|---|---|
| MinkowskiNet | 5 fps | 200 ms | none | 167 ms over |
| Optimised sparse kernels | 18 fps | 55.6 ms | system, 3.6x | 22 ms over |
| Fast-LiDARNet | 47 fps | 21.3 ms | algorithm, 2.6x more (9.4x total) | 12 ms to spare |
LLMs outgrow GPU memory
| Precision | Bytes per param | Weights | GPUs |
|---|---|---|---|
| FP32 | 4 | 2,120 GB | 27 |
| FP16 or BF16 | 2 | 1,060 GB | 14 |
| INT8 | 1 | 530 GB | 7 |
| INT4 | 0.5 | 265 GB | 4 |
| Model | Year | Parameters | FP16 weights |
|---|---|---|---|
| Transformer | 2017 | 0.05 B | 0.1 GB |
| GPT-2 | 2019 | 1.5 B | 3 GB |
| Megatron-LM | 2019 | 8.3 B | 16.6 GB |
| GPT-3 | 2020 | 175 B | 350 GB |
| MT-NLG | 2022 | 530 B | 1,060 GB |
Cloud GPUs and the memory wall
| Part | Dense FP16 TOPS | Bandwidth GB/s | Power W | Memory GB | TOPS/W |
|---|---|---|---|---|---|
| P100 (2016), PCIe | 18.7 | 732 | 250 | 16 | 0.075 |
| V100 SXM (2017) | 125 | 900 | 300 | 32 | 0.42 |
| A100 SXM (2020) | 312 | 2,039 | 400 | 80 | 0.78 |
| H100 SXM (2022) | 990 | 3,430 | 700 | 96 | 1.41 |
| B100 SXM (2024) | 1,750 | 8,192 | 700 | 192 | 2.5 |
Growth from P100 to B100
- Dense FP16 TOPS
- 18.7 to 1,750, about 94x
- Board power
- 250 to 700 W, 2.8x
- Memory capacity
- 16 to 192 GB, 12x
- Memory bandwidth
- 732 to 8,192 GB/s, about 11x
- TOPS per watt
- 0.075 to 2.5, about 33x
Memory wall on an H100
- H100 compute
- 990 x 10^12 dense FP16 ops/s
- H100 bandwidth
- 3.35 x 10^12 bytes/s
- Break-even intensity
- 990 / 3.35 = 295 ops per byte
- LLM decoding, batch 1
- 2 ops per 2-byte weight = 1 op per byte
- Trend (Gholami et al.)
- FLOPS 3.0x, DRAM bandwidth 1.6x, interconnect 1.4x per 2 years
Edge accelerator families
| Family | Performance | Power | TOPS/W | Shape |
|---|---|---|---|---|
| Snapdragon 845 to 8 Gen 1 | 3 to 52 TOPS, 17x | 9 to 10 W, 1.1x | 0.33 to 5.2 | fixed thermal lid |
| Apple A11 to A17 Pro | 0.6 to 35 TOPS, 58x | 6 to 8 W, about 1x | 0.075 to 2.63 (A15) | fixed thermal lid |
| Jetson Nano to AGX Orin 64GB | 0.5 to 275 TOPS, 550x | 10 to 60 W, 6x | 0.05 to 4.6 | ladder of power tiers |
Which tier for which power budget
- Microcontroller
- milliwatts (STM32F746 about 360 mW)
- Phone SoC
- about 10 W, fixed by heat
- Robot module (Jetson)
- 10 to 60 W, configurable
- Cloud GPU
- 250 to 700 W
The microcontroller tier
A microcontroller is a compact integrated circuit for embedded systems that puts a processor, memory and I/O peripherals on a single chip: no OS, no DRAM, milliwatts and kilobytes. Part 03: Hardware gap
| Board | Core | MHz | DMIPS | mW | SRAM | Flash |
|---|---|---|---|---|---|---|
| Arduino Zero | Cortex-M0+ | 48 | 45 | 6 | 32 kB | 256 kB |
| Arduino Due | Cortex-M3 | 84 | 158 | 26 | 96 kB | 512 kB |
| STM32F407VG | Cortex-M4 | 168 | 210 | 132 | 192 kB | 1 MB |
| STM32F746NG | Cortex-M7 | 216 | 462 | 360 | 320 kB | 1 MB |
DMIPS arithmetic
- Arm DMIPS/MHz
- M0+ 0.95, M3 and M4 1.25, M7 2.14
- STM32F746
- 2.14 x 216 = 462
- STM32F407
- 1.25 x 168 = 210
- MCU TOPS upper bound
- 216 MHz x 2 MACs x 2 ops = 0.0009 TOPS
Cloud, mobile and tiny AI
| Tier | Device | Activation memory | Weight storage | Memory vs tiny | Storage vs tiny |
|---|---|---|---|---|---|
| Cloud AI | A100-class GPU | 80 GB HBM | about 1 TB or more | 250,000x | about 1,000,000x |
| Mobile AI | Flagship phone | 4 GB DRAM | 256 GB flash | 12,500x | 256,000x |
| Tiny AI | STM32F746 | 320 kB SRAM | 1 MB flash | 1x | 1x |
Ratios to have ready
- Cloud to mobile, activation
- 80 GB / 4 GB = 20x
- Mobile to tiny, activation
- 4 GB / 320 kB = 12,500x
- Cloud to tiny, activation
- 80 x 10^6 kB / 320 kB = 250,000x
- Mobile to tiny, storage
- 256 GB / 1 MB = 256,000x
- 530 B FP16 against tiny
- 1.06 TB / 1 MB = 1,060,000x
| Quantity | Size | Cloud, 80 GB | Mobile, 256 GB and 4 GB | Tiny, 1 MB and 320 kB |
|---|---|---|---|---|
| FP32 weights | 102.2 MB | yes, 783x headroom | yes | no, 102x over flash |
| INT8 weights | 25.6 MB | yes | yes | no, 25.6x over flash |
| Peak activation | 7.2 MB | yes | yes | no, 22.5x over SRAM |
Slide errata
Answer with the corrected fact, and name the slide version if a question depends on it. Part 01, Part 02, Part 03
What the slides get wrong
- Slide 2
- "without loosing accuracy" should read "without losing accuracy".
- Slide 5
- Legend bubbles 2M to 64M are parameter counts, not MACs. "ResNetXt" means ResNeXt. The chart is Figure 2 of the Once-for-All paper, credited to Deng et al.
- Slide 6
- "ResNetXt-50" and "ResNetXt-101" are again ResNeXt.
- Slide 10
- The citation is the EfficientViT backbone paper (Cai et al., ICCV 2023). The 48.9x result is Zhang, Cai and Han, arXiv 2402.05008. "EfficientVIT" should read EfficientViT.
- Slide 15
- The x axis skips 2019 with uneven spacing, so slopes are distorted. "MegatronLM" is Megatron-LM (8.3 B). The 2017 Transformer has 65 M (base) and 213 M (big), rounded to 0.05 B.
- Slide 18
- "Memory Bendwidth" means Bandwidth. H100 96 GB is an OEM SXM5 SKU, and 3,430 GB/s matches neither the 80 GB (3.35 TB/s) nor the 96 GB part. P100 values are the PCIe card (SXM2 was 300 W). B100 1,750 is the dense half of 3.5 PFLOPS sparse.
- Slide 20
- No A17 power is plotted, so its TOPS/W cannot be computed. The A13 6 TOPS is NanoReview's figure; the A17 Pro 35 TOPS comes from Apple's 2023 event.
- Slide 21
- Standard TX2 has 8 GB (4 GB is the TX2 4GB variant). The series mixes precisions: Nano and TX2 in FP16, Xavier INT8, Orin sparse INT8. The caption credits NanoReview but the URL is Connect Tech.
- Slide 22
- "on a single" is missing "chip". "KB" should be kB. STM32F407VG has 192 kB SRAM plus 4 kB backup, not 196 kB. Due's 158 DMIPS implies 1.88 DMIPS/MHz; Arm rates the M3 at 1.25 (105).
- Slide 23
- Cloud and mobile columns were updated from MCUNet's 2020 table (V100 16 GB, iPhone 11 more than 64 GB) to 80 GB and 256 GB; the tiny column is unchanged, so the columns are not one dated snapshot.
- Slide 24
- The MIT course is 6.5940, "TinyML and Efficient Deep Learning Computing" (Fall 2024 page), now "TinyML and Efficient AI Computing". The slide omits year and URL.