Majid Al-RaimiFull guide

COE 592Lecture 00Full guide

Course overview

The whole lecture on one page, taught concept by concept. Work through the parts in order, mark each concept once you understand it, and open the slide chips when you want the original slides.

Parts
2
Concepts
8
Slides
12
Reading
48 min
Understood
0/8 concepts

Part 01: The instructor and edge AI in practice

Who teaches the course, and what his lab's projects show about running vision and security models on UAVs, farm robots, vehicles and IoT devices.

5 concepts, slides 1-8

Why this part matters

This part is the course's motivation written as a portfolio. Every image on slides 3 to 8 is the same problem you will face in the research project: a model sitting next to a sensor, a budget of weight, power, memory and latency, and a requirement to stay correct when the world is wet, noisy or hostile.

Read these slides as budget problems rather than as a list of the instructor's papers. That reading habit is what the rest of COE 592 builds on, because pruning, quantization, neural architecture search, distillation and TinyML are all answers to one question: how much model can this platform afford, and how do we make a good model fit inside that? The bio matters less for exams than the three words in the lab's name, which turn out to be the syllabus in miniature.

By the end you can

  1. Expand SERV and map each word to a later course theme.
  2. Name the sensor modalities on a UAV or vehicle inspection rig and explain why payload limits force small on-board models or data offloading.
  3. Read a detector's size class (n, s, m, l, x) from Ultralytics tables and compute the parameter, FLOP and mAP trade-off between two sizes.
  4. Read a detection figure critically: separate confidence from correctness, and a saliency map from a metric.
  5. State why an IoT intrusion detector must be lightweight and what generalizable means.
  6. Explain why roadside and rolling-stock inspection are real-time, safety-critical edge inference problems.

The course is COE 592, Machine Learning on Embedded Systems, offered in term T261 and taught by Dr. Abdul Jabbar Siddiqui of the Department of Computer Engineering. Of everything on the title and bio slides, one line carries the most information: he leads the SERV Lab, and SERV stands for Secure, Efficient, Robust Vision. Those three adjectives are the map of the course before the syllabus arrives in Part 02.

Instructor profile, as stated on slide 2

Name
Dr. Abdul Jabbar Siddiqui
Roles
Assistant Professor, Computer Engineering, KFUPM; Research Scholar, SDAIA-KFUPM JRC for AI; Lead of SERV Lab
Degrees
PhD, University of Ottawa (2021); MASc, University of Ottawa (2015); BSc, KFUPM (2012)
Previously
National Research Council of Canada; NSERC TRANSIT Network
Research interests
Edge AI, Visual Intelligence and Analytics, Adversarial machine learning, AI for Cybersecurity, Intelligent transportation systems, UAV-based Inspection and Monitoring, Internet of Things

Reading SERV as a syllabus

Look at the interests list again and notice that every item is an instance of one question: can a model run near the sensor, inside a budget of memory, compute, energy and latency, and still be trusted? Edge AI and IoT name the setting. UAV inspection and transportation systems name the platforms. Adversarial machine learning and AI for cybersecurity name the threats. MIT 6.5940, whose course description slide 11's takeaways echo almost word for word, opens with the same framing, that deep networks "demand extraordinary levels of computation, hindering its deployment on everyday devices" (Han, MIT 6.5940).

LetterWordCourse themeWhere you meet it in this part
SSecureAdversarial machine learning and intrusion detectionPNet-IDS on slide 6
EEfficientCompression, pruning, quantization, nano detectorsEcoWeedNet on slide 4, PNet-IDS on slide 6
RRobustBehaviour under rain, noise, blur and domain shiftYOLO-RAW on slide 5
VVisionCameras, thermal imagers and LiDAR as the inputEvery figure from slide 3 to slide 8
SERV mapped onto the course
A camera on three legs. Secure, Efficient and Robust each hold up one corner; remove any leg and the deployment tips over.

Efficient is the theme that gets the most lecture time, because it is where the arithmetic lives: how many parameters, how many operations, how many milliseconds. Robust is what happens to that efficient model when rain, noise or a new camera changes the input distribution. Secure is the adversarial case, where the change in the input is chosen on purpose: Goodfellow, Shlens and Szegedy defined adversarial examples as inputs formed by applying "small but intentionally worst-case perturbations", and traced the vulnerability to the models' linear behaviour (Goodfellow et al., 2015). A model that is only efficient can be fooled; a model that is only robust may not fit on the drone.

Recall

What do the letters of SERV stand for, and which course theme does each map to?

Secure (adversarial machine learning, intrusion detection), Efficient (compression, quantization, nano detectors), Robust (rain, noise, blur, domain shift), Vision (the input modality across all of them).

Quick check

What does the SERV in SERV Lab stand for?

Slide 3 is a wall of figures with no captions, so start by describing what is visibly there. A small unmanned helicopter sits on gravel with its camera payload circled (panel A). Beside it is a rugged field case holding a ground control station (panel B). A snowy corridor with a railway line running through it has been rebuilt in 3D (panel C). Two road vehicles carry roof rigs with four labelled sensors: Visual Camera, Thermal Camera, GPS receiver and LiDAR. Three top-down strips of the same track corridor follow, two with the rails traced in blue and encroaching vegetation segmented in red, one with the vegetation segmented in purple. A photogrammetric point cloud of the corridor appears once in true colour and once coloured by elevation. Finally a ROS-style dashboard shows six panes: an IR camera looking left and right, RGB cameras looking right, forward and left, and a LiDAR output.

Read together, this appears to be a rail corridor inspection platform: a Multi-sensor platform mounted on a UAV and on vehicles, mapping the track and the vegetation around it. The slide attributes no paper and no result, so treat everything beyond the visible labels as a description of the problem class. The instructor did co-author a review of exactly this field, a literature survey of automated machine vision inspection systems in railways, prepared at the National Research Council of Canada for Transport Canada (Siddiqui, Mammeri and Liu, 2021), which is the closest citable link for the theme.

Four sensors, four different questions

Each sensor on the roof rig answers a question the others cannot. A visual camera sees colour and texture, which is what separates green vegetation from grey ballast in the orthomosaic strips. A Thermal camera images emitted infrared instead of reflected light, so it sees temperature: a hot bearing or a person on the track at night shows up with no visible illumination at all. LiDAR is, in NOAA's definition, a remote sensing method that uses light in the form of a pulsed laser to measure ranges, and those ranges become the 3D Point cloud in which clearance and elevation can be measured. GPS contributes no picture of the scene, only where and when each frame was taken, which is what lets overlapping aerial photos be stitched into a georeferenced strip and the segmentation overlays registered onto it.

SensorMeasuresGood atWeak atData rate pressure
Visual (RGB) cameraReflected visible lightColour and texture: vegetation against rail, paint, signsDarkness, rain, glareHighest: full-colour frames at video rate
Thermal (IR) cameraEmitted infrared, so temperatureHot bearings, people and animals at nightNo colour, lower resolutionModerate: fewer pixels per frame
LiDARLaser pulse time of flight, so rangeGeometry, clearance, elevation, 3D point cloudsSparse at long range, cost, weightHigh: millions of points per second
GPS receiverPosition and timeGeoreferencing every frame and pointNothing about the scene itselfNegligible
The four modalities on the slide 3 rigs

Why the payload cannot carry much

The reason this slide belongs in an embedded machine learning course is the platform, not the sensors. A UAV lifts a fixed mass, and every gram of camera, LiDAR or computer is a gram less battery. Its flight time is set by that battery, and the compute it can run is set by what the battery can feed without shortening the flight. A car rig has more of all three, but still nothing like a data center. So there are exactly two ways to turn the sensor streams into an inspection result: shrink the model until it runs on board, which is the E of SERV, or ship the raw data somewhere that can process it.

Three sensors feed one small chip on the platform. The weight, power and compute bars fill almost to the limit but must not overflow it.

Shipping the raw data does not scale. Satyanarayanan's edge computing argument gives the numbers: if 12,000 users each stream 1080p video, the ingress link into the cloud needs about 100 Gbit/s, and a million users would need 8.5 Tbit/s. The cumulative demand is "considerably lower if the raw data is analyzed on cloudlets" near the sensors and only the extracted information travels onward (Satyanarayanan, 2017). Five cameras on one vehicle are a very small version of the same problem.

Worked example

How much video does a five-camera rig produce?

  1. Assume a bit rate per camera

    Take five 1080p cameras at an assumed 8 Mbit/s each after compression. This is an illustrative figure, not one from the slide.
  2. Total stream

    5 × 8 = 40 Mbit/s leaving the rig continuously.
  3. One hour of driving

    40 × 3600 = 144,000 Mbit, which is 18 GB per hour.
  4. Upload over a rural link

    At a 5 Mbit/s cellular uplink, 144,000 / 5 = 28,800 s, that is 8 hours to upload one hour of driving.
  5. Result

    The rig cannot upload what it records. It must detect on board and upload detections, which is why the model on the chip has to fit the platform's budget.

Recall

Name the four sensors labelled on the slide 3 vehicle rigs and one thing each measures.

Visual camera (colour and texture from reflected light), thermal camera (emitted infrared, so temperature), LiDAR (pulsed-laser range, giving 3D points), GPS receiver (position and time for georeferencing).

A farm robot drives between rows of cotton seedlings and must find the weeds without stopping. That is an Object detection task, and the model that does it rides on the robot: no cloud, no rack, just whatever compute the vehicle's battery can feed. Slide 4 shows the instructor's answer, EcoWeedNet, described on the slide as lightweight, computationally and energy efficient, and aimed at sustainable, low-carbon, consumer-electronics-based Precision agriculture.

Farm device
consumer electronics, vehicle or robot

Drives the field and carries the camera.

captured images
Smart farm images
cotton seedlings and weeds

Frames taken in the field, in sun and shadow.

images in
EcoWeedNet
lightweight detector

A small network that runs within the robot's budget.

boxes
Weed detections
automated output

Boxes the robot can act on.

The EcoWeedNet pipeline as drawn on slide 4

Efficient: what a nano detector is

The comparison on slide 4 is against YOLO11n and YOLO12n. In the YOLO family the suffix is a size class: n (nano) is the smallest, then s, m, l and x. The classes share one architecture and differ in width and depth, so the table below is the cleanest picture of what "lightweight" costs and buys. The figures are Ultralytics' own, on COCO validation at 640 px.

ModelParametersFLOPsmAP50-95CPU ONNX latencyT4 TensorRT latency
YOLO11n2.6M6.5B39.556.1 ms1.5 ms
YOLO12n2.6M7.5B40.6not listed1.64 ms
YOLO11m20.1M68.1B51.5183.2 ms4.7 ms
YOLO11x56.9M195.3B54.7462.8 ms11.3 ms
YOLO size classes (Ultralytics documentation)

Worked example

Nano versus medium

  1. Parameters

    20.1 / 2.6 ≈ 7.7x fewer parameters in YOLO11n than YOLO11m.
  2. Compute

    68.1 / 6.5 ≈ 10.5x fewer floating point operations per image.
  3. Latency

    CPU latency 183.2 / 56.1 ≈ 3.3x lower. The ratio is smaller than the FLOP ratio because small models leave part of the hardware idle.
  4. Price

    51.5 − 39.5 = 12.0 mAP50-95 points lower.
  5. Result

    A nano model trades a tenth of the compute for twelve points of accuracy. On a robot that must run at frame rate on a battery, that is usually the right trade, and it is the trade EcoWeedNet is designed to improve on.

The EcoWeedNet paper reports 95.2% mAP@0.5 on the CottonWeedDet12 dataset with about 4.21% of YOLOv4's parameters and 6.59% of its GFLOPs, using two parameter-free attention modules (Khater et al., 2025). Whether a Lightweight model is good enough is always this pair of numbers: an accuracy metric next to a cost metric, never one alone.

What a GradCAM++ heatmap does and does not prove

The slide's Figure 5 lays GradCAM++ heatmaps from YOLO11n, YOLO12n and EcoWeedNet side by side on seedling images, and its caption says EcoWeedNet showed better emphasis and focus on weed-relevant regions. GradCAM++ computes a weighted combination of the positive partial derivatives of the last convolutional layer's feature maps with respect to a class score (Chattopadhyay et al., 2018). In plain terms it answers "which pixels pushed this score up". It does not answer "is this score right". A model can attend to the weed and still misclassify it, or attend to the soil and get lucky. The caption's claim is a qualitative observation about attention; the 95.2% is the quantitative claim about correctness, and they are different kinds of evidence.

Robust: the same detector in the rain

Slide 5 moves from a farm to a rainy sky. The task is to tell a UAV from a bird in weather that smears both into grey blobs. The baseline is YOLOv5-m, the medium class of an earlier family (21.2M parameters, 49.0B FLOPs in the Ultralytics YOLOv5 table). The instructor's model is YOLO-RAW, published in IEEE Transactions on Intelligent Transportation Systems and evaluated on three adverse test sets, rainy, additive Gaussian noise and motion blurred (Munir et al., 2025). The slide never expands RAW, and neither does the paper's title or abstract, so this course does not expand it either.

SceneObjectYOLOv5-mYOLO-RAW
Rain skyGull 1 (top left)no boxBird 0.50
Rain skyGull 2 (centre)UAV 0.29 (false positive)Bird 0.63
Rain skyGull 3no boxBird 0.54
Rain skyGull 4no boxBird 0.73
Rain skyGull 5 (bottom)Bird 0.35Bird 0.37 (approximate)
Rain fieldReal droneUAV 0.90UAV 0.93
Every label printed on slide 5, with the same gull numbered the same way in both panels

Read the second row first. YOLOv5-m looks at the centre gull and reports UAV 0.29: a False positive, and one that a drone-detection system would act on. Its only correct box is Bird 0.35. YOLO-RAW boxes all five birds with confidence scores from about 0.37 to 0.73 and reports no drone in the sky. In the bottom row, where a real drone flies over a field, both models find it, 0.90 against 0.93. The gain is not on the easy target; it is on the ambiguous ones.

A detector reports every box above a confidence threshold. Ultralytics' default is 0.25, and its documentation says raising it "can help reduce false positives" (Ultralytics predict mode). So a natural question is whether the baseline's mistake could be fixed by tuning the threshold instead of changing the model. Drag the slider and watch what each setting costs.

SimulatorConfidence threshold on the slide 5 scene
Scene
0.25
YOLOv5-m
UAV 0.29Bird 0.35
YOLO-RAW
Bird 0.50Bird 0.63Bird 0.54Bird 0.73Bird 0.37~
ModelShownTPFPFNPrecisionRecall
YOLOv5-m21140.500.20
YOLO-RAW55001.001.00

At the default 0.25 YOLOv5-m reports a false UAV at 0.29 and one bird; YOLO-RAW reports all five birds and no false positives.

Boxes and confidences are the ones printed on slide 5; the fifth YOLO-RAW label is partly occluded and marked ~. Teal boxes match the true class, light boxes do not. Ground truth is 5 birds.

At 0.30 the false UAV disappears and YOLOv5-m loses nothing else, but it still sees one bird in five. At 0.60 YOLOv5-m already reports nothing and YOLO-RAW has lost three of its five correct birds. The threshold only slices whatever confidence mass the model produced. YOLO-RAW is better because it moved that mass onto the right classes in the first place, not because it was thresholded differently.

precision=TPTP+FP,recall=TPTP+FN\text{precision} = \frac{TP}{TP + FP}, \qquad \text{recall} = \frac{TP}{TP + FN}
The two numbers a threshold trades against each other

Recall

In the slide 5 rain scene, which model produced a false positive, what was it, and at what confidence?

YOLOv5-m labelled a gull as UAV 0.29. YOLO-RAW labelled the same bird Bird 0.63.

Recall

What does a GradCAM++ heatmap show and what does it not show?

Where in the image the positive gradients of the class score concentrated, that is which regions drove the score. It does not show whether the prediction is correct; that needs mAP, precision or recall on held-out data.

Recall

Roughly how many times fewer parameters and FLOPs does YOLO11n have than YOLO11m, and at what mAP cost?

About 7.7x fewer parameters (2.6M against 20.1M) and 10.5x fewer FLOPs (6.5B against 68.1B), for 12.0 mAP50-95 points (39.5 against 51.5).

Quick check

Why would a farm robot's designer pick a nano YOLO variant over a medium one?

Quick check

In the slide 5 rain scene, which model labelled a gull as a UAV?

Quick check

What does a GradCAM++ heatmap over a seedling image show?

A home or factory gateway forwards packets from hundreds of Internet of Things devices, and it has a fraction of a phone's compute. Somewhere in that box a decision has to be made, per packet or per flow, about whether the traffic is an attack. That is an Intrusion detection system, and slide 6 is the first page of the instructor's PNet-IDS paper, a lightweight and generalizable Convolutional neural network for intrusion detection in the Internet of Things.

The paper as printed on slide 6

Title
PNet-IDS: A Lightweight and Generalizable Convolutional Neural Network for Intrusion Detection in Internet of Things
Authors
Auwal Sani Iliyasu, Abdul Jabbar Siddiqui (Member, IEEE), Houbing Song (Fellow, IEEE), Fahad Jibrin Abdu
Venue
IEEE Access, open access research article
Timeline
Received 6 May 2025, accepted 22 May 2025, published 2 June 2025, current version 19 June 2025
DOI
10.1109/ACCESS.2025.3575705
Funding
KFUPM Deanship of Research and the IRC for Intelligent Secure Systems, grant INSS2309

Why an intrusion detector must be lightweight

The abstract states the problem in one sentence: deep learning detectors perform well, but their "high computational and storage requirements make these impractical for most IoT devices" (Iliyasu et al., 2025). NIST's guide to intrusion detection and prevention lists four classes of system, and the relevant one here is network-based: a sensor that watches traffic on a segment and analyses it for suspicious activity (NIST SP 800-94). On an IoT gateway that sensor is the gateway itself, so the model competes for the same small CPU that forwards the packets, and it has to keep up with the link, which network engineers call running at line rate.

Worked example

Line-rate arithmetic on a gateway (illustrative numbers)

  1. Packets per second

    A 100 Mbit/s link carrying 500-byte packets delivers 100×106/(500×8)=25,000100 \times 10^{6} / (500 \times 8) = 25{,}000 packets per second. A flow-level IDS aggregates packets before scoring, so this per-packet figure is the upper bound on inspection cost.
  2. Time budget per packet

    1 / 25,000 = 40 µs. Any per-packet decision slower than that falls behind the link.
  3. A heavy model

    If one decision needs 10 MFLOPs and the gateway CPU sustains 1 GFLOP/s, one decision takes 10 ms, about 250x too slow.
  4. A light model

    At 0.02 MFLOPs per decision the same CPU answers in 20 µs, inside the budget.
  5. Result

    The deadline is set by the link, not by the designer. Reducing FLOPs is what moves a detector from impractical to deployable, which is why the paper counts FLOPs, parameters and model size as first-class results.

That is the Efficient leg. The Robust leg is the word generalizable in the title. An IDS trained on one dataset meets, in deployment, traffic and attacks it never saw, a distribution shift. PNet-IDS addresses this with a knowledge distillation framework, in which a small student network is trained to reproduce a larger teacher, and reports that robustness against distribution shifts in network traffic is enhanced as a result. The benchmarks are BoT-IoT and CIC-IDS2017, and the claimed wins are reduced parameter count, reduced FLOPs and reduced model size while maintaining high accuracy and precision (Iliyasu et al., 2025).

Recall

Give two reasons an IoT intrusion detector must be lightweight, and say what generalizable means in PNet-IDS.

It runs on gateways with little compute and storage, and it must decide at line rate, between packet arrivals. Generalizable means keeping accuracy under distribution shift, on traffic or attacks unlike the training set, which PNet-IDS supports with knowledge distillation.

Quick check

Why must an IoT intrusion detector such as PNet-IDS be lightweight?

The last two slides are photographs with no captions, and both show a decision that has to be made before the world moves on. Slide 7 pairs a low-speed autonomous shuttle with a pedestrian crossing in front of it and a schematic of a T-intersection where a red arrow, a vehicle turning left, crosses a blue arrow, a pedestrian in the crosswalk. Slide 8 shows grayscale undercarriage images of rail cars: one component circled as defective or missing, and a crack pulled out in a zoom box.

The left-turn pedestrian conflict

The schematic is the left-turn pedestrian conflict, a named pattern in Intelligent transportation systems safety analysis. A left-turning driver scans for oncoming vehicles, accepts a gap, and turns into a crosswalk where a pedestrian is already crossing. FHWA research on left-turn phasing found that protected phasing reduced vehicle-to-vehicle injury crashes but did not significantly reduce vehicle-to-pedestrian crashes, and that under permissive phasing the turning driver must pick a gap in oncoming traffic and yield to the crosswalk at the same time, which prior studies flag as a strong conflict potential (FHWA, HRT-18-044). For an autonomous shuttle the problem is the same one the human driver has, and the detector that watches the crosswalk must answer before the vehicle enters it.

Looking under the train

Slide 8 belongs to Automated visual inspection of rolling stock: cameras placed by the track or under it photograph each wagon as it passes, and a model flags anything that does not look like an intact component. The instructor reviewed exactly this field for Transport Canada while at NRC (Siddiqui, Mammeri and Liu, 2021). The slide itself carries no citation, so the images are described here as an example of the problem class, not attributed to a particular system.

A wayside camera sweeps a passing wagon, boxes a crack, and pulls it out into a zoom for the inspector.

Physics sets the deadline

What unites the shuttle and the wagon is that the time available for Inference is not a design preference. It is set by how fast the scene changes. Satyanarayanan's remark that "the speed of light is an obvious physical limit on latency" (Satyanarayanan, 2017) is the reason a cloud round trip cannot be the answer here: the distance the vehicle covers during the latency of the reply is the safety margin lost.

Worked example

Metres lost per millisecond of latency

  1. The shuttle

    At 10 m/s a vehicle covers 1 m every 100 ms. A 200 ms cloud round trip is 2 m of travel before the detector's answer arrives.
  2. The wagon

    A wagon passing a wayside camera at 15 m/s moves 1.5 m every 100 ms. At 30 frames per second the model has about 33 ms per frame before the next one arrives and the wagon has moved another 0.5 m.
  3. Result

    In both cases the budget is tens of milliseconds, and the network alone can spend that. The model runs on an Embedded system next to the camera, which means it must be small enough to fit there.
ScenarioSensorDecisionDeadline driverCost of a miss
Shuttle at a T-intersectionForward camera on the vehicleYield, brake or proceed on a left turnVehicle and pedestrian speedA pedestrian is struck
Rolling-stock undercarriageWayside or under-track cameraFlag a cracked or missing componentTrain speed past the cameraA defect leaves the yard undetected
Two safety-critical edge inference problems

Recall

Why is a cloud round trip unacceptable for the shuttle and the wagon inspection, in one sentence each?

The shuttle moves about 2 m during a 200 ms round trip, which is the margin between yielding and striking a pedestrian. The wagon moves out of the camera's view within tens of milliseconds, so the frame must be analysed before the next one arrives.

Recap

If you remember nothing else

  • SERV: Secure (adversarial ML, IDS), Efficient (compression, nano models), Robust (rain, noise, blur, shift), Vision.
  • Inspection rigs carry RGB, thermal, LiDAR and GPS; a UAV's weight, battery and compute cap what the on-board model can be.
  • Analyze near the sensor, ship the extracted information: 12,000 1080p streams already need 100 Gbit/s into a cloud.
  • A nano YOLO is the smallest scale: YOLO11n has 2.6M parameters and 6.5B FLOPs against 20.1M and 68.1B for YOLO11m, for 12 mAP points.
  • EcoWeedNet: 95.2% mAP@0.5 on CottonWeedDet12 with about 4.21% of YOLOv4's parameters and 6.59% of its GFLOPs.
  • GradCAM++ shows where the class score's evidence sits; it does not certify correctness.
  • YOLOv5-m produced a false UAV at 0.29 in rain; YOLO-RAW boxed the birds at 0.37 to 0.73. A threshold slices confidence mass; a better model moves it. RAW is not expanded.
  • PNet-IDS cuts FLOPs, parameters and size so an IoT gateway can detect intrusions in real time, and uses knowledge distillation to hold up under traffic distribution shift.
  • Shuttle intersections and undercarriage inspection are deadline-bound: distance travelled per millisecond of latency decides where inference runs.

Sources

Part 02: Course takeaways and getting started

What the course promises you will be able to do, the five pillars of efficient deep learning on constrained devices, and what to do in week one.

3 concepts, slides 9-12

Why this part matters

This part is the course contract. Every later lecture, from efficiency metrics through pruning, quantization, architecture search and deployment, is one turn of a five-step loop that slide 11 draws as a neural network funnelling down into a small development board.

Your exams and your research project will ask you to do two things with that loop: place a technique in it (is knowledge distillation an accelerate step or a compare step?), and compute whether a given model fits a given device. Both start here. The part also covers what the instructor asks of you in week one, and what the deck deliberately leaves to Blackboard.

By the end you can

  1. State the five course takeaways and reorder them as a measure, optimize, compare, explore, deploy loop.
  2. Compute a model's weight size from parameter count and bit width, and judge whether it fits a device tier.
  3. Explain why accuracy alone cannot select a model for a constrained device.
  4. Distinguish on-device inference from on-device training by their memory needs.
  5. Set up a week-one plan: survey, prerequisites check, device tier, candidate model, metrics log.

On the first evening each student says three things out loud: their name, their educational background (previous studies), and their work background (current or previous work, industry or occupation). Then everyone opens the Pre-course survey at https://forms.cloud.microsoft/r/5ssu8gkcS9, also printed as a QR code on slide 10, and answers questions about background, skills and expectations.

It is tempting to file both under administration, but they serve the same purpose as takeaway 1 later in this part: you cannot tune what you have not measured. A graduate course on efficient deep learning draws people from computer engineering, electrical engineering, computer science and industry. Some have trained convolutional networks for years and never flashed a microcontroller; others write firmware for embedded systems daily and have never opened PyTorch. The instructor uses the room's answers to set depth and pacing, so an honest answer about a gap is worth more to you than an impressive one.

What is worth refreshing before week two

The deck does not list prerequisites, so treat the following as a hedged suggestion drawn from what the course will ask you to do rather than as a slide claim. The topics ahead (compression, pruning, quantization, deployment to boards) match the syllabus of MIT's 6.5940 TinyML and Efficient Deep Learning Computing, which lists an introductory machine learning subject and a computation structures subject as its prerequisites. Translated to tools, four skills will carry most of the weight.

Skills that will probably matter, and why

Python
Every framework, script and notebook in efficient deep learning assumes it. The PyTorch beginner tutorial itself assumes basic familiarity with Python.
PyTorch
The compression and quantization literature (Deep Compression, MCUNet, TinyTL) ships PyTorch code; you will read and modify such code, not write frameworks.
Basic CNNs
Convolutions, pooling and fully connected layers are what get pruned and quantized. The free Deep Learning book by Goodfellow, Bengio and Courville covers them.
C for microcontrollers
LiteRT for Microcontrollers is written in C++ 17 for 32-bit platforms with no operating system, so deploying to a board means touching C or C++ code.

None of these needs to be deep on day one. If you can train a small classifier in PyTorch, explain what a convolution computes, and compile and flash a blinking-LED program in C, you can follow the first weeks and fill in the rest as the loop demands it.

Recall

What three things does the icebreaker ask each student to share?

Name, educational background (previous studies), and work background (current or previous work, industry or occupation).

Take a small image classifier with 3 million parameters, trained in PyTorch and saved in 32-bit floating point. Its weights alone occupy 12 MB. A Raspberry Pi Pico has 2 MB of flash and 264 KB of SRAM. The model does not fit, and no amount of accuracy changes that. Everything this course teaches is what you do next.

Slide 11 states the promise as five takeaways. You will know the key Efficiency metrics for deep learning computing. You will understand how to speed up Inference and Training of neural networks on resource-constrained platforms. You will analyse the tradeoffs between different optimization techniques. You will be aware of recent research trends and industry practice. And you will get hands-on experience compressing and implementing deep learning models on resource-constrained devices. The infographic relabels the same five as Measure efficiency, Accelerate, Analyse trade-offs, Track the frontier, and Build and deploy, with a five-word footer: Measure, Optimize, Compare, Explore, Deploy.

A network funnels into one spine; the five pillars light up in order, the board at the bottom powers on, and a loop returns to pillar 1

Read the five as a loop, not a list

The order on the slide is not arbitrary. It is the order in which you actually work. You measure a baseline: size, memory, compute, latency, energy. You apply one optimization, say pruning. You compare the new numbers against the baseline and note what the step cost in accuracy. You look to the research literature and to industry for the next technique. You deploy on the real board to check that the paper numbers survive contact with hardware, because measured latency and energy routinely differ from analytical estimates. Then you measure again, and the loop turns.

PillarVerbQuestion you askDeep Compression example
1 Measure efficiencyMeasureHow big, how slow, how hungry is the baseline?AlexNet: 240 MB, 61 M parameters, 32-bit weights
2 AccelerateOptimizeWhich technique shrinks the cost that hurts most?Prune 9x fewer connections, quantize to 5 bits, Huffman code
3 Analyse trade-offsCompareWhat did each step cost in accuracy or effort?35x smaller with no accuracy loss, after retraining
4 Track the frontierExploreWhat has research or industry done since?MCUNet, TinyTL and on-device training under 256 KB
5 Build and deployDeployDoes it really run on the target hardware?3x to 4x layerwise speedup and 3x to 7x energy efficiency measured on hardware
The five pillars as one loop, with Deep Compression (Han, Mao and Dally, 2016) as the running example

Deep Compression is the canonical worked loop. Han, Mao and Dally started from AlexNet at 240 MB and VGG-16 at 552 MB (measure), pruned redundant connections, quantized the survivors from 32 to as few as 5 bits and Huffman coded the result (optimize), reported 35x and 49x reductions with no loss of accuracy (compare), and measured 3x to 4x layerwise speedup and 3x to 7x better energy efficiency on real hardware (deploy). The companion paper on learning both weights and connections had already shown pruning alone cutting AlexNet from 61 M to 6.7 M parameters. Everything after 2016 (MCUNet, TinyTL, on-device training under 256 KB) is the explore step feeding the next turn.

A preview of the efficiency metrics

Slide 11 names the metrics without listing them; the list below previews later lectures and is not slide content. The ordering follows the MIT 6.5940 sequence, which opens with the basics of efficient deep learning before pruning, quantization, neural architecture search, distillation and on-device training. Learn the names now so the loop has something concrete to measure.

MetricWhat it countsTypical unit
Parameter countHow many learnable weights and biases the model hasmillions (M)
Model sizeBytes needed to store the weights: parameters times bits per weight over 8MB or KB
Peak memoryThe most memory needed at any moment: weights plus the largest activation tensorsMB or KB
ComputeMultiply-accumulate operations (MACs) or floating-point operations (FLOPs) per inferenceMMACs or GFLOPs
LatencyTime from one input to its outputms
ThroughputInferences completed per secondinferences/s
EnergyJoules per inference, or average power while runningmJ or mW
Efficiency metrics you will meet (preview, not slide content)

The first two are the easiest to compute by hand, and the calculation is one you will repeat all term. A weight stored in b bits takes b / 8 bytes, so the size of the weights is the parameter count times bits per weight divided by eight.

size (bytes)=Nparams×b8\text{size (bytes)} = \frac{N_{\text{params}} \times b}{8}
Weight storage for N parameters at b bits each; MB here means one million bytes

Worked example

A 3 million parameter model at four bit widths

  1. Start from float32

    3,000,000 × 32 / 8 = 12,000,000 bytes, about 12 MB. This is what PyTorch saves by default.
  2. Halve the bits, halve the bytes

    At 16 bits the same weights take 6 MB; at 8 bits, 3 MB; at 4 bits, 1.5 MB. Nothing else changed: the parameter count is identical.
  3. Test against two boards

    The Raspberry Pi Pico has 2 MB of flash, so only the 4-bit version fits. The Arduino Nano 33 BLE Sense Rev2 has 1 MB, so none of the four fit, and quantization alone is not enough.
  4. Result

    Bit width divides size, but the target device decides whether the result is small enough. Below 1 MB this model also needs pruning or a different architecture.
Bits per weightCalculationSizeFits Pico flash (2 MB)?Fits Nano 33 BLE Sense flash (1 MB)?
32-bit3,000,000 × 32 / 812,000,000 B (12 MB)NoNo
16-bit3,000,000 × 16 / 86,000,000 B (6 MB)NoNo
8-bit3,000,000 × 8 / 83,000,000 B (3 MB)NoNo
4-bit3,000,000 × 4 / 81,500,000 B (1.5 MB)YesNo
3 M parameters at four bit widths, tested against two microcontroller flash budgets
The same 3 M parameters at 32, 16, 8 and 4 bits; only the 4-bit bar slides inside the 2 MB flash line

Try it yourself. The calculator applies the formula for any parameter count and bit width and checks the result against generic device tiers. The device figures are datasheet capacities for one example board per tier (Arduino, Raspberry Pi and NVIDIA product pages), not what is free once the runtime and your code are loaded.

SimulatorDoes it fit?
Target device

Microcontroller, example board: Arduino Nano 33 BLE Sense Rev2 (nRF52840)

3.00 M
Bits per weight
100 KB
Weight size12.00 MB3.00 M × 32 ÷ 8
Activations100 KByour estimate, not computed
Flash, weights storedDoes not fit
needs 12.00 MB (weights only)of 1.05 MB
SRAM, weights read from flashFits
needs 116 KB (activations plus a 16 KB allowance for runtime and stack)of 262 KB
SRAM, weights copied inDoes not fit
needs 12.12 MB (weights, activations and the same allowance)of 262 KB

Activations, the runtime and your application code also need memory. These figures are datasheet capacities, not what is free. Model sizes use MB as 10^6 bytes; the microcontroller banks use the datasheet binary figures (1 MiB flash, 256 KiB SRAM), which are slightly larger.

Three misconceptions the loop corrects

Inference stacks weights and one set of activations; training adds stored activations, gradients and optimizer state and spills past the SRAM line

Quick check

You compare pruning and quantization on the same network and report size, latency and accuracy for each. Which takeaway is this?

Quick check

A network has 12 million parameters stored as 8-bit integers. What is its weight size?

Quick check

Why can accuracy alone not choose a model for a microcontroller?

Quick check

Which memory cost appears in on-device training but not in on-device inference?

Recall

Name the five course takeaways in the order of the loop, then say what comes after the fifth.

Measure efficiency, accelerate (optimize), analyse trade-offs (compare), track the frontier (explore), build and deploy. After deploying you measure the deployed model again, and the loop turns.

Recall

A 3 M parameter model with 8-bit weights: how many bytes, and does it fit 1 MB of flash?

3,000,000 × 8 / 8 = 3,000,000 bytes, about 3 MB. It does not fit 1 MB of flash; even at 4 bits it is 1.5 MB.

Recall

Why is on-device training harder than on-device inference?

Training must store activations for the backward pass, gradients and optimizer state, so peak memory is far higher. TinyTL shows activations, not parameters, are the bottleneck, and LiteRT for Microcontrollers does not support on-device training.

Slide 12 contains three words: "View on Blackboard". That is the whole syllabus section of the deck. It does not give grading weights, exam dates, project deliverables, a textbook, an office location or office hours. Anything you read elsewhere claiming to know those for this course is not coming from these slides.

What is known so far (meeting details from the course registry, not this deck)

Meetings
Monday and Wednesday, 20:10 to 21:25
Room
Building 59, Room 2004
Syllabus, grading, deliverables
On Blackboard, not in this deck
Textbook, office, office hours
Not stated in this deck; check Blackboard

Plan the project before the syllabus tells you to

Takeaway 5 promises hands-on experience compressing and implementing models on constrained devices, and the MIT counterpart of this course ends in an open-ended design project. Read that as a strong hint that a hardware Deployment is coming, whatever Blackboard eventually calls it. The students who struggle with such projects are rarely short of techniques; they are short of a baseline to compare against and a board that arrives on time. Both are fixed in week one.

  1. Pick a device tier first. The memory budget of the tier decides which models are even candidates, so the device is an input to model selection, not an output of it.
  2. Pick a candidate model that already fits, or is within one or two optimizations of fitting, and record its parameter count and float32 size with the formula from the previous concept.
  3. Start a metrics log before the first optimization. A baseline recorded after you have started changing things is not a baseline.
TierExample boardMemoryWhat it is typically asked to run
MicrocontrollerArduino Nano 33 BLE Sense Rev2 (nRF52840)1 MB flash, 256 KB SRAMKeyword spotting, gesture and anomaly detection, tiny image classifiers
MicrocontrollerRaspberry Pi Pico and Pico 22 MB flash and 264 KB SRAM; Pico 2 has 4 MB and 520 KBSame class, a little more headroom
Single-board computerRaspberry Pi 4 Model B1 to 8 GB LPDDR4 RAMQuantized MobileNet-class vision, small detectors, a Linux runtime
GPU moduleJetson Orin Nano Super Developer Kit8 GB LPDDR5 at 102 GB/s, 67 INT8 TOPSReal-time detection, small transformers, on-device fine-tuning
One example board per tier, with datasheet capacities (Arduino, Raspberry Pi and NVIDIA product pages)

The last column is a guide, not a datasheet claim. MCUNet reached over 70% ImageNet top-1 accuracy on a commercial microcontroller, so the microcontroller tier is more capable than its numbers suggest, but only after the whole loop has been applied to a purpose-built architecture. A model chosen for a GPU module and then squeezed onto a microcontroller is the hardest path in the course.

WhenWhat to record
Week 1Target device tier and its flash, SRAM or RAM budget from the datasheet
Week 1Candidate model, its parameter count and float32 weight size
Week 1Baseline accuracy on a held-out set you will not change later
First run on hardwareLatency per inference and peak memory as the runtime reports them
Every optimizationThe same four numbers again, next to the baseline, plus what it cost
A starter metrics log

Keep the log as a plain table with one row per experiment and the same columns every time: accuracy, parameter count, size, peak memory, latency, energy if you can measure it, and a one-line note of what changed. This is the compare step made permanent, and it is the table an examiner or a reviewer will want to see.

Quick check

Where is the COE 592 grading scheme published, according to the deck?

Recall

Why does the deck not tell you the grading scheme, and what should you do about it?

Slide 12 says only "View on Blackboard". The syllabus, assessments, textbook and office hours live on Blackboard, so check there; the site will add assessments once the syllabus is transcribed.

Recall

What three numbers should be in your metrics log before you run the first optimization?

Baseline accuracy, weight size (from parameter count and bit width) and latency on the target hardware, with peak memory as soon as the runtime reports it.

Recap

If you remember nothing else

  • Introduce yourself with name, education and work background, and fill in the pre-course survey so the course can be calibrated to the room.
  • Five takeaways: measure efficiency, accelerate, analyse trade-offs, track the frontier, build and deploy. Read them as a loop, not a list.
  • Weight size in bytes equals parameters times bits per weight divided by 8: 3 M parameters is 12 MB at 32-bit and 1.5 MB at 4-bit.
  • Size, peak memory, compute, latency, throughput and energy are separate metrics and can disagree.
  • Compression trades accuracy and effort; Deep Compression's 35x came with retraining, not for free.
  • Training on device must hold activations and gradients, far more than inference, and LiteRT for Microcontrollers does not support it.
  • The syllabus, grading and deliverables are on Blackboard, not in this deck. Class: Mon and Wed 20:10 to 21:25, Building 59, Room 2004.
  • Pick a device tier and a candidate model in week one, and start a metrics log before the first optimization.

Sources