Viva colour cheat sheet · one page per topic, blocks stacked one below the other · A4. Print with "Background graphics" enabled.
1

Deep Learning foundations

MAI590/DEN790 · losses, optimisers, backprop, CNN arithmetic, detection metrics, LSTM, attention
formulanumbers to quotetrap / critiquesay this in the exam

Softmax & cross-entropy

$$p_k=\frac{e^{s_k}}{\sum_j e^{s_j}},\qquad L=-\log p_y,\qquad \frac{\partial L}{\partial s_j}=p_j-\mathbb 1[j=y]$$

Init loss $=\log C$ ($\ln10=2.30$). Example (3.2, 5.1, −1.7): $p=(0.13,0.87,0.00)$, $L=2.04$.

Multiclass SVM (hinge)

$$L_i=\sum_{j\neq y}\max(0,\,s_j-s_y+1)$$

Cat example $=2.9$; init $=C-1$. Zero once margins hold (stops learning); softmax never stops. $2W$ keeps $L=0$ → need $\lambda R(W)$.

Optimisers

$$\text{SGD: }w\leftarrow w-\alpha\nabla L\qquad\text{Momentum: }v\leftarrow\rho v+\nabla L,\ w\leftarrow w-\alpha v$$ $$\text{Adam: }m\leftarrow\beta_1m+(1-\beta_1)g,\ v\leftarrow\beta_2v+(1-\beta_2)g^2,\ \hat m=\tfrac{m}{1-\beta_1^t},\ \hat v=\tfrac{v}{1-\beta_2^t},\ w\leftarrow w-\alpha\tfrac{\hat m}{\sqrt{\hat v}+\epsilon}$$

$\beta_1$ 0.9, $\beta_2$ 0.999, start lr $10^{-3}$ or $5\cdot10^{-4}$. Numerical gradient check: $(1.25322-1.25347)/10^{-4}=-2.5$.

Backprop rules

downstream = local × upstream. add distributes, mul swaps, max routes. $q=x+y,\ f=qz$: $\partial f/\partial x=z=-4$.

$$y=xW:\quad \frac{\partial L}{\partial x}=\frac{\partial L}{\partial y}W^T,\qquad \frac{\partial L}{\partial W}=x^T\frac{\partial L}{\partial y}$$

Sigmoid: $\sigma'=\sigma(1-\sigma)=0.73\cdot0.27=0.20$; $dw=[-0.2,-0.39,0.2]$, $dx=[0.39,-0.59]$.

Convolution arithmetic

$$W'=\Big\lfloor\frac{W-K+2P}{S}\Big\rfloor+1,\qquad \#\text{params}=C_{out}(C_{in}K^2+1),\qquad \text{MACs}=C_{in}K^2\cdot W'H'C_{out}$$

3×32×32, 10×(5×5), S1 P2 → 10×32×32, 760 params, 768,000 MACs. "Same" $P=(K-1)/2$. Receptive field $1+L(K-1)$. VGG: three 3×3 = one 7×7, $27C^2$ vs $49C^2$. ResNet $H(x)=F(x)+x$.

Numbers to quote

ILSVRC top-5AlexNet 8L · VGG 7.3% · ResNet-152 3.57%CIFAR-1050k/10k, 32×32×3 = 3072Max-pool 2×2/2[[1,1,2,4],[5,6,7,8],[3,2,1,0],[1,2,3,4]] → [[6,8],[3,4]]Dropoutp 0.5; inverted: ÷(1−p) in trainingConv 7×7 backward∂L/∂b = 8; ∂L/∂W = [[20,10,2],[5,18,16],[15,10,4]]Transformer sizes12L/213M · GPT-2 48L/1.5B · GPT-3 96L/175B

Detection metrics

$$\mathrm{IoU}=\frac{|P\cap G|}{|P\cup G|},\quad P=\frac{TP}{TP+FP},\quad R=\frac{TP}{TP+FN},\quad AP=\sum_i\Delta R_iP_i,\quad \mathrm{mAP}=\tfrac1C\sum AP_c$$

TP: class ok and IoU ≥ τ. Duplicate = FP. No TN. Course table: Cat 0.6667, Dog 0.5, Bicycle 0.3333 → mAP 0.5000. Bicycle IoU 0.74 fails τ 0.75.

RNN → LSTM

$$h_t=\tanh(W_{hh}h_{t-1}+W_{xh}x_t)\quad\Rightarrow\quad \prod_t \mathrm{diag}(\tanh')W^T\ \text{vanishes/explodes}$$ $$\begin{pmatrix}i\\f\\o\\g\end{pmatrix}=\begin{pmatrix}\sigma\\\sigma\\\sigma\\\tanh\end{pmatrix}W\begin{pmatrix}h_{t-1}\\x_t\end{pmatrix},\ c_t=f\odot c_{t-1}+i\odot g,\ h_t=o\odot\tanh c_t$$

$W$: $4h\times(h+d)$. $\partial c_t/\partial c_{t-1}=\mathrm{diag}(f)$ → "uninterrupted flow, like ResNet". Clip for exploding.

Attention & ViT

$$Q=XW_Q,\ K=XW_K,\ V=XW_V,\qquad Y=\mathrm{softmax}\!\Big(\frac{QK^T}{\sqrt D}\Big)V$$

$\sqrt D$: $\mathrm{Var}(q\cdot k)=D$ → avoid softmax saturation. Permutation-equivariant → positional encoding. Mask future with $-\infty$. Block = MHSA → +res → LN → MLP(D→4D→D) → +res → LN; 6 matmuls; $O(N^2)$. ViT: $N=HW/P^2$ patches (224/16 → 196 tokens of 768), CLS token, low inductive bias → needs big data.

Say this in the exam

"Training minimises a loss by gradient descent; backprop is just the chain rule on the computational graph. CNNs share small filters across positions; LSTMs keep a cell state whose gradient passes through an element-wise gate; attention lets every token read every other token in one step at $O(N^2)$ cost." Overfitting = low train / high test error; underfitting = both high and "cannot be fixed by more epochs".

Viva cheat sheet · Deep Learningpage 1 of 7
2

Intelligent Robots foundations

MAI675/DEN775 · PID, Hough, RANSAC, camera, YOLO/MIO, Kalman, ROS 2

The loop and the pipeline

Sense→Compute→Actuate⟳image→grey→edges (Sobel/Canny)→ROI→Hough→filter L/R→centre, error→PID→servo / wheels

PID

$$u=K_pe+K_i\!\int\! e\,dt+K_d\frac{de}{dt}\qquad u[n]=K_pe[n]+K_i\sum e[i]\Delta t+K_d\frac{e[n]-e[n-1]}{\Delta t}$$

$e=SP-PV$. P only → steady-state error (0 output at $e=0$); I removes it; D damps, amplifies noise. Tune $K_p$, then $K_d$, then $K_i$ (0.0001). $K_p=125/3500=0.0357$. Example (2.0,0.5,0.1), $\Delta t$ 0.1, $e=-0.48$, $e_{prev}=-0.30$, $\sum=-0.12$: $-0.96-0.06-0.18=\mathbf{-1.20}$.

Hough & RANSAC

$$\rho=x\cos\theta+y\sin\theta,\quad m=-\cot\theta,\ b=\rho/\sin\theta\qquad S=\frac{\log(1-P)}{\log(1-p^k)}$$

Normal form: vertical lines finite. (3,3),(4,3),(5,3) → $A(3,90°)=3$. RANSAC $P$ 0.99, $p$ 0.5: $k$=2→17, 3→35, 4→72. Reject lane parabola $|a|\ge0.003$.

Camera & stereo

$$\lambda\begin{bmatrix}u\\v\\1\end{bmatrix}=K[R|t]\begin{bmatrix}X\\Y\\Z\\1\end{bmatrix},\ K=\begin{bmatrix}f_x&s&c_x\\0&f_y&c_y\\0&0&1\end{bmatrix},\ X=\frac{(u-c_x)Z}{f_x},\ Z=\frac{f_xB}{d}$$

(700,400), $Z$ 10, $f$ 800, $c$ (640,360) → (0.75, 0.5, 10). Stereo $d=80$, $f_x$ 795, $B$ 0.2 → $Z=1.9875$ m. One image: 2 eq, 3 unknowns. Quality = reprojection error.

YOLO & MIO

Output $S\times S\times(5B+C)$: 7×7×30. Loss: $\sqrt w,\sqrt h$; $\lambda_{noobj}=0.5$; NMS IoU > 0.5. MIO: in lane if $x_L(y)\le x\le x_R(y)$, $x(y)=(y-b)/m$; MIO = argmax $y_{bottom}$. Three-car example: Car 3 (x 500) out of lane → Car 2 (290 > 260). FCW: tracks (confirm [2 3], delete 5), closest in lane; $d=1.2v+v^2/(2\cdot0.4\cdot9.8)$ → 24.8 m at 10 m/s.

Kalman filter · five scalar equations and their origin

$$\underbrace{\mu_p=\mu+v\Delta t+\tfrac12a\Delta t^2}_{\text{mechanics}}\quad \underbrace{p_p=p+q}_{\text{variances add}}\quad \underbrace{K=\frac{p_p}{p_p+r}}_{\text{confidence}}\quad \underbrace{\mu=\mu_p+K(z-\mu_p)}_{\text{Bayes mean}}\quad \underbrace{p=(1-K)p_p}_{\text{Bayes variance}}$$

Derivation: $N(\mu_p,p)\times N(z,r)$ → $\frac1{\sigma^2}=\frac1p+\frac1r$, $\mu=\frac{r\mu_p+pz}{p+r}$; set $K=\frac{p}{p+r}$. $K\to1$ trust sensor, $K\to0$ trust prediction; $0

Fusion: $z_f=\frac{\sum z_i/r_i}{\sum 1/r_i},\ r_f=\frac1{\sum1/r_i}$; (0.9,1.1),(1,4) → 0.94, 0.8; prior 10.1 → $K=0.927$, $x=0.87$, $p=0.74$. $r_i$ never changes during filtering.

Numbers to quote

Ultrasonicd = t·0.034/2 cm (2000 µs → 34 cm)L293DIN 10 fwd, 01 rev, 00 stop; EN = PWMLine bar I²Cbar 0x3E, robot 81, request 240Grey0.299R+0.587G+0.114B; (120,200,80) → 162Sobel ex.Gx 275, Gy 145 → M 311, θ 27.8°HoughLinesPrho 2, θ π/180, thr 50, minLen 10, gap 30BC net24/36/48 (5×5 s2), 64/64, FC 1254/1254/256/1, lr 1e-4ROS 2Jazzy · colcon · /thing_on Bool q10 · frame "map"

Say this in the exam

"Bang-bang chooses a direction; PID chooses how much." "Hough votes in $(\rho,\theta)$ because slope is infinite for vertical lines." "RANSAC keeps the model most points agree with; least squares is pulled by outliers." "The Kalman filter is a recursive Bayesian estimator: predict widens, update shrinks." "MIO comes from confirmed tracks because detections flicker."

Traps

Pull-up button pressed = LOW. Never delay(), use millis(). Pooling has no parameters; PID I-term needs anti-windup in practice. Behaviour cloning is regression (ELU + regression layer), not softmax. Gazebo = world, RViz = belief.

Viva cheat sheet · Intelligent Robotspage 2 of 7
3

Edge AI foundations

AIRE325 / MAI633 · quantisation, regression on MCUs, energy, TinyML, protocols, state charts

Quantisation

$$x_q=\mathrm{round}\!\Big(\frac{x}{s}\Big)+z,\qquad x\approx s(x_q-z),\qquad s=\frac{x_{max}-x_{min}}{255}$$ $$y_q=\frac{s_as_x}{s_y}(a_q-z_a)(x_q-z_x)+\frac{s_b}{s_y}b_q+z_y$$

Error ≤ $s/2$, variance $s^2/12$. Weights per-channel, activations per-tensor, bias int32 with $s_b=s_as_x$. PTQ (observe min/max) vs QAT (fake-quant nodes). MCU: int32 accumulator, shift >> 7.

Numbers to quote

Slide examplex 4.0 → x_q 102; a_q 64; b_q 13; y_q 112.8 → 11.05 (float 11.0)Exercises3.2/0.04 → 80 · 0.05(100−128) = −1.4 · 6.375/255 = 0.025Accuracy dropMobileNet 71.9→71.0 · KWS 94.0→93.8 · CIFAR 82.4→82.3Uno2 KB SRAM, 32 KB flash; int8+PROGMEM ≈ 4× neurons (50→200)Nano 33 BLE Sense256 KB / 1 MB; IMU, mic, T/H, pressureGesture lab119×6 samples, thr 2.5 g, 50-15-2, 600 ep, arena 8 KBKWS lab2500 ms @ 16 kHz, MFCC, Edge Impulse

Regression on a microcontroller

$$\hat\beta=(X^TX)^{-1}X^Ty\qquad m=\frac{n\sum xy-\sum x\sum y}{n\sum x^2-(\sum x)^2},\ c=\bar y-m\bar x$$

w = !(~X*X)*~X*y (~ transpose, ! Gauss–Jordan inverse, pivot < 1e-6 → abort). (1,2),(2,2.5),(3,3.5): $m=4.5/6=0.75$, $c=1.17$. Quadratic (1,2),(2,3),(3,5): $\beta=[2,-0.5,0.5]$. $\begin{bmatrix}2&1\\5&3\end{bmatrix}^{-1}=\begin{bmatrix}3&-1\\-5&2\end{bmatrix}$. "Linear in parameters, nonlinear in features."

Energy budgets

$$\text{uptime}=\frac{\text{capacity}}{\text{inf/day}\times\text{mAh/inf}},\qquad S(t)=e^{-\lambda t},\ \lambda=\tfrac1{\text{mean}},\qquad E[N]=\frac{T}{E[T]}$$

1000 mAh, 50 mAh every 2 h → 1.67 d. 1200 mAh, 40 mAh, 5 d → every 4 h. Uno 500 mAh: awake 50 mA → 10 h; asleep 0.1 mA → 30 days; inference 45 mA·0.8 s = 0.01 mAh. $\lambda=1/12$: $S(2)$ 0.846, $S(8.3)$ 0.5, $S(24)$ 0.135; run if $S<0.3$ (t > 14.4 h). Adaptive interval $60TE/B$: 180 s @5000, 900 s @1000. Morning 7 min + night 22.5 min → 67 events → 2010 mAh.

TinyML flow

train (PC)→prune · cluster · quantise→.tflite→xxd → model.h→TFLite Micro:ErrorReporterGetModel + versionOpResolver<N>tensor arenaInterpreterAllocateTensorsinput → Invoke → output

Raspberry Pi: Interpreter → allocate_tensors → set_tensor → invoke → get_tensor; detection input [1,320,320,3] uint8, outputs boxes/classes/scores, thr 0.3, EfficientDet-Lite0. Arduino only if model < 20 KB. Hierarchy: Arduino wakes Pi at 10–30 cm.

Protocols

UARTI²CSPI
wires3 (Rx,Tx,GND)2 (SDA,SCL)4 + n (SCK,MOSI,MISO,SS)
clocknone (baud)mastermaster
duplexfullhalffull
speed9600/115200100k/400k/3.4M~10 MHz
notesRS-232 ±3–25 V, MAX2327-bit → 128 addr; START SDA↓ while SCL high; ACK SDA low; 4.7 kΩ pull-upsSS low selects; byte per 8 clocks

State charts (method)

classify devices (in/out, dig/ana)→output string→count states→outputs/state→transitions (PB, After t, guard)

Ride OFF→ON 50%→OFF after 20 s. Fan Idle→30%→60%→100%. Traffic Green 120 → Yellow 30 → 3 blinks (6 states) → Red 120; pedestrian PB only in Green with > 30 s left. Python: Enum + loop; (value+1) % 4.

Say this in the exam

"Quantisation stores each weight as an integer plus a shared scale and zero-point; integer inference needs only an int32 accumulator and a rescale." "Idle current dominates the battery: sleep and wake on interrupts." "$S(t)$ is a survival probability, not a density." "Least squares has a closed form, $(X^TX)^{-1}X^Ty$, so a microcontroller can learn without gradient descent."

Viva cheat sheet · Edge AIpage 3 of 7
4

Computer Vision foundations

MAI621/ECE621 · histograms, Otsu, morphology, Fourier filters, spatial filters, Canny, restoration, tracking

Images, neighbours, histograms

f = i·r (illumination × reflectance); bits = M·N·k (1024² × 8 = 8,388,608). N4 / ND / N8; m-connectivity removes multiple paths. D4 = |Δx|+|Δy|, D8 = max, De = √: (1,2)→(4,6) = 7, 4, 5. Two-pass labelling with equivalence table. Point ops: negative (L−1)−r, log c·log(1+r), power s = c·rᵞ (γ<1 brightens), slicing.

$$p(g)=\frac{h(g)}{RC},\quad s_k=\mathrm{round}\big((L-1)\textstyle\sum_{j\le k}p_j\big):\ h=[5,4,0,0,2,1,3,0,4,1],\ L=10\ \to\ s=2,4,4,4,5,5,7,7,9,9\ (\text{"almost, not completely, flat"})$$

Otsu & segmentation

$$P_1=\sum_{i\le k}p_i,\ m(k)=\sum_{i\le k}ip_i,\ m_G=\sum ip_i,\quad \sigma_B^2(k)=\frac{(m_GP_1-m)^2}{P_1(1-P_1)}=P_1P_2(m_1-m_2)^2,\quad k^*=\arg\max,\ \eta=\sigma_B^2/\sigma_G^2$$

Freq 8,7,2,6,9,4 (N 36): m_G 2.3611; σ²_B = 1.593, 2.564, 2.629, 2.142, 0.870 → k* = 2, η ≈ 0.84. Limitation: histogram only, no spatial info. Iterative T = ½(m₁+m₂); Niblack T = m + k·s. Split/merge (std < 5): 4 quadrants pass, R1∪R3 and R2∪R4 merge → 2 regions. K-means 7 points: C₁(1,1), C₂(5,7) → it.1 {1,2,3}/{4–7} → it.2 point 3 moves → (1.25,1.5), (3.9,5.1) → it.3 stable.

Morphology

$$A\ominus B=\{z:(B)_z\subseteq A\}\ \text{(fit, shrink)}\qquad A\oplus B=\{z:(\hat B)_z\cap A\ne\emptyset\}\ \text{(hit, grow)}\qquad (A\ominus B)^c=A^c\oplus\hat B$$

Opening = erode→dilate (salt, thin joints); closing = dilate→erode (pepper, holes). 3×3 block: erode by 3×3 ones → centre only; by vertical line → middle row; dilate by 3×3 ones → full 5×5; by cross → plus with zero corners. Boundary β = A − (A⊖B); fill X_k = (X_{k−1}⊕B) ∩ Aᶜ; coins: binarise → erode → label → count.

Fourier domain & frequency filters

$$F(u,v)=\sum_x\sum_yf\,e^{-j2\pi(ux/M+vy/N)},\ f=\tfrac1{MN}\sum_u\sum_vF\,e^{+j2\pi(\cdot)};\quad f\star h\leftrightarrow FH\ (\text{pad }2M\times2N,\ (-1)^{x+y}\text{ centres})$$ $$H_{ILPF}=\mathbb 1[D\le D_0]\ (\text{rings}),\ H_{BLPF}=\frac1{1+(D/D_0)^{2n}},\ H_{GLPF}=e^{-D^2/2D_0^2}\ (\text{no ringing}),\ H_{HP}=1-H_{LP},\ H_{Lap}=-4\pi^2D^2,\ g=f-\nabla^2f$$

Unsharp g = f + k(f − f_LP). Homomorphic: ln f = ln i + ln r, H = (γ_H−γ_L)[1−e^{−cD²/D₀²}]+γ_L (0.25, 2, 1, 80). Notch pairs ±(u_k,v_k) kill stripes. Sampling: fs > 2·f_max else aliasing. Mask weights sum 1 → DC passes (low-pass); sum 0 → edge detector; [0 −1 0; −1 5 −1; 0 −1 0] emphasises edges.

Spatial filters & edges

$$g(x,y)=\sum_{s=-a}^{a}\sum_{t=-b}^{b}w(s,t)f(x+s,y+t);\quad \nabla^2f=f(x\!+\!1,y)+f(x\!-\!1,y)+f(x,y\!+\!1)+f(x,y\!-\!1)-4f;\quad g_x=(z_7\!+\!2z_8\!+\!z_9)-(z_1\!+\!2z_2\!+\!z_3)$$

Box 3×3 of 106,104,99,95,100,108,98,90,85 → 98.33; median = 5th of the 9 sorted values (slide: 0,0,1,1,1,2,2,2,4 → 1). Borders: discard (512→510), zero-pad (false edges), replicate. Second derivative: thin double edges, fine detail. Canny: Gaussian 5×5/159 (→ 41) → Sobel G_x −191, G_y −181, |G| 263, θ 134° → NMS (263 vs 7, 255 → kept) → hysteresis T_H 200 / T_L 50 (weak kept if 8-connected to strong).

Restoration

$$g=h\star f+\eta;\quad \hat f=g-\frac{\sigma_\eta^2}{\sigma_L^2}(g-m_L);\quad \hat F=\Big[\frac1H\frac{|H|^2}{|H|^2+K}\Big]G;\quad I^{t+1}=I^t+\lambda\!\sum_{N,S,E,W}\!c_d\nabla_dI,\ c=e^{-(|\nabla I|/K)^2}$$

Contraharmonic Q>0 pepper, Q<0 salt; median for salt-and-pepper. Local filter μ 75.14, σ² 6259.5, v² 400: 186→178.9, 95→93.7, 36→38.5. Diffusion 186 with N255 S157 E212 W208, K 30, λ 1/7: c = .005, .392, .472, .584 → 188.0 (edge kept).

Video: background, GMM/EM, tracking

$$\gamma(z_{nk})=\frac{\pi_k\mathcal N(x_n|\mu_k,\Sigma_k)}{\sum_j\pi_j\mathcal N(x_n|\mu_j,\Sigma_j)}\ (\text{E}),\quad \mu_k=\tfrac1{N_k}\sum_n\gamma_{nk}x_n,\ \pi_k=\tfrac{N_k}N\ (\text{M});\qquad P(X_t|y_{0:t})\propto P(y_t|X_t)\!\int\!P(X_t|X_{t-1})P(X_{t-1}|y_{0:t-1})$$

Frame differencing = boundaries + ghosts; temporal median or Stauffer–Grimson GMM background. E-step example x=1, N(0,1)/N(4,1): γ₁ = 0.982. Tracking = detection + prediction; Markov on state and observation. Kalman: linear-Gaussian, mean + covariance, one object. Particle filter: weighted samples, clutter, multimodal. Too strong dynamics → ignores data; too strong observation → repeated detection; drift.

Numbers to quote

Otsuk* = 2, σ²_B 2.629, η .84Equalisation0→2, 1→4, 4→5, 6→7, 8→9K-means(1.25,1.5), (3.9,5.1) after 3 it.Sobel/Canny−191, −181, 263; T 200/50Diffusion188.0Wiener-local178.9 / 93.7 / 38.5Laplacian H−4π²D² (−39.48 at D=1)Homomorphicγ_L .25, γ_H 2, c 1, D₀ 80PadP ≥ A+C−1, 2M×2N

Say this in the exam

"Otsu maximises between-class variance from the histogram alone." "Erosion keeps a pixel only where the structuring element fits; dilation wherever it hits; they are duals." "Filtering in space is multiplication in frequency, but DFT convolution is circular, so we pad." "Canny: smooth, gradient, thin by non-maximum suppression, link by hysteresis." "EM alternates responsibilities and parameter updates; a Kalman filter predicts then corrects."

Traps

Ideal low-pass rings (sinc kernel). Zero padding of borders creates false edges. Otsu fails on unimodal or unevenly lit images. Second derivative amplifies noise: smooth first. 8-connectivity merges diagonal blobs. Mask sums: 1 passes DC, 0 rejects it.

Viva cheat sheet · Computer Visionpage 4 of 7
5

Analysis and Computing foundations

MATLAB for signals and CT images · polynomials · linear algebra · Fourier and sampling
formulanumbers to quotetrap / critiquesay this in the exam

MATLAB essentials

Names: letter first, no specials, < 32 chars, case-sensitive. A(row,col) rows split by ; · A(1,:) row · A(1,1:2:5) cols 1,3,5 · .* ./ .^ element-wise, * matrix (inner dims agree) · zeros ones eye rand · for i=1:n … end, if … elseif … else … end, odd test fix(x/2)~=x/2 · index starts at 1 → a(n+1).

Course matrix A = [12 10 13 15 16; 22 45 65 1 0; 22 33 41 23 45; 21 30 12 6 2; 1 0 0 1 7]: sum(A) = [78 118 131 46 70], sum(sum(A)) = 443, diag = [12 45 41 6 7], trace 111. [3,5].*[4,8] = [12,40].

Rectangle rule & loops

$$\int_a^b f\,dt\approx\sum_k f(t_k)\,\Delta t\qquad \int_0^\pi\sin t\,dt:\ \Delta t=0.5\to1.98,\ \Delta t=0.1\to1.9995,\ \text{exact }2$$

[0 2 10 20 3 15] with +1 if > 10 else −1 → [−1 1 9 21 2 16]. Plot: t=0:0.01:2*pi (0:1:2π gives only 7 points).

Images in MATLAB

sprintf names CT_001…012→imread → uint8 → double→hist_im: find(Y==i), Hist(i+1)→volume histogram (sum of slices)→valley T = 95→Z = Y if Y<95 else 255→imwrite(Z./255)

uint8 0–255 (uint16 65,535); toy 4×4 histogram [5 4 3 4]. medfilt2(X,[3 3]) odd window → middle of 9: {0,20,0,127,112,100,128,135,0} → 100; {0,5,8,60,99,99,109,125,155} → 99; ECG 1-D [1 3]. 3-D: V(:,:,k), isosurface(…,15), daspect([1 1 .4]), plot3, trisurf.

Polynomials & fitting

Coefficients highest power first: x³+4x²+9x+16 → [1 4 9 16]. roots ↔ poly; conv = multiply ([1 2 3 4]⋆[1 4 9 16] = [1 6 20 50 75 84 64]); deconv; polyder → [3 8 9]; polyint → [0.25 1.333 4.5 16 0].

$$\text{polyfit: }\min_p\sum_i(y_i-P(x_i))^2\ \Rightarrow\ (V^TV)p=V^Ty\quad(\text{41 pts}\to p\approx[0.965,\,0.140,\,4.969]);\qquad\text{spline passes exactly through the data}$$

Linear algebra

$$x=A\backslash b\ (\text{LU, not inv});\quad A^{-1}=\frac1{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix};\quad A^{-1}=\frac{\mathrm{adj}A}{|A|},\ C_{ij}=(-1)^{i+j}M_{ij};\quad x_{LS}=(A^TA)^{-1}A^Tb;\quad Av=\lambda v$$

Solution exists iff rank(A) = rank([A b]) = r; unique if r = n (det ≠ 0); infinite if r < n (pinv, rref); A x = 0 non-trivial iff rank < n. Over-determined consistent → exact; inconsistent → least squares. Slide system A = [1 2 3; 4 5 6; 7 8 0], y = [366; 804; 351] → x = [25; 22; 99] (y/A is a bug). [3 −4; 6 −8] singular. Circuit A = [1 −1 1; −1 1 −1; 4 2 0; 0 2 5], b = [0;0;8;9] → i = [1, 2, 1] A. Chemical CO₂ + H₂O → O₂ + C₆H₁₂O₆: null space t·[6 6 6 1]. dot = projection, cross = moment.

Fourier series, transform, modulation, sampling

$$x(t)=a_0+\sum_n\big[a_n\cos(2\pi nf_0t)+b_n\sin(2\pi nf_0t)\big],\ a_0=\tfrac1T\!\int_T\!x,\ a_n=\tfrac2T\!\int_T\!x\cos,\ b_n=\tfrac2T\!\int_T\!x\sin;\quad \pm1\text{ square: }a_0=0,\ b_n=\tfrac{4}{\pi n}\ (n\text{ odd})$$ $$X(f)=\!\int\! x(t)e^{-j2\pi ft}dt;\ \ \mathrm{rect}_\tau\leftrightarrow\tau\,\mathrm{sinc}(f\tau)\ (\text{zeros }n/\tau);\ \ e^{-at}u(t)\leftrightarrow\tfrac1{a+j2\pi f};\ \ e^{-a|t|}\leftrightarrow\tfrac{2a}{a^2+(2\pi f)^2};\ \ x\cos(2\pi f_0t)\leftrightarrow\tfrac12[X(f\!-\!f_0)+X(f\!+\!f_0)]$$ $$x_s=x\cdot\textstyle\sum_k\delta(t-kT_s)\ \leftrightarrow\ f_s\sum_kX(f-kf_s):\ \text{copies at }kf_s\pm f_m\ \Rightarrow\ \boxed{f_s\ge2f_m}\ \text{(200 Hz → 400 Hz; 150 → 300)};\ \text{RC: }\tfrac{V_o}{V_{in}}=\tfrac1{1+sRC}$$

Series: periodic only ("main limitation"); transform: periodic and aperiodic ("main advantage"). syms t; fourier(exp(-t^2)) = √π e^{−w²/4}; fourier(exp(-abs(t))) = 2/(1+w²).

Numbers to quote

sum(sum(A))443 · trace 111∫sin, Δt .5 / .11.98 / 1.9995 (exact 2)Lung thresholdT = 95 (valley of 12-slice histogram)Medians100 and 99polyfit[0.965 0.140 4.969] ≈ x²+5A\y[25 22 99]Circuiti = [1 2 1] ASquare waveb₁ 1.273, b₃ 0.424, b₅ 0.255; even 0Nyquistfs ≥ 2fmvar([2 4 6])4 (N−1 denominator)

Say this in the exam

"Backslash is Gaussian elimination (LU); the inverse is only for theory." "A solution exists when b lies in the column space: rank(A) = rank([A b])." "Fitting minimises squared error and need not touch the points; interpolation must." "A square wave has only odd harmonics falling as 1/k, so its edges need infinite bandwidth." "Sampling replicates the spectrum every fs; keep the copies apart with fs ≥ 2fm, then low-pass to recover."

Traps

MATLAB indices start at 1 (Hist(i+1), a(n+1)). uint8 saturates at 255: convert to double. medfilt2 needs an odd window. y/A ≠ A\y. det ≈ 0 means numerically singular. Variable names are case-sensitive (items vs Items bug). Fourier series only for periodic signals.

Viva cheat sheet · Analysis and Computingpage 5 of 7
6

Pattern Recognition and Machine Learning

MAI540 · data preparation, regression, confusion matrix, ROC/PR, kNN, SVM/softmax, neural networks

Data preparation

$$\text{TF-IDF}=tf\log\tfrac{N}{df};\quad z=\tfrac{x-\mu}{\sigma}\ (|z|>3);\quad IQR=Q_3-Q_1,\ \text{fences }Q_1-1.5IQR,\ Q_3+1.5IQR;\quad x'=\tfrac{x-\min}{\max-\min};\quad x_{rob}=\tfrac{x-\mathrm{med}}{IQR};\quad IR=\tfrac{N_{maj}}{N_{min}}$$

Cyclical: sin/cos(2πx/T). Drop |r| > 0.95 (0.85 in exam). PCA: standardise → covariance → eigenvectors → top k (max variance, numeric only). Fit scalers on training data only. SMOTE {900,100}→{900,900}. Ten bias types; reject-option classification. Robust example (70−50)/20 = 1.0. IQR exam [3,4,5,5,6,6,7,8,9,28]: Q1 5, Q3 8, fences 0.5/12.5 → 28 outlier. Question C: missing = (7300+6200)/2 = 6750; Age (x−25)/20; Salary (x−5000)/2500; one-hot Dept/Region.

Regression & loss

$$A=\frac{\sum y\sum x^2-\sum x\sum xy}{n\sum x^2-(\sum x)^2},\quad B=\frac{n\sum xy-\sum x\sum y}{n\sum x^2-(\sum x)^2};\quad SE\to0\text{ (cancels)},\ SAE,\ SSE,\ MSE=\tfrac{SSE}{N},\ RMSE=\sqrt{MSE}\ \text{(use this)}$$

Errors −5, 0, 5 → SE 0, SAE 10, SSE 50, MSE 16.67, RMSE 4.08. Cats (n 15): Σx 550, Σy 39, Σxy 1882, Σx² 27352 → den 107780, B 0.0629, A 0.293, MSE 0.344, RMSE 0.59. Bookstore 6 pts: Σx 30, Σy 39, Σx² 160, Σxy 211 → A −1.5, B 1.6. Life expectancy: male 75 − 0.5·cig + veg (73, 81), female 80 − … (73, 84). R² = 1 − Σ(y−ŷ)²/Σ(y−ȳ)².

Fit quality & cross-validation

MPE (training error) and PV (train/test gap): good = low/low; overfit = low/high "dangerous, undetectable"; underfit = high/low (detectable). Test set 30% random (wrong way = last rows, test MSE 0.608); k-fold MSE = mean of folds (5-fold: .127 .985 .074 .517 .819 → 0.504); LOOCV n folds (bookstore errors .5, 2, 1, 0, 1.14 → mean 0.928, std 0.749). Test set cheap/high variance; LOOCV expensive/no waste; k-fold between. Stratified: 10% positives → 2 per fold of 20. Learning curve: gap = overfit, both low = underfit.

Confusion matrix (rows = human, cols = AI)

$$Acc=\tfrac{TP+TN}{all},\ P=\tfrac{TP}{TP+FP},\ R=TPR=\tfrac{TP}{TP+FN},\ Spec=\tfrac{TN}{TN+FP},\ FPR=\tfrac{FP}{FP+TN},\ F_1=\tfrac{2PR}{P+R},\ IoU=\tfrac{|A\cap B|}{|A\cup B|}>0.5$$

Precision when false alarms cost (spam, law, chemo); recall when misses cost (cancer, fraud, airport, lawsuits). TP1 FN0 FP2 TN7 → P 33%, R 100%, Spec 78%. TP80 FP20 FN10 → .80/.89/F1 .84. 150/165 = 90.9%. 3-class: P 85.7/78.4/77.8, R 75/80/84, Acc 80%. Models A (.826/.95 → spam) vs B (.766/.98 → cancer). PE2: Prec C 65/88 = .739, Rec B .55, Acc 340/463 = .734. Macro (equal classes) · micro (pooled) · weighted. OvR n, OvO n(n−1)/2 (4 → 4, 6).

ROC & PR

Sort scores desc, threshold at each (ties together), plot (FPR, TPR). (0,0) all −, (1,1) all +, (0,1) ideal, diagonal random. AUC 1 perfect, 0.5 random, 0.7 = 70% chance + ranks above −; exactly 1 → suspect a bug. Deck example .95+ .93+ .87− .85(−−+) .76− .53+ .43− .25+ → (0,.2)(0,.4)(.2,.4)(.6,.6)(.8,.6)(.8,.8)(1,.8)(1,1), AUC ≈ 0.56. PR: x recall, y precision, ignores TN → rare positives. PE2 Q19 (P 6, N 2): (0,1/6)(0,1/3)(.5,.5)(.5,2/3)(1,2/3)(1,5/6)(1,1).

kNN, linear classifier, losses

$$f=Wx+b\ (W\ 10\times3072);\quad L^{SVM}_i=\sum_{j\ne y}\max(0,s_j-s_y+1);\quad L^{soft}_i=-\log\frac{e^{s_y}}{\sum_je^{s_j}};\quad L=\tfrac1N\sum L_i+\lambda R(W);\quad W\leftarrow W-\alpha\nabla L$$

NN train O(1), predict O(N); L1 Σ|Δ|, L2 √Σ(Δ)²; k and distance on validation, test once; never on pixels (same L2 for shifted/tinted; 4, 16, 64 points). 4-pixel scores cat −96.8, dog 437.9, ship 61.95. Hinge: 2.9, 0, 12.9 → 5.27; init C−1; stops at margin. Softmax [3.2, 5.1, −1.7] → 24.5/164/0.18 → .13/.87/0 → 2.04; init ln C (2.3); never stops. Minibatch 32–256; Adam; step ×0.1 at 30/60/90; cosine ½α₀(1+cos(tπ/T)). 2-layer f = W₂max(0,W₁x), 3072→100→10 (307,200 + 1,000); without ReLU → linear. σ(0) = 0.5.

Numbers to quote

Memory10k×100 float32 = 4 MB; 1000 RGB 224² = 150 MBIris150 × 4, 3 classes; 105/45 split; kNN k=3 acc .955Poly([2,3],2)[1 2 3 4 6 9]Cats fitA .293, B .0629, RMSE .59; 5-fold .504CIFAR-1050k/10k, 32×32×3 = 3072Losseshinge 5.27 mean; softmax 2.04; init C−1 / ln CPE2 Q22a 2.386, b 1.923, RMSE .9093-fold workbook1.896, .252, .282 → .81

Say this in the exam

"Accuracy lies under imbalance; report precision and recall and say which error is expensive." "Fit the scaler on training data only, or the test set leaks." "Overfitting is low training error with a large gap; it is dangerous because you only see it after deployment." "Hinge stops once the margin is met; softmax keeps pushing." "Without a non-linearity, two layers collapse into one."

Traps

SE cancels. Never tune on the test set. The "wrong way" test split is unshuffled. Cosine similarity really ranges −1 to 1 (key says 0 to 1). Deck-03 regression slide prints MAE .39/MSE .29 but its table gives .51/.335. Lag-1 example slide says .35 (correct .25). Count neural-net layers by weight layers.

Viva cheat sheet · Pattern Recognitionpage 6 of 7
7

Probability and Random Processes

events, Bayes, CT pixel classification, random variables, distributions, counting, histogram fits, Parzen

Events & Bayes

$$\Pr(B\mid A)=\tfrac{\Pr(AB)}{\Pr(A)};\quad A\perp B\iff\Pr(AB)=\Pr(A)\Pr(B);\quad \Pr(H\mid E)=\frac{\Pr(E\mid H)\Pr(H)}{\sum_j\Pr(E\mid B_j)\Pr(B_j)}\ \text{(prior · likelihood → posterior)}$$

Axioms ≥ 0, Pr(S) = 1, disjoint add. Disjoint ≠ independent; not transitive. Cond. indep.: Pr(AB|C) = Pr(A|C)Pr(B|C) (1/25 example). Drug table: .5, .5, .45 → dependent; Pr(B|A) = .9. Simpson: 10% vs 5%, 95% vs 50%, overall 10.8% vs 45.9%. Rain .64 → .875. Coin HHH: .512/.125 → .3185 → .804. Disease 1%/99%/5%: Pr(+) .0594 → Pr(D|+) .1667. Machines: .029 → .517.

Bayes classifies CT pixels

$$\text{decide }k_1\iff p(q|k_1)P(k_1)>p(q|k_2)P(k_2);\quad E_I=\tfrac{P(k_1)}2\operatorname{erfc}\tfrac{T-\mu_1}{\sqrt2\sigma_1};\quad E_{II}=P(k_2)-\tfrac{P(k_2)}2\operatorname{erfc}\tfrac{T-\mu_2}{\sqrt2\sigma_2};\quad \operatorname{erfc}(x)=\tfrac2{\sqrt\pi}\!\int_x^\infty\! e^{-t^2}dt$$

Lung μ 89.4747, σ² 787.9435 (σ 28.07), 46870 px; chest μ 208.4971, σ² 794.1484 (σ 28.18), 86472 px; priors .3515/.6485. Peaks .0142; weighted .00499/.00918; cross ≈ 145. q = 140: 9.89e−4 vs 4.79e−4 → lung. T = 148: .006516 + .010315 = .016831; optimum T = 145, .016284. erfc(0) = 1, erfc(−x) = 2 − erfc(x), erfc(z) = 2[1 − Φ(z√2)]. MATLAB: p_lung(q+1), [min_error T_Opt] = min(Total_Error).

Random variables, E, Var

$$E[X]=\sum xP(x)=\int xf;\quad \operatorname{Var}=E[X^2]-(E[X])^2;\quad F(x)=\int_{-\infty}^x f;\quad \hat f=\tfrac{\text{count}}{nh}$$

Die 3.5, E[X²] 91/6, Var 35/12. Two coins (heads) E 1; two dice E 7, Var 35/6. PMF .2/.5/.3 → 1.1, 1.7, Var .49. f = 2x: F = x², P(.2–.6) = .32, E 2/3, Var 1/18. f = 3x²: 3/4, 3/5, 3/80. Exp: E = 1/λ by parts. Three dice |S| 216: P(5) 6/216, P(10) 27/216, P(3) 1/216, P(4) 3/216. Histogram h = 1, n 20: .15 .25 .30 .20 .10. Image [5 4 3 4] via find(Y==i), Hist(i+1).

Distributions

lawpdf / pmfmeanvarlecture numbers
Bernoulli pp^x(1−p)^{1−x}pp(1−p)p .3: var .21
Binomial n,pC(n,x)p^x(1−p)^{n−x}npnp(1−p)100,.4: 40, 24, P(40) .0812
Poisson λλ^x e^{−λ}/x!λλλ 4: P(2) .1465, P(4) .1954, P(≤3) .4335
Normal μ,σ²e^{−(x−μ)²/2σ²}/√(2πσ²)μσ²N(70,25): .9772, .6826, .0013; Φ(1) .8413, Φ(1.96) .975, Φ(2) .9772, Φ(3) .9987
Uniform a,b1/(b−a)(a+b)/2(b−a)²/12U(0,10): .3, .4, .3, 5, 8.33
Exponential λλe^{−λx}, F = 1−e^{−λx}1/λ1/λ²λ 2: .6321, .1353, .2326, .5, .25
Gamma k,λλ^k x^{k−1}e^{−λx}/Γ(k); Erlang F = 1−e^{−λx}Σ(λx)ⁿ/n!k/λk/λ²3,2: F(1) .3233, F(.5) .0802, diff .2431, 1.5, .75
Beta α,βx^{α−1}(1−x)^{β−1}/Bα/(α+β)αβ/((α+β)²(α+β+1))2,5: 2/7, .0255, mode .2

Counting & density estimation

$$n_1n_2\cdots n_k;\quad P^n_r=\tfrac{n!}{(n-r)!};\qquad p(q)=\frac1n\sum_i\frac1V\varphi\Big(\frac{q-q_i}{h}\Big),\ V=h^d;\qquad k=\tfrac{m^2}{v},\ \theta=\tfrac vm;\quad \alpha=\mu\big(\tfrac{\mu(1-\mu)}{\sigma^2}-1\big)$$

2 coins 4, 3 coins 8, 2 dice 36, 3 dice 216, M&Ms 12; P(4,3) 24; 5! 120. Histogram over 136 slices: histcounts, normalise, L2 = ‖p_h − p_m‖. Parametric L2: exp .0613 < beta .0706 < Gauss .0726 < gamma .0811 < uniform .0942. Parzen: window ≥ 0, area 1; h .1 noisy, 1 smooth; conv(hist, kernel); L2 .018–.021 (triweight .01818 best, Gaussian σ3 .02118). Square-window 15 pts: counts 3,3,4,2,3,3,3,1,2,2,2,2,1,1,1 → /15 → normalise 33/15 → count/33. Kernels: Epanechnikov ¾(1−u²) (optimal, AMISE), quartic 15/16(1−u²)², triweight 35/32(1−u²)³, cosine π/4·cos(πu/2), triangular 1−|u|.

Say this in the exam

"Bayes turns a likelihood into a posterior; the denominator is the same for every class, so I compare likelihood × prior." "The optimal threshold is where the weighted densities cross; the error is the two tails, written with erfc." "Variance is E[X²] minus the square of the mean." "Poisson is the binomial with n → ∞ and np fixed." "Parzen puts a little window on every sample; h trades noise for smoothness."

Traps

Disjoint events are never independent. A 99%-sensitive test on a 1% disease is right one time in six. MATLAB indexes from 1 (q+1). Lecture-3 slide writes (μ₂−T) in Type II; the code (and the correct form) uses (T−μ₂). The Parzen slide counts come from the sketch, not a literal |q−qᵢ| ≤ 0.5 count. Gamma "scale" θ = 1/λ.

Viva cheat sheet · Probabilitypage 7 of 7