Init loss $=\log C$ ($\ln10=2.30$). Example (3.2, 5.1, −1.7): $p=(0.13,0.87,0.00)$, $L=2.04$.
Cat example $=2.9$; init $=C-1$. Zero once margins hold (stops learning); softmax never stops. $2W$ keeps $L=0$ → need $\lambda R(W)$.
$\beta_1$ 0.9, $\beta_2$ 0.999, start lr $10^{-3}$ or $5\cdot10^{-4}$. Numerical gradient check: $(1.25322-1.25347)/10^{-4}=-2.5$.
downstream = local × upstream. add distributes, mul swaps, max routes. $q=x+y,\ f=qz$: $\partial f/\partial x=z=-4$.
$$y=xW:\quad \frac{\partial L}{\partial x}=\frac{\partial L}{\partial y}W^T,\qquad \frac{\partial L}{\partial W}=x^T\frac{\partial L}{\partial y}$$Sigmoid: $\sigma'=\sigma(1-\sigma)=0.73\cdot0.27=0.20$; $dw=[-0.2,-0.39,0.2]$, $dx=[0.39,-0.59]$.
3×32×32, 10×(5×5), S1 P2 → 10×32×32, 760 params, 768,000 MACs. "Same" $P=(K-1)/2$. Receptive field $1+L(K-1)$. VGG: three 3×3 = one 7×7, $27C^2$ vs $49C^2$. ResNet $H(x)=F(x)+x$.
TP: class ok and IoU ≥ τ. Duplicate = FP. No TN. Course table: Cat 0.6667, Dog 0.5, Bicycle 0.3333 → mAP 0.5000. Bicycle IoU 0.74 fails τ 0.75.
$W$: $4h\times(h+d)$. $\partial c_t/\partial c_{t-1}=\mathrm{diag}(f)$ → "uninterrupted flow, like ResNet". Clip for exploding.
$\sqrt D$: $\mathrm{Var}(q\cdot k)=D$ → avoid softmax saturation. Permutation-equivariant → positional encoding. Mask future with $-\infty$. Block = MHSA → +res → LN → MLP(D→4D→D) → +res → LN; 6 matmuls; $O(N^2)$. ViT: $N=HW/P^2$ patches (224/16 → 196 tokens of 768), CLS token, low inductive bias → needs big data.
"Training minimises a loss by gradient descent; backprop is just the chain rule on the computational graph. CNNs share small filters across positions; LSTMs keep a cell state whose gradient passes through an element-wise gate; attention lets every token read every other token in one step at $O(N^2)$ cost." Overfitting = low train / high test error; underfitting = both high and "cannot be fixed by more epochs".
$e=SP-PV$. P only → steady-state error (0 output at $e=0$); I removes it; D damps, amplifies noise. Tune $K_p$, then $K_d$, then $K_i$ (0.0001). $K_p=125/3500=0.0357$. Example (2.0,0.5,0.1), $\Delta t$ 0.1, $e=-0.48$, $e_{prev}=-0.30$, $\sum=-0.12$: $-0.96-0.06-0.18=\mathbf{-1.20}$.
Normal form: vertical lines finite. (3,3),(4,3),(5,3) → $A(3,90°)=3$. RANSAC $P$ 0.99, $p$ 0.5: $k$=2→17, 3→35, 4→72. Reject lane parabola $|a|\ge0.003$.
(700,400), $Z$ 10, $f$ 800, $c$ (640,360) → (0.75, 0.5, 10). Stereo $d=80$, $f_x$ 795, $B$ 0.2 → $Z=1.9875$ m. One image: 2 eq, 3 unknowns. Quality = reprojection error.
Output $S\times S\times(5B+C)$: 7×7×30. Loss: $\sqrt w,\sqrt h$; $\lambda_{noobj}=0.5$; NMS IoU > 0.5. MIO: in lane if $x_L(y)\le x\le x_R(y)$, $x(y)=(y-b)/m$; MIO = argmax $y_{bottom}$. Three-car example: Car 3 (x 500) out of lane → Car 2 (290 > 260). FCW: tracks (confirm [2 3], delete 5), closest in lane; $d=1.2v+v^2/(2\cdot0.4\cdot9.8)$ → 24.8 m at 10 m/s.
Derivation: $N(\mu_p,p)\times N(z,r)$ → $\frac1{\sigma^2}=\frac1p+\frac1r$, $\mu=\frac{r\mu_p+pz}{p+r}$; set $K=\frac{p}{p+r}$. $K\to1$ trust sensor, $K\to0$ trust prediction; $0 Fusion: $z_f=\frac{\sum z_i/r_i}{\sum 1/r_i},\ r_f=\frac1{\sum1/r_i}$; (0.9,1.1),(1,4) → 0.94, 0.8; prior 10.1 → $K=0.927$, $x=0.87$, $p=0.74$. $r_i$ never changes during filtering.
"Bang-bang chooses a direction; PID chooses how much." "Hough votes in $(\rho,\theta)$ because slope is infinite for vertical lines." "RANSAC keeps the model most points agree with; least squares is pulled by outliers." "The Kalman filter is a recursive Bayesian estimator: predict widens, update shrinks." "MIO comes from confirmed tracks because detections flicker."
Pull-up button pressed = LOW. Never delay(), use millis(). Pooling has no parameters; PID I-term needs anti-windup in practice. Behaviour cloning is regression (ELU + regression layer), not softmax. Gazebo = world, RViz = belief.
Error ≤ $s/2$, variance $s^2/12$. Weights per-channel, activations per-tensor, bias int32 with $s_b=s_as_x$. PTQ (observe min/max) vs QAT (fake-quant nodes). MCU: int32 accumulator, shift >> 7.
w = !(~X*X)*~X*y (~ transpose, ! Gauss–Jordan inverse, pivot < 1e-6 → abort). (1,2),(2,2.5),(3,3.5): $m=4.5/6=0.75$, $c=1.17$. Quadratic (1,2),(2,3),(3,5): $\beta=[2,-0.5,0.5]$. $\begin{bmatrix}2&1\\5&3\end{bmatrix}^{-1}=\begin{bmatrix}3&-1\\-5&2\end{bmatrix}$. "Linear in parameters, nonlinear in features."
1000 mAh, 50 mAh every 2 h → 1.67 d. 1200 mAh, 40 mAh, 5 d → every 4 h. Uno 500 mAh: awake 50 mA → 10 h; asleep 0.1 mA → 30 days; inference 45 mA·0.8 s = 0.01 mAh. $\lambda=1/12$: $S(2)$ 0.846, $S(8.3)$ 0.5, $S(24)$ 0.135; run if $S<0.3$ (t > 14.4 h). Adaptive interval $60TE/B$: 180 s @5000, 900 s @1000. Morning 7 min + night 22.5 min → 67 events → 2010 mAh.
Raspberry Pi: Interpreter → allocate_tensors → set_tensor → invoke → get_tensor; detection input [1,320,320,3] uint8, outputs boxes/classes/scores, thr 0.3, EfficientDet-Lite0. Arduino only if model < 20 KB. Hierarchy: Arduino wakes Pi at 10–30 cm.
| UART | I²C | SPI | |
|---|---|---|---|
| wires | 3 (Rx,Tx,GND) | 2 (SDA,SCL) | 4 + n (SCK,MOSI,MISO,SS) |
| clock | none (baud) | master | master |
| duplex | full | half | full |
| speed | 9600/115200 | 100k/400k/3.4M | ~10 MHz |
| notes | RS-232 ±3–25 V, MAX232 | 7-bit → 128 addr; START SDA↓ while SCL high; ACK SDA low; 4.7 kΩ pull-ups | SS low selects; byte per 8 clocks |
Ride OFF→ON 50%→OFF after 20 s. Fan Idle→30%→60%→100%. Traffic Green 120 → Yellow 30 → 3 blinks (6 states) → Red 120; pedestrian PB only in Green with > 30 s left. Python: Enum + loop; (value+1) % 4.
"Quantisation stores each weight as an integer plus a shared scale and zero-point; integer inference needs only an int32 accumulator and a rescale." "Idle current dominates the battery: sleep and wake on interrupts." "$S(t)$ is a survival probability, not a density." "Least squares has a closed form, $(X^TX)^{-1}X^Ty$, so a microcontroller can learn without gradient descent."
f = i·r (illumination × reflectance); bits = M·N·k (1024² × 8 = 8,388,608). N4 / ND / N8; m-connectivity removes multiple paths. D4 = |Δx|+|Δy|, D8 = max, De = √: (1,2)→(4,6) = 7, 4, 5. Two-pass labelling with equivalence table. Point ops: negative (L−1)−r, log c·log(1+r), power s = c·rᵞ (γ<1 brightens), slicing.
$$p(g)=\frac{h(g)}{RC},\quad s_k=\mathrm{round}\big((L-1)\textstyle\sum_{j\le k}p_j\big):\ h=[5,4,0,0,2,1,3,0,4,1],\ L=10\ \to\ s=2,4,4,4,5,5,7,7,9,9\ (\text{"almost, not completely, flat"})$$Freq 8,7,2,6,9,4 (N 36): m_G 2.3611; σ²_B = 1.593, 2.564, 2.629, 2.142, 0.870 → k* = 2, η ≈ 0.84. Limitation: histogram only, no spatial info. Iterative T = ½(m₁+m₂); Niblack T = m + k·s. Split/merge (std < 5): 4 quadrants pass, R1∪R3 and R2∪R4 merge → 2 regions. K-means 7 points: C₁(1,1), C₂(5,7) → it.1 {1,2,3}/{4–7} → it.2 point 3 moves → (1.25,1.5), (3.9,5.1) → it.3 stable.
Opening = erode→dilate (salt, thin joints); closing = dilate→erode (pepper, holes). 3×3 block: erode by 3×3 ones → centre only; by vertical line → middle row; dilate by 3×3 ones → full 5×5; by cross → plus with zero corners. Boundary β = A − (A⊖B); fill X_k = (X_{k−1}⊕B) ∩ Aᶜ; coins: binarise → erode → label → count.
Unsharp g = f + k(f − f_LP). Homomorphic: ln f = ln i + ln r, H = (γ_H−γ_L)[1−e^{−cD²/D₀²}]+γ_L (0.25, 2, 1, 80). Notch pairs ±(u_k,v_k) kill stripes. Sampling: fs > 2·f_max else aliasing. Mask weights sum 1 → DC passes (low-pass); sum 0 → edge detector; [0 −1 0; −1 5 −1; 0 −1 0] emphasises edges.
Box 3×3 of 106,104,99,95,100,108,98,90,85 → 98.33; median = 5th of the 9 sorted values (slide: 0,0,1,1,1,2,2,2,4 → 1). Borders: discard (512→510), zero-pad (false edges), replicate. Second derivative: thin double edges, fine detail. Canny: Gaussian 5×5/159 (→ 41) → Sobel G_x −191, G_y −181, |G| 263, θ 134° → NMS (263 vs 7, 255 → kept) → hysteresis T_H 200 / T_L 50 (weak kept if 8-connected to strong).
Contraharmonic Q>0 pepper, Q<0 salt; median for salt-and-pepper. Local filter μ 75.14, σ² 6259.5, v² 400: 186→178.9, 95→93.7, 36→38.5. Diffusion 186 with N255 S157 E212 W208, K 30, λ 1/7: c = .005, .392, .472, .584 → 188.0 (edge kept).
Frame differencing = boundaries + ghosts; temporal median or Stauffer–Grimson GMM background. E-step example x=1, N(0,1)/N(4,1): γ₁ = 0.982. Tracking = detection + prediction; Markov on state and observation. Kalman: linear-Gaussian, mean + covariance, one object. Particle filter: weighted samples, clutter, multimodal. Too strong dynamics → ignores data; too strong observation → repeated detection; drift.
"Otsu maximises between-class variance from the histogram alone." "Erosion keeps a pixel only where the structuring element fits; dilation wherever it hits; they are duals." "Filtering in space is multiplication in frequency, but DFT convolution is circular, so we pad." "Canny: smooth, gradient, thin by non-maximum suppression, link by hysteresis." "EM alternates responsibilities and parameter updates; a Kalman filter predicts then corrects."
Ideal low-pass rings (sinc kernel). Zero padding of borders creates false edges. Otsu fails on unimodal or unevenly lit images. Second derivative amplifies noise: smooth first. 8-connectivity merges diagonal blobs. Mask sums: 1 passes DC, 0 rejects it.
Names: letter first, no specials, < 32 chars, case-sensitive. A(row,col) rows split by ; · A(1,:) row · A(1,1:2:5) cols 1,3,5 · .* ./ .^ element-wise, * matrix (inner dims agree) · zeros ones eye rand · for i=1:n … end, if … elseif … else … end, odd test fix(x/2)~=x/2 · index starts at 1 → a(n+1).
Course matrix A = [12 10 13 15 16; 22 45 65 1 0; 22 33 41 23 45; 21 30 12 6 2; 1 0 0 1 7]: sum(A) = [78 118 131 46 70], sum(sum(A)) = 443, diag = [12 45 41 6 7], trace 111. [3,5].*[4,8] = [12,40].
[0 2 10 20 3 15] with +1 if > 10 else −1 → [−1 1 9 21 2 16]. Plot: t=0:0.01:2*pi (0:1:2π gives only 7 points).
uint8 0–255 (uint16 65,535); toy 4×4 histogram [5 4 3 4]. medfilt2(X,[3 3]) odd window → middle of 9: {0,20,0,127,112,100,128,135,0} → 100; {0,5,8,60,99,99,109,125,155} → 99; ECG 1-D [1 3]. 3-D: V(:,:,k), isosurface(…,15), daspect([1 1 .4]), plot3, trisurf.
Coefficients highest power first: x³+4x²+9x+16 → [1 4 9 16]. roots ↔ poly; conv = multiply ([1 2 3 4]⋆[1 4 9 16] = [1 6 20 50 75 84 64]); deconv; polyder → [3 8 9]; polyint → [0.25 1.333 4.5 16 0].
$$\text{polyfit: }\min_p\sum_i(y_i-P(x_i))^2\ \Rightarrow\ (V^TV)p=V^Ty\quad(\text{41 pts}\to p\approx[0.965,\,0.140,\,4.969]);\qquad\text{spline passes exactly through the data}$$Solution exists iff rank(A) = rank([A b]) = r; unique if r = n (det ≠ 0); infinite if r < n (pinv, rref); A x = 0 non-trivial iff rank < n. Over-determined consistent → exact; inconsistent → least squares. Slide system A = [1 2 3; 4 5 6; 7 8 0], y = [366; 804; 351] → x = [25; 22; 99] (y/A is a bug). [3 −4; 6 −8] singular. Circuit A = [1 −1 1; −1 1 −1; 4 2 0; 0 2 5], b = [0;0;8;9] → i = [1, 2, 1] A. Chemical CO₂ + H₂O → O₂ + C₆H₁₂O₆: null space t·[6 6 6 1]. dot = projection, cross = moment.
Series: periodic only ("main limitation"); transform: periodic and aperiodic ("main advantage"). syms t; fourier(exp(-t^2)) = √π e^{−w²/4}; fourier(exp(-abs(t))) = 2/(1+w²).
"Backslash is Gaussian elimination (LU); the inverse is only for theory." "A solution exists when b lies in the column space: rank(A) = rank([A b])." "Fitting minimises squared error and need not touch the points; interpolation must." "A square wave has only odd harmonics falling as 1/k, so its edges need infinite bandwidth." "Sampling replicates the spectrum every fs; keep the copies apart with fs ≥ 2fm, then low-pass to recover."
MATLAB indices start at 1 (Hist(i+1), a(n+1)). uint8 saturates at 255: convert to double. medfilt2 needs an odd window. y/A ≠ A\y. det ≈ 0 means numerically singular. Variable names are case-sensitive (items vs Items bug). Fourier series only for periodic signals.
Cyclical: sin/cos(2πx/T). Drop |r| > 0.95 (0.85 in exam). PCA: standardise → covariance → eigenvectors → top k (max variance, numeric only). Fit scalers on training data only. SMOTE {900,100}→{900,900}. Ten bias types; reject-option classification. Robust example (70−50)/20 = 1.0. IQR exam [3,4,5,5,6,6,7,8,9,28]: Q1 5, Q3 8, fences 0.5/12.5 → 28 outlier. Question C: missing = (7300+6200)/2 = 6750; Age (x−25)/20; Salary (x−5000)/2500; one-hot Dept/Region.
Errors −5, 0, 5 → SE 0, SAE 10, SSE 50, MSE 16.67, RMSE 4.08. Cats (n 15): Σx 550, Σy 39, Σxy 1882, Σx² 27352 → den 107780, B 0.0629, A 0.293, MSE 0.344, RMSE 0.59. Bookstore 6 pts: Σx 30, Σy 39, Σx² 160, Σxy 211 → A −1.5, B 1.6. Life expectancy: male 75 − 0.5·cig + veg (73, 81), female 80 − … (73, 84). R² = 1 − Σ(y−ŷ)²/Σ(y−ȳ)².
MPE (training error) and PV (train/test gap): good = low/low; overfit = low/high "dangerous, undetectable"; underfit = high/low (detectable). Test set 30% random (wrong way = last rows, test MSE 0.608); k-fold MSE = mean of folds (5-fold: .127 .985 .074 .517 .819 → 0.504); LOOCV n folds (bookstore errors .5, 2, 1, 0, 1.14 → mean 0.928, std 0.749). Test set cheap/high variance; LOOCV expensive/no waste; k-fold between. Stratified: 10% positives → 2 per fold of 20. Learning curve: gap = overfit, both low = underfit.
Precision when false alarms cost (spam, law, chemo); recall when misses cost (cancer, fraud, airport, lawsuits). TP1 FN0 FP2 TN7 → P 33%, R 100%, Spec 78%. TP80 FP20 FN10 → .80/.89/F1 .84. 150/165 = 90.9%. 3-class: P 85.7/78.4/77.8, R 75/80/84, Acc 80%. Models A (.826/.95 → spam) vs B (.766/.98 → cancer). PE2: Prec C 65/88 = .739, Rec B .55, Acc 340/463 = .734. Macro (equal classes) · micro (pooled) · weighted. OvR n, OvO n(n−1)/2 (4 → 4, 6).
Sort scores desc, threshold at each (ties together), plot (FPR, TPR). (0,0) all −, (1,1) all +, (0,1) ideal, diagonal random. AUC 1 perfect, 0.5 random, 0.7 = 70% chance + ranks above −; exactly 1 → suspect a bug. Deck example .95+ .93+ .87− .85(−−+) .76− .53+ .43− .25+ → (0,.2)(0,.4)(.2,.4)(.6,.6)(.8,.6)(.8,.8)(1,.8)(1,1), AUC ≈ 0.56. PR: x recall, y precision, ignores TN → rare positives. PE2 Q19 (P 6, N 2): (0,1/6)(0,1/3)(.5,.5)(.5,2/3)(1,2/3)(1,5/6)(1,1).
NN train O(1), predict O(N); L1 Σ|Δ|, L2 √Σ(Δ)²; k and distance on validation, test once; never on pixels (same L2 for shifted/tinted; 4, 16, 64 points). 4-pixel scores cat −96.8, dog 437.9, ship 61.95. Hinge: 2.9, 0, 12.9 → 5.27; init C−1; stops at margin. Softmax [3.2, 5.1, −1.7] → 24.5/164/0.18 → .13/.87/0 → 2.04; init ln C (2.3); never stops. Minibatch 32–256; Adam; step ×0.1 at 30/60/90; cosine ½α₀(1+cos(tπ/T)). 2-layer f = W₂max(0,W₁x), 3072→100→10 (307,200 + 1,000); without ReLU → linear. σ(0) = 0.5.
"Accuracy lies under imbalance; report precision and recall and say which error is expensive." "Fit the scaler on training data only, or the test set leaks." "Overfitting is low training error with a large gap; it is dangerous because you only see it after deployment." "Hinge stops once the margin is met; softmax keeps pushing." "Without a non-linearity, two layers collapse into one."
SE cancels. Never tune on the test set. The "wrong way" test split is unshuffled. Cosine similarity really ranges −1 to 1 (key says 0 to 1). Deck-03 regression slide prints MAE .39/MSE .29 but its table gives .51/.335. Lag-1 example slide says .35 (correct .25). Count neural-net layers by weight layers.
Axioms ≥ 0, Pr(S) = 1, disjoint add. Disjoint ≠ independent; not transitive. Cond. indep.: Pr(AB|C) = Pr(A|C)Pr(B|C) (1/25 example). Drug table: .5, .5, .45 → dependent; Pr(B|A) = .9. Simpson: 10% vs 5%, 95% vs 50%, overall 10.8% vs 45.9%. Rain .64 → .875. Coin HHH: .512/.125 → .3185 → .804. Disease 1%/99%/5%: Pr(+) .0594 → Pr(D|+) .1667. Machines: .029 → .517.
Lung μ 89.4747, σ² 787.9435 (σ 28.07), 46870 px; chest μ 208.4971, σ² 794.1484 (σ 28.18), 86472 px; priors .3515/.6485. Peaks .0142; weighted .00499/.00918; cross ≈ 145. q = 140: 9.89e−4 vs 4.79e−4 → lung. T = 148: .006516 + .010315 = .016831; optimum T = 145, .016284. erfc(0) = 1, erfc(−x) = 2 − erfc(x), erfc(z) = 2[1 − Φ(z√2)]. MATLAB: p_lung(q+1), [min_error T_Opt] = min(Total_Error).
Die 3.5, E[X²] 91/6, Var 35/12. Two coins (heads) E 1; two dice E 7, Var 35/6. PMF .2/.5/.3 → 1.1, 1.7, Var .49. f = 2x: F = x², P(.2–.6) = .32, E 2/3, Var 1/18. f = 3x²: 3/4, 3/5, 3/80. Exp: E = 1/λ by parts. Three dice |S| 216: P(5) 6/216, P(10) 27/216, P(3) 1/216, P(4) 3/216. Histogram h = 1, n 20: .15 .25 .30 .20 .10. Image [5 4 3 4] via find(Y==i), Hist(i+1).
| law | pdf / pmf | mean | var | lecture numbers |
|---|---|---|---|---|
| Bernoulli p | p^x(1−p)^{1−x} | p | p(1−p) | p .3: var .21 |
| Binomial n,p | C(n,x)p^x(1−p)^{n−x} | np | np(1−p) | 100,.4: 40, 24, P(40) .0812 |
| Poisson λ | λ^x e^{−λ}/x! | λ | λ | λ 4: P(2) .1465, P(4) .1954, P(≤3) .4335 |
| Normal μ,σ² | e^{−(x−μ)²/2σ²}/√(2πσ²) | μ | σ² | N(70,25): .9772, .6826, .0013; Φ(1) .8413, Φ(1.96) .975, Φ(2) .9772, Φ(3) .9987 |
| Uniform a,b | 1/(b−a) | (a+b)/2 | (b−a)²/12 | U(0,10): .3, .4, .3, 5, 8.33 |
| Exponential λ | λe^{−λx}, F = 1−e^{−λx} | 1/λ | 1/λ² | λ 2: .6321, .1353, .2326, .5, .25 |
| Gamma k,λ | λ^k x^{k−1}e^{−λx}/Γ(k); Erlang F = 1−e^{−λx}Σ(λx)ⁿ/n! | k/λ | k/λ² | 3,2: F(1) .3233, F(.5) .0802, diff .2431, 1.5, .75 |
| Beta α,β | x^{α−1}(1−x)^{β−1}/B | α/(α+β) | αβ/((α+β)²(α+β+1)) | 2,5: 2/7, .0255, mode .2 |
2 coins 4, 3 coins 8, 2 dice 36, 3 dice 216, M&Ms 12; P(4,3) 24; 5! 120. Histogram over 136 slices: histcounts, normalise, L2 = ‖p_h − p_m‖. Parametric L2: exp .0613 < beta .0706 < Gauss .0726 < gamma .0811 < uniform .0942. Parzen: window ≥ 0, area 1; h .1 noisy, 1 smooth; conv(hist, kernel); L2 .018–.021 (triweight .01818 best, Gaussian σ3 .02118). Square-window 15 pts: counts 3,3,4,2,3,3,3,1,2,2,2,2,1,1,1 → /15 → normalise 33/15 → count/33. Kernels: Epanechnikov ¾(1−u²) (optimal, AMISE), quartic 15/16(1−u²)², triweight 35/32(1−u²)³, cosine π/4·cos(πu/2), triangular 1−|u|.
"Bayes turns a likelihood into a posterior; the denominator is the same for every class, so I compare likelihood × prior." "The optimal threshold is where the weighted densities cross; the error is the two tails, written with erfc." "Variance is E[X²] minus the square of the mean." "Poisson is the binomial with n → ∞ and np fixed." "Parzen puts a little window on every sample; h trades noise for smoothness."
Disjoint events are never independent. A 99%-sensitive test on a 1% disease is right one time in six. MATLAB indexes from 1 (q+1). Lecture-3 slide writes (μ₂−T) in Type II; the code (and the correct form) uses (T−μ₂). The Parzen slide counts come from the sketch, not a literal |q−qᵢ| ≤ 0.5 count. Gamma "scale" θ = 1/λ.