Unit circle: each sample x is put at angle t·x
Centroid = φ̂(t). Target point = (e−t²/2, 0). The line between the two points is the error.
Samples
Centroid φ̂(t)
Target φ(t)
Error
Path of φ̂ for t ∈ [0, 3]
Sample histogram
The red curve is the N(0,1) density.
φ̂(t) as t changes
The dots are the 17 knots in the code (linspace(0, 3, 17)).
Target e−t²/2
Re φ̂ = mean cos(tx)
Im φ̂ = mean sin(tx)
- This plot shows the raw values, not multiplied by the weights wk. It shows how large the difference is at each t.
- The purple shading is only the difference between Re φ̂ and the target. It does not include the Im error, and it has no weights.
- Taylor expansion: φ(t) ≈ 1 + iμt − (σ² + μ²)·t²/2 + O(t³). A small t shows only the mean and the variance. Only a larger t shows the shape, for example skewness and kurtosis.
Contribution of each knot to the loss
N · wk · err(tk). The sum of all bars = SIGReg. The gray curve is the shape of the window e−t²/2.
- Vertical axis: the top = max(1, highest bar) × 1.1. So when the top is 1.1, every bar is less than 1.
- Gray curve: it only shows how the weight decreases as t increases. It is scaled to the panel height and does not use the vertical scale of the bars.
- Why the bars go up and then down: each bar = err(t) × weight. err(t) increases with t (at t = 0 all points are at (1, 0), so err is always 0). The weight wk = trapezoid coefficient (Δt at the two ends, 2Δt in the middle) × e−t²/2 decreases with t (at t = 3 it is only about 0.011). The product is high in the middle and low at the two sides.
- A perfect N(0,1) also has this shape: with only N samples, the centroid moves randomly. The expected noise is E[N·err(t)] = 1 − e−t², so a bar from pure noise ≈ 2Δt · e−t²/2 · (1 − e−t²). The maximum is at t = √ln3 ≈ 1.05, and this value does not depend on N.
- So: the position of the peak does not always mean that there is a problem near t ≈ 1. The bar heights depend on the samples. When you click "Resample", the heights change a lot, but the shape stays almost the same. For N(0,1), the sum of all bars (SIGReg) is about 1.0 on average.
- Compare the distributions: Collapsed, Too wide, and Shifted have large differences at small and medium t. Their peak is on the left, and their bars are very high. For Uniform, Bimodal, and Laplace, the differences appear only at larger t, so the peak moves to the right. But the window is small on the right, so these differences have much less weight in the loss.
- What the window does: SIGReg mainly penalizes low-order errors (mean and variance). It penalizes higher-order differences in shape less, because the estimate is noisier at large t.