Two experiments. Part 1 shows distinct float values collapsing onto shared INT8 codes. Part 2 shows how many scales you share across a weight matrix (per-tensor, per-row, per-column or block-wise) changes what an outlier costs you.
Enter a range, generate 100 random points, and watch distinct float values collapse onto shared INT8 codes. Add clipping to trade outlier accuracy for finer resolution.
Symmetric: scale = max(|min|, |max|) / 127, zero-point 0, codes -127..127. Asymmetric: scale = (max - min) / 255, zero-point shifts the range onto -128..127. Dequantized value = (code - zero-point) * scale. With no clipping the quantization range is the range you enter, not the min/max of the random points. Clipping shrinks that range: the step size gets smaller (finer resolution for most points), but any point outside the clip range saturates to the nearest edge code and takes a large error. Percentile clipping derives the range from the points themselves (in symmetric mode from their absolute values). The tight cluster + one outlier option mimics a real tensor where one extreme value stretches the range and squeezes everything else, which is the case clipping is meant to fix.
A 16 × 32 weight matrix of typical values (standard deviation 1), quantized to INT8. Per-tensor uses one scale for everything. Per-row and per-column use one per row or column. Block-wise uses one per group of consecutive weights along each row. More scales shrink the blast radius of an outlier, but each scale is extra data to store.
Bits per weight assumes one 16-bit scale per group, plus an 8-bit zero-point per group in asymmetric mode, on top of the 8 bits for each weight. Block-wise here means groups of consecutive weights within a row, so a block of 32 equals per-row in this 32-column matrix. No clipping or calibration is applied: each group's range comes straight from its own min and max, and rounding is to nearest. Which axis is a "row" depends on layout. PyTorch's Linear layer stores weights as [out_features, in_features], so per-row there is per-output-channel. The data is random, not a real model's weights: it shows the mechanism, not a benchmark.