INT8 quantization playground

Two experiments. Part 1 shows distinct float values collapsing onto shared INT8 codes. Part 2 shows how many scales you share across a weight matrix (per-tensor, per-row, per-column or block-wise) changes what an outlier costs you.

1. Where values collapse

Enter a range, generate 100 random points, and watch distinct float values collapse onto shared INT8 codes. Add clipping to trade outlier accuracy for finer resolution.

Where individual values collapse

Top: the original values, each one distinct. Bottom: each INT8 code is a column, and every point that lands in it stacks up. A tall stack is many different floats reduced to one identical number.
Scroll over the plot to zoom at the cursor, drag to pan.

Biggest collisions

Symmetric: scale = max(|min|, |max|) / 127, zero-point 0, codes -127..127. Asymmetric: scale = (max - min) / 255, zero-point shifts the range onto -128..127. Dequantized value = (code - zero-point) * scale. With no clipping the quantization range is the range you enter, not the min/max of the random points. Clipping shrinks that range: the step size gets smaller (finer resolution for most points), but any point outside the clip range saturates to the nearest edge code and takes a large error. Percentile clipping derives the range from the points themselves (in symmetric mode from their absolute values). The tight cluster + one outlier option mimics a real tensor where one extreme value stretches the range and squeezes everything else, which is the case clipping is meant to fix.

2. Granularity lab: how many weights share one scale?

A 16 × 32 weight matrix of typical values (standard deviation 1), quantized to INT8. Per-tensor uses one scale for everything. Per-row and per-column use one per row or column. Block-wise uses one per group of consecutive weights along each row. More scales shrink the blast radius of an outlier, but each scale is extra data to store.

Original weights

Teal is negative, coral is positive, and the colour saturates at 3 standard deviations. Amber outlines mark the injected outliers. Hover any cell for its exact values.

Absolute error after the round trip

All four side by side

Click a row to select that granularity.

Bits per weight assumes one 16-bit scale per group, plus an 8-bit zero-point per group in asymmetric mode, on top of the 8 bits for each weight. Block-wise here means groups of consecutive weights within a row, so a block of 32 equals per-row in this 32-column matrix. No clipping or calibration is applied: each group's range comes straight from its own min and max, and rounding is to nearest. Which axis is a "row" depends on layout. PyTorch's Linear layer stores weights as [out_features, in_features], so per-row there is per-output-channel. The data is random, not a real model's weights: it shows the mechanism, not a benchmark.