T27.AI

Blog

One outlier, 23 zeros: a course module on the OCP MX block

2026-10-06 · 8 min read

[one block of 32, chosen to show the flush, no model and no accuracy number; ocp_mx.t27 is merged into t27 but not on the site yet; lesson 33 is a recording, not a browser player; software only] The t27 course has three new lessons on the Microscaling (MX) formats of the Open Compute Project, where 32 weights share one scale byte. One weight of 3.3 among 31 small ones sets that scale, and in MXFP4 23 of the 30 non-zero small ones become zero; MXFP6 E2M3 loses 4, MXINT8 2, MXFP6 E3M2 and MXFP8 E4M3 none. The numbers come from a spec with 49 tests, and changing one exponent bias from 15 to 16 fails exactly one of them.

Three-panel engraved illustration for: One outlier, 23 zeros: a course module on the OCP MX block
View the complete triptych at full size
#t27#NumberFormats#MutationTesting

The t27 course now has 33 lessons. Its new eleventh module, "AI numbers: the MX block", is three lessons on the Microscaling (MX) formats of the Open Compute Project: weights stored in 4, 6 or 8 bits, with one shared scale for every 32 of them. Lesson 31 explains the shared scale and compiles its spec in your browser. Lesson 32 runs that spec on a real machine. Lesson 33 shows what one large weight does to the other 31 weights in its block, 30 of them non-zero and one exactly zero: in MXFP4, 23 of the 30 non-zero ones become zero. Every number in this post is printed by a command in a recording or asserted by a test in a spec.

t27c on ocp_mx.t27 -- OCP MX v1.0, native

t27c 0.4.0 and Zig 0.16.0 on the t27c lab: the test that counts the flushed weights, 49 of 49 tests pass, then one exponent bias is changed from 15 to 16 and exactly one test fails; git restores the spec. 32.7 s, shown in full; the prompt and the typing are staged, every byte printed is real. Open the recording on its own page.

One byte of scale for 32 weights

An MX block is 32 elements and one shared scale. Each element is a small number: FP4 E2M1, FP6 E2M3 or E3M2, FP8 E4M3 or E5M2, or an 8-bit integer. The scale is one byte called E8M0: eight exponent bits, no sign, no mantissa. Code e means 2 to the power e minus 127, code 255 means NaN, and there is no zero. A block of MXFP4 is 32 x 4 bits plus 8, or 4.25 bits per weight.

Lesson 31 opens e8m0.t27, our spec of that byte, in the player. The t27 compiler runs as WebAssembly in the page, emits all seven backends and runs the spec's checks: 17 of 27 pass, none fails, and 10 are skipped, each with the reason the browser cannot run it. Lesson 32 is the same spec on a real machine, where nothing is skipped: 18 of 18 tests pass in Zig, and 9 invariants are proved while the spec compiles, so compiling is the check.

t27c on e8m0.t27 -- the MX shared scale byte, native

t27c 0.4.0 and Zig 0.16.0 on the t27c lab: the constants of e8m0.t27, the test report (18 pass, 9 invariants proved at compile time), the Zig tests, and the Verilog the compiler writes for the NaN check. 19.9 s, shown in full. Open the recording on its own page.

One outlier, format by format

Section 6.3 of the MX specification sets the scale from the largest magnitude in the block: its power of two, minus the largest exponent the element can show. One value decides the scale for all 32. Our worked block has 31 weights between -0.17 and 0.17, one of them exactly zero, and one outlier, 3.3. ocp_mx.t27 converts the block to each element type and counts the non-zero inputs that come back as zero. Every count, scale and error below is an assert in the spec; the errors are asserted exactly, in units of 2^-24, and shown here as decimals.

ElementBits per weightShared scaleFlushed to zero (of 31 non-zero inputs)Sum of absolute errors
MXFP4 E2M14.252^-1232.5610
MXFP6 E2M36.252^-140.5625
MXFP6 E3M26.252^-300.3218
MXINT88.252^120.2451
MXFP8 E4M38.252^-700.1064

In MXFP4 the outlier itself does not survive either: 3.3 over the scale 2^-1 is 6.6, past FP4's largest value of 6.0, so it is clamped and stored as 3.0. At six bits, range beats precision on this block. E3M2, with one more exponent bit and one less mantissa bit, flushes nothing and has the smaller error, because the outlier pushes the small weights down into E2M3's subnormals. The spec asserts that comparison too.

The codes the spec expects for every element of the block did not come from the spec. They were computed by a separate implementation of section 6.3 in exact rational arithmetic (Python fractions), and the spec checks them bit for bit for MXFP4, both MXFP6 types, MXFP8 E4M3 and MXINT8.

This is one block, built to show the effect. It says nothing about a whole model, where the share of blocks that hold such an outlier decides how much it costs. Lesson 29, earlier in the course, measures formats on a real language model; this module does not.

A spec that catches its own mistake

The recording ends by planting a bug: the exponent bias of E5M2 is changed from 15 to 16. Exactly one test fails, bias_puts_the_min_normal_at_two_to_one_minus_bias, and git restores the file. In review, 11 such mutants were tried, and each made at least one test fail; the pull request lists them. Review also found that the first version generated Zig without errors but did not compile it: 20 errors, from array indexes and signed division. "Generates" and "compiles" are different claims, and the spec now makes both. The compiler defect behind it is filed as t27#6867.

Where else to learn this

Before writing this, we read the pages of twelve other resources that teach or implement these formats, on 6 October 2026. The table puts them in ten rows. "Counted" means the page states how many small values one outlier turns into zero.

ResourceKindMX block and E8M0 scaleZeros from one outlierRuns whereTests behind it
A Visual Guide to Quantization, M. Grootendorst (2024)illustrated articleno: INT8 and INT4, no FP8, FP4 or MXshown in a picture, not countedstatic figures—
Introducing NVFP4, NVIDIA (2025)articlecompares MXFP4 with NVFP4 and the error of their scale bytesnot countedstatic—
MXFP4 in Transformers, Hugging Face (2025)article and docsMXFP4, blocks of 32, no worked blocknoneeds a GPUbenchmarks, no tests named
Quantization Fundamentals and Quantization in Depth, DeepLearning.AIvideo coursesnot named on the course pagesnocode examplesnot stated
TinyML and Efficient Deep Learning, MIT 6.5940 (Fall 2024)university coursenot named on the course page; slides not readnot on the course pageColab labsnot stated
OCP MX v1.0 and arXiv:2310.10537standard and paperyes, the normative definitiongives the rule, no worked blockPDFno test vectors published
microsoft/microxcalinglibraryyes: MXFP8, MXFP6, MXFP4, MXINT8 (and INT4, INT2), blocks of 32noPyTorch and CUDAyes, in Python
graphcore-research/gfloatlibraryyes: OCP MX element and block formatsnoPython, notebooksyes, cross-checked against torchao
Understanding MXFP4 Quantization, K. Sharma (2025)interactive pageMXFP4 only, with its own scale rulesays an outlier costs precision, not countedyour browser, your numbersnone cited
Floating Point Conversion Calculator, sw23interactive pageE8M0 and every MX element type, one value at a timeno blockyour browser, shareable linksyes: test vectors and CI
This module, t27 lessons 31 to 33course moduleyes: five element types, blocks of 32, E8M0counted: 23, 4, 0, 2 and 0 of 31 non-zero inputslesson 31 in your browser; lesson 33 a recordingyes: 49 tests in the spec

Each of these does something this module does not. A Visual Guide to Quantization covers the whole field, from post-training quantization to GGUF, with patient pictures. NVIDIA and Hugging Face show real models running at these precisions, with accuracy and memory figures; nothing here touches a model. The two DeepLearning.AI courses teach with video and guided coding, and MIT's course is a full semester. microxcaling runs MX inside real PyTorch models. sw23's calculator takes every MX element type apart bit by bit, with rounding and overflow modes and test vectors behind it, and Understanding MXFP4 Quantization lets you type your own numbers and watch the codes change.

gfloat is the fair comparison for "derived from the standard and tested": a readable Python reference over many formats, MX blocks included, cross-checked against a second library. What we did not find anywhere is the count this module is built on: how many of a block's small weights one outlier sends to zero, element type by element type. The standard gives the rule but no worked block, the articles describe the effect in words or as one error figure, and the libraries can encode such a block, but no page we read shows the count. That is a narrow claim, and it is the only one we make. Nor do we run the block's tests in your browser yet: lesson 33 is a recording.

One difference matters if you compare numbers across tools. Understanding MXFP4 Quantization picks the scale as the ceiling of log2(max / 6), and its full visualizer uses blocks of 16. Section 6.3 of the OCP specification takes the largest power of two not above the block's maximum, divided by the largest power of two the element can hold, and its blocks are 32. The specification allows other conversions, so neither is wrong. But for our block, with its outlier of 3.3, the first rule gives a scale of 2^0 and the second 2^-1, so the two tools give different codes for the same weights.

For people who build MX

ocp_mx.t27 is one file. It decodes every code of FP4 E2M1 (16), FP6 E2M3 and E3M2 (64 each), INT8, and OFP8 E4M3 and E5M2 (256 each), NaN and Inf included; encodes with round to nearest, ties to even, saturating; and runs the block conversion of section 6.3. Each source is cited next to the code it justifies. Its 49 tests assert exact integers: codes, scales, counts, and errors in units of 2^-24. The named code points and the five block-code tables can be run against a conversion in a kernel or in hardware. If one disagrees with your reading of the specification, we want that report, as an issue in gHashTag/t27.

Until this spec, the t27 format catalog behind arXiv:2606.09686 listed mxfp8, mxfp6 and mxfp4 and cited OCP MX v1.0, but no spec in t27 decoded an MX element or ran the block conversion (t27#6827).

What this does not show

Try it

What this does not settle

Receipts

Work with me

Want this kind of check on your own design?

I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.