Blog
[one block of 32, chosen to show the flush, no model and no accuracy number; ocp_mx.t27 is merged into t27 but not on the site yet; lesson 33 is a recording, not a browser player; software only] The t27 course has three new lessons on the Microscaling (MX) formats of the Open Compute Project, where 32 weights share one scale byte. One weight of 3.3 among 31 small ones sets that scale, and in MXFP4 23 of the 30 non-zero small ones become zero; MXFP6 E2M3 loses 4, MXINT8 2, MXFP6 E3M2 and MXFP8 E4M3 none. The numbers come from a spec with 49 tests, and changing one exponent bias from 15 to 16 fails exactly one of them.

The t27 course now has 33 lessons. Its new eleventh module, "AI numbers: the MX block", is three lessons on the Microscaling (MX) formats of the Open Compute Project: weights stored in 4, 6 or 8 bits, with one shared scale for every 32 of them. Lesson 31 explains the shared scale and compiles its spec in your browser. Lesson 32 runs that spec on a real machine. Lesson 33 shows what one large weight does to the other 31 weights in its block, 30 of them non-zero and one exactly zero: in MXFP4, 23 of the 30 non-zero ones become zero. Every number in this post is printed by a command in a recording or asserted by a test in a spec.
t27c on ocp_mx.t27 -- OCP MX v1.0, native
An MX block is 32 elements and one shared scale. Each element is a small number: FP4 E2M1, FP6 E2M3 or E3M2, FP8 E4M3 or E5M2, or an 8-bit integer. The scale is one byte called E8M0: eight exponent bits, no sign, no mantissa. Code e means 2 to the power e minus 127, code 255 means NaN, and there is no zero. A block of MXFP4 is 32 x 4 bits plus 8, or 4.25 bits per weight.
Lesson 31 opens e8m0.t27, our spec of that byte, in the player. The t27 compiler runs as WebAssembly in the page, emits all seven backends and runs the spec's checks: 17 of 27 pass, none fails, and 10 are skipped, each with the reason the browser cannot run it. Lesson 32 is the same spec on a real machine, where nothing is skipped: 18 of 18 tests pass in Zig, and 9 invariants are proved while the spec compiles, so compiling is the check.
t27c on e8m0.t27 -- the MX shared scale byte, native
Section 6.3 of the MX specification sets the scale from the largest magnitude in the block: its power of two, minus the largest exponent the element can show. One value decides the scale for all 32. Our worked block has 31 weights between -0.17 and 0.17, one of them exactly zero, and one outlier, 3.3. ocp_mx.t27 converts the block to each element type and counts the non-zero inputs that come back as zero. Every count, scale and error below is an assert in the spec; the errors are asserted exactly, in units of 2^-24, and shown here as decimals.
| Element | Bits per weight | Shared scale | Flushed to zero (of 31 non-zero inputs) | Sum of absolute errors |
|---|---|---|---|---|
| MXFP4 E2M1 | 4.25 | 2^-1 | 23 | 2.5610 |
| MXFP6 E2M3 | 6.25 | 2^-1 | 4 | 0.5625 |
| MXFP6 E3M2 | 6.25 | 2^-3 | 0 | 0.3218 |
| MXINT8 | 8.25 | 2^1 | 2 | 0.2451 |
| MXFP8 E4M3 | 8.25 | 2^-7 | 0 | 0.1064 |
In MXFP4 the outlier itself does not survive either: 3.3 over the scale 2^-1 is 6.6, past FP4's largest value of 6.0, so it is clamped and stored as 3.0. At six bits, range beats precision on this block. E3M2, with one more exponent bit and one less mantissa bit, flushes nothing and has the smaller error, because the outlier pushes the small weights down into E2M3's subnormals. The spec asserts that comparison too.
The codes the spec expects for every element of the block did not come from the spec. They were computed by a separate implementation of section 6.3 in exact rational arithmetic (Python fractions), and the spec checks them bit for bit for MXFP4, both MXFP6 types, MXFP8 E4M3 and MXINT8.
This is one block, built to show the effect. It says nothing about a whole model, where the share of blocks that hold such an outlier decides how much it costs. Lesson 29, earlier in the course, measures formats on a real language model; this module does not.
The recording ends by planting a bug: the exponent bias of E5M2 is changed from 15 to 16. Exactly one test fails, bias_puts_the_min_normal_at_two_to_one_minus_bias, and git restores the file. In review, 11 such mutants were tried, and each made at least one test fail; the pull request lists them. Review also found that the first version generated Zig without errors but did not compile it: 20 errors, from array indexes and signed division. "Generates" and "compiles" are different claims, and the spec now makes both. The compiler defect behind it is filed as t27#6867.
Before writing this, we read the pages of twelve other resources that teach or implement these formats, on 6 October 2026. The table puts them in ten rows. "Counted" means the page states how many small values one outlier turns into zero.
| Resource | Kind | MX block and E8M0 scale | Zeros from one outlier | Runs where | Tests behind it |
|---|---|---|---|---|---|
| A Visual Guide to Quantization, M. Grootendorst (2024) | illustrated article | no: INT8 and INT4, no FP8, FP4 or MX | shown in a picture, not counted | static figures | — |
| Introducing NVFP4, NVIDIA (2025) | article | compares MXFP4 with NVFP4 and the error of their scale bytes | not counted | static | — |
| MXFP4 in Transformers, Hugging Face (2025) | article and docs | MXFP4, blocks of 32, no worked block | no | needs a GPU | benchmarks, no tests named |
| Quantization Fundamentals and Quantization in Depth, DeepLearning.AI | video courses | not named on the course pages | no | code examples | not stated |
| TinyML and Efficient Deep Learning, MIT 6.5940 (Fall 2024) | university course | not named on the course page; slides not read | not on the course page | Colab labs | not stated |
| OCP MX v1.0 and arXiv:2310.10537 | standard and paper | yes, the normative definition | gives the rule, no worked block | no test vectors published | |
| microsoft/microxcaling | library | yes: MXFP8, MXFP6, MXFP4, MXINT8 (and INT4, INT2), blocks of 32 | no | PyTorch and CUDA | yes, in Python |
| graphcore-research/gfloat | library | yes: OCP MX element and block formats | no | Python, notebooks | yes, cross-checked against torchao |
| Understanding MXFP4 Quantization, K. Sharma (2025) | interactive page | MXFP4 only, with its own scale rule | says an outlier costs precision, not counted | your browser, your numbers | none cited |
| Floating Point Conversion Calculator, sw23 | interactive page | E8M0 and every MX element type, one value at a time | no block | your browser, shareable links | yes: test vectors and CI |
| This module, t27 lessons 31 to 33 | course module | yes: five element types, blocks of 32, E8M0 | counted: 23, 4, 0, 2 and 0 of 31 non-zero inputs | lesson 31 in your browser; lesson 33 a recording | yes: 49 tests in the spec |
Each of these does something this module does not. A Visual Guide to Quantization covers the whole field, from post-training quantization to GGUF, with patient pictures. NVIDIA and Hugging Face show real models running at these precisions, with accuracy and memory figures; nothing here touches a model. The two DeepLearning.AI courses teach with video and guided coding, and MIT's course is a full semester. microxcaling runs MX inside real PyTorch models. sw23's calculator takes every MX element type apart bit by bit, with rounding and overflow modes and test vectors behind it, and Understanding MXFP4 Quantization lets you type your own numbers and watch the codes change.
gfloat is the fair comparison for "derived from the standard and tested": a readable Python reference over many formats, MX blocks included, cross-checked against a second library. What we did not find anywhere is the count this module is built on: how many of a block's small weights one outlier sends to zero, element type by element type. The standard gives the rule but no worked block, the articles describe the effect in words or as one error figure, and the libraries can encode such a block, but no page we read shows the count. That is a narrow claim, and it is the only one we make. Nor do we run the block's tests in your browser yet: lesson 33 is a recording.
One difference matters if you compare numbers across tools. Understanding MXFP4 Quantization picks the scale as the ceiling of log2(max / 6), and its full visualizer uses blocks of 16. Section 6.3 of the OCP specification takes the largest power of two not above the block's maximum, divided by the largest power of two the element can hold, and its blocks are 32. The specification allows other conversions, so neither is wrong. But for our block, with its outlier of 3.3, the first rule gives a scale of 2^0 and the second 2^-1, so the two tools give different codes for the same weights.
ocp_mx.t27 is one file. It decodes every code of FP4 E2M1 (16), FP6 E2M3 and E3M2 (64 each), INT8, and OFP8 E4M3 and E5M2 (256 each), NaN and Inf included; encodes with round to nearest, ties to even, saturating; and runs the block conversion of section 6.3. Each source is cited next to the code it justifies. Its 49 tests assert exact integers: codes, scales, counts, and errors in units of 2^-24. The named code points and the five block-code tables can be run against a conversion in a kernel or in hardware. If one disagrees with your reading of the specification, we want that report, as an issue in gHashTag/t27.
Until this spec, the t27 format catalog behind arXiv:2606.09686 listed mxfp8, mxfp6 and mxfp4 and cited OCP MX v1.0, but no spec in t27 decoded an MX element or ran the block conversion (t27#6827).
ocp_mx.t27 was merged into t27 on 6 October 2026 (t27#6828, squash commit 322cc77d7). The ocp_mx recording was made before that, on the branch commit 1a3786657; the file is byte-identical in both.ocp_mx.t27 is not on the site yet, so the page does not run its tests (the site's evaluator passes all 49 when run on the file).Work with me
I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.