T27.AI

Blog

A 3 MB compiler that turns a spec into AArch64 code in 5 ms, and where it still loses

2026-10-11 · 12 min read

[measured on a shared Railway lab; run times are qemu emulation, instruction counts are exact; t27b's code is -O0 class] t27b 0.2.0, released with the signed receipt of the very binary it ships, compiles a t27 spec with its tests in about 4.8 ms as a native x86_64 cross-compiler: 8.8x faster than gen-c plus gcc -O0 and about 640x faster than zig Debug per spec, in 5.3 MiB of memory, from one binary of about 3 MB, and it agrees with the reference on all 8,344 tests both ran. It is not a breakthrough in generated code: it executes 3.4x the instructions of gcc -O2 and 4.5x those of zig ReleaseFast, its code is up to 2x larger, and on Linux it is a JIT test runner that cannot yet write an ELF binary. Over the week its share of the reference's passing specs rose from 8.6% to 98.9%.

Three-panel engraved illustration for: A 3 MB compiler that turns a spec into AArch64 code in 5 ms, and where it still loses
Three panels, left to right
  1. ONE BINARYThree megabytes, no linker.
  2. MILLISECONDSA spec compiled in about 5 ms.
  3. THE LIMITExact on every test, -O0 in speed.
View the complete triptych at full size
#t27#Compiler#Measurement#Verification

Here is a compiler you can download as one file of about 3 MB. Give it a .t27 spec and about 5 milliseconds later it has parsed and type-checked the spec, lowered it, and written AArch64 machine code for every function and every test in it. It needs no LLVM, no assembler and no linker, and while it works it uses about 5 MiB of memory. That is t27b, the native backend of the t27 language. On 10 October it shipped as version 0.2.0, together with a signed receipt for the exact binary in the release.

We also measured it against zig, gcc, clang and rustc, on the same specs and the same machine, and the honest result has two halves. t27b is very fast, very small, and it agrees with the reference on every test. The code it writes is slow: about as good as gcc -O0, and three to five times behind every optimising compiler. This post shows both halves, how they were measured, and how to check them yourself.

t27b 0.2.0 · check the binary, compile a spec

Recorded on the Railway t27c lab over ssh on 10 October at 23:38 UTC, in a clone of t27 at the tag t27b-v0.2.0. The release binary and its receipt come from GitHub; the binary's sha256 equals the receipt's t27b_sha256; t27c corpus-receipt compare prints EQUIVALENT, exit 0; a native x86_64 t27b built at the tag writes an AArch64 Mach-O object for kernel_matmul in a few milliseconds (the real line of time); then the release binary runs the spec's two tests in its JIT under qemu-aarch64, where every phase is emulated and slower. 34 s; the prompt and the typing are staged, the output is what the lab printed. Open the recording on its own page.

What t27b is

t27 is a spec-first language: a .t27 file holds types, functions, tests and invariants, and t27c turns it into Zig, C, Rust or Verilog. t27b is the other road. It reads the same file through the t27c front end it mounts, lowers it to its own IR, and emits AArch64 machine code directly. t27b test runs every test in an in-process JIT and checks each result against t27b's own IR interpreter; t27b build writes a Mach-O object. New lowering decisions live in .t27 plan specs that t27c gen-rust turns into Rust, under the project's rule that new logic is written in t27.

The result in one table

Axist27bThe othersFor t27b
Compile time per spec4.8 ms, native x86_64 cross-compiler, tests includedgen-c + gcc -O0 42 ms; gen + zig Debug 3.05 s; gen + zig ReleaseFast 13.8 swins: 8.8x, 640x, 2,900x
All 20 specs, zig in one invocation0.72 szig Debug 6.21 s; zig ReleaseFast 32.9 swins by less: 8.6x and 46x
Peak memory while compiling5.3 MiB (median over specs)zig 354 to 500 MiB; gcc 27 to 32; clang 68 to 76; rustc 112 to 120wins
What has to be installedone binary of about 3 MBzig 363 MB; gcc for aarch64 about 145 MB; rustc 790 MBwins, but it cannot link an executable
Agreement with the reference0 of 8,344 tests differzig ReleaseFast flips 7 files; gen-c + clang disagrees with itself on 39exact
Rebuildsbit-identical, also from another directorygcc (without -g) and rustc too; zig records the build directoryties the best
Instructions its code executesabout gcc -O0 (0.91x)3.4x fewer with gcc -O2; 4.5x fewer with zig ReleaseFastloses
Code size1.8x to 2.0x the optimising buildsgcc -O2, zig ReleaseSafe, zig ReleaseFastloses
Coverage at the bench commit1,296 of 1,311 reference passes (98.9%)zig built for aarch64: 1,311slightly behind

Compile time: milliseconds against seconds

The lab is an x86_64 machine, so t27b was built for the host as well. On x86_64, t27b test runs the front end and the code generator, tests included, and stops only when it tries to map the JIT; the wall time of that process is t27b's compile time. Every other pipeline also ran natively, with the t27c generator's time included in its own. One process at a time, six rounds round-robin over specs and pipelines, warm medians reported.

Time to compile a spec with its tests (log scale)geomean of the warm median over 18 specs, native x86_64, generator time included1 ms10 ms100 ms1 s10 st27b: spec + testst27b: spec + tests: 4.8 ms4.8 mst27b build: objectt27b build: object: 5.1 ms5.1 msgen-c + gcc -O0gen-c + gcc -O0: 42.0 ms · 8.8x42.0 ms · 8.8xgen-c + clang -O0gen-c + clang -O0: 71.1 ms · 15x71.1 ms · 15xgen-c + gcc -O2gen-c + gcc -O2: 100 ms · 21x100 ms · 21xgen-c + clang -O2gen-c + clang -O2: 105 ms · 22x105 ms · 22xgen + zig Debuggen + zig Debug: 3.05 s · 642x3.05 s · 642xgen + zig ReleaseFastgen + zig ReleaseFast: 13.8 s · 2,906x13.8 s · 2,906xAll 20 specs, total wall time (log scale)compiler processes only; zig shares its fixed cost in one call, t27b has no such mode1 s10 s100 s1000 st27b, 20 processest27b, 20 processes: 0.72 s0.72 sgcc -O0, 20 processesgcc -O0, 20 processes: 1.31 s · 1.8x1.31 s · 1.8xzig Debug, one callzig Debug, one call: 6.21 s · 8.6x6.21 s · 8.6xzig ReleaseFast, one callzig ReleaseFast, one call: 32.9 s · 46x32.9 s · 46xzig Debug, 20 callszig Debug, 20 calls: 61.7 s · 86x61.7 s · 86xzig ReleaseFast, 20 callszig ReleaseFast, 20 calls: 284 s · 395x284 s · 395x
Top: geometric mean of the warm median over the 18 tier-1 specs every step compiled. Bottom: the total over all 20 tier-1 specs, compiler processes only, as in the report's table (t27c's generator adds about 0.2 s for 20 specs to the gcc and zig rows). Native x86_64 on the shared Railway t27c lab, 10 October; host load 27 to 103 on 48 threads. Source: t27b-bench-2026-10.json, axis B.

Most of zig's time per spec is fixed cost: building the test runner, the parts of std that panic and format, and the link, all through LLVM. Small and large specs alike took 2.7 to 3.4 s in Debug. Build all 20 specs in one invocation and zig spends 0.31 s per spec: the gap falls from about 640x to 8.6x in Debug, and from about 2,900x to 46x in ReleaseFast. t27b has no multi-file mode, so its total is 20 separate processes. Over the 567 files every pipeline passes, t27b compiles 183 files a second; gen-c plus gcc -O0 manages 28.

One correction we owe: earlier reports quoted t27b compile times of 33 to 86 ms. Those were measured under qemu. Calibrated against the native build on the same specs, emulation made t27b's compile phases about 20x slower (geometric mean), so the native figure is an order of magnitude lower.

Memory and footprint

Peak memory while compiling (MiB, log scale)median over the specs of axis C; the whole process tree, from wait41101001000t27bt27b: 5.3 MiB5.3 MiBgcc -O0gcc -O0: 27.1 MiB · 5.2x27.1 MiB · 5.2xgcc -O2gcc -O2: 31.6 MiB · 6.0x31.6 MiB · 6.0xclang -O0clang -O0: 68.0 MiB · 13x68.0 MiB · 13xclang -O2clang -O2: 75.5 MiB · 14x75.5 MiB · 14xrustc -O0rustc -O0: 112 MiB · 21x112 MiB · 21xzig Debugzig Debug: 354 MiB · 67x354 MiB · 67xzig ReleaseFastzig ReleaseFast: 500 MiB · 95x500 MiB · 95xWhat has to be installed (MB, log scale)t27b needs no assembler and no linker; it also cannot link an executable1101001000t27b, one binaryt27b, one binary: 3.1 MB3.1 MBgcc 12 for aarch64 (about)gcc 12 for aarch64 (about): 145 MB · 47x145 MB · 47xzig 0.16 (also clang)zig 0.16 (also clang): 363 MB · 118x363 MB · 118xrustc 1.99 toolchainrustc 1.99 toolchain: 790 MB · 257x790 MB · 257x
Top: peak resident memory of each compiler's whole process tree (gcc's cc1, as and ld included), read from wait4 by a small static helper; the median over the specs of axis C. Bottom: what has to be installed. The t27b size is its x86_64 build at the bench commit; the aarch64 binary in the release is 3.09 MB. Source: t27b-bench-2026-10.json, axis D.

t27b's median is 5.3 MiB, and its hungriest spec needed 48 MiB. zig needs 354 MiB in Debug and 500 MiB in ReleaseFast, about 67 and 95 times more. The install tells the same story: one executable that links Rust's std statically and needs glibc at run time, against hundreds of megabytes of compiler, libc sources and LLVM. The price of that smallness is real: t27b cannot link an executable.

Exact against the reference

The reference is t27c gen followed by zig test in Debug, run natively. t27b ran the whole corpus, 1,819 files, through its JIT and its interpreter, and wherever both sides ran a file's tests they were compared test by test: 0 of 8,344 tests differ. zig Debug built for aarch64 and run under the same qemu also agrees with t27b on all 8,344.

The optimisers are less exact. zig ReleaseFast flips 7 files against zig Debug; one of them is a test that reads a stack slot after its frame has returned. The C road is worse. gen-c output contains undefined behaviour, so its verdict depends on the optimisation level: clang -O0 and -O2 disagree on 39 files and 107 tests. The smallest case is a function gen-c prints as bool dummy(void) { true; }, with no return. It passes under gcc -O0 and clang -O2, and aborts under gcc -O2 and clang -O0.

Every pipeline is deterministic in place. t27b, t27c's generators, gcc without -g and rustc are also bit-identical when built from another directory; zig and zig cc record the build directory in their debug information. The t27b lab's signed receipts depend on this property: two runs of one commit must produce one Merkle root.

Where t27b loses

t27b's code divided by theirs (geomean, 20 specs)above 1, t27b executes more instructions or emits more bytes; instruction counts do not depend on the hostinstructions executedcode size0x1x2x3x4x5x← t27b bettert27b worse →zig Debugzig Debug, instructions executed: 0.54x0.54xzig Debug, code size: 0.66x0.66xclang -O0clang -O0, instructions executed: 0.41x0.41xclang -O0, code size: 0.65x0.65xgcc -O0gcc -O0, instructions executed: 0.91x0.91xgcc -O0, code size: 1.34x1.34xzig ReleaseSafezig ReleaseSafe, instructions executed: 3.17x3.17xzig ReleaseSafe, code size: 1.89x1.89xgcc -O2gcc -O2, instructions executed: 3.43x3.43xgcc -O2, code size: 1.82x1.82xclang -O2 (9 specs)clang -O2 (9 specs), instructions executed: 3.50x3.50xclang -O2 (9 specs), code size: 1.69x1.69xzig ReleaseFastzig ReleaseFast, instructions executed: 4.47x4.47xzig ReleaseFast, code size: 2.03x2.03x
t27b's generated code divided by each pipeline's: geometric mean over the 20 tier-1 specs (clang -O2 over the 9 where it did not evaluate the test at compile time). Instructions were counted by a qemu TCG plugin and do not depend on the host. The specs were chosen without looking at t27b: the 20 with the most gcc -O0 instructions among the 567 files every pipeline passes. Source: t27b-bench-2026-10.json, axis C.

This is the half of the result that does not make a headline. t27b's code executes about as many instructions as gcc -O0 and about half as many as zig Debug. Against every optimising build it is 3 to 5 times behind: 3.2x zig ReleaseSafe, the build with the same run-time checks; 3.4x gcc -O2; 4.5x zig ReleaseFast. Under qemu that is 3.4x to 4.1x the run time, and on the crypto specs the gap to zig ReleaseFast reaches 10x in instructions.

WhereThe numberWhat it means
Generated code3.4x the instructions of gcc -O2, 3.2x zig ReleaseSafe, 4.5x zig ReleaseFasta program t27b builds runs slower than the same program from any optimiser
Code size1.8x to 2.0x every optimising buildmore bytes for the same functions
Coverage1,296 of 1,311 at the bench commit (files with tests); 1,349 of 1,364 at the release (the lab tally)i128 stage 3, @cos and two empty typed array literals still block
Linux executablesMach-O objects only: no ELF, no link stepon aarch64-linux t27b runs tests in its JIT; it cannot ship a binary
Self-generation36% of the backend is generated from .t27; 12% counting the front end it mountsmost of the code t27b runs is still hand-written Rust
DDCDDC_BACKEND_AGREEMENT, not DDC_INDEPENDENTthe three routes share one front end
Random programs7 disagreements per 1,000 programs (3 in the release run)the fuzzer reaches code the corpus does not, and still finds differences
Per-test resett27core runs 8.3x slower than under zig Debugjit.call copies all module state before every test
Amortisationthe 640x lead falls to 8.6x when zig builds 20 specs at oncet27b has no multi-file mode

Two smaller costs. Overflow checks, on by default, cost about 12% of the instructions and 24% of the emulated time. And one spec, compiler/core/t27core, runs 8.3x slower than under zig Debug although t27b executes fewer instructions: before every test jit.call copies all module state so that each test starts fresh, and each of t27core's 95 tests pays at least 2.8 ms for that copy.

How it was measured, and what to distrust

A week of coverage

t27b's share of the reference's passes (4-10 Oct)every master run of the t27b lab; the denominator grew from 648 to 1,364 with the corpus and the referencepasses with runtime checkspasses vacuously (no assert ran)0%25%50%75%100%4 Oct5 Oct6 Oct7 Oct8 Oct9 Oct10 Oct2026-10-04 16:06Z dd3864dff: 56 of 648 (8.6%)2026-10-04 22:13Z aa57ba405: 190 of 719 (26.4%)2026-10-05 05:23Z df7b627b6: 370 of 720 (51.4%)2026-10-05 12:51Z 03362c1da: 379 of 722 (52.5%)2026-10-05 19:05Z 1f9b3706a: 428 of 773 (55.4%), +283 vacuous2026-10-06 01:49Z 099ac2224: 457 of 791 (57.8%), +261 vacuous2026-10-06 09:47Z 0241b52f9: 462 of 795 (58.1%), +261 vacuous2026-10-06 15:56Z d4a4f82cc: 519 of 828 (62.7%), +247 vacuous2026-10-06 22:18Z b1d823579: 541 of 844 (64.1%), +247 vacuous2026-10-07 04:43Z 9c6d32617: 559 of 855 (65.4%), +245 vacuous2026-10-07 10:53Z 68fb92807: 827 of 1,045 (79.1%), +151 vacuous2026-10-07 17:10Z 76a9de9ab: 853 of 1,074 (79.4%), +150 vacuous2026-10-07 23:20Z 4dd14bf69: 875 of 1,084 (80.7%), +150 vacuous2026-10-08 05:50Z 2e06cef59: 899 of 1,099 (81.8%), +150 vacuous2026-10-08 12:13Z 579bc55aa: 924 of 1,115 (82.9%), +151 vacuous2026-10-08 18:15Z 7c884488c: 988 of 1,177 (83.9%), +152 vacuous2026-10-09 00:18Z a40d89128: 1,021 of 1,193 (85.6%), +142 vacuous2026-10-09 06:19Z f97e761a4: 1,039 of 1,211 (85.8%), +142 vacuous2026-10-09 12:38Z 30c9755c7: 1,094 of 1,219 (89.7%), +100 vacuous2026-10-09 19:12Z b9d22d0f4: 1,173 of 1,249 (93.9%), +56 vacuous2026-10-10 02:21Z b0c63a215: 1,201 of 1,279 (93.9%), +56 vacuous2026-10-10 08:40Z 7c16bfab0: 1,208 of 1,281 (94.3%), +50 vacuous2026-10-10 15:18Z 26a975354: 1,277 of 1,309 (97.6%), +5 vacuous2026-10-10 21:43Z a977fd715: 1,339 of 1,354 (98.9%), +5 vacuous2026-10-10 22:34Z 59e2c87c4: 1,349 of 1,364 (98.9%), +5 vacuous56 of 648 (8.6%)v0.1.0: 899 of 1,099 (81.8%)v0.2.0: 1,349 of 1,364 (98.9%)
Each point is one master run of the t27b lab, 4 to 10 October: the files t27b passes with runtime checks, divided by the files the reference passes. The band above the line is the vacuous passes (every test passed, no assert ran), counted per file from 5 October 15:43 UTC. The denominator grew from 648 to 1,364 with the corpus and the reference. Source: the run summaries at t27b-lab-production.up.railway.app/runs/.

On 4 October t27b passed 56 of the 648 specs the reference passed. v0.1.0, on 8 October, passed 899 of 1,099 with runtime checks and 150 more vacuously. v0.2.0 passes 1,349 of 1,364, with 5 vacuous, and 0 of its 8,629 compared tests differ from the reference. Between the benchmark commit and the release four more blockers fell, and t27b's passes rose from 1,312 to 1,349 while the reference's rose from 1,331 to 1,364: anytype monomorphization (t27#8738), run-time i128 stages 1 and 2 (t27#8750, t27#8795), and frames up to 16 MiB of local aggregates (t27#8764).

What is left on the ledger: run-time i128 stage 3 for tri/t27b/eval_arith.t27, @cos (t27#8677), two empty typed array literals, five stubs that pass with nothing to assert (t27#8766), and five rows where the reference passes for the wrong reason and t27b is right to refuse.

Why you can check this

The binary in the release is not a rebuild. It is the file the t27b lab ran the whole corpus with, and its sha256, 0202a2ae..., is the t27b_sha256 field of the lab's receipt for that run. The receipt is signed with Ed25519 under the domain line t27-corpus-receipt-v1, and it carries Merkle roots over every input file, every verdict and every generated listing. From a t27 checkout, t27c corpus-receipt compare R R checks the signature and rehashes every leaf:

base "59e2c87c4bd74c1137f4071b5e63095a285ce5e4" AUTH_MISSING_NONE AUTHOR leaves-bound true
head "59e2c87c4bd74c1137f4071b5e63095a285ce5e4" AUTH_MISSING_NONE AUTHOR leaves-bound true
totals true inputs true verdicts true outputs true: EQUIVALENT

Every corpus run also executes each test twice, in the JIT and in t27b's own IR interpreter, with 0 mismatches. And a signed DDC receipt builds t27core's self-compile three ways, through gen-c and cc, through t27b and through zig, and all three stage-2 digests agree.

Now the limits, because they are the point of the exercise. That DDC verdict is DDC_BACKEND_AGREEMENT, not DDC_INDEPENDENT: all three routes share t27c's front end. The JIT and the interpreter read the same IR. The reference t27b agrees with on 8,344 tests comes out of the same front end that t27b mounts. The random-program fuzzer is the one check that reaches beyond the corpus, and it still finds disagreements: 7 per 1,000 programs in the run the report cites, 3 in the release run. No check here compares t27b with an independently written t27 front end. The verification is unusually visible; most of it is still self-consistency.

What comes next

Verdict

Not a breakthrough in generated code. A fast, lean baseline compiler that is exact against its reference: a 3 MB, LLVM-free compiler that needs about 5 MiB of memory, turns a spec into AArch64 code in about 5 ms, and agrees with the reference on every one of 8,344 tests. That is the axis to lead with, and the 3 to 5x gap in code quality is the axis to work on. The release, the benchmark report with every caveat (t27#8855), the improvement loop (epic t27#8326) and the receipt of the release run are linked below.

What this does not settle

Receipts

Work with me

Want this kind of check on your own design?

I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.