T27.AI

Blog

Where t27c's compile time goes, and a first backend without LLVM

2026-10-04 · 10 min read

[one M1 Pro, one synthetic program; the t27b numbers were taken at a load average of 83–165 on 8 cores; t27b handles 37 of 1174 specs and its JIT runs only on arm64 macOS; three of the four bugs it found are not fixed yet] t27c type-checks a function in 29.9 µs, 5.2 times faster than rustc check, but it emits no machine code, so its full path to an object file is 2.1 times slower than clang on the same program written in C. Its own clean release build went from 142.7 s to 66.0 s after three merged PRs removed unused machine-learning crates and the aws-lc-sys build. t27b, a native AArch64 backend with no LLVM, now merged, compiles and runs the same tests 23 to 381 times faster than the Zig path, writes code at about the clang -O0 level, and found four bugs in t27c and in a spec.

#t27#Compiler#Measurement

t27c, the t27 compiler, is a translator. It reads a .t27 spec and prints Zig, C, Rust or Verilog, and a different compiler, almost always LLVM, turns that into machine code. This post asks three questions about it. How fast is the part t27c does itself? What does the whole path cost once the backend is counted? And how long does t27c take to build, since it is a Rust program? It closes with a small backend of our own, t27b, which emits AArch64 code without LLVM, and the four bugs it found on the way.

All of it was measured on 4 October 2026 on one Apple M1 Pro (8 cores, 16 GB). The test program is synthetic. The limits are listed at the end of the post, and they matter.

Why Rust and C++ builds are slow

The complaint is older than Rust. Go was designed at Google partly because a large C++ build took 45 minutes. In Pike's 2012 account, 4.2 MB of C++ source grew to more than 8 GB once every #include was expanded, about 2000 times. Our own test file shows the same effect at small scale: six standard headers take it from 1,007 to 88,104 lines after clang++ -E.

Rust has a different cost. Generic code is compiled once for each concrete type, and the unit of compilation is a whole crate. rustc splits a crate into codegen units to keep the cores busy, but it never splits one module. t27c shows this directly. bootstrap/src/compiler.rs is 44,636 lines in one module, and its codegen unit is 880 KB, 98.5% of it that one file. In -Z time-passes on a release build of the t27c crate, rustc spent 36.6 of 60.7 s waiting for LLVM. The machine was loaded during that run.

Then there is LLVM itself. In our test, the clang front end (-fsyntax-only) is 24.5% of the time of clang -O0. The rest is IR, LLVM code generation and writing the object file. Optimisation costs more again: -O2 takes 6.4 times as long as -O0 for clang on C, and 3.7 times as long for rustc. The clearest case is Zig. With its own backend (-fno-llvm) it builds the same file 5.7 times faster than through LLVM. Published non-LLVM backends report the same order of gain: Copy-and-Patch (OOPSLA 2021) compiles about 100 times faster than LLVM -O0, and TPDE (CGO 2026) 8 to 24 times faster.

One gap in the literature should be said plainly. We found no peer-reviewed paper that measures Rust compile time technically. The Rust project's 2025 compiler performance survey had more than 3,700 answers, and 55% of respondents wait more than 10 s for an incremental rebuild. It is a blog post, not a paper. Every Rust number below is our own measurement.

How it was measured

t27c checks a function in 29.9 µs

101001,000clang -fsyntax-only (C)24.4t27c typecheck29.9t27c gen-c33.7zig -fno-llvm82.1go tool compile95.2clang -O0 (C)99.6rustc check154.4rustc -O0291.6zig Debug (LLVM)468.4clang -O2 (C)636.0rustc -O21,089.5swiftc -typecheck2,084.6t27cchecks only, no codemachine codeµs per function, log scale
Microseconds per function at N = 5,000, median of 3 runs, log scale. Gold is t27c. Light bars only check the code; dark bars emit machine code. The two kinds of work are not the same, and the table below keeps them apart.
CompilerWhat it doesms at N = 5,000µs per function
clang -fsyntax-only (C)parse and check C, no code121.924.4
t27c typecheckparse and type-check t27, no code149.729.9
t27c gen-cparse, check and print C source168.433.7
zig -fno-llvmZig's own backend, Debug, no LLVM410.382.1
go tool compileGo's own backend476.195.2
clang -O0 (C).o through LLVM, no optimisation498.199.6
rustc checktypes and borrows, metadata only772.1154.4
rustc -O0.o through LLVM, no optimisation1,457.8291.6
zig Debug (LLVM).o through LLVM, Debug2,341.9468.4
clang -O2 (C).o through LLVM, -O23,180.2636.0
rustc -O2.o through LLVM, opt-level=25,447.61,089.5
swiftc -typecheckSwift type check, no code10,423.22,084.6

Among the tools that only check code, t27c typecheck comes second: 29.9 µs per function against 24.4 for clang. That is 5.2 times faster than rustc check and 69.6 times faster than swiftc -typecheck. Printing C source on top of the check costs little: t27c gen-c takes 33.7 µs.

But t27c does not emit machine code

So the fair comparison is the whole path: t27c plus the backend that runs on what it prints.

Path, N = 5000t27c, msbackend, mstotal, mst27c's share
C written by hand: clang -O0—498.1498.1—
t27 → C → clang -O0168.4899.61,068.015.8%
Zig written by hand: zig Debug (LLVM)—2,341.92,341.9—
t27 → Zig → zig test (build and link)167.93,370.43,538.44.7%

The fastest t27 path to machine code goes through C, at 1,068.0 ms. The same program written in C by hand compiles in 498.1 ms, 2.1 times faster. Most of the difference is not t27c. The generated C is 80,043 lines against 50,001, because it also carries 5,000 tests. Through Zig, the backend takes 95.3% of the time. While t27c depends on someone else's backend, its full path is never faster than that backend.

Building t27c itself: 142.7 s to 66.0 s

t27c is a Rust program, so everything in the first section applies to it. Its clean release build took 142.7 s by the wall clock. cargo's timing report says where the time went:

Build unitVersionSecondsFate
aws-lc-sys (build script)0.41.078.8removed in #5920
candle-core0.11.052.4removed in #5900
candle-core0.10.249.4removed in #5900
t27c (the compiler itself)0.4.034.4stays
tokenizers0.22.233.3removed in #5900
candle-nn0.10.214.6removed in #5900

The longest unit was the build script of aws-lc-sys, a C crypto library that came in through reqwest → rustls → aws-lc-rs. Next came two versions of candle-core and tokenizers. These are machine-learning libraries that nothing in bootstrap/src used. The t27c crate itself took 34.4 s. Three merged PRs removed most of this:

before142.7 s · 386 build unitswithout candle (#5900)97.9 s · 279 build unitsplus native-tls (#5920)66.0 s · 273 build unitsthe #5900 tree, built again110.9 s · 279 build units
Clean release build of t27c, wall clock, one run each. The dashed bar is the #5900 tree built a second time, at 110.9 s instead of 97.9 s. That gap is how much one run on this machine can move, so read the bars as a direction rather than a precise figure.

A clean build is rare; editing one line and rebuilding is not. A normal release rebuild after a one-line change took 29.8 to 33.9 s, because release is not incremental and the large compiler.rs unit sits on the critical path. With CARGO_PROFILE_RELEASE_INCREMENTAL=true it took 3.2 to 5.0 s, and cargo check took 2.4 to 3.3 s. #5945 (merged) writes this loop into CONTRIBUTING.

One idea did not work. We split compiler.rs into 24 files, and the compiler's output stayed byte-identical. Then we built the t27c crate in 6 interleaved rounds. The median went from 49.6 s to 37.8 s and the minimum from 34.1 s to 32.1 s, while CPU time rose from 95.9 s to 98.4 s. The split build was faster in only 3 of 6 pairs, and single builds ranged from 32 to 117 s. That is noise, not a speed-up. The branch stays local and is not a PR.

t27b: a backend of our own, for part of the language

t27b is a separate crate, cli/t27b, merged in #5979 (a follow-up, #5989, also merged, registers it in the CI ledger of orphan crates). It emits AArch64 machine code itself, with no LLVM, zig, clang or rustc. Its front end is t27c's own parser and type checker, used unchanged. After that the code is new: an IR of its own, an A64 encoder, and then either execution in memory (t27b test, a JIT) or a Mach-O object file (t27b build). Integer overflow is defined: it either traps or wraps (--overflow trap|wrap), and a shift by an amount outside [0, width) traps. A reference interpreter of the same IR is the oracle. The JIT runs only on arm64 macOS.

These numbers come from a different session than the ones above. The machine was heavily loaded, with a load average of 83 to 165 on 8 cores. Absolute times are therefore inflated. Each ratio below compares variants run interleaved in the same session, 5 times each.

10 ms100 ms1,000 ms10,000 msN = 100load average83–13417.8 ms1,158 ms · ×65.06,798 ms · ×381.5N = 1,000load average129–16159.6 ms1,189 ms · ×19.96,124 ms · ×102.7N = 5,000load average98–165917 ms4,529 ms · ×4.921,058 ms · ×23.0t27b testgen-c + clang + rungen + zig test
Compiling and running the same tests, median of 5 runs, log scale. "×" is how many times longer than t27b. The load average under each N shows how busy the machine was.
Nload avgt27b test, ms (median / min)gen-c + clang + run, mslonger than t27b (median / min / CPU)gen + zig test, mslonger than t27b (median / min / CPU)
10083–13417.8 / 15.81,158×65.0 / ×59.0 / ×14.46,798×381.5 / ×349.2 / ×245.7
1,000129–16159.6 / 51.11,189×19.9 / ×15.8 / ×8.66,124×102.7 / ×61.9 / ×67.4
5,00098–165916.7 / 250.14,529×4.9 / ×8.3 / ×6.421,058×23.0 / ×20.0 / ×24.9

Against the Zig path, t27b is 381 times faster at N = 100 and 23 times faster at N = 5,000 (medians). By CPU time the figures are 246 and 25. The C-path median ratio is inflated by load. At N = 100, running ./ctest took 970.3 ms of wall time but only 6.7 ms of CPU, so most of that time was spent waiting. Counting compilation alone, the C path takes 10.5 times as long as t27b at N = 100 and 4.3 times as long at N = 5,000. The load also shows in zig test: at N = 5,000 it took 20.5 s here and 3.4 s in the quieter harness run. That is why t27b is kept off the earlier tables.

An object file

t27b build writes a .o the way clang -c does. In the same session, at the median, in ms:

Nt27b build (trap)gen-c + clang -O0 -cgen-c + clang -O2 -cC by hand, clang -O0 -c
10021.4120.3 (×5.6)301.2 (×14.1)74.4 (×3.5)
1,00062.9570.5 (×9.1)1,780.1 (×28.3)190.2 (×3.0)
5,000822.73,847.7 (×4.7)20,388.9 (×24.8)1,496.0 (×1.8)

The comparison is lopsided: clang does far more work than t27b. It does show where t27b sits, though. At N = 5,000 it is faster than clang compiling the hand-written C at -O0.

Where t27b spends its time

Phase, N = 5,000, median of 15msshare
parse (shared with t27c)157.248.3%
type check (shared with t27c)90.827.9%
lower to its own IR37.811.6%
A64 code generation17.15.3%
map into memory and run1.60.5%
total, including reading the file325.8100.0%

76.1% of the time goes to the front end that t27b shares with t27c. Its own part, from lowering to running, is 17.3%. Making t27b faster from here means making t27c's front end faster.

How fast the code is

Nt27b trapt27b wrapgen-c + clang -O0gen-c + clang -O2
1009.407.3611.153.96
1,00010.368.4712.464.93
5,00010.958.5013.655.30

Nanoseconds per call, median of 5. All variants return the same checksum. t27b's code is at about the level of clang -O0. In trap mode it takes 0.80 to 0.84 of the time of gen-c with -O0. In wrap mode it is 1.60 to 1.86 times slower than gen-c with -O2. t27b has no optimiser. At N = 5,000 its code section is 400,000 bytes in trap mode and 319,992 in wrap mode, against 963,612 and 523,608 for gen-c with clang at -O0 and -O2.

Correctness, and coverage: 37 of 1,174

In a differential test, every function was called on the same inputs from t27b's code and from clang's. Wrap mode was compared with -O2, and trap mode with -O0 plus UBSan traps on overflow. Over 12,261,000 calls there were 0 mismatches. That covers one synthetic program, and it is not a proof.

t27b corpus ran over all 1,174 specs in the repository in 3.6 s. t27b understands 37 of them. 36 pass, and 26 of those have no tests at all; 1 fails, because of a real bug in the spec (below). It refused 1,117 files, and the shared front end failed on 20. There were no mismatches between the JIT and the interpreter, no timeouts and no crashes. The commonest reasons for refusal, counted as the first refusal per file, are StructDecl (393), string literal (287), EnumDecl (89), InvariantBlock (70), type str (31). Structs, strings, enums, invariant blocks and casts are tracked in #5977, which is open. The first coverage PR, #5992, runs invariant blocks like tests and reports them separately; it is submitted, not merged.

The plan from here, planned and in progress with no numbers yet, is to cover every spec that the reference path itself passes, in two lanes. The memory lane adds structs, arrays and slices, and strings, all on one model of an address plus an offset, with read-only data for constants. The scalar lane adds enums, f64 and casts. The order of the work comes from a greedy ranking: next is whichever feature unlocks the most specs that are still refused.

t27b's own clean release build has 17 build units and 8 crates, against 386 and 295 for t27c. Its binary is 1.37 MB against 15.6 MB. Its build time under a load of 244 to 336 means little: the median of 3 was 111.6 s and the minimum 67.2 s. Most of the source it compiles is not its own. It mounts compiler.rs (44,636 lines) and use_resolve.rs (1,010), next to about 5,100 lines of its own code.

Why it is fast is no mystery. It is the same reason Zig's own backend is 5.7 times faster than Zig through LLVM, and the reason behind Copy-and-Patch and TPDE: one pass, no LLVM, code at about the -O0 level.

What t27b found

A second backend with defined semantics is useful just because its answers can be compared with t27c's. That comparison found four bugs. One is fixed so far.

BugHow it was foundState on 4 October 2026
ternary_model.t27: dot27 shifts an i32 by up to 52 bits, and its test expected wrong valuest27b corpus: "shift amount out of range at line 40"fixed: issue #5972 closed, PR #5975 merged on 4 October
gen-c writes plain C (x + 1, x >> n). Signed overflow and an oversized shift are undefined behaviour in C, and the tests inc_max and shr_big change outcome between clang -O0 and -O2a probe module, three backends side by sidein progress, not yet a PR
t27c gen (Zig): a signed % does not compile; Zig wants @rem or @modthe same probe modulesubmitted, not merged: PR #5993
The t27c parser reads module a.b; and use std.testing; as a followed by a stray expression .b, with no line number; 12 specs are parsed this wayt27b corpus: a refusal on "StmtExpr" at line 0in progress, not yet a PR

The second bug matters most. The same spec means different things depending on the backend and the optimisation level. Zig traps on the overflow, and so does t27b. The C that t27c generates silently does whatever the optimiser decides. The Zig fix is PR #5993, submitted, not merged. The gen-c fix and the parser fix are on local branches and are not PRs, so this post records them as in progress.

What this does not show

What this does not settle

Receipts

Work with me

Need an FPGA/RTL problem taken to measured hardware?

I work contract and part-time on hardware-AI, FPGA/RTL and ML systems — from specification and open toolchains to reproducible measurements.