Blog
[one M1 Pro, one synthetic program; the t27b numbers were taken at a load average of 83–165 on 8 cores; t27b handles 37 of 1174 specs and its JIT runs only on arm64 macOS; three of the four bugs it found are not fixed yet] t27c type-checks a function in 29.9 µs, 5.2 times faster than rustc check, but it emits no machine code, so its full path to an object file is 2.1 times slower than clang on the same program written in C. Its own clean release build went from 142.7 s to 66.0 s after three merged PRs removed unused machine-learning crates and the aws-lc-sys build. t27b, a native AArch64 backend with no LLVM, now merged, compiles and runs the same tests 23 to 381 times faster than the Zig path, writes code at about the clang -O0 level, and found four bugs in t27c and in a spec.
t27c, the t27 compiler, is a translator. It reads a .t27 spec and prints Zig, C, Rust or Verilog, and a different compiler, almost always LLVM, turns that into machine code. This post asks three questions about it. How fast is the part t27c does itself? What does the whole path cost once the backend is counted? And how long does t27c take to build, since it is a Rust program? It closes with a small backend of our own, t27b, which emits AArch64 code without LLVM, and the four bugs it found on the way.
All of it was measured on 4 October 2026 on one Apple M1 Pro (8 cores, 16 GB). The test program is synthetic. The limits are listed at the end of the post, and they matter.
The complaint is older than Rust. Go was designed at Google partly because a large C++ build took 45 minutes. In Pike's 2012 account, 4.2 MB of C++ source grew to more than 8 GB once every #include was expanded, about 2000 times. Our own test file shows the same effect at small scale: six standard headers take it from 1,007 to 88,104 lines after clang++ -E.
Rust has a different cost. Generic code is compiled once for each concrete type, and the unit of compilation is a whole crate. rustc splits a crate into codegen units to keep the cores busy, but it never splits one module. t27c shows this directly. bootstrap/src/compiler.rs is 44,636 lines in one module, and its codegen unit is 880 KB, 98.5% of it that one file. In -Z time-passes on a release build of the t27c crate, rustc spent 36.6 of 60.7 s waiting for LLVM. The machine was loaded during that run.
Then there is LLVM itself. In our test, the clang front end (-fsyntax-only) is 24.5% of the time of clang -O0. The rest is IR, LLVM code generation and writing the object file. Optimisation costs more again: -O2 takes 6.4 times as long as -O0 for clang on C, and 3.7 times as long for rustc. The clearest case is Zig. With its own backend (-fno-llvm) it builds the same file 5.7 times faster than through LLVM. Published non-LLVM backends report the same order of gain: Copy-and-Patch (OOPSLA 2021) compiles about 100 times faster than LLVM -O0, and TPDE (CGO 2026) 8 to 24 times faster.
One gap in the literature should be said plainly. We found no peer-reviewed paper that measures Rust compile time technically. The Rust project's 2025 compiler performance survey had more than 3,700 answers, and 55% of respondents wait more than 10 s for an incremental rebuild. It is a blog post, not a paper. Every Rust number below is our own measurement.
N independent functions, each an 8-iteration loop over a u32. The t27 version also gives each function its own test block. The same program is written by hand in C, C++, Rust, Zig, Go and Swift, for N = 100, 1000 and 5000.| Compiler | What it does | ms at N = 5,000 | µs per function |
|---|---|---|---|
clang -fsyntax-only (C) | parse and check C, no code | 121.9 | 24.4 |
t27c typecheck | parse and type-check t27, no code | 149.7 | 29.9 |
t27c gen-c | parse, check and print C source | 168.4 | 33.7 |
zig -fno-llvm | Zig's own backend, Debug, no LLVM | 410.3 | 82.1 |
go tool compile | Go's own backend | 476.1 | 95.2 |
clang -O0 (C) | .o through LLVM, no optimisation | 498.1 | 99.6 |
rustc check | types and borrows, metadata only | 772.1 | 154.4 |
rustc -O0 | .o through LLVM, no optimisation | 1,457.8 | 291.6 |
zig Debug (LLVM) | .o through LLVM, Debug | 2,341.9 | 468.4 |
clang -O2 (C) | .o through LLVM, -O2 | 3,180.2 | 636.0 |
rustc -O2 | .o through LLVM, opt-level=2 | 5,447.6 | 1,089.5 |
swiftc -typecheck | Swift type check, no code | 10,423.2 | 2,084.6 |
Among the tools that only check code, t27c typecheck comes second: 29.9 µs per function against 24.4 for clang. That is 5.2 times faster than rustc check and 69.6 times faster than swiftc -typecheck. Printing C source on top of the check costs little: t27c gen-c takes 33.7 µs.
So the fair comparison is the whole path: t27c plus the backend that runs on what it prints.
| Path, N = 5000 | t27c, ms | backend, ms | total, ms | t27c's share |
|---|---|---|---|---|
| C written by hand: clang -O0 | — | 498.1 | 498.1 | — |
| t27 → C → clang -O0 | 168.4 | 899.6 | 1,068.0 | 15.8% |
| Zig written by hand: zig Debug (LLVM) | — | 2,341.9 | 2,341.9 | — |
| t27 → Zig → zig test (build and link) | 167.9 | 3,370.4 | 3,538.4 | 4.7% |
The fastest t27 path to machine code goes through C, at 1,068.0 ms. The same program written in C by hand compiles in 498.1 ms, 2.1 times faster. Most of the difference is not t27c. The generated C is 80,043 lines against 50,001, because it also carries 5,000 tests. Through Zig, the backend takes 95.3% of the time. While t27c depends on someone else's backend, its full path is never faster than that backend.
t27c is a Rust program, so everything in the first section applies to it. Its clean release build took 142.7 s by the wall clock. cargo's timing report says where the time went:
| Build unit | Version | Seconds | Fate |
|---|---|---|---|
aws-lc-sys (build script) | 0.41.0 | 78.8 | removed in #5920 |
candle-core | 0.11.0 | 52.4 | removed in #5900 |
candle-core | 0.10.2 | 49.4 | removed in #5900 |
t27c (the compiler itself) | 0.4.0 | 34.4 | stays |
tokenizers | 0.22.2 | 33.3 | removed in #5900 |
candle-nn | 0.10.2 | 14.6 | removed in #5900 |
The longest unit was the build script of aws-lc-sys, a C crypto library that came in through reqwest → rustls → aws-lc-rs. Next came two versions of candle-core and tokenizers. These are machine-learning libraries that nothing in bootstrap/src used. The t27c crate itself took 34.4 s. Three merged PRs removed most of this:
candle-core and candle-nn. The build went from 142.7 s to 97.9 s, and from 386 to 279 build units. The set of failing t27c suite tests was the same before and after.reqwest, tokio, axum, hyper, jsonwebtoken) behind the cargo features net and server. A default build now compiles 49 crates instead of 212.reqwest to native-tls, so the aws-lc-sys build disappears. The build with the network stack went from 97.9 s to 66.0 s, 54% below where it started.A clean build is rare; editing one line and rebuilding is not. A normal release rebuild after a one-line change took 29.8 to 33.9 s, because release is not incremental and the large compiler.rs unit sits on the critical path. With CARGO_PROFILE_RELEASE_INCREMENTAL=true it took 3.2 to 5.0 s, and cargo check took 2.4 to 3.3 s. #5945 (merged) writes this loop into CONTRIBUTING.
One idea did not work. We split compiler.rs into 24 files, and the compiler's output stayed byte-identical. Then we built the t27c crate in 6 interleaved rounds. The median went from 49.6 s to 37.8 s and the minimum from 34.1 s to 32.1 s, while CPU time rose from 95.9 s to 98.4 s. The split build was faster in only 3 of 6 pairs, and single builds ranged from 32 to 117 s. That is noise, not a speed-up. The branch stays local and is not a PR.
t27b is a separate crate, cli/t27b, merged in #5979 (a follow-up, #5989, also merged, registers it in the CI ledger of orphan crates). It emits AArch64 machine code itself, with no LLVM, zig, clang or rustc. Its front end is t27c's own parser and type checker, used unchanged. After that the code is new: an IR of its own, an A64 encoder, and then either execution in memory (t27b test, a JIT) or a Mach-O object file (t27b build). Integer overflow is defined: it either traps or wraps (--overflow trap|wrap), and a shift by an amount outside [0, width) traps. A reference interpreter of the same IR is the oracle. The JIT runs only on arm64 macOS.
These numbers come from a different session than the ones above. The machine was heavily loaded, with a load average of 83 to 165 on 8 cores. Absolute times are therefore inflated. Each ratio below compares variants run interleaved in the same session, 5 times each.
| N | load avg | t27b test, ms (median / min) | gen-c + clang + run, ms | longer than t27b (median / min / CPU) | gen + zig test, ms | longer than t27b (median / min / CPU) |
|---|---|---|---|---|---|---|
| 100 | 83–134 | 17.8 / 15.8 | 1,158 | ×65.0 / ×59.0 / ×14.4 | 6,798 | ×381.5 / ×349.2 / ×245.7 |
| 1,000 | 129–161 | 59.6 / 51.1 | 1,189 | ×19.9 / ×15.8 / ×8.6 | 6,124 | ×102.7 / ×61.9 / ×67.4 |
| 5,000 | 98–165 | 916.7 / 250.1 | 4,529 | ×4.9 / ×8.3 / ×6.4 | 21,058 | ×23.0 / ×20.0 / ×24.9 |
Against the Zig path, t27b is 381 times faster at N = 100 and 23 times faster at N = 5,000 (medians). By CPU time the figures are 246 and 25. The C-path median ratio is inflated by load. At N = 100, running ./ctest took 970.3 ms of wall time but only 6.7 ms of CPU, so most of that time was spent waiting. Counting compilation alone, the C path takes 10.5 times as long as t27b at N = 100 and 4.3 times as long at N = 5,000. The load also shows in zig test: at N = 5,000 it took 20.5 s here and 3.4 s in the quieter harness run. That is why t27b is kept off the earlier tables.
t27b build writes a .o the way clang -c does. In the same session, at the median, in ms:
| N | t27b build (trap) | gen-c + clang -O0 -c | gen-c + clang -O2 -c | C by hand, clang -O0 -c |
|---|---|---|---|---|
| 100 | 21.4 | 120.3 (×5.6) | 301.2 (×14.1) | 74.4 (×3.5) |
| 1,000 | 62.9 | 570.5 (×9.1) | 1,780.1 (×28.3) | 190.2 (×3.0) |
| 5,000 | 822.7 | 3,847.7 (×4.7) | 20,388.9 (×24.8) | 1,496.0 (×1.8) |
The comparison is lopsided: clang does far more work than t27b. It does show where t27b sits, though. At N = 5,000 it is faster than clang compiling the hand-written C at -O0.
| Phase, N = 5,000, median of 15 | ms | share |
|---|---|---|
| parse (shared with t27c) | 157.2 | 48.3% |
| type check (shared with t27c) | 90.8 | 27.9% |
| lower to its own IR | 37.8 | 11.6% |
| A64 code generation | 17.1 | 5.3% |
| map into memory and run | 1.6 | 0.5% |
| total, including reading the file | 325.8 | 100.0% |
76.1% of the time goes to the front end that t27b shares with t27c. Its own part, from lowering to running, is 17.3%. Making t27b faster from here means making t27c's front end faster.
| N | t27b trap | t27b wrap | gen-c + clang -O0 | gen-c + clang -O2 |
|---|---|---|---|---|
| 100 | 9.40 | 7.36 | 11.15 | 3.96 |
| 1,000 | 10.36 | 8.47 | 12.46 | 4.93 |
| 5,000 | 10.95 | 8.50 | 13.65 | 5.30 |
Nanoseconds per call, median of 5. All variants return the same checksum. t27b's code is at about the level of clang -O0. In trap mode it takes 0.80 to 0.84 of the time of gen-c with -O0. In wrap mode it is 1.60 to 1.86 times slower than gen-c with -O2. t27b has no optimiser. At N = 5,000 its code section is 400,000 bytes in trap mode and 319,992 in wrap mode, against 963,612 and 523,608 for gen-c with clang at -O0 and -O2.
In a differential test, every function was called on the same inputs from t27b's code and from clang's. Wrap mode was compared with -O2, and trap mode with -O0 plus UBSan traps on overflow. Over 12,261,000 calls there were 0 mismatches. That covers one synthetic program, and it is not a proof.
t27b corpus ran over all 1,174 specs in the repository in 3.6 s. t27b understands 37 of them. 36 pass, and 26 of those have no tests at all; 1 fails, because of a real bug in the spec (below). It refused 1,117 files, and the shared front end failed on 20. There were no mismatches between the JIT and the interpreter, no timeouts and no crashes. The commonest reasons for refusal, counted as the first refusal per file, are StructDecl (393), string literal (287), EnumDecl (89), InvariantBlock (70), type str (31). Structs, strings, enums, invariant blocks and casts are tracked in #5977, which is open. The first coverage PR, #5992, runs invariant blocks like tests and reports them separately; it is submitted, not merged.
The plan from here, planned and in progress with no numbers yet, is to cover every spec that the reference path itself passes, in two lanes. The memory lane adds structs, arrays and slices, and strings, all on one model of an address plus an offset, with read-only data for constants. The scalar lane adds enums, f64 and casts. The order of the work comes from a greedy ranking: next is whichever feature unlocks the most specs that are still refused.
t27b's own clean release build has 17 build units and 8 crates, against 386 and 295 for t27c. Its binary is 1.37 MB against 15.6 MB. Its build time under a load of 244 to 336 means little: the median of 3 was 111.6 s and the minimum 67.2 s. Most of the source it compiles is not its own. It mounts compiler.rs (44,636 lines) and use_resolve.rs (1,010), next to about 5,100 lines of its own code.
Why it is fast is no mystery. It is the same reason Zig's own backend is 5.7 times faster than Zig through LLVM, and the reason behind Copy-and-Patch and TPDE: one pass, no LLVM, code at about the -O0 level.
A second backend with defined semantics is useful just because its answers can be compared with t27c's. That comparison found four bugs. One is fixed so far.
| Bug | How it was found | State on 4 October 2026 |
|---|---|---|
ternary_model.t27: dot27 shifts an i32 by up to 52 bits, and its test expected wrong values | t27b corpus: "shift amount out of range at line 40" | fixed: issue #5972 closed, PR #5975 merged on 4 October |
gen-c writes plain C (x + 1, x >> n). Signed overflow and an oversized shift are undefined behaviour in C, and the tests inc_max and shr_big change outcome between clang -O0 and -O2 | a probe module, three backends side by side | in progress, not yet a PR |
t27c gen (Zig): a signed % does not compile; Zig wants @rem or @mod | the same probe module | submitted, not merged: PR #5993 |
The t27c parser reads module a.b; and use std.testing; as a followed by a stray expression .b, with no line number; 12 specs are parsed this way | t27b corpus: a refusal on "StmtExpr" at line 0 | in progress, not yet a PR |
The second bug matters most. The same spec means different things depending on the backend and the optimisation level. Zig traps on the overflow, and so does t27b. The C that t27c generates silently does whatever the optimiser decides. The Zig fix is PR #5993, submitted, not merged. The gen-c fix and the parser fix are on local branches and are not PRs, so this post records them as in progress.
u32. Real Rust and C++ lose most of their time to generics, templates, macros and imports, and this program has almost none.clang -O2.Work with me
I work contract and part-time on hardware-AI, FPGA/RTL and ML systems — from specification and open toolchains to reproducible measurements.