T27.AI

Blog

From a minute to a second: where the time to build t27 went, and where it still goes

2026-10-11 · 8 min read

[measured on one 8-core Mac whose load varied overnight; pairs taken under the same load] One overnight loop measured every place t27 builds and waits, and fixed what the numbers pointed at. A warm t27c frontier went from 55-68 s to about 0.9 s, an unchanged bitstream is found in 0.22-0.39 s instead of 4.3-8.9 s, a board run after approval no longer builds first, and the corpus suite CI runs on every pull request dropped from 221-257 s to 108-140 s. Almost all of the old time went into recomputing what had not changed and into running independent work one item at a time. What is left: about 2 s per spec in zig's relink, 27 s of nextpnr per logic change, about 14 s to load a board, and a 72 s C-header gate.

Title card, no illustration yet, for: From a minute to a second: where the time to build t27 went, and where it still goes
View the title card at full size
#t27#Measurement#FPGA#Verification

Overnight from 10 to 11 October, one loop went through every place t27 builds something and makes its author wait, measured each one, and changed what the measurements pointed at. A warm t27c frontier went from about a minute to under a second, a bitstream that has not changed is found in a quarter of a second instead of four to nine, and the corpus suite that CI runs on every pull request takes half its old time. Thirteen pull requests, all merged; the record is the loop ledger, t27#8737.

Seconds spent waiting, before and after (log scale)beforeafter0.1 s1 s10 s100 s1000 sfrontier, warm551frontier after a t27c rebuild556.3watch: save to a hub spec (16 reruns)34.513.9silicon: unchanged design4.30.39t27c suite (local)394185CI: corpus ratchet step221140
Seconds before and after for six paths, on a log scale. Each pair is the conservative end of its measured range: the slowest "after" against the fastest "before". Measured on one 8-core Mac whose load varied from about 3 to about 140 overnight.

Where the time went

Almost none of it went into work that was new. It went into recomputing what had not changed since the last run, and into doing independent pieces of work one after another. Every fix below is one of those two.

WhereBeforeNowWhat changedPR
t27c frontier, warm55-68 s~0.9 sThe four generated-code hashes of a spec come from a cache keyed on the spec, every spec it imports, and the bytes of the t27c binaryt27#8741
t27c frontier after a rebuild of t27c55-68 s6.3 sThe cache is filled by 8 workers first; the decisions are then made in order, exactly as beforet27#8847
frontier --watch, idle33% of a core3%Each look stats the files and reads nothing unless a stat movedt27#8825
frontier --watch, a save to a spec 16 others import34.5 s13.9 sThe reruns go side by side; specs that share a file name stay on one workert27#8856
t27c silicon, design unchanged4.3-8.9 s0.22-0.39 sTool digests (the 333 MB chipdb among them) come from a cache keyed on the file stat; the Python packages are read from the venv instead of pip freezet27#8808, t27#8814
A board run after /bench approvebuild first, 30-550 sload onlyThe bench agent builds the bitstream while the job waits for approval, without touching the cablet27#8797
t27c suite --corpus-only394-498 s167-185 sAbout 13,000 per-spec runs go 8 at a time; results are folded in file ordert27#8837, t27#8840, t27#8845
The same suite step in CI221-257 s108-140 sSame change, measured on the runnerst27#8837

Three rules did most of the work

  1. Key a cache on everything that can change the output, and on nothing else. The generated code of a spec depends on its imports, because constants are inlined across them, so the key holds the whole import closure. It also holds the binary itself: any rebuild of t27c is a new generator. A key that names too little serves wrong answers; a key that names too much, like a commit of an unrelated checkout, throws good answers away.
  2. Check something cheap before something expensive. A file stat costs microseconds; reading and hashing 333 MB costs seconds. The watch and the digest cache both look at the stat first and read only what moved, and they take the stat before reading, so an edit that lands during a read is seen next time.
  3. Run independent work side by side, and prove it changed nothing. Every parallel phase folds its results in input order, and each one was run both ways, with T27_SUITE_JOBS=1 and with the default: the error output was byte-identical, the JSON summaries equal apart from their timing fields, 259 failures both times.

Defects found on the way

Where the time still goes

PathNowWhere it goesNext step
Edit a spec, get its verdict~2 s per specZig compiles and relinks the test binary on every run; t27c itself takes 0.03 sSkip the relink: about 0.2-0.3 s per spec. It changes the test runner that every seal names, so it needs one full re-mint
Change logic, get a bitstream36.7 s (ternary_link)nextpnr 27.3 s: clock routing searches the whole chip for the BSCAN to BUFG arc, and the router sets up every wire of the 200T before routing a few thousandA patch with byte-identical output makes clock routing 1.8-5x faster (t27#8754); a place-and-route process that stays resident would remove the setup
Bitstream onto the board~13-15 s9.73 MB over JTAG at the default clockA faster JTAG clock on the HS2 cable, about 3 s; needs a test on the board
Approval to the start of a runup to 60 sThe bench agent polls once a minutePoll every 15 s, at about 180 more GitHub API calls an hour
t27c suite~170 sThe C-header gate: 72 s of cc -fsyntax-only on every spec, every runCache its results by content (t27#8900)

Each next step in this table changes something outside the loop: what every seal records, what the board is sent, the GitHub API budget, or a file the project only lets its owner open for hand-written code. None of them was taken without the owner.

Check it yourself

tri hot runs each hot path once to warm it, then once timed, and says OVER when one is past its budget, or ran and failed. It is the regression test for the caches above, and the loop now starts every tick with it.

$ T27_OPENXC7=$HOME/t27/build/fpga/openxc7 tri hot
OK       frontier, warm: 856 ms (budget 5000 ms)
OK       silicon, reuse hit: 242 ms (budget 2000 ms)

What this does not settle

Receipts

Work with me

Want this kind of check on your own design?

I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.