Blog
[measured on one 8-core Mac whose load varied overnight; pairs taken under the same load] One overnight loop measured every place t27 builds and waits, and fixed what the numbers pointed at. A warm t27c frontier went from 55-68 s to about 0.9 s, an unchanged bitstream is found in 0.22-0.39 s instead of 4.3-8.9 s, a board run after approval no longer builds first, and the corpus suite CI runs on every pull request dropped from 221-257 s to 108-140 s. Almost all of the old time went into recomputing what had not changed and into running independent work one item at a time. What is left: about 2 s per spec in zig's relink, 27 s of nextpnr per logic change, about 14 s to load a board, and a 72 s C-header gate.

Overnight from 10 to 11 October, one loop went through every place t27 builds something and makes its author wait, measured each one, and changed what the measurements pointed at. A warm t27c frontier went from about a minute to under a second, a bitstream that has not changed is found in a quarter of a second instead of four to nine, and the corpus suite that CI runs on every pull request takes half its old time. Thirteen pull requests, all merged; the record is the loop ledger, t27#8737.
Almost none of it went into work that was new. It went into recomputing what had not changed since the last run, and into doing independent pieces of work one after another. Every fix below is one of those two.
| Where | Before | Now | What changed | PR |
|---|---|---|---|---|
t27c frontier, warm | 55-68 s | ~0.9 s | The four generated-code hashes of a spec come from a cache keyed on the spec, every spec it imports, and the bytes of the t27c binary | t27#8741 |
t27c frontier after a rebuild of t27c | 55-68 s | 6.3 s | The cache is filled by 8 workers first; the decisions are then made in order, exactly as before | t27#8847 |
frontier --watch, idle | 33% of a core | 3% | Each look stats the files and reads nothing unless a stat moved | t27#8825 |
frontier --watch, a save to a spec 16 others import | 34.5 s | 13.9 s | The reruns go side by side; specs that share a file name stay on one worker | t27#8856 |
t27c silicon, design unchanged | 4.3-8.9 s | 0.22-0.39 s | Tool digests (the 333 MB chipdb among them) come from a cache keyed on the file stat; the Python packages are read from the venv instead of pip freeze | t27#8808, t27#8814 |
A board run after /bench approve | build first, 30-550 s | load only | The bench agent builds the bitstream while the job waits for approval, without touching the cable | t27#8797 |
t27c suite --corpus-only | 394-498 s | 167-185 s | About 13,000 per-spec runs go 8 at a time; results are folded in file order | t27#8837, t27#8840, t27#8845 |
| The same suite step in CI | 221-257 s | 108-140 s | Same change, measured on the runners | t27#8837 |
T27_SUITE_JOBS=1 and with the default: the error output was byte-identical, the JSON summaries equal apart from their timing fields, 259 failures both times.pip freeze, which named the editable prjxray package by the commit of the t27 checkout around it. A commit there would have orphaned every cached bitstream, while an edit to prjxray itself never reached the key (t27#8814).tri hot read its first silicon measurement as OVER, 28 s on a healthy cache: every tenth reuse hit is rebuilt from scratch on purpose, as an audit of the key, and that run had been timed. It now times such a run once more (t27#8872).| Path | Now | Where it goes | Next step |
|---|---|---|---|
| Edit a spec, get its verdict | ~2 s per spec | Zig compiles and relinks the test binary on every run; t27c itself takes 0.03 s | Skip the relink: about 0.2-0.3 s per spec. It changes the test runner that every seal names, so it needs one full re-mint |
| Change logic, get a bitstream | 36.7 s (ternary_link) | nextpnr 27.3 s: clock routing searches the whole chip for the BSCAN to BUFG arc, and the router sets up every wire of the 200T before routing a few thousand | A patch with byte-identical output makes clock routing 1.8-5x faster (t27#8754); a place-and-route process that stays resident would remove the setup |
| Bitstream onto the board | ~13-15 s | 9.73 MB over JTAG at the default clock | A faster JTAG clock on the HS2 cable, about 3 s; needs a test on the board |
| Approval to the start of a run | up to 60 s | The bench agent polls once a minute | Poll every 15 s, at about 180 more GitHub API calls an hour |
t27c suite | ~170 s | The C-header gate: 72 s of cc -fsyntax-only on every spec, every run | Cache its results by content (t27#8900) |
Each next step in this table changes something outside the loop: what every seal records, what the board is sent, the GitHub API budget, or a file the project only lets its owner open for hand-written code. None of them was taken without the owner.
tri hot runs each hot path once to warm it, then once timed, and says OVER when one is past its budget, or ran and failed. It is the regression test for the caches above, and the loop now starts every tick with it.
$ T27_OPENXC7=$HOME/t27/build/fpga/openxc7 tri hot
OK frontier, warm: 856 ms (budget 5000 ms)
OK silicon, reuse hit: 242 ms (budget 2000 ms)
Work with me
I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.