T27.AI

Blog

The FPGA flow, layer by layer: what a rewrite from specs is worth

2026-10-03 · 5 min read

[one design, one laptop at a load of 12.6 on 8 cpus; L1 and L2 figures are Amdahl ceilings, not results; hours a year rest on stated assumptions] We timed every layer of one real XC7A200T build through openXC7. Synthesis took 12.7 s, place and route 70.1 s, FASM to frames 33.9 s and frames to bitstream 0.2 s, 116.9 s in all. With the two layers already rebuilt from t27 specs, the flow takes 83.5 s, 1.40 times faster, and the bitstream is byte-identical. Place and route is now 60% of the build: if it took no time the flow would be 8.75 times faster, while an instant synthesis step caps out at 1.65 times. The numbers are in a recorded terminal session. The tools are not in a tri release yet.

#FPGA#OpenToolchain#t27

The previous post rebuilt the back half of openXC7 from t27 specs. FASM to frames and frames to bitstream now come from the specs, byte for byte. That raises the next question: what would it be worth to do the same for every layer of the flow? To answer it we timed every layer of one real build and put the numbers next to each other.

One build, every layer

The design is trinet_node_v2_ax7203 for the xc7a200tfbg484-2 on an AX7203 board, 121,587 lines of FASM. tri devkit flow --build runs the openXC7 flow with every step timed. It then runs the t27 replacements on the same FASM, three times each, and compares their output with openXC7's byte for byte. The times are wall seconds on one laptop at a load of 12.6 on 8 cpus.

LayeropenXC7 toolNowSharet27Same bytes
L1 Synthesisyosys12.69 s10.9 %not yet–
L2 Place & routenextpnr-xilinx70.09 s60.0 %not yet–
L3 FASM → framesfasm2frames.py33.86 s29.0 %bitwalk --fasm 0.42 syes
L4 Frames → .bitxc7frames2bit0.22 s0.2 %bitwalk --write 0.25 syes
Whole flow116.86 s83.45 s (1.40×)

tri devkit · the FPGA flow, layer by layer

The recorded run: tri devkit flow --build and tri devkit impact. Every byte printed is real and arrives when it did; the prompt and the typing are staged, and silences over 2 s are shortened, with a note in the title bar while that happens. Open the recording on its own page.

The rewrite so far saves 33.4 s on this build, all of it in L3. L4 is a tie: bitwalk --write took 0.25 s against 0.22 s for xc7frames2bit. The bitstream is the same file either way, so this is a faster route to the same result, not a better result.

Where the rest of the time goes

After L3, place and route is 60.0 % of the build and synthesis is 10.9 %. Amdahl's law gives the most a rewrite of either could ever give: make the layer take 0 s and see what is left.

If this took 0 sFlowAgainst openXC7 today
L1 Synthesis83.5 s → 70.8 s1.65×
L2 Place & route83.5 s → 13.4 s8.75×

So the next lever is place and route, by a wide margin. A spec-driven synthesis step, even an instant one, caps the whole flow at 1.65×. These are ceilings, not plans. nextpnr-xilinx is a large, mature C++ program, and the realistic path is work inside it (we already send patches upstream) plus spec-checked timing and routing rules, not a rewrite from zero.

What it is worth in hours

tri devkit impact turns the saving into hours. At an assumed 20 builds a day for one person over 230 working days, 33.4 s a build is 42.7 hours a year of waiting. The builds and the people are assumptions, so the page at t27.ai/#/devkit lets you put in your own.

Can I run it?

Not yet. The run above used a local build of tri devkit, and bitwalk is built from an open pull request, gHashTag/t27#5609. Neither is in a tri release, so there is no install command to give here. Both are to ship in the tri release archives with a one-line install; that work is tracked in github.com/gHashTag/trinity/issues/1272, and this post will get the command when a release has it.

What this does not show

What this does not settle

Receipts

Work with me

Need an FPGA/RTL problem taken to measured hardware?

I work contract and part-time on hardware-AI, FPGA/RTL and ML systems — from specification and open toolchains to reproducible measurements.