T27.AI

Blog

The back half of openXC7, from a spec: FASM to frames to bitstream

2026-10-03 · 8 min read

[one design on one board, equivalence not a better bitstream; timings on one busy laptop, ratios are rough; the fpga-assembler speed figure is its README's, not ours] openXC7 turns FASM into a Xilinx 7-series bitstream with two prjxray tools, fasm2frames in Python and xc7frames2bit in C++. We rebuilt that half so every bit position, frame address and packet word comes from three t27 specs. Against the two tools it matches byte for byte on 22 of 22 FASM files and 16 of 16 frame files, and it is about 6 to 14 times faster on small real designs, about 85 times on a 121,587-line one, and up to about 200 times on synthetic test files. The comparison found a silent bit-wrap in fasm2frames (f4pga/prjxray#2574), a missing part check in xc7frames2bit (#2573), and a bug of our own on kintex7. On an AX7203 board (XC7A200T), a .bit written this way for a 121,587-line design was byte-identical to openXC7's, loaded into SRAM, and returned 403,200 of 403,200 tagged receipts.

#FPGA#OpenToolchain#Verification#t27

openXC7 builds a Xilinx 7-series bitstream in two halves. The first half is yosys and nextpnr: Verilog in, FASM out. FASM is a text list of features such as CLBLM_R_X3Y0.SLICEM_X0.ALUT.INIT[63:0] = 64'h.... The second half turns that list into a bitstream. prjxray's fasm2frames (Python) looks up every feature's bits and writes configuration frames, and xc7frames2bit (C++) wraps the frames in configuration packets. We rewrote the second half so that every bit position, frame address and packet word comes from a t27 spec, and compared it byte for byte with the two tools it replaces.

openXC7 todaythis branch (gHashTag/t27#5609)FASMfasm2framesprjxray · Python.framesxc7frames2bitprjxray · C++.bitFASMbitwalk --fasmframes.t27.framesbitwalk --writepackets.t27 · far.t27.bit= byte-identical 22/22= byte-identical 16/16yosys + nextpnr produce the FASM; openFPGALoader loads the .bit
The second half of the flow. Each gold box is a driver generated from t27 specs. The dashed lines are the corpus checks against the prjxray tool above it.

What the specs hold

The driver, bitwalk, is Rust. t27c gen-rust generates its rule functions from these specs. The FASM grammar, the feature lookup, file I/O and the text output are written by hand; the bit arithmetic is not. Each spec section cites the prjxray file and line it reproduces (prjxray c9f02d857, prjxray-db 517d66a).

Measured

CheckCorpusResult
FASM to frames, against fasm2frames22 files. 13 real designs: 7 from nextpnr on xc7a100t, and 6 Vivado bitstreams on xc7a35t read back with bit2fasm. 7 synthetic files, 2 of them on kintex7 xc7k325t. 2 files that must be refused22/22 byte-identical. Both refuse files are refused for the same reason
FASM to frames, against fpga-assembler's reference framesThe 5 reference-parity cases in lromor/fpga-assembler#495/5, including the kintex7 pseudo-PIP case and the 83-bit RXCDR_CFG value
Frames to .bit, against xc7frames2bit16 frame files on xc7a35t, xc7a100t and xc7a200t16/16 byte-identical
CRC, against Vivado6 Vivado bitstreamsAll 12 CRC words reproduce
Frame ECC, against Vivado32,520 frames, 752 of them with data0 wrong
Mutation gate20 defects seeded into the three specs, one at a time20/20 caught. A defect counts as caught only if the spec tests and the corpus sweep both fail

The mutation gate earned its place again while this post was being written. Adding kintex7 support changed how the driver counts bits outside a tile, and one seeded defect survived: a tile window one word too long. The spec tests caught it, but the corpus sweep did not, because the new count absorbed exactly the bit the defect moved. The count is now split in two, so a tile that starts inside its frame can never have its own bits outside it, and the defect is caught again.

Speed

0.10 s1.0 s10.0 s100.0 sOne real designxc7a100t, nextpnr, 1,785 lines1.6 s0.20 s · 8×Synthetic test filexc7a100t, 1,049 lines19.3 s0.14 s · 138×Synthetic test filekintex7 xc7k325t, 1,123 lines71.8 s0.34 s · 211×All 22 corpus filesone after another150.7 s3.3 s · 45×fasm2frames (prjxray, Python)bitwalk --fasm (t27 specs)seconds, log scale
Seconds per file, log scale. fasm2frames was timed when each corpus file was built; bitwalk is the best of 3 runs on the same laptop. The machine was busy with other jobs in both cases, so read the ratios as rough.
Filefasm2framesbitwalk --fasmRatio
One real design, xc7a100t, nextpnr, 1,785 lines1.6 s0.20 s8×
Synthetic test file, xc7a100t, 1,049 lines19.3 s0.14 s138×
Synthetic test file, kintex7 xc7k325t, 1,123 lines71.8 s0.34 s211×
All 22 corpus files, one after another150.7 s3.3 s45×

The synthetic files are no longer than a real design, about 1,100 lines against 1,785, yet fasm2frames takes 12 to 45 times longer on them. We have not profiled why, so we make no claim about the cause. We also re-timed fasm2frames on three files at the same load as our runs: 3.4 s, 32.5 s and 104.4 s, which makes the ratios larger, not smaller. The figure keeps the lower ones.

What else does this job

ToolLanguageStepWhere the bit rules live
fasm2frames (prjxray)PythonFASM to framesPython code over prjxray-db
xc7frames2bit (prjxray)C++frames to .bitC++ code
fpga-assembler (lromor)C++FASM to .bitC++ code over prjxray-db. Its README reports about 10 times the speed of fasm2frames; we did not time it here
bitwalk (gHashTag/t27#5609)Rust generated from t27, plus a hand-written driverbotht27 specs over prjxray-db

Found on the way

A differential is only useful when the two sides disagree. Three times they did, and each time the tool we compared against had the defect.

One defect was ours. The kintex7 tile grid has four tiles that start 2 words below their frame, and our driver read the offset as unsigned and stopped. The spec now says how such a tile's bits are placed (seg_shift), with a test.

On the board

Files that match are a claim about files. To check the chain on hardware we built one real design for the ALINX AX7203 (XC7A200T): trinet node 0, 121,587 FASM lines, 20,230 frames. On that FASM bitwalk was again byte-identical to openXC7 for both the frames and the .bit, and took about 0.8 s where fasm2frames took about 68 s, on a host that was busy the whole time. openFPGALoader then loaded bitwalk's .bit into SRAM over JTAG, with no flash write, and the FPGA reported DONE.

Run on the AX7203Result
SRAM load of bitwalk's .bitdone 1 in 16.7 s
All 42 ternary matrices of the model, each answer tagged under the node keyreceipts verified (tag) : 403200/403200, 33,792/33,792 rows bit-exact
One layer with int8 activationsreceipts verified (tag) : 51840/51840, 320/320 rows bit-exact

tri x7-board · t27 back half on the AX7203

A recorded session: tri x7-board compare, load and receipts, run again for the recording. Every byte printed is real and arrives when it did; the prompt and the typing are staged, and silences over 2 s are shortened, with a note in the title bar while that happens. That session re-timed fasm2frames at 33.3 s against 0.43 s, on a machine with a load of 23 on 8 cpus. Open the recording on its own page.

Because the two .bit files are identical, this shows nothing the openXC7 file would not have shown. The claim is equivalence on one design on one board, not a better bitstream. The front half, yosys and nextpnr-xilinx, is openXC7's and unchanged.

What this does not show

What this does not settle

Receipts

Work with me

Want this kind of check on your own design?

I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.