Blog
[one design on one board, equivalence not a better bitstream; timings on one busy laptop, ratios are rough; the fpga-assembler speed figure is its README's, not ours] openXC7 turns FASM into a Xilinx 7-series bitstream with two prjxray tools, fasm2frames in Python and xc7frames2bit in C++. We rebuilt that half so every bit position, frame address and packet word comes from three t27 specs. Against the two tools it matches byte for byte on 22 of 22 FASM files and 16 of 16 frame files, and it is about 6 to 14 times faster on small real designs, about 85 times on a 121,587-line one, and up to about 200 times on synthetic test files. The comparison found a silent bit-wrap in fasm2frames (f4pga/prjxray#2574), a missing part check in xc7frames2bit (#2573), and a bug of our own on kintex7. On an AX7203 board (XC7A200T), a .bit written this way for a 121,587-line design was byte-identical to openXC7's, loaded into SRAM, and returned 403,200 of 403,200 tagged receipts.
openXC7 builds a Xilinx 7-series bitstream in two halves. The first half is yosys and nextpnr: Verilog in, FASM out. FASM is a text list of features such as CLBLM_R_X3Y0.SLICEM_X0.ALUT.INIT[63:0] = 64'h.... The second half turns that list into a bitstream. prjxray's fasm2frames (Python) looks up every feature's bits and writes configuration frames, and xc7frames2bit (C++) wraps the frames in configuration packets. We rewrote the second half so that every bit position, frame address and packet word comes from a t27 spec, and compared it byte for byte with the two tools it replaces.
packets.t27: the configuration packet stream. Headers, registers, commands, the CRC and the write sequence xc7frames2bit uses.far.t27: the frame address walk for xc7a35t, xc7a100t and xc7a200t, with the part picked by IDCODE.frames.t27: the frame ECC, and where a database bit lands in a frame. That covers the tile's word offset, the alias shift of a *_SING tile, a tile that starts below its frame, and what happens to a bit that falls outside.The driver, bitwalk, is Rust. t27c gen-rust generates its rule functions from these specs. The FASM grammar, the feature lookup, file I/O and the text output are written by hand; the bit arithmetic is not. Each spec section cites the prjxray file and line it reproduces (prjxray c9f02d857, prjxray-db 517d66a).
| Check | Corpus | Result |
|---|---|---|
FASM to frames, against fasm2frames | 22 files. 13 real designs: 7 from nextpnr on xc7a100t, and 6 Vivado bitstreams on xc7a35t read back with bit2fasm. 7 synthetic files, 2 of them on kintex7 xc7k325t. 2 files that must be refused | 22/22 byte-identical. Both refuse files are refused for the same reason |
| FASM to frames, against fpga-assembler's reference frames | The 5 reference-parity cases in lromor/fpga-assembler#49 | 5/5, including the kintex7 pseudo-PIP case and the 83-bit RXCDR_CFG value |
Frames to .bit, against xc7frames2bit | 16 frame files on xc7a35t, xc7a100t and xc7a200t | 16/16 byte-identical |
| CRC, against Vivado | 6 Vivado bitstreams | All 12 CRC words reproduce |
| Frame ECC, against Vivado | 32,520 frames, 752 of them with data | 0 wrong |
| Mutation gate | 20 defects seeded into the three specs, one at a time | 20/20 caught. A defect counts as caught only if the spec tests and the corpus sweep both fail |
The mutation gate earned its place again while this post was being written. Adding kintex7 support changed how the driver counts bits outside a tile, and one seeded defect survived: a tile window one word too long. The spec tests caught it, but the corpus sweep did not, because the new count absorbed exactly the bit the defect moved. The count is now split in two, so a tile that starts inside its frame can never have its own bits outside it, and the defect is caught again.
| File | fasm2frames | bitwalk --fasm | Ratio |
|---|---|---|---|
| One real design, xc7a100t, nextpnr, 1,785 lines | 1.6 s | 0.20 s | 8× |
| Synthetic test file, xc7a100t, 1,049 lines | 19.3 s | 0.14 s | 138× |
| Synthetic test file, kintex7 xc7k325t, 1,123 lines | 71.8 s | 0.34 s | 211× |
| All 22 corpus files, one after another | 150.7 s | 3.3 s | 45× |
The synthetic files are no longer than a real design, about 1,100 lines against 1,785, yet fasm2frames takes 12 to 45 times longer on them. We have not profiled why, so we make no claim about the cause. We also re-timed fasm2frames on three files at the same load as our runs: 3.4 s, 32.5 s and 104.4 s, which makes the ratios larger, not smaller. The figure keeps the lower ones.
| Tool | Language | Step | Where the bit rules live |
|---|---|---|---|
fasm2frames (prjxray) | Python | FASM to frames | Python code over prjxray-db |
xc7frames2bit (prjxray) | C++ | frames to .bit | C++ code |
| fpga-assembler (lromor) | C++ | FASM to .bit | C++ code over prjxray-db. Its README reports about 10 times the speed of fasm2frames; we did not time it here |
bitwalk (gHashTag/t27#5609) | Rust generated from t27, plus a hand-written driver | both | t27 specs over prjxray-db |
A differential is only useful when the two sides disagree. Three times they did, and each time the tool we compared against had the defect.
fasm2frames does not refuse a feature of a site that a *_SING tile lacks. On a bottom SING tile the bits wrap into words 99-100 of the frame, which belong to another tile. On a top SING tile they are dropped. The exit code is 0 in both cases, and one line of FASM reproduces it. Reported as f4pga/prjxray#2574. The same happens on kintex7: one synthetic file has 337 wrapped and 356 dropped bits. bitwalk writes the same bytes by default, which is why the 22/22 above holds; --strict refuses the file and names the line.xc7frames2bit accepts frames for the wrong part: xc7a100t frames with an xc7a200t part exit 0 and write 192 extra frames. bitwalk refuses them. Reported in f4pga/prjxray#2573, with fixes in openXC7/prjxray#27 and #28.RXCDR_CFG value, fpga-assembler's own frames differ from the reference: 18 bits missing, 5 extra. bitwalk matches the reference. Our reading of the source: the parser packs a long binary literal into 64-bit words from its most significant end, so for 83 bits the first word holds only the low 19 bits. Emulating that packing gives exactly the 18 missing and 5 extra bits.One defect was ours. The kintex7 tile grid has four tiles that start 2 words below their frame, and our driver read the offset as unsigned and stopped. The spec now says how such a tile's bits are placed (seg_shift), with a test.
Files that match are a claim about files. To check the chain on hardware we built one real design for the ALINX AX7203 (XC7A200T): trinet node 0, 121,587 FASM lines, 20,230 frames. On that FASM bitwalk was again byte-identical to openXC7 for both the frames and the .bit, and took about 0.8 s where fasm2frames took about 68 s, on a host that was busy the whole time. openFPGALoader then loaded bitwalk's .bit into SRAM over JTAG, with no flash write, and the FPGA reported DONE.
| Run on the AX7203 | Result |
|---|---|
| SRAM load of bitwalk's .bit | done 1 in 16.7 s |
| All 42 ternary matrices of the model, each answer tagged under the node key | receipts verified (tag) : 403200/403200, 33,792/33,792 rows bit-exact |
| One layer with int8 activations | receipts verified (tag) : 51840/51840, 320/320 rows bit-exact |
tri x7-board · t27 back half on the AX7203
tri x7-board compare, load and receipts, run again for the recording. Every byte printed is real and arrives when it did; the prompt and the typing are staged, and silences over 2 s are shortened, with a note in the title bar while that happens. That session re-timed fasm2frames at 33.3 s against 0.43 s, on a machine with a load of 23 on 8 cpus. Open the recording on its own page.Because the two .bit files are identical, this shows nothing the openXC7 file would not have shown. The claim is equivalence on one design on one board, not a better bitstream. The front half, yosys and nextpnr-xilinx, is openXC7's and unchanged.
Work with me
I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.