T27.AI

Blog

Three bits, each one necessary: the MMCM's performance clock on silicon

2026-10-04 · 6 min read

[one clock path on one AX7203 board; LEDs read by eye, no counter readback; built with a nextpnr-xilinx 115 commits behind main, with one merged fix applied locally; the faster loop was timed on one small design] An MMCM output can reach a regional clock buffer through a performance-clock path, and two rows in Project X-Ray's database say which bits switch that path on. With the rows on prjxray-db master, a counter clocked through it does not run on an XC7A200T. With the rows we proposed in openXC7/prjxray-db#30, it runs, and the two bitstreams differ in exactly three bits. Clearing any one of the three stops the counter again, so each is necessary. nextpnr-xilinx main still pins a database with the old rows. Along the way, the edit-to-board loop went from about 104–111 s to about 19–28 s, mostly from a router flag, the t27-spec assembler and a faster JTAG clock.

#FPGA#OpenToolchain#Verification

What was measured

An MMCM in a Xilinx 7-series clock-management tile can hand an output to the regional clock buffers through a performance-clock path. The path is two muxes: one in the CMT tile (CLK_PERF0 to CLK_PERF3) and one in the HCLK_CMT tile above it (PERFCLK0 to PERFCLK3). Project X-Ray's database has rows that say which configuration bits select each input. For the path used here, the rows on prjxray-db master are short. CLK_PERF2.CLKOUT3 is 29_994 alone, and PERFCLK2.MUXED2 is 28_176 29_133. The rows we proposed in openXC7/prjxray-db#30, at commit c030ed6, add 28_994 and 29_1022 to the first and 28_180 to the second.

Those rows came from a population. cavearr ran a campaign of 1083 Vivado specimens, and we re-derived its numbers with our own bit reader before writing the rows. A population says which bits move together, but it does not say a design runs. This post is the board half.

The probe

The board is an ALINX AX7203 (xc7a200tfbg484-2). The 200 MHz board clock goes through IBUFDS and BUFG into an MMCM. The VCO runs at 1000 MHz and CLKOUT3_DIVIDE = 10, so CLKOUT3 is 100 MHz. CLKOUT3 goes through CLK_PERF2 and PERFCLK2 <- MUXED2 into a BUFR in bypass mode, and the BUFR clocks a 27-bit counter. A 28-bit counter on the BUFG is the reference. The top bit of each counter drives an LED, so both LEDs should toggle every 1.34 s, since 2^27 / 100 MHz = 2^28 / 200 MHz. A third LED blinks fast while the MMCM reports LOCKED.

The counter sits in ordinary fabric, not in the I/O column, so the clock has to leave the BUFR through the horizontal clock spine. An earlier post used the same trick for its one-bit A/B, and this probe also depends on the ENABLE_BUFFER bit that post verified.

The A/B: one route, two databases, three bits

One place-and-route run produced one FASM file, and fasm2frames assembled it twice. The first time was against prjxray-db master 517d66a, and the second against the same checkout with the eight files of c030ed6 copied in. Past the sync word, the two bitstreams differ in exactly three bits, and all three are PERF bits: 28_994 and 29_1022 in the CMT tile, and 28_180 in the HCLK_CMT tile.

rowsreference LEDBUFR counter LEDLOCKED
master 517d66ablinks, ~1.3 ssteadylocked
c030ed6blinks, ~1.3 sblinks in step with the referencelocked

The LEDs on this board are active-low, so "steady" means a counter frozen at its reset value. With the master rows, the MMCM locks and the reference runs, but no clock reaches the BUFR.

Knockouts: each bit is necessary

The counter runs with all three bits set, but that does not show each one is needed, since one could be a passenger. So we took the working frames, cleared exactly one of the three bits, and wrote a bitstream for each, with no new place-and-route. Past the sync word, each knockout differs from the working bitstream in three bytes: the bit itself and the frame's ECC word.

bit clearedrow it belongs toBUFR counter LED
none (c030ed6 rows)blinks in step with the reference
28_994CLK_PERF2.CLKOUT3steady
29_1022CLK_PERF2.CLKOUT3steady
28_180PERFCLK2.MUXED2 (used bit)steady
none, re-flashed after the knockoutsblinks in step with the reference

Clearing any one bit stops the counter, and the working bitstream still ran when flashed again afterwards, so the board did not change state in between. 28_180 is a bit that our own earlier version of the rows left out. It is the HCLK "used" bit, (26+p)_180 for PERFCLKp. We found it was missing only when our check started counting every bit in a fixed set, instead of only the bits that some row mentioned.

perf2.py diff · knockout · the three CLK_PERF2 bits

The recorded run, 2026-10-04 at 02:53 UTC+7, after the board runs: perf2.py diff names the three bits that separate the two A/B bitstreams, and no master row uses any of them. perf2.py knockout rebuilds the three knockout frames files, each byte-identical to the one that was flashed, and each knockout .bit differs from the working one in 3 bytes past the sync word. The LEDs are not in the recording. Every byte printed is real; the prompt and the typing are staged, and silences over 2 s are shortened. Open the recording on its own page.

The decode agrees with the board

The database is read in two directions: fasm2frames encodes features into bits, and bit2fasm decodes bits back into features. We decoded both bitstreams with both databases and kept the PERF features of the two tiles.

bitstreamdecoded with masterdecoded with c030ed6
master rows (counter dead)CLK_PERF2.CLKOUT3, PERFCLK2.MUXED2none
c030ed6 rows (counter running)CLK_PERF2.CLKOUT3, PERFCLK2.MUXED2CLK_PERF2.CLKOUT3, PERFCLK2.MUXED2

With the master rows, the dead bitstream decodes as if the path were configured. Each master row is a subset of the real one, and all of its bits are present. A decoder built on those rows therefore reports a working clock path that the silicon does not have. With c030ed6, the decode matches the board in both cases.

What the open toolchain ships today

prjxray-db#30 is open. nextpnr-xilinx main pins prjxray-db at 6b8695e, which is cavearr's prjxray-db#13, merged on 7 September. At that commit, CLK_PERF2.CLKOUT3 is 29_994 and PERFCLK2.MUXED2 is 28_176 29_133, the same rows as master 517d66a, which left the counter dead here. We read the rows at that commit. We did not build nextpnr-xilinx main or run the probe through it.

How the build was set up, and what was stale

The router did not choose this path. Left alone, nextpnr routes CLKOUT3 through CLK_PERF0, which has no rows on master at all. A pip blacklist removed the alternatives, so the route had to take CLK_PERF2 <- CLKOUT3 and PERFCLK2 <- MUXED2.

We placed and routed with classic nextpnr-xilinx. The local himbaechel chip database for the xc7a200t had been generated from a db without the PERF rows, so those pips had no configuration bits and were dropped.

The classic binary was stale. It was built from b608fd2c, which is 115 commits behind openXC7/nextpnr-xilinx main. We found that out only after the first board runs, through two symptoms:

The lesson we wrote down is to check git log HEAD..origin/main before blaming a tool. An earlier draft of this report said classic nextpnr-xilinx had never received a fix that it had in fact received, and the fix was ours.

A faster loop

Each knockout needs a flash and a look at the LEDs, so the time from an edit to the board matters. At the start, one cycle took about 104–111 s. By the end, it took about 19–28 s. Synthesis is not in the table, because it was not timed here.

stepbeforeafterwhat changed
place and route49 s13.2 s--router router1 instead of router2
FASM to frames37–39 s0.43–1.72 sbitwalk, the t27-spec assembler, instead of fasm2frames
frames to bitstream0.8 s0.8 sunchanged (xc7frames2bit)
flash to SRAM17.5–22.6 s4.5–12.2 sopenFPGALoader --freq 30000000
total~104–111 s~19–28 s

Router2's log says it spent 38.1 s, but the per-net times it reports add up to about 1.5 s, so most of that time went to something other than routing nets. Router1 finished routing in 2.53 s. The two routes differ in 225 lines of FASM but take the same PERF path. We flashed the router1 bitstream, and the two LEDs blink together. Both routes report a maximum frequency above 270 MHz on both clocks, against the default 12 MHz target, so timing was never under pressure.

bitwalk wrote frames byte-identical to fasm2frames for both databases. The knockouts skipped place-and-route entirely: a frame edit, frames to bitstream, then a flash, which takes about 6–14 s. At 30 MHz, the flash took anywhere from 4.5 to 12.2 s for bitstreams of the same size, and we do not know why.

One more cycle was recorded after the board work. Its place-and-route took 5.95 s, not 13.2 s, and its frames-to-bitstream step took 2.11 s, not 0.8 s. Other jobs kept the laptop at a load average of about 110 the whole time, so a single timing here moves by that much, and neither number is a better estimate than the other.

perf2.py loop --flash · place and route to SRAM, timed

The recorded run, 2026-10-04 at 02:54 UTC+7: perf2.py loop --flash places and routes the probe with router1, assembles it with bitwalk and xc7frames2bit, checks that the result is identical past the sync word to the bitstream that ran on the board, and loads it into SRAM. In this recording: place and route 5.95 s (router1 reports 0.78 s of it), FASM to frames 1.14 s, frames to bitstream 2.11 s, SRAM load 9.85 s, 19.06 s in total, with other jobs holding the laptop at a load average of about 110 on 8 cores. Synthesis is not in it. Every byte printed is real; the prompt and the typing are staged, and silences over 2 s are shortened. Open the recording on its own page.

The previous post said nothing there pointed to a faster place-and-route. This is not one either, in the sense that post meant: it is a router flag, measured on a design with a few hundred wires. We did not measure whether router1 holds up on the 121,587-line design from that post, which takes 70.1 s to place and route.

What this does not show

What this does not settle

Receipts

Work with me

Want this kind of check on your own design?

I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.