Blog
[measured on the 391-item strict t27 bench: 0/150 flat vs 3/150 masked on the same 30 hardest items; 65% of flat first compile errors are undeclared identifiers; strict loop closed at 30 rounds / 8083 samples = 4/391 pass, 283/391 compile; three decode constraints, zero parameters; benched checkpoint = t27-domain finetune of tern_tc, the package ships the base] tern_tc, a 9M ternary model sized to XC7A200T block RAM, cannot learn the names of a scope -- so instead of scaling it we constrained its decoder: a scope mask at every decode step, diversity by mask rule rather than temperature (15/30 vs 3/30 union compile), and a repeat-n-gram ban after 1517/2177 blocked candidates turned out to be repetition cascades. The result is not a solver: 1.0% of the strict set passes. It is a draft model for a verifier -- 72% of the strict set gets a compiling body, tri tc-draft returns the first test-verified one or an honest refusal.
tern_tc is a 9M-parameter ternary model (weights in {-1, 0, +1}), six layers, width 320, 8K vocabulary, sized to the block RAM of an XC7A200T. Its job is to fill in function bodies in t27 specs, where the spec's own tests judge the answer. On the strict half of that bench -- 391 functions whose tests catch both constant mutants -- flat sampling scored 0 passes in 150 samples on the hardest 30 items, and 65% of the first compile errors were 'use of undeclared identifier'. The model does not know the names, and at 9M parameters it cannot. This post is about what we did instead of scaling it: we constrained the decoder. Provenance, so nothing hides: every number below comes from one checkpoint -- this architecture after a t27-domain finetune, the base model continued for about three epochs (22M tokens) over the t27 spec corpus itself. The published package ships the base checkpoint and its card claims no quality; the finetuned bench checkpoint has since been published separately -- playra/tern-tc-9m-t27, with the same kind of honest card.
The C engine speaks a step-io protocol: it emits one token, then accepts one mask line ('.' free, '+' allow-list, '-' ban-list). An external masker -- the same code that knows the spec's scopes -- computes which identifiers are grammatical at that position: the function's own signature, sibling functions, module constants declared before it, and names the body itself has already declared. Inside a string or comment, and after a dot, the mask lifts. Mid-identifier, only continuations that lead to an allowed name survive. 'Declare before use' stops being a hope and becomes a property of the decode. The ban is -1e30 on the logits, before argmax and before top-p, so greedy and sampling obey it alike.
On the same 30 items where flat generation was 0/150, masked generation scored 3 passes in 150 samples -- 2.0%, the model's first nonzero result on this bench -- and a generate-judge-retry loop turned that into 3 of 30 items solved at 149 samples, each verified by the spec's own tests.
The loop needs different samples each round. The conventional lever is temperature. We measured both levers over the same historical artifacts: three temperatures (0.7/0.8/1.1) at one mask rule union-cover 3 of 30 items at compile; three mask rules at one temperature union-cover 15 of 30. Every pass in that set went to temperature 0.7. Temperature shuffles the same mistakes; changing the rule changes the legal set itself.
When the full strict loop stalled at 2 passes in 20 rounds, the post-mortem found why the candidates never compiled: 1517 of 2177 blocked_codegen candidates were repetition cascades -- r.unshift(r.unshift(... over and over. The model walks into a loop at temperature >= 0.7 and never walks out. The fix is the same discipline as the mask, aimed at the second failure mode: ban any n-gram the continuation already contains (-1e30 before sampling, --no-repeat 4). Both late passes in the final run arrived after that ban went in. Scope masking, rule diversity and repeat-banning are three constraints on one decoder, and none of them cost a parameter.
The full loop closed at 30 rounds, 8083 samples, about 21 per item. Four passes out of 391 (1.0%) -- rounds 3, 7, 26 and 27 -- and each went to a different arm (s0.7, w0.7, w0.8, m0.7), which is the rule-diversity claim made concrete. The other number matters more: 283 of 391 items (72%) produced at least one compiling body. The masked 9M model can speak the language of nearly three quarters of the strict set; the wall is semantics, not syntax. Its role is a draft model for a verifier, packaged as one command: tri tc-draft takes a spec and a function signature, pays the prefill once for N masked samples, judges each with t27c, prints the first body that passes the spec's own tests, and exits 1 with a status table when nothing does. No unverified body is ever presented as an answer.
For scale, the 100M fp model on the same bench scores 7.4% pass@10 while its ternary twin scores 2.3% -- both after the same t27 finetune, so the comparison is apples to apples -- and the gap grew from 1.5x to 3.2x as data quadrupled. Ternarity is expensive for code; that result is published too, and this post is not an argument against it -- it is the honest account of what a board-sized ternary model is actually for.
The weights that ship are the base checkpoint: the package, its model card, the bench charts and the CI contract are generated from t27 specs that carry their own tests, and a public verify workflow re-hashes the package, rebuilds the C engine, reproduces the board generation receipt (1 3 204 276 405 659 85 1516) and checks bit parity with PyTorch -- green against the bytes on the hub today. The bench numbers in this post are from the finetuned checkpoint described above, not from the shipped base. That finetune is published: https://huggingface.co/playra/tern-tc-9m-t27 -- same C engine, same tokenizer, one command sh scripts/verify.sh checks the hashes and reproduces the cross-engine receipt (C == PyTorch CPU == MPS); its card carries the same numbers as this post, with the protocol for each.
Work with me
I audit RTL and build independent, bit-exact models, then take the result through synthesis and, when useful, onto an Artix-7 board. The first conformance module is free.