A neuron in chunks
You will learn
How a neuron with more than 27 inputs adds up its dot products chunk by chunk, and why a test with zero chunks matters.
One packed chunk holds 27 trits, so a wider neuron needs several. neuronN in bitnet_neuron_nchunk.t27 takes up to 8 chunks of activations and weights, adds dot27 for nchunks pairs starting at chunk 0, and quantizes the total. The browser skips all 11 tests because it does not know assert_eq yet; the native t27c runs all 11, all pass, none vacuous. The recording changes the loop bound on line 36 from nchunks to 1, so the neuron always reads exactly one chunk. The other neuron tests fill every chunk alike, so one chunk lands on the same side of the threshold. Exactly one test fails, neuron_zero_chunks, which asks for no chunks and expects Z. Every byte in the recording was printed by the command; only the typing is staged.
Try it
In the recording, find the loop line before and after the change; then in the spec frame find neuron_two_chunks and explain why the same bug does not break it.

t27c on the t27c lab (Railway), spec at t27 5ff0ec512: 11 tests pass natively; stopping the chunk loop after one chunk fails exactly one test, neuron_zero_chunks; git restores the spec.
specs/ternary/bitnet_neuron_nchunk.t27
module BitnetNeuronN;
// Sign-only ternary multiply of two packed trits {N=0b00, Z=0b01, P=0b10}.
fn tmul(ta: u8, tb: u8) -> i8 {
if (ta == 1) { return 0; }
if (tb == 1) { return 0; }
if (ta == tb) { return 1; }
return -1;
}
// 27-trit ternary dot product of two 54-bit packed vectors, via a real loop
// (gen-verilog local-decl hoist, #1741). Result in [-27, +27].
fn dot27(a: u64, b: u64) -> i16 {
var acc : i16 = 0;
var i : u32 = 0;
while (i < 27) {
var ta : u8 = ((a >> (i << 1)) & 3) as u8;
var tb : u8 = ((b >> (i << 1)) & 3) as u8;
acc = acc + tmul(ta, tb) as i16;
i = i + 1;
}
return acc;
}
// Ternary activation re-quantizer: v > +t -> P(2), v < -t -> N(0), else Z(1).
fn quantize(v: i16, threshold: i16) -> u8 {
if (v > threshold) { return 2; }
if (v < -threshold) { return 0; }
return 1;
}
// A BitNet neuron over an arbitrary chunk count: loop the ternary dot product
// across the first `nchunks` (activation, weight) chunk pairs of the packed
// arrays, then re-ternarize with the threshold. Packed-array element indexing
// (#1748) and the loop (#1741) make this fully parameterized. Cross-checked
// against a reference over random packed inputs in tests/bitnet_neuron_nchunk.rs.
pub fn neuronN(acts: [8]u64, weights: [8]u64, nchunks: u32, threshold: i16) -> u8 {
var acc : i16 = 0;
var c : u32 = 0;
while (c < nchunks) {
acc = acc + dot27(acts[c], weights[c]);
c = c + 1;
}
return quantize(acc, threshold);
}
test dot27_all_n { assert_eq(dot27(0, 0), 27); }
test dot27_all_p_x_n { assert_eq(dot27(12009599006321322, 0), -27); }
test dot27_all_z { assert_eq(dot27(6004799503160661, 6004799503160661), 0); }
test quantize_pos { assert_eq(quantize(100, 10), 2); }
test quantize_neg { assert_eq(quantize(-100, 10), 0); }
test quantize_band { assert_eq(quantize(5, 10), 1); }
// Full-neuron coverage over packed-array inputs. Chunk constants: P (+1 lanes) =
// 12009599006321322, N (-1 lanes) = 0, Z (0 lanes) = 6004799503160661. Uses the
// [N]Type{...} array-literal syntax.
test neuron_all_p_x_p { assert_eq(neuronN([8]u64{12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322}, [8]u64{12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322}, 8, 10), 2); }
test neuron_all_p_x_n { assert_eq(neuronN([8]u64{12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 0, 0, 0, 0}, [8]u64{0, 0, 0, 0, 0, 0, 0, 0}, 4, 10), 0); }
test neuron_all_z { assert_eq(neuronN([8]u64{6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661}, [8]u64{6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661, 6004799503160661}, 8, 10), 1); }
test neuron_two_chunks { assert_eq(neuronN([8]u64{12009599006321322, 12009599006321322, 0, 0, 0, 0, 0, 0}, [8]u64{12009599006321322, 12009599006321322, 0, 0, 0, 0, 0, 0}, 2, 10), 2); }
test neuron_zero_chunks { assert_eq(neuronN([8]u64{12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322}, [8]u64{12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322, 12009599006321322}, 0, 10), 1); }
// W699: the hardware boundary, with a width the EMITTER can produce.
//
// W698 accepted `[8]u64` while the port emitter still sized entry ports with
// `type_to_width`, whose last arm is `_ => 32`. This parameter became
// `input wire [31:0]` -- a silent 16x narrowing, and the banner, the census,
// the corpus column and yosys all reported success. It was retracted the same
// wave.
//
// W699 replaced that call with `entry_port_width`, which returns None rather
// than a plausible number and makes the entry point refuse LOUDLY in the
// generated source. Verified end to end: this parameter now emits
// `input wire [511:0]`, and `on_comb` and the function it forwards to both
// take [511:0], so nothing is truncated between the boundary and the body.
fn on_comb(acts: [8]u64, weights: [8]u64, nchunks: u32, threshold: i16) -> u8 { return neuronN(acts, weights, nchunks, threshold); }
endmodule
All lessons
Module 1 · Lab: our own research
A number format of our own, an honest scoreboard, and a model's tables multiplied on the board.
Module 2 · AI numbers: the MX block
How AI chips keep weights in a few bits: one shared scale per block, the scale byte itself, and what one outlier does to its neighbours.
Module 3 · Ternary weights
Weights that are only minus, zero or plus a scale, the five rules a ternary alphabet must pass, and a test pass that checked nothing.
Module 4 · The Ternary Network Float
A rule the compiler enforces before any test runs, and a 17-bit float whose exponent is four balanced trits.
Module 5 · Arithmetic on signed numbers
Multiply two signed numbers, add them when their signs differ, and do both at once in a multiply-accumulate.
Module 6 · Parts of a neuron
A ReLU that bends at zero, a power of two for softmax, and an argmax that names the answer.
Module 7 · Learning from a mistake
A loss that prices a wrong guess in bits, one step that moves a weight against its gradient, and the hidden layer that XOR needs.
Module 8 · BitNet: ternary networks
A threshold that squeezes a sum back to three values, one neuron that becomes a different function when its weights change, and a neuron that reads its inputs 27 trits at a time.
Module 9 · The ternary MAC as a chip
The 27-trit dot product as wires with no register, the same sum added into a register on every clock, and a small whole network to close the course.