A loss in bits
You will learn
How the loss -log2(p) prices a guess in bits, and why a perfect guess must cost exactly 0.
Recording pending: it waits on tri test and tri mutate plant from gHashTag/t27#7400, the two commands the recording runs, and until then the widget below is a placeholder that shows no run. Training starts with a number for how wrong a guess was. gft_nll.t27 takes p, the probability the model gave the true class, and returns -log2(p): p = 1.0 costs 0, p = 0.5 costs 1 bit, and p = 0.25 costs 2. on_comb is neg(log2v(p)), and neg flips the sign bit, except that it returns 0 for 0. One header line still says the output is log2(x); the code returns its negative. The browser skips all 3 tests because its runner does not know assert_eq yet; the native t27c runs all 3, all pass, none vacuous. The recording makes neg return 65536 for 0, a zero with the sign bit set, and exactly one test fails, perfect. Every byte in the recording was printed by the command; only the typing is staged.
Try it
In the recording, find the changed line and the test that fails; then in the spec frame find the header line that says the output is log2(x) and compare it with on_comb.

Recording pending: waits on tri test and tri mutate plant from gHashTag/t27#7400. Until then this page is a placeholder and shows no run.
specs/ternary/gft_nll.t27
module GftNll;
// #1764 + GF-T: GF-T negative-log-likelihood (cross-entropy loss for a one-hot
// label) -- given the softmax probability p of the true class, returns the loss
// -log2(p) as a positive GF-T16. Composes the verified gft_log2 primitive with a
// sign flip. log2(x) = (o-40) + log2(1+m/512); the fractional part is
// a Q16 quartic in m, then the integer+fraction real value is normalized
// (fixed->GF-T, flat priority-encoder -> yosys-synthesizable) into a signed GF-T16.
// Accuracy: <=0.008 abs vs true log2 (output-quantization limited). x<=0 saturates.
//
// Input: x positive GF-T16 (u32). Output: log2(x) as signed GF-T16 (u32).
// arithmetic floor-shift by 9 (the poly has negative coefficients, and the
// generated Verilog `>>` on a signed reg is LOGICAL -> would fill 0 for negatives;
// this reproduces Python's floor(x/512) using only non-negative shifts).
fn asr9(v: i32) -> i32 {
if (v >= 0) { return v >> 9; }
return 0 - (((0 - v) + 511) >> 9);
}
// round(65536 * log2(1+m/512)) for m in [0,511], Q Horner with rounded shifts.
fn log2_frac(m: i32) -> i32 {
var p : i32 = 0 - 5528;
p = asr9(p * m + 256) + 21224;
p = asr9(p * m + 256) + (0 - 44443);
p = asr9(p * m + 256) + 94274;
p = asr9(p * m + 256);
return p;
}
// encode a positive Q16.16 magnitude `mag` (value = mag/2^16) as a GF-T16 with the
// given sign. Flat left-normalization so the leading 1 lands at bit 30.
fn encode(sign: i32, mag: i32) -> u32 {
if (mag == 0) { return 0; }
var acc : i32 = mag;
var e : i32 = 0;
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
if (acc < 1073741824) { acc = acc << 1; e = e + 1; }
// acc now in [2^30, 2^31): mantissa = round((acc - 2^30) / 2^21), off = 54 - e.
var mant : i32 = ((acc - 1073741824) + 1048576) >> 21;
var off : i32 = 54 - e;
if (mant >= 512) { mant = 0; off = off + 1; }
if (off < 1) { off = 1; mant = 0; }
if (off > 80) { off = 80; mant = 511; }
return ((sign << 16) | (off << 9) | mant) as u32;
}
fn log2v(x: u32) -> u32 {
if (x == 0) { return ((1 << 16) | (80 << 9) | 511) as u32; }
if ((x >> 16) == 1) { return ((1 << 16) | (80 << 9) | 511) as u32; }
var o : i32 = ((x >> 9) & 127) as i32;
var m : i32 = (x & 511) as i32;
var intpart : i32 = o - 40;
var frac : i32 = log2_frac(m);
var val : i32 = (intpart << 16) + frac;
if (val == 0) { return 0; }
if (val < 0) { return encode(1, 0 - val); }
return encode(0, val);
}
fn neg(v: u32) -> u32 {
if (v == 0) { return 0; }
return v ^ 65536;
}
// cross-entropy loss for one class: -log2(p), p a probability in (0,1].
fn on_comb(p: u32) -> u32 {
return neg(log2v(p));
}
// p = 1.0 -> loss 0 (perfect prediction).
test perfect { assert_eq(on_comb(20480), 0); }
// p = 0.5 -> loss -log2(0.5) = +1.0 (0x5000).
test half { assert_eq(on_comb(19968), 20480); }
// p = 0.25 -> loss 2.0 (0x5200).
test quarter { assert_eq(on_comb(19456), 20992); }
endmodule
All lessons
Module 1 · Lab: our own research
A number format of our own, an honest scoreboard, and a model's tables multiplied on the board.
Module 2 · AI numbers: the MX block
How AI chips keep weights in a few bits: one shared scale per block, the scale byte itself, and what one outlier does to its neighbours.
Module 3 · Ternary weights
Weights that are only minus, zero or plus a scale, the five rules a ternary alphabet must pass, and a test pass that checked nothing.
Module 4 · The Ternary Network Float
A rule the compiler enforces before any test runs, and a 17-bit float whose exponent is four balanced trits.
Module 5 · Arithmetic on signed numbers
Multiply two signed numbers, add them when their signs differ, and do both at once in a multiply-accumulate.
Module 6 · Parts of a neuron
A ReLU that bends at zero, a power of two for softmax, and an argmax that names the answer.
Module 7 · Learning from a mistake
A loss that prices a wrong guess in bits, one step that moves a weight against its gradient, and the hidden layer that XOR needs.
Module 8 · BitNet: ternary networks
A threshold that squeezes a sum back to three values, one neuron that becomes a different function when its weights change, and a neuron that reads its inputs 27 trits at a time.
Module 9 · The ternary MAC as a chip
The 27-trit dot product as wires with no register, the same sum added into a register on every clock, and a small whole network to close the course.