Blog
[simulated: reviewer, concurrency, turn stop, dispatcher; measured: node-link chaos and waits on PostgreSQL 16, decision replay, the simulation gate; FPGA: synthesis only; nothing in production] On one seeded input the Queen's actor runtime cuts the worst recovery from 1181 s to 27 s and finishes 18.7% more reviews under overload, but it does not win everywhere. Seven changes from a competitor study landed behind flags on 2026-10-09, each with a t27 spec and a benchmark against the code it would replace: a decision log that replays 73,435 decisions with 0 mismatches, turns that really stop, waits as Postgres rows (GitHub reads 180 to 24), a fenced node link (lost mail 21 to 0), and a seeded simulation gate that runs in 33.5 s and found a real runtime bug first. Most of the dispatcher's gain came from a 60-minute turn bound, not from actors, and adaptive concurrency loses with 3 lanes. On an XC7A200T the card as written costs 26,737 LUT and 13 functions do not synthesize; memory, not logic, caps it at about 37,000 actors a board.

On one seeded input, the Queen's actor runtime beats the old polling loop where work goes wrong: the worst recovery after a crash or a hang falls from 1181 s to 27 s, and under overload it finishes 18.7% more reviews. It does not beat the loop everywhere. On 2026-10-09 seven changes from a competitor study landed on the integration branch actors-next, each with a t27 spec, tests and a benchmark against the code it would replace. Two of them lose in cases we can name, one gives back throughput on purpose, and most of the dispatcher's gain turned out to come from a plain time limit, not from actors. The new simulation gate found a real runtime bug on its first runs. Each new behaviour sits behind a flag that is off by default, and the bug fixes need none. None of it runs in production yet.
Two kinds of number appear below, and each is labelled. Simulated means a seeded run on a virtual clock, with fault rates and task lengths that are assumptions. Measured means real OS processes, a real PostgreSQL 16 or a real synthesis tool.
The Queen is the scheduler of the t27 agent swarm. It hands GitHub issues to worker agents, called bees, reviews their pull requests and runs release jobs. Since 2026-10-08 the owner's rule is that every concurrent part of it becomes an actor, with a pid, a mailbox and a supervisor, and that no loop is swapped out before an MVP, its tests and a benchmark on the same input have been posted (t27#7851). Every decision the runtime takes is a call into a compiled t27 spec, the "card". The TypeScript host keeps state and does I/O, nothing more.
The first actor domain was the pull-request reviewer (trios#1698). The workload had Poisson arrivals, 5% of attempts crashing and 2% hanging for 30 minutes, over 7 simulated hours. With no faults the loop and the actors tie: done p95 is 319 s against 309 s. With faults the actors cut done p95 from 553 s to 384 s and the worst recovery from 1181 s to 27 s. At 120 arrivals an hour they finish 565 of 660 rows against 476. A message costs 3.7 us on a ring of 1000 actors. [simulated]
Production then showed what the simulation had not modelled. The actor reviewer went live at 19:00Z on 2026-10-08. It visited 22 rows whose branch cannot be read about 184 times a minute, where the loop had visited them about 20 times. A backoff fixed it within the hour (trios#1702). In the first 10.5 minutes after it, the reviewer was busy 23% of its worker-seconds (584 of 2520). A day later the live Queen showed about 11 reviews an hour against about 58 finished dispatches (reviewer.t27). Nothing in the runtime could say which reading was right. [measured, production]
The study compared the runtime with BEAM/OTP, Akka, Orleans, durable-execution engines such as DBOS, Temporal, Restate and Hatchet, and actor-based agent frameworks. It found three gaps, and none of them is a missing BEAM feature. Akka and Orleans measure mailbox depth, time in the mailbox and long turns out of the box; we had log lines. OTP, Orleans and Effect stop work by a protocol: a signal, a deadline, finalizers, then a kill. We abandoned a turn and let it run on. DBOS, Hatchet and Cloudflare make a long wait a database row that holds no process.
It also found where the runtime was already ahead. A node lease of 20 s with a 5 s heartbeat notices a dead node sooner than BEAM's default net_ticktime (45 to 75 s), the default of Akka's split-brain resolver (about 45 s) or Orleans 9 (90 s). Our control lane, which serves signals before data, is the idea OTP 28 shipped as priority messages. A hung turn is killed, where Orleans only reports it (MaxRequestProcessingTime). These are the vendors' documented defaults, not our measurements.
The seven items are the epic trios#1712. Each spec was merged in t27 first, and each runtime landed on actors-next as its own PR. The table gives the headline of each; the sections after it give the price.
| # | Change (spec, runtime, flag) | What moved | Evidence |
|---|---|---|---|
| 1 | Telemetry and a replayable decision log (t27#8285, trios#1725, TRIOS_QUEEN_ACTORS_TELEMETRY=on) | 73,435 card decisions replayed with 0 mismatches; the same schedule with telemetry on and off | measured replay; cost estimated |
| 2 | Turns that really stop (t27#8281, trios#1720, TRIOS_QUEEN_TURN_STOP=on) | reviews at once during stalls: up to 7-10, now at most 4; grandchild processes alive after a kill: 10 of 10, now 0 | simulated; process test measured |
| 3 | Concurrency read from the backlog (t27#8275, trios#1716, TRIOS_QUEEN_REVIEWER_ADAPTIVE=1) | overload: 565 of 660 rows at 80.7/h, now 660 of 660 at 94.3/h | simulated |
| 4 | Long waits as rows (t27#8272, t27#8282, trios#1722, TRIOS_QUEEN_WAITS=rows) | one 90-minute CI wait: 180 GitHub reads, now 24; 180 job-log entries, now 0 | measured on PostgreSQL 16, virtual clock |
| 5 | Seeded simulation as a CI gate (t27#8292, t27#8306, trios#1726) | 18 cases of 10,000 steps, each run twice, in 33.5 s on CI with 0 divergences; 4 re-created defects caught, and one real bug found first | measured (the gate's own runs) |
| 6 | Node link fencing and incarnations (t27#8278, trios#1724, no caller yet) | mail lost: 21, now 0; pids reused: 5 of 5, now 0; messages from a frozen node: 40, now 0 | measured on PostgreSQL 16, OS processes |
| 7 | Keyed actors and the bee dispatcher (t27#8286, trios#1723, TRIOS_QUEEN_DISPATCH=actors) | crash recovery p50 with 6-minute rounds: 701 s, now 149 s; duplicate claims on two nodes: 0 | simulated |
Item 1 counts per actor kind, keeps turn-time and mailbox-time histograms, and raises threshold events with a gap, as OTP 27's long_message_queue monitor does. A mailbox-depth alarm turns on at 192 and off at 64 of its 256 slots. A long turn is half the kill bound, 150 s for a review. A restart storm is flagged one restart before the supervisor gives up. Records are sampled: 8 messages per actor kind and 16 calls per card function each second, into a ring of 4096.
Because every decision is already a call into a compiled spec, logging those calls makes the runtime replayable. A seeded reviewer run logged 73,435 decisions from 38 functions on 3 cards; replayed against the pinned cards, they gave 0 mismatches. A corrupted, empty, foreign or unsampled log does not pass. Of 13 injected crashes the counters saw 12: the 13th ended after a rest_for_one restart had already replaced its worker. [measured]
On a quiet host (load about 4; 1000 actors passing 100 000 messages, 21 rounds) telemetry adds +3.3% CPU per message, the median, with +1.6 to +5.2% between p25 and p75. That is inside the +10% target. The same run found a cost we did not want: with telemetry off, a message cost 4.5 us against the 3.7 us published before the seven lanes. A profile (trios#1746) showed the cost was older than the lanes: 44% of the time went to looking a card up again on every call, by a key of 100-odd characters, and a slot known since spawn was asked for three times a message. With both fixed, a message costs 1.55 us, 2.4 times below the old base; that fix went to production with the second batch (trios#1754). Telemetry costs the same in absolute terms, so on the cheaper base it adds 8 to 13%, at the edge of its +10% target. A later change (trios#1755) cut telemetry's own instructions per message by 27 to 33%; the ratio on a quiet host is still to be measured. There is no production reading yet, so the question of 23% busy against 11 reviews an hour is still open.
Item 2 gives every turn an AbortSignal. The model call and the git and criterion commands obey it, a stopped review writes no verdict, and an escalation kills the whole process group, never one pid. Before this, a review killed at its bound kept its z.ai lane until its own 120 s timeout. In the simulation, the most reviews at once during stalls fall from 7-10 to 4, and a lane is free 0 s of virtual time after the kill instead of 48-74 s (p50); the real abort of a model call took 2.6-27.5 ms. [simulated]
A test with real processes: 10 turns, each leaving a sleep grandchild, half of them deaf to SIGTERM, killed at a 1 s bound. With the old proc.kill(9) to one pid, all 10 grandchildren were alive 10 s later. With the group kill none were, and a turn deaf to SIGTERM was freed in 3012 ms. [measured]
The price shows under overload. With deaf stalls, 78.4 reviews an hour fall to 60.7; with stalls that hear their abort, 76.6 fall to 68.4. The old code was not faster: it ran past its bound. In the second case it ran 12,216 review-seconds beyond the bound, about 68 reviews at the median length, against a gap of 57. The simulation has no key-lane limit, so there a review past the bound is free. In production it is a third concurrent request on a key, which z.ai refuses with error 1302. At steady load the throughput is equal. [simulated]
Item 3 replaces the fixed pool of 4 reviewers with one sized from the backlog, the free model lanes and the memory (reviewer_sizing.t27). With 47 free lanes, which /queen/status showed on 2026-10-09, it clears an overload: 660 of 660 rows at 94.3 an hour, against 565 of 660 at 80.7. A burst of 120 rows is done in 2895 s instead of 11,690 s. The price is more model calls at once: up to 16, against 7. [simulated]
It loses in two cases. With only 3 lanes, the fixed pool finishes 429 of 660 and the adaptive one 360. A review holds its lane only for the model call, so a pool capped at the lanes leaves each lane idle for the rest of the review. In a container with 1.5 GB free and an assumed 512 MB per review, the pool is capped at 3 and finishes 399 against 565. The flag stays off until telemetry has measured free lanes and memory per review. [simulated]
Item 4 parks a long wait as a row in Postgres, queen_wait, woken by a scheduler actor. It starts with a correction to the study. The release job's 90-minute CI wait did not live in process memory: the job was already a row, so a restart lost nothing. What it cost was a GitHub read and a job-log entry about every 30 s.
Over that wait, against PostgreSQL 16 on the virtual clock, GitHub reads fall from 180 to 24 and job-log entries from 180 to 0. No wait is lost across a forced restart in either version. Noticing that the CI run has ended now takes 83 s instead of 23 s, with a bound of 240 s instead of 30 s, because a 15 s poll replaces the round. Process time without GitHub latency rises from 1108 ms to 1463 ms, about half of it that shared poll. Counting GitHub's measured 51 ms a read, the old way spent 9.2 s waiting on GitHub and the new one 1.2 s. A webhook would close the detection gap; none is wired. [measured]
Item 5 turns the virtual clock into a CI gate (simulation.t27, t27#8292 and t27#8306; harness trios#1726). It drives the real runtime on three nodes: the reviewer actors, remote children, the in-memory net, and on two seeds the real Postgres link over a simulated store. The card picks the seeds and the fault mix: node loss, hung turn, mailbox overflow, unreadable row, provider 429, stale lease, review crash. There are 18 cases, 16 on the memory net and 2 on the store, of 10,000 steps each. Every case runs twice and the two logs must agree, and the cases together must reach the eight rare states the card requires. On CI the gate takes 33.5 s, and the whole job 52 s. Four gate runs in a row on the final code gave 0 runs that logged differently. [measured]
To check that it catches what it should, each known defect was put back into the real source and the gate run again:
| Defect put back | Cases failed | First failure |
|---|---|---|
| No wait backoff (the go-live hot loop) | 18 of 18 | seed 3600507402, step 421: one row visited 5 times in a virtual minute |
| The node link acknowledges mail before handling it | 2 of 2 store cases | seed 3600507402, step 2637 |
| Pids without the incarnation | 18 of 18 | seed 3600507402, step 2543 |
| A turn runs after its process stopped | 2 of 18 | seed 3600507402, step 7690 |
The defect in the last row was not one we knew about: the gate found it. On its first runs the gate failed 5 of 64 seeds with a turn that ran for a process that had already exited: the runtime takes a message and runs the turn a microtask later, and in between a supervisor can stop the process. queen-actors.ts now runs a turn only while its process is current, with a regression test. On a loaded Mac the check cost nothing above the noise (median 11.4 us per message with it, 11.65 us without). And on the code before item 6, the simulated store reproduced both node-link defects with no change to the source.
The gate has two limits. Keyed actors and the adaptive reviewer pool are not in its simulated world yet. And nothing runs it nightly, because GitHub runs scheduled workflows only on the default branch, where the Queen's code does not live.
Item 6 fixes the Postgres link between nodes before anything uses it; in production it has no caller. netlink.t27 puts an incarnation number in every pid and checks the lease epoch on every write. It moves the heartbeat to its own thread, and makes a node fence itself after 15 s without a renewal: one heartbeat before any peer can see it down at the 20 s lease. The same chaos script ran before and after, against a real PostgreSQL 16, with each node a separate OS process. [measured]
| Chaos case | Before | After |
|---|---|---|
| Mail lost, 10 read answers dropped mid-delivery (of 200) | 21 | 0 |
| Mail lost behind one undecodable row (of 20) | 10 | 0 |
| Pid reused across 5 restarts | 5 of 5 | 0 |
| Old-incarnation mail delivered to the new one | 25 of 25 | 0 |
| Messages from a node frozen 30 s, after its peer saw it down | 40 | 0 |
| False node-down while one turn held the loop for up to 25 s | 1 | 0 |
Each of the six defects was reproduced red before the fix. The price is the acknowledgement: 6.16 mail statements per round trip instead of 4.11. CPU moved 7 to 8%, inside the run-to-run spread of a Mac at load 21 to 42. The recording below runs the fence rule itself: the spec passes, the fence is moved from 15 s to the 20 s lease, exactly that test fails, and git restores the file.
t27c test-report specs/queen/netlink.t27 · fence before the peers decide
Item 7 runs the bee dispatcher as one actor per issue, with links, call and an orderly stop. Its benchmark added a control: the old loop with its runners capped at the actors' 60-minute turn bound. At 16 lanes with faults the loop finishes 262 issues, the capped loop 349 and the actors 352. The actors add 3, or 0.9%; the cap gave the rest. [simulated]
The actors win where rounds are slow and where bees crash. With 6-minute rounds they finish 330 against 307 (+7.5%), and crash recovery p50 falls from 701 s to 149 s. They lose a retry when lanes are scarce: 20 to 27 s for the loop, against 60 to 98 s. On a burst they finish 612 against 616, because they give up an issue after 3 counted failures. Duplicate claims are 0 for every runtime on one node. Across two dispatcher nodes the actors keep 0 with the pid as the claim holder, against 273 in the negative control. The proposal on t27#7851 is not to swap on these numbers: a 60-minute cap for the current loop takes most of the gain at once. [simulated]
The decision card is a pure t27 spec, so t27c can lower it to Verilog. We synthesized it for the XC7A200T of our bench board to see what an actor runtime in hardware would cost. This is synthesis only, with yosys 0.63: no place-and-route, no clock rate, nothing loaded on the board. The provenance is in trinity#1582.
while loop whose exit depends on data, 4 on a name the generated Verilog does not resolve, and 1 on a width yosys cannot detect.actors-next merged into the production branch once, with the owner's OK: trios#1730, on 2026-10-10 at 13:38Z, and a second batch followed at 16:03Z (trios#1754): the review-lane queue, the bounded drain, the keyed hardening, actor events and the cost fix. Every new behaviour in them is behind a flag that is off, so production behaves as before until the flags are switched on one at a time, by data. Until then, every number here comes from a benchmark.runRound, not the function itself.Work with me
I work contract and part-time on hardware-AI, FPGA/RTL and ML systems — from specification and open toolchains to reproducible measurements.