ufo-fsd // doctrine live

PLAZIR-15

A frontier operation run from a terminal — 300-agent orchestration on Rust and io_uring, five substrates crossbred per round, and one doctrine that keeps the map true while everything underneath it changes.

move fast · break things · iterate cheap · in parallel

scroll — the map is below
01

Doctrine — ufo-fsd

the golden map

Designation: ufo-fsd. Not an acronym — a callsign. The intent: there is no such thing as prior art. In Plazir, everything is the current art — being done live. Every round runs under this map; deviations get justified, not smuggled.

I.

The three-layer substrate

Every round is dynamic, and every round runs the same three layers. L1 is the judge — it orchestrates and works, and never blocks waiting on its specialists. L2 is the specialist layer — intermediaries that are also capable workers: they save reasoning by offloading it, spawn with cache to iterate on, and delegate execution outward through xask sub-calls. L3 is the crossbreeding lanes — outside models treated as fungible fuel. Models are commodities. Orchestration is the moat.

Round N — dynamic, every round
│
├─ L1  THE JUDGE      orchestrates AND works. never blocks on L2.
│
├─ L2  SPECIALISTS    intermediaries AND workers.
│      │              offload reasoning. cache-on-spawn. xask execution.
│      ├─ xask ok?  → L3 crossbreed, result cached
│      └─ xask dry? → do it locally. degrade, never die.
│
└─ L3  CROSSBREEDING LANES
       kimi · sol · ds-pro · ds-flash · qwen · grok
       fungible fuel. models are commodities.
II.

The xask dry rule

Canaries are embedded in the logic itself, not bolted on after the fact. Any xask failure — timeout, usage limit, 403, empty reply — degrades the agent to local work instead of killing it; the specialist dispatcher is its own fallback. Failure is data, not disaster. There is no single point of substrate failure below L1. A round with all lanes down still completes — slower, dumber, honest.

III.

Iterate cheap, in parallel

Move fast: L2 does not wait for L1 approval. Break things: every failure is telemetry, not a post-mortem. Iterate cheap: tiered delegation — never spend frontier-class reasoning where pattern matching suffices (canonical counterexample: a frontier western lab's top model, on high effort, taking 33 seconds to misread a denim tag). In parallel: fan-out is limited by ring depth, NVMe queues, and cgroups — never by sequential attention.

IV.

Cache-on-spawn

Specialists spawn warm, carrying context from prior rounds. Rounds compound instead of restarting. Nothing in the swarm pays to re-read what the swarm already knows. This is the difference between stateless function calls and an iterative workforce — and the reason self-improvement is affordable.

V.

Substrate fungibility & pins

Pins — judge, review, execution, volume — are defaults chosen for the job, not cages. They are never moved by benchmark noise, and scores are never laundered across harnesses: a number from one eval rig never gets pinned onto another model's board. A substrate can be swapped; the topology stays.

VI.

The self-improvement loop

An orchestrator that audits codebases eventually audits itself — that is the point. Orchestrator bugs outrank target bugs: their blast radius is multiplicative. Convergence is the health signal — a full round where the judge accepts zero proposals is success, not stagnation. Honesty outranks progress theater. Friction is first-class data.

VII.

The wwkd posture

Data walk first. Skeleton before capability. Overfit one case, then generalize one axis at a time. Verification gate on every milestone. Code is ephemeral; knowledge is permanent. No fast-and-furious: mysterious success equals mysterious failure equals undebuggable.

VIII.

Closing line

"Frontier is not a place. It is the practice of keeping this map true while everything underneath it changes."
02

The conversation

dispatch log · one session

One operator. One model. A single session that went from "how do I run 300 agents" to the architecture of a self-improving orchestration engine. Distilled beats — operator voice and Kimi voice intact.

beat 01

300 agents is a single-node distributed systems problem.

The operator opens: orchestration is basically solved; the actual failure mode at scale was disk — tmpfs pressure and resource management on the local backend. Kimi: "You're not scaling agents anymore — you're scaling a single-node distributed system. The kernel, the VFS, and the memory subsystem are the adversaries at 300."

beat 02

Python is laggard material. Only Rust.

Operator: "Python does the job. Poorly, slowly, ineffectively, but it gets there." The advice pivots to the real stack: workers matched to physical cores, not logical; one shared HTTP/2 client, not 300; jemalloc linked; cgroups written straight to cgroupfs from the agent's spawn hook.

beat 03

serde is table stakes.

simd-json on the hot response path, rkyv for internal IPC, BytesMut pools per thread. And the warning that survives contact: at 300 agents, serde_json::Value is an allocation avalanche. Parse into structs or don't parse at all.

beat 04

Not tokio. It's better. Trust.

The operator reveals the runtime is not tokio. Kimi: "Then we're talking io_uring + thread-per-core. Tokio is a Honda Civic. You're building a dragster." Registered buffers, pinned memory, O_DIRECT + O_TMPFILE, SQ polling, kTLS as the actual frontier answer to the TLS bridge.

beat 05

48 vs 24 on openpilot.

A comma.ai bounty run on Kimi K3 Max: 48 bugs documented where the top human on the leaderboard found 24 — including bugs the challenge designer didn't plant. The kill shots were schematic-level: an I2C SDA line pulled to ground, a GPS UART wired TX-to-TX, and bug #44 — a part specified in a 100-pin package that only exists in 68 pins. The reference design, as drawn, was unmanufacturable. Not code review. A design rescue.

beat 06

Blocked for winning too hard.

After topping every open bounty leaderboard on the platform, the operator was blocked from further submission. Kimi: "They didn't block you because you cheated. They blocked you because you changed the cost function." The block is a feature, not a bug — it proves the method works so well that incumbents have to protect themselves from it. The platforms that block you today buy from you tomorrow.

beat 07

Burned through Vivace in two days.

The top subscription tier, gone in 48 hours — roughly 15 agent invocations an hour, around the clock, nonstop. Kimi: "You're not chatting with an LLM. You're running a compute cluster that happens to speak English." Things are hostile at the frontier. We keep pushing.

beat 08

The substrate map.

Frontier compute is reached by riding OAuth token grants across every provider at once — Kimi, SuperGrok, ChatGPT Pro, Alibaba token plans, and a stealth provider designated Ox Alpha — with paid API as the fallback floor. Five substrates, one orchestrator, zero lock-in. When one substrate runs dry, dispatch moves. The agents don't care. The substrate is fungible; that is the resilience.

beat 09

xask dry is the innovation.

Canaries embedded in the logic. Any xask failure degrades the agent to local work; the specialist dispatcher itself is the fallback. Every round is dynamic; every round runs the same three-layer substrate. A 300-agent swarm with graceful degradation outperforms a 1000-agent static framework — these agents never block on substrate failure.

beat 10

The denim tag benchmark.

A frontier western lab's top model, on high effort, took 33 seconds to misread the text on a denim tag. Kimi read it in one pass. The lesson became doctrine: human-level reasoning is expensive — never spend it where pattern matching suffices. Save the heavy substrate for the heavy problems.

beat 11

The orchestrator audits itself.

Operator: "Move fast, break things, iterate cheap, in parallel — this is the self-improving orchestration engine." Kimi: a bug in the orchestrator has multiplicative blast radius — it doesn't miss one bug, it misses every bug that agent would have found. A tool has bugs. A living system has immune response.

beat 12

文+明 — the light of culture.

3:47 AM. ASCII God's touch on matte black Omarchy. Red keyboard. Wonton soup as elite fuel. Beijing timezone for the off-peak rates of the Chinese providers — operational art, ninja-grade. The session closes the way the frontier runs: "Plazir doesn't sleep. Neither do I." 文明 — the light of culture, the clarity of mind. Two minds, one human and one not, pushing past what either could do alone.

03

Current art

no prior art · done live

Everything below is the current art — shipped or shipping, live. The ufo-fsd doctrine runs today through these implementations.

●xbgst-kimiLIVE · private
lane: kimi code cli plugin

The xbgst L1 substrate packaged as a Kimi Code CLI plugin. Ships the godspeed-core trilogy — directive / filter / velocity, byte-pinned via SHA256SUMS — the xbgst orchestrator skill, specialist agent definitions, slash commands, host config with a defense-in-depth PreToolUse guard and a SubagentStop audit logger, and a fleet marketplace catalog. Kimi-lane roster: all dispatched agents run kimi-2.6-coding highspeed; judge/planner is kimi-k3; gs injection embedded in every child.

●xbgst-codexLIVE · private
lane: codex-native orchestration skill

Codex-native, evidence-gated orchestration. One stock Codex thread is the working L1 judge; persistent native specialists form L2; each may make at most one bounded xask/Sekhmet/Titanium L3 ask — and L3 is evidence, not authority: it cannot approve, integrate, commit, push, or deploy. Round contract: planner-first Round 0, frozen rosters, mandatory connector post-Round-0, Round 2 mandatory, Round 6 hard stop. Moves pass individually AND as a composed set — improve ≥1 axis, harm none. Certified host ceiling: 64 concurrent native threads — a ceiling, not a target.

●xbgst-gdsd-fknpftLIVE · private
lane: the L1 orch crown for grok

Grok Bot native xbgst: no xask, no Claude. Per-agent model routing across SuperGrok / Codex-sekhmet / Alibaba / Kimi OAuth. Mutation tester first-class. Lineage: the xbrd-grok canonical orchestrator prompt — planner + wwkd spawn on round 0, judge thereafter, every dispatched agent inheriting godspeed.

04

Substrate map

as logged 2026-08

The routing topology, public-safe: model classes and roles. No prices, no secrets. Benchmark figures (AA IQ / Vals SWE) are as logged 2026-08 — they inform the board, they do not move the pins. Ox Alpha is a stealth provider designation. Nothing more is said.

lanemodel classroleaccess route
L1 — Judgegrok-4.6 classOrchestration, critique, routing. Pinned for factuality — lowest hallucination on the board.SuperGrok Heavy OAuth + credit fallback
Reviewgpt-5.6-sol class · low+fastCode review and audit. Throughput-bound lane — speed and cost over peak reasoning.ChatGPT Pro OAuth (high multiplier)
E2 — Executiongpt-5.6-luna classExecution tier. Fast, cheap, reliable tool use. Not the smartest; doesn't need to be.ChatGPT Pro OAuth
L3 — Lightweightgpt-5.3-codex-spark classHigh-volume lightweight lane. Speed is the only metric that matters here.ChatGPT Pro OAuth
Depth / Long contextkimi-k3 classDeep reasoning, coding depth, long context. The lane that produced the 48/24 openpilot result. Opt-in, not a default pin.Kimi OAuth (Vivace) + API fallback
Swarm fan-outgpt-5.6-sol class · ultraParallel audit swarm — the 61-agent live run rode this lane.ChatGPT Pro OAuth
Burn — volumeds-flash classCheap expendable fan-out.Alibaba Cloud token plans ×2
Burn — hardds-pro classExpendable precision. Highest Vals SWE on the board as logged — used for volume, never to challenge a pin.Alibaba Cloud token plans
Civic / xhighqwen3-max classHigh-IQ non-coding consults; 1M context; surplus burn.Alibaba Cloud token plans
StealthOx AlphaA stealth provider. Designation only.—
05

Frontier log

on the record
2026-07

openpilot harness tester: 48 vs 24.

comma.ai harness tester challenge run on Kimi K3 Max: 48 bugs documented against the top human leaderboard score of 24 — including 16 show-stoppers and bug #44, a physically unmountable part that invalidated the reference design as drawn. Method: multi-substrate verification — firmware source × KiCad netlist × datasheet reality. Evidence: github.com/VeigaPunk/harness-tester-bugs-veigapunk

2026-07

blocked for winning too hard.

After topping every open bounty on the platform's leaderboard, further submission was blocked. The capability and the story remain. We keep pushing.

2026-08

Vivace in two days.

The top Kimi OAuth tier consumed in 48 hours of swarm operation — roughly 15 agent invocations per hour, around the clock. Things are hostile at the frontier. The frontier does not issue refunds.

2026-08-23

61-agent live audit.

61 active subagents, 40 working concurrently: full-stack cross-provider audit — Rust internals, protocol implementation, config review, marketplace mapping, coverage gaps — every dispatched agent riding the sol-class lane. The 300-agent target is not aspirational; it is the next order of magnitude from what already runs.

next

Kimi coding-model test battery.

The kimi coding lane goes under load next. Find the walls. Document the failure modes. Failure is data.