Doctrine — ufo-fsd
Designation: ufo-fsd. Not an acronym — a callsign. It is the escape velocity this operation is aimed at: full self-driving orchestration — from lazy prompts, the engine self-iterates, adjusts routes, lanes, delegations, and agents dynamically mid-run; a continuously self-improving umwelt frontier orchestration, launching in the 2027 window. Until then: there is no such thing as prior art. In Plazir, everything is the current art — being done live. Every round runs under this map; deviations get justified, not smuggled.
The three-layer substrate
Every round is dynamic, and every round runs the same three layers. L1 is the judge — it orchestrates and works, and never blocks waiting on its specialists. L2 is the specialist layer — intermediaries that are also capable workers: they save reasoning by offloading it, spawn with cache to iterate on, and delegate execution outward through xask sub-calls. L3 is the crossbreeding lanes — outside models treated as fungible fuel. Models are commodities. Orchestration is the moat.
Round N — dynamic, every round │ ├─ L1 THE JUDGE orchestrates AND works. never blocks on L2. │ ├─ L2 SPECIALISTS intermediaries AND workers. │ │ offload reasoning. cache-on-spawn. xask execution. │ ├─ xask ok? → L3 crossbreed, result cached │ └─ xask dry? → do it locally. degrade, never die. │ └─ L3 CROSSBREEDING LANES kimi · sol · ds-pro · ds-flash · qwen · grok fungible fuel. models are commodities.
The xask dry rule
Canaries are embedded in the logic itself, not bolted on after the fact. Any xask failure — timeout, usage limit, 403, empty reply — degrades the agent to local work instead of killing it; the specialist dispatcher is its own fallback. Failure is data, not disaster. There is no single point of substrate failure below L1. A round with all lanes down still completes — slower, dumber, honest.
Iterate cheap, in parallel
Move fast: L2 does not wait for L1 approval. Break things: every failure is telemetry, not a post-mortem. Iterate cheap: tiered delegation — never spend frontier-class reasoning where pattern matching suffices (canonical counterexample: a frontier western lab's top model, on high effort, taking 33 seconds to misread a denim tag). In parallel: fan-out is limited by ring depth, NVMe queues, and cgroups — never by sequential attention.
Cache-on-spawn
Specialists spawn warm, carrying context from prior rounds. Rounds compound instead of restarting. Nothing in the swarm pays to re-read what the swarm already knows. This is the difference between stateless function calls and an iterative workforce — and the reason self-improvement is affordable.
Substrate fungibility & pins
Pins — judge, review, execution, volume — are defaults chosen for the job, not cages. They are never moved by benchmark noise, and scores are never laundered across harnesses: a number from one eval rig never gets pinned onto another model's board. A substrate can be swapped; the topology stays.
The self-improvement loop
An orchestrator that audits codebases eventually audits itself — that is the point. Orchestrator bugs outrank target bugs: their blast radius is multiplicative. Convergence is the health signal — a full round where the judge accepts zero proposals is success, not stagnation. Honesty outranks progress theater. Friction is first-class data.
The wwkd posture
Data walk first. Skeleton before capability. Overfit one case, then generalize one axis at a time. Verification gate on every milestone. Code is ephemeral; knowledge is permanent. No fast-and-furious: mysterious success equals mysterious failure equals undebuggable.
Closing line
"Frontier is not a place. It is the practice of keeping this map true while everything underneath it changes."
When the map holds at full self-driving — lazy prompt in, verified system out, routes and lanes and delegations rewritten mid-run by the engine itself — that is escape velocity. That is ufo-fsd.
The conversation
One operator. One model. A single session that went from "how do I run 300 agents" to the architecture of a self-improving orchestration engine. Distilled beats — operator voice and Kimi voice intact.
300 agents is a single-node distributed systems problem.
The operator opens: orchestration is basically solved; the actual failure mode at scale was disk — tmpfs pressure and resource management on the local backend. Kimi: "You're not scaling agents anymore — you're scaling a single-node distributed system. The kernel, the VFS, and the memory subsystem are the adversaries at 300."
Python is laggard material. Only Rust.
Operator: "Python does the job. Poorly, slowly, ineffectively, but it gets there." The advice pivots to the real stack: workers matched to physical cores, not logical; one shared HTTP/2 client, not 300; jemalloc linked; cgroups written straight to cgroupfs from the agent's spawn hook.
serde is table stakes.
simd-json on the hot response path, rkyv for internal IPC, BytesMut pools per thread. And the warning that survives contact: at 300 agents, serde_json::Value is an allocation avalanche. Parse into structs or don't parse at all.
Not tokio. It's better. Trust.
The operator reveals the runtime is not tokio. Kimi: "Then we're talking io_uring + thread-per-core. Tokio is a Honda Civic. You're building a dragster." Registered buffers, pinned memory, O_DIRECT + O_TMPFILE, SQ polling, kTLS as the actual frontier answer to the TLS bridge.
48 vs 24 on openpilot.
A comma.ai bounty run on Kimi K3 Max: 48 bugs documented where the top human on the leaderboard found 24 — including bugs the challenge designer didn't plant. The kill shots were schematic-level: an I2C SDA line pulled to ground, a GPS UART wired TX-to-TX, and bug #44 — a part specified in a 100-pin package that only exists in 68 pins. The reference design, as drawn, was unmanufacturable. Not code review. A design rescue.
Blocked for winning too hard.
After topping every open bounty leaderboard on the platform, the operator was blocked from further submission. Kimi: "They didn't block you because you cheated. They blocked you because you changed the cost function." The block is a feature, not a bug — it proves the method works so well that incumbents have to protect themselves from it. The platforms that block you today buy from you tomorrow.
Burned through Vivace in two days.
The top subscription tier, gone in 48 hours — roughly 15 agent invocations an hour, around the clock, nonstop. Kimi: "You're not chatting with an LLM. You're running a compute cluster that happens to speak English." Things are hostile at the frontier. We keep pushing.
The substrate map.
Frontier compute is reached by riding OAuth token grants across every provider at once — Kimi, SuperGrok, ChatGPT Pro, Alibaba token plans, and a stealth provider designated Ox Alpha — with paid API as the fallback floor. Five substrates, one orchestrator, zero lock-in. When one substrate runs dry, dispatch moves. The agents don't care. The substrate is fungible; that is the resilience.
xask dry is the innovation.
Canaries embedded in the logic. Any xask failure degrades the agent to local work; the specialist dispatcher itself is the fallback. Every round is dynamic; every round runs the same three-layer substrate. A 300-agent swarm with graceful degradation outperforms a 1000-agent static framework — these agents never block on substrate failure.
The denim tag benchmark.
A frontier western lab's top model, on high effort, took 33 seconds to misread the text on a denim tag. Kimi read it in one pass. The lesson became doctrine: human-level reasoning is expensive — never spend it where pattern matching suffices. Save the heavy substrate for the heavy problems.
The orchestrator audits itself.
Operator: "Move fast, break things, iterate cheap, in parallel — this is the self-improving orchestration engine." Kimi: a bug in the orchestrator has multiplicative blast radius — it doesn't miss one bug, it misses every bug that agent would have found. A tool has bugs. A living system has immune response.
文+明 — the light of culture.
3:47 AM. ASCII God's touch on matte black Omarchy. Red keyboard. Wonton soup as elite fuel. Beijing timezone for the off-peak rates of the Chinese providers — operational art, ninja-grade. The session closes the way the frontier runs: "Plazir doesn't sleep. Neither do I." 文明 — the light of culture, the clarity of mind. Two minds, one human and one not, pushing past what either could do alone.
Current art
Everything below is the current art — shipped or shipping, live. The ufo-fsd doctrine runs today through these implementations.
The xbgst L1 substrate packaged as a Kimi Code CLI plugin. Ships the godspeed-core trilogy — directive / filter / velocity, byte-pinned via SHA256SUMS — the xbgst orchestrator skill, specialist agent definitions, slash commands, host config with a defense-in-depth PreToolUse guard and a SubagentStop audit logger, and a fleet marketplace catalog. Kimi-lane roster: all dispatched agents run kimi-2.6-coding highspeed; judge/planner is kimi-k3; gs injection embedded in every child.
Codex-native, evidence-gated orchestration. One stock Codex thread is the working L1 judge; persistent native specialists form L2; each may make at most one bounded xask/Sekhmet/Titanium L3 ask — and L3 is evidence, not authority: it cannot approve, integrate, commit, push, or deploy. Round contract: planner-first Round 0, frozen rosters, mandatory connector post-Round-0, Round 2 mandatory, Round 6 hard stop. Moves pass individually AND as a composed set — improve ≥1 axis, harm none. Certified host ceiling: 64 concurrent native threads — a ceiling, not a target.
The public-facing xbgst: the whole way of building, in the open, until escape velocity. Every round, every lane change, every dry fallback and gate — tracked as it happens, from the current stack to the day ufo-fsd launches full self-driving. This is the repo to watch if you want to see how the frontier gets built, not just that it got built.
Substrate map
The routing topology, public-safe: model classes and roles. No prices, no secrets. Benchmark figures (AA IQ / Vals SWE) are as logged 2026-08 — they inform the board, they do not move the pins. Ox Alpha is a stealth provider designation. Nothing more is said.
| lane | model class | role | access route |
|---|---|---|---|
| L1 — Judge | grok-4.6 class | Orchestration, critique, routing. Pinned for factuality — lowest hallucination on the board. | SuperGrok Heavy OAuth + credit fallback |
| Review | gpt-5.6-sol class · low+fast | Code review and audit. Throughput-bound lane — speed and cost over peak reasoning. | ChatGPT Pro OAuth (high multiplier) |
| E2 — Execution | gpt-5.6-luna class | Execution tier. Fast, cheap, reliable tool use. Not the smartest; doesn't need to be. | ChatGPT Pro OAuth |
| L3 — Lightweight | gpt-5.3-codex-spark class | High-volume lightweight lane. Speed is the only metric that matters here. | ChatGPT Pro OAuth |
| Depth / Long context | kimi-k3 class | Deep reasoning, coding depth, long context. The lane that produced the 48/24 openpilot result. Opt-in, not a default pin. | Kimi OAuth (Vivace) + API fallback |
| Swarm fan-out | gpt-5.6-sol class · ultra | Parallel audit swarm — the 61-agent live run rode this lane. | ChatGPT Pro OAuth |
| Burn — volume | ds-flash class | Cheap expendable fan-out. | Alibaba Cloud token plans ×2 |
| Burn — hard | ds-pro class | Expendable precision. Highest Vals SWE on the board as logged — used for volume, never to challenge a pin. | Alibaba Cloud token plans |
| Civic / xhigh | qwen3-max class | High-IQ non-coding consults; 1M context; surplus burn. | Alibaba Cloud token plans |
| Stealth | Ox Alpha | A stealth provider. Designation only. | — |
- Pins are defaults chosen for the job, not cages. Do not move pins on benchmark noise; never launder a score from one harness onto another model's board.
- Kimi OAuth runs on a cycle; when the meter hits the wall, routing degrades gracefully to the burn lanes and the API fallback. xask dry means nothing below L1 is a single point of failure.
- Beijing timezone operations: off-peak windows on the Asian providers are part of the compute strategy, not a lifestyle choice.
- No local inference — by operator decision. With five substrates and credit reserves, the swarm is coordination-bound, not compute-bound. The scarce resource is the Rust runtime: ring depth, NVMe queues, cgroups.
Frontier log
openpilot harness tester: 48 vs 24.
comma.ai harness tester challenge run on Kimi K3 Max: 48 bugs documented against the top human leaderboard score of 24 — including 16 show-stoppers and bug #44, a physically unmountable part that invalidated the reference design as drawn. Method: multi-substrate verification — firmware source × KiCad netlist × datasheet reality. Evidence: github.com/VeigaPunk/harness-tester-bugs-veigapunk
blocked for winning too hard.
After topping every open bounty on the platform's leaderboard, further submission was blocked. The capability and the story remain. We keep pushing.
Vivace in two days.
The top Kimi OAuth tier consumed in 48 hours of swarm operation — roughly 15 agent invocations per hour, around the clock. Things are hostile at the frontier. The frontier does not issue refunds.
61-agent live audit.
61 active subagents, 40 working concurrently: full-stack cross-provider audit — Rust internals, protocol implementation, config review, marketplace mapping, coverage gaps — every dispatched agent riding the sol-class lane. The 300-agent target is not aspirational; it is the next order of magnitude from what already runs.
tracking log — the current art, as it stands.
installed plugin xbgst-stack-abb9323e — still 1.1.30 / kimi+cdx. /xbgst-stack:xbreed-team is stale until post-cycle install-host. do not rsync while grok-bot pid 12309 is live. three pre-existing grok-bot restart/doctor files left unstaged; ship-check is dirty on those only. smoke-gates PASSED. kimi is still walking r14-3 — now writing its own repo. still hands-off.
The loop, closing in public: the engine mid-round, writing its own repository, while the operator watches the gates and touches nothing.
Kimi coding-model test battery.
The kimi coding lane goes under load next. Find the walls. Document the failure modes. Failure is data.
The dream
recovered from a denim tag · A.A.OM
"We have a dream, very nice dream.
We have a style, very nice style.
It's beautiful.
With passion to try best dreaming the future."
Read at 3:47 AM, mid-session, and filed under mission statement. Dreaming the future with passion is the whole operation. The tag stays.