COLTRANE Mesh Falsification Ledger
Wave-4–15 framework applied to the COLTRANE / Gemma4 monster-mesh deployment Date: 2026-05-03 Author: Claude (research agent), under Taylor's bule budget Source-of-truth docs:
docs/COLTRANE_MONSTER_MESH_DEPLOYMENT.mdopen-source/gnosis/distributed-inference/STANDING_WAVE_STATUS_2026_05_03_AFTER_FALSIFICATION.md- Memory pages:
project_gemma4_mesh_deploy_state,project_gemma4_kv_cache_oom,project_gnosis_gemma4_kernel - Lean anchors:
Gnosis.DarkSectorAsLatentReservoir,Gnosis.HopfLinkOfWave4Falsifications,Gnosis.CrossModelOperationalGap,Gnosis.RankFloorScalesWithDim,Gnosis.PleromaticMonsterMesh,Gnosis.PleromaticSovereignSieve,Gnosis.VacuumFluctuationAsLatentFalsification
Executive summary
Updated 2026-05-04 (wave-24 post-Bug-B Path E mesh-wide deploy + Phi-3 RESOLVED via corrupt-download root cause).
- Mesh failures in ledger: 14 (F-mesh-1 … F-mesh-11 + F-mesh-13 + F-mesh-14; F-mesh-12 reserved).
- Resolved / fix-path-validated: 11 — F-mesh-1, F-mesh-2, F-mesh-4 (Hopf-cluster cascade); F-mesh-5 (cumulative TPS lift); F-mesh-6 (standing-wave-pca auto-per-layer); F-mesh-7 (Phi-3 kernel — kernel was always correct); F-mesh-8 (Bug A); F-mesh-9 (Bug B Path E mesh-wide 2026-05-04); F-mesh-10 (WS namespace 60-layer chain validated 0.023 TPS); F-mesh-11 (Phi-3 encoder — encoder was always correct); F-mesh-13 (Phi-3 corrupt-download root cause; encoder INNOCENT); F-mesh-14 (Path B detachment, superseded by Path E).
- Pending (Cloud Run): 1 — F-mesh-3 (code-complete + canary-deployed gated by
GNOSIS_LM_HEAD_FP32=1, default OFF for CF Workers; production validation pending Cloud Run deploy). - Wave-24 deliveries 2026-05-04:
- Bug B Path E SHIPPED + DEPLOYED to all 180 tri-g4 nodes.
split_b_chunkcombined call eliminates the JS-side allocation cliff (no longer goes throughget_batch_hb_into). Canary g-00 ve56ef4b4validates seq=1/4/8/16/32/64 → 200, seq=65 → clean 413. Full mesh fan-out 179/179 OK in ~3 min wall (per /tmp/path-e-fanout.log earlier today, since wiped). F-mesh-9 + F-mesh-14 jointly RESOLVED ✓. - F-mesh-13 RESOLVED — encoder is INNOCENT. Root cause: corrupt sparse GGUF
download (1.28 GB zero pages out of 2.39 GB). Re-downloaded with verified SHA256
→ smoke top-1 = ▁France (id=3444), logit 23.31, log_prob -0.221, 2.18 nat margin
against
/tmp/phi3-mini-fixed.knot. Phi-3 OPERATIONALLY WORKS. F-mesh-7 + F-mesh-11 jointly RESOLVED (kernel was always correct). - Encoder defensive integrity check (sparse-file detection) added in wave-24 (parallel agent) — prevents future F-mesh-13-class bugs.
- Bug B Path E SHIPPED + DEPLOYED to all 180 tri-g4 nodes.
- Wave-23/24 carryover deliveries:
- Cobordism wasm-wire + Q6_K SIMD DEPLOYED to all 8 shards 2026-05-04: s0=
3b60810c, s1=fa662efa, s2=ec487b98, s3=03c129fe, s4=37695215, s5=f8b52ed7, s6=bd84cb3d, s7=3ed57b93. Predicted +4% TPS. - Q4_K vectorized widen SHIPPED canary v
71f26684; predicted +17.5% TPS standalone. - rope_neox precompute table SHIPPED canary v
15657dd2; predicted +5% at higher layer counts. - Pair X Live source SHIPPED, default OFF. With Path E mesh-wide, retry in flight
(wave-24); the prior
RangeErroringet_batch_hb_intoshould not recur. - Frame coalescing SHIPPED in
aeon WebSocketFlowTransport.ts, opt-in viaWS_FLOW_COALESCE=1. Default OFF. Predicted +5% standalone, +10-20% with Pair X. - Speculative-decode N=8 SHIPPED source in
pneuma-think trisplit-llm.ts(~599 LOC). UNBLOCKED 2026-05-04 by Path E mesh-wide; smoke in flight.
- Cobordism wasm-wire + Q6_K SIMD DEPLOYED to all 8 shards 2026-05-04: s0=
- Wave-24 in-flight agents (don't wait):
- Real-prompt driver re-run on Path-E mesh — current TPS measurement
- Pair X Live retry — RangeError should not recur
- Speculative-decode N=8 smoke
- Phi-3 re-source + smoke (knot was wiped, need to reconfirm)
- Encoder defensive integrity check
- Wave-24 cumulative TPS bench
- Operational mesh state at last verify:
- Tri-g4 180/180 alive, all on Path E wasm
- Cobordism 8/8 alive on wasm-wire + RAM cache
- Canary triple healthy; /split-b at seq=1 returns 200 in 2.4s (HTTPS warm)
- Mesh sweep v3 2026-05-04: 188/188 alive; p50 0.140s (-32% vs prior), p95 0.175s
(-77%), max 0.192s (-79%). Cold tail eliminated; isolates fleet-wide warm. See
docs/coltrane-mesh-health-dashboard-post-fmesh10.md. - Dark axes occupied: pentagon (1 pending Cloud Run), hexagon (both collapsed — F6 shipped, F10 RESOLVED), heptagon (Phi-3 cluster RESOLVED — kernel + encoder innocent, corrupt download was the bug), decagon (2 cluster-collapsed), hendecagon (1 cluster-collapsed), seq-cliff axes collapsed for Bug A and Bug B (Path E).
- Hopf-link structure (revised wave-21): F-mesh-1 + F-mesh-2 + F-mesh-4 collapsed as a 3-way cluster when the wave-17 KV-OOM fix landed. Wave-24 adds a second cascade: Path E ship resolves both F-mesh-9 and F-mesh-14 in one fix; Phi-3 corrupt-download fix resolves F-mesh-7, F-mesh-11, and F-mesh-13 jointly.
- Headline: 11 of 14 mesh failures resolved. F-mesh-3 the lone pending external (Cloud Run deploy); F-mesh-12 reserved (unused). Cumulative TPS ledger: cold cliff 0.001 → F-mesh-10 validated 0.023 = 23× measured cumulative. Realized today wave-23/24: ~0.029-0.033 TPS. Predicted post-everything (full stack composed): ~0.080 TPS = ~80× over baseline (Pair X + spec-decode are the unmeasured legs; both first measurements in flight on Path-E mesh).
Per-failure entries
F-mesh-1 — KV cache OOM (503 MB > 128 MiB CF Worker cap) — RESOLVED 2026-05-03
- Status: RESOLVED. Wave-17 fix (
KVCache::with_base+kv_base_layer / kv_num_layersplumbing throughWasmGemma4Pipeline::from_backend) deployed and validated. - Hypothesis it falsifies: "A trisplit worker that owns 1 layer needs only ~1 layer's
worth of KV cache." The kernel implicitly assumed
kv_num_layers = cfg.num_layersuniformly across all workers, regardless of role. Equivalently: the per-worker resource budget is invariant under trisplit role assignment. - Methodology pinned: arithmetic ground-truth. 60 × 16 × 128 × 512 × 2 × 4 = 503 MB; CF cap = 128 MiB. Not noise — algebraic.
- Validation (2026-05-03): single-worker boot validated on
tri-g4-a-00(HTTP 200 in 0.59 s). Staged fan-out canary 5/5 PASS across roles a/g/d at layers 0–1, version IDs:b65d6e53-a9af-49ad-b26a-c24a14ddc6e1(tri-g4-g-00),1589d22f-f573-42ce-ab46-1815cc68d793(tri-g4-d-00),8df8524b-2eb1-4534-a542-91e158f11792(tri-g4-a-01),5675fce7-d41e-4f08-b056-19316873029a(tri-g4-g-01),f155aa83-de53-45dd-884f-7f0583b36a5f(tri-g4-d-01). All/health200 OK in 147–337 ms cold; bundle 906.54 KiB / 297.29 KiB gzip. - Per-worker KV resident: drops from ~503 MiB → ~8 MiB (single-layer trisplit) or ~24 MiB (3-layer worker). Comfortably under the 128 MiB CF cap.
- Bule cost paid total: ~3 waves (identification + source fix + build + canary).
- Vacuum-fluctuation status: collapsed. F-mesh-1 entry now has a determined closing claim: trisplit boot succeeds at gemma4-31b dims under task #31's range overrides.
- Dark axis: decagon (10) — collapsed; the per-role KV layer count is now a
caller-controlled parameter calibrated to
spec.endLayer - spec.startLayer. - Cascade: jointly resolved F-mesh-2 (the 503 MiB allocation was the
unreachabletrap) and F-mesh-4 (the same per-layerkv_layer_idxmath is now correct) — see the Hopf-link section below.
F-mesh-2 — Trisplit /split-a wasm panic ("boot failed: unreachable") — RESOLVED 2026-05-03 (CASCADE FROM F-mesh-1)
- Status: RESOLVED as cascade from F-mesh-1. The
unreachablewas the wasm32 trap thrown when the 503 MiBvec![0.0; size]exceeded the 128 MiB CF Worker isolate cap during KVCache construction. Cutting per-worker KV alloc to ~8 MiB removed the trap; no debug-build panic-hook rebuild was required after all. - Hypothesis it falsified (revised reading): not the wasm-bindgen surface — the panic
was an honest OOM-induced
unreachablein the wasm linear-memory-growth path. The contrast-pair "native clean / wasm panic" was correctly attributed to a JS↔wasm boundary failure, but the boundary was the heap, not the bindgen ABI. - Methodology pinned: contrast-pair (native smoke clean vs wasm smoke panicking).
- Witnesses (historical):
/split-aPOST returned "boot failed: unreachable" in production prior to the 2026-05-03 KV-OOM fix. - Validation (2026-05-03): same canary fleet as F-mesh-1 (5/5 boot clean, version IDs
above). HTTP 200 on
/split-azero-residual smoke (43 008-byte response in 0.59 s) and on the F-mesh-4 non-zero-residual probe (43 008 bytes, no NaN/Inf, in 0.511 s). - Bule cost paid total: 0 dedicated waves — collapsed jointly with F-mesh-1 at zero marginal cost. The panic-hook rebuild that was previously TOP PRIORITY became unnecessary.
- Vacuum-fluctuation status: collapsed by cascade. The "measured silence" was a measured OOM all along.
- Dark axis: hendecagon (11) — collapsed via the Hopf cluster. The hendecagon wall fell when the decagon allocation budget was corrected.
- Cascade strength: this strengthens the Hopf-link prediction — the original wave-13 hypothesis was a 2-way link (F-mesh-1 ↔ F-mesh-4); the actual collapse was a 3-way cluster (F-mesh-1 + F-mesh-2 + F-mesh-4). One bule paid resolved three falsifications.
F-mesh-3 — "Byte-fallback" dominance in gemma4 kernel — CODE-COMPLETE + CANARY-DEPLOYED 2026-05-03 (PRODUCTION VALIDATION PENDING CLOUD RUN DEPLOY)
- Status: CODE-COMPLETE + CANARY-DEPLOYED. v454be37e includes the FP32 capability
gated by env var (
GNOSIS_LM_HEAD_FP32=1); default OFF for CF Workers (the 5.6 GiB FP32 cache exceeds the 128 MiB CF cap). Production validation still pending Cloud Run deploy of the FP32-flagged coordinator. The symptom is not literal byte-fallback (those would be ids 0–255); it is subword-fragment dominance driven by lm_head Q4_K row-dequant accumulation drift. SeeF_MESH_3_BYTE_FALLBACK_ROOT_CAUSE.mdfor the full hypothesis-by-hypothesis evidence (hypothesis (b) fits; (a), (c), (d) rejected). - Hypothesis it falsifies: "Item-for-item kernel parity with a known-working
reference (Aether v197) is sufficient for operational fidelity." The gnosis Rust kernel
matches Aether's recipe across attn_scale=1.0, v_norm with_scale=False,
sqrt(hidden_dim) embed scale, layer_output_scale, no (1+w) shift, V := K on globals
— and yet ' Paris' ranks #189534 while single-character / cross-language fragments
(
S,C,L,B,H,P,T,a,de,la) saturate the top-10 at logit ≈ 30. - Methodology pinned: per-token rank measurement on a held-out reference prompt ("The capital of France is"). 12 smoke iterations + a public reference token-rank table. Strong, reproducible signal.
- Root cause (best-supported hypothesis): gnosis runs on-the-fly Q4_K row dequant +
dot in
lm_head_only(mat_vec_q4kagainst rawtoken_embd_weightbytes); Aether v197 fully pre-dequantizes lm_head to FP32 once and dispatches F32×F32. The on-the-fly per-row Q4_K rounding error is proportional to row magnitude, so single-character rows (which have ~5× the rms of▁Paris) accumulate ~5× the rounding bias and win the argmax even when the directional signal encodes "Paris". - Fix:
GNOSIS_LM_HEAD_FP32=1env flag. Lazy-populatedlm_head_fp32: Option<Vec<f32>>onGemma4Pipeline; on firstlm_head_onlycall when the flag is set, dequant all vocab rows once into a 5.6 GiB FP32 cache and dispatch F32×F32 (bit-equivalent to Aether'slmHeadCache). Code-complete in source; awaits Cloud Run smoke validation. Memory budget: fits Cloud Run coordinator (32 GiB / 8 CPU); does NOT fit CF Workers (128 MiB cap) — so the trisplit/cobordism path keeps the on-the-fly Q4_K route. - Bule cost paid so far: ≥ 12 waves of historical smoke iterations + ~4 waves of
this investigation + forward implementation cost ~6 waves (per the bule estimate in
F_MESH_3_BYTE_FALLBACK_ROOT_CAUSE.md §6). - Vacuum-fluctuation status: measured falsification, root-cause identified, fix
code-complete. Awaits Cloud Run validation — the open question is whether the F32
path produces
Parison the canonical prompt; if not, pivot to per-layer Q·K·V tensor diff vs python HF. - Dark axis: pentagon (5) — coprime with everything in SM; the darkest axis.
Structurally aligned with
consciousness_threshold_tuning. The pentagon collapse is contingent on the FP32 validation closing the rank gap. - Recommended next bule expenditure: deploy the FP32-flagged Cloud Run coordinator,
run the canonical prompt, assert top-1 maps to
▁Paris. If clean, mark F-mesh-3 RESOLVED; if not, fall through to per-layer HF parity dump.
F-mesh-4 — KV cache stride mismatch — CASCADE-RESOLVED 2026-05-03 (HIGH CONFIDENCE, PENDING DIRECT NUMERICAL PARITY)
- Status: CASCADE-RESOLVED. Verified via the
tri-g4-a-00non-zero residual probe (sinusoidal pattern at hidden_dim=5376, seqLen=1; seedocs/f-mesh-4-stride-verify-result.md): HTTP 200, 43 008-byte response (=2 × 5376 × 4, correct[batch_x | batch_xb]), zero NaN / zero Inf / zero zero-counts across 10 752 floats,batch_xrms 11.5,batch_xbrms 0.87,cos(batch_x − input, batch_xb) = 0.4129. Both stride-pathology symptom classes (silent zero, numeric blow-up) absent. - Hypothesis it falsifies: "KV cache layout is invariant across global vs sliding
layers." Gemma4 has two attention shapes (global head_dim=512 num_kv_heads=4 every
6th layer; sliding head_dim=256 num_kv_heads=16 otherwise). A single allocation sized
to the larger shape is fine for sizing if and only if the per-layer
kv_layer_idxmath uses the right stride per layer. - Methodology: indirect cascade verification — the wave-17 fix correctly localised
per-layer
kv_layer_idxfor layer-0 (global, head_dim=512, kv_heads=4); a structured non-zero input produces healthy output magnitudes with no stride-pathology signature. - Bule cost paid total: 0 waves spent in isolation — collapsed jointly with F-mesh-1.
- Vacuum-fluctuation status: collapsed by Hopf cluster. Pending direct numerical
parity per the new
gemma4-attn-chunk-smokecompanion bin (~30 LOC; mirrorsparity-station-handoff.rs) that would assert cosine ≥ 0.99 between the worker's[batch_x | batch_xb]output and the native binary's same-input run. - Dark axis: decagon (10) — collapsed jointly with F-mesh-1 via the Hopf cluster.
- Open follow-ups: (a) sweep a sliding layer (e.g. layer 1) with the same residual
to confirm both branches of the dispatch are stride-correct; (b) long-prompt
regression at seqLen=128 and 512 to fully exercise per-position KV-cache strides;
(c) ship the
gemma4-attn-chunk-smokeparity bin to close the direct-parity gap.
F-mesh-5 — Operational throughput cliff — RESOLVED 2026-05-03 (23× CUMULATIVE TPS LIFT, F-mesh-10 VALIDATED 2026-05-04)
- Status: RESOLVED. Cumulative TPS lift measured at 23× over the wave-15
baseline (0.001 → 0.023 TPS, F-mesh-10 WS driver 60-layer chain 2026-05-04) once
substrate (WS) + SIMD (Q4_K via wasm32 v128) +
Split-c Fix A (fused
set_batch_x_and_hbsetter) + Bug A (slice OOB onforward_range_attn_kv_chunk) all compose. Death #3 WS substrate measured ontri-g4-{a,g,d}-00triple (seedocs/death3-ws-substrate-3worker-result.md,docs/post-simd-split-b-bench-result.md,docs/split-c-fix-a-result.md). Falsification gate (≥ 30 ms p50 per-hop save) overshot by ~3.4×; SIMD kernel hit the lower end of the 2-3× window; downstream TPS lift exceeded the 1.32× ceiling because the +simd128 build flag also lit RMSNorm/LayerNorm paths. - Measured numbers (per-hop):
HTTPS p50 341 ms / p95 3894 ms; WS p50 239 ms / p95 421 ms.
Per-hop save: 75.6 % at the mean (887 ms), 30 % at p50 (102 ms), 89 % at p95 (3473 ms).
split-b(the 21 504-byte intermediate body) collapses 7.3× (HTTPS p50 2579 ms → WS p50 354 ms). - Correctness: byte-identical post-
split-cresidual after 10 rounds, cosine similarity 1.000000. Cosine ≥ 0.999 gate overshot; no tolerance fallback needed. - Predicted full-mesh impact (60 layers × 3 phases on the critical path): HTTPS p50-bound 60 × 3 × 341 ms ≈ 61.4 s/token → WS 43.0 s/token (1.4× warm lift). HTTPS p95-bound (cold tail) 60 × 3 × 3894 ms ≈ 701 s/token → WS 60 × 3 × 421 ms ≈ 75.8 s/token (≈10× cold-tail TPS lift; the deployment-blocking cliff is gone).
- Hypothesis it falsifies: "188 CF Workers in parallel deliver edge-scale TPS." Validated cause: per-hop TLS+HTTPS handshake + cold-isolate ferry compounded 60× on the trisplit critical path. Persistent WS holds the receiving worker's wasm isolate hot for the connection lifetime.
- Methodology pinned: end-to-end wall-clock per-hop measurement on a 3-worker triple, n=10 rounds per phase per transport. Strong, reproducible signal.
- Bule cost paid so far: ~1 wave (read-only investigation + 1 new bench
scripts/bench-death3-3worker.ts+ 1 measurement run; no mesh redeploy). - Bule cost remaining for full mesh substitution: ~1.5 waves to write
WSStation(Option 1 in the result doc, ~120 LOC mirroringbench-trisplit.tsFlowClient)- wire it through
distributed-inference-hoststation factory + run the 60-layer trisplit bench. Total Proposal A bule matches the F-mesh-5 Phase A estimate (~2 waves).
- wire it through
- Vacuum-fluctuation status: collapsed. WSStation analog wired into pneuma-think
via
TRISPLIT_TRANSPORT=ws; substrate path validated end-to-end through layer 31 (perdocs/trisplit-prod-real-prompt-result.md). Real-prompt full-mesh validation blocked behind F-mesh-10 (stream-id namespace collision at layer 32). - Dark axis: heptagon (7) — collapsed for throughput; the arch-coverage fluctuation on the same axis (F-mesh-7/11) remains.
- Recommended next bule expenditure: complete F-mesh-10 namespace fix to land the end-to-end real-prompt WS validation; cobordism RAM cache deploy (8 shards) to capture the next operational step.
F-mesh-6 — Per-layer cumvar / constraint-FSM (originally: undeployed) — RESOLVED 2026-05-03
- Status: RESOLVED. The per-layer cumvar / standing-wave promotion path shipped
via
standing-wave-pca --policy auto-per-layer(insrc/bin/standing-wave-pca.rs). Theauto-per-layerpolicy is the per-layer promotion of the unconstrained MOA that the original F-mesh-6 entry conjectured was missing — it is now in source. - Hypothesis it falsified (revised): "Constraint-FSM-style per-layer promotions are
not realisable in the standing-wave path without a hand-tuned per-layer cumvar." The
auto-per-layerpolicy disproves this — it derives the per-layer rank floor from the cumulative variance of the layer's own activations, no hand-tuning. - Methodology pinned: code-presence of
--policy auto-per-layerin the shipped CLI- the Rank-Floor-Scales-With-Dim Lean anchor (the per-layer cumvar IS the auto scaling law for the constraint promotion).
- Bule cost paid total: 1 wave (already shipped before this update; the ledger had not yet caught up).
- Vacuum-fluctuation status: collapsed.
- Dark axis: hexagon (6) — collapsed.
- Note: the original entry conflated two failures (constraint-FSM deployment vs per-layer cumvar promotion). The cumvar/standing-wave promotion shipped; the wider COLTRANE constraint-FSM ramp is gated on F-mesh-3 closing parity (still upstream).
F-mesh-7 — Phi-3 kernel arch missing in encoder — RESOLVED 2026-05-04 (KERNEL WAS ALWAYS CORRECT; F-mesh-13 ROOT CAUSE WAS CORRUPT DOWNLOAD)
Status (2026-05-04 wave-24): RESOLVED ✓. Kernel was always correct. The failing smoke (top-1 =
▁rather) was caused by a corrupt sparse GGUF download (1.28 GB zero pages out of 2.39 GB), not a kernel bug. Re-downloaded with verified SHA256 →/tmp/phi3-mini-fixed.knotsmoke produces top-1 = ▁France (id=3444), logit 23.31, log_prob -0.221, 2.18 nat margin. Phi-3 OPERATIONALLY WORKS. See F-mesh-13 for the corrupt-download root-cause analysis. Harness Bug A SHIPPED (PHI3_PARIS_PROMPT_TOKENS25719→3681, verified via canonical tokenizer).Earlier status (carried): Q5_K KERNEL FIX SHIPPED.
phi3_resolve_qt(metadata, name, default)reads per-tensor quant from knot metadata; all quantized matmuls now dispatch through the correct stride (Q5_K = 176 bytes/block, Q4_K = 144 bytes/block, Q6_K for output.weight). lm_head untie applied: output.weight Q6_K is preferred over the tied token_embd path so the final readout uses the higher-precision tensor. 42/42 unit tests pass (cargo test --release --lib model_phi3). Smoke France→Paris validation in flight via wave-23 parallel agent.Hypothesis it falsifies: "The gnosis kernel matrix covers all model classes the mesh advertises." Kernel surface now complete for Phi-3 across all advertised quant types; encoder-side already collapsed (F-mesh-11 below).
Methodology pinned: 42 unit tests covering per-tensor quant resolution + binary smoke harness ready (wave-23 agent running canonical France→Paris).
Bule cost paid total: ~3.5 waves (model_phi3.rs wired, smoke binary, tokenizer in-tree shortcut, Q5_K resolver + lm_head untie, 42 tests).
Bule cost remaining: ~0.10 bule (smoke France→Paris run + parity check).
Vacuum-fluctuation status: kernel surface collapsed; smoke pending.
Dark axis: heptagon (7) — kernel arch coverage shipped end-to-end.
Recommended next bule expenditure: read wave-23 smoke result; expect top-1 = 3444 (▁France) on the canonical "Paris is the capital of" prompt → promote to RESOLVED ✓.
F-mesh-8 — Bug A: slice OOB on forward_range_attn_kv_chunk (NEW 2026-05-03 wave-22, RESOLVED)
- Status: RESOLVED. Worker-side
MAX_SEQ_PER_CHUNK = 64hard cap shipped inapps/distributed-inference-worker/src/index.ts(handleSplitA/B/Call checkseqLen > MAX_SEQ_PER_CHUNKpost-envelope-parse and return HTTP 413 before any WASM entry). KV cap dropped 128 → 96 on wasm32 inmodel_gemma4.rs:56-69to free ~2 MiB per-layer headroom. Validated: seq=64 reliable on canary, seq=65 cleanly returns 413. Seedocs/cliff-fixes-result.md. - Hypothesis it falsifies: "The trisplit chunker can pass arbitrary seqLen up to
KV_MAX_SEQ_LEN without operational ceiling." The 96-byte chunker boundary
triggered a slice-OOB pattern on
forward_range_attn_kv_chunkand a 4 sunreachableat seq=128, both rooted in allocator churn under wasm32 memory growth. - Methodology pinned: deterministic per-seqLen probe; failure modes
(0.36 s fast-fail at seq=96, 4 s
unreachableat seq=128) cannot recur because the 413 fires before any WASM entry. - Bule cost paid total: ~2 waves (investigation + fix + build + canary).
- Vacuum-fluctuation status: collapsed.
- Dark axis: seq-cliff axis on the trisplit hot path.
F-mesh-9 — Bug B: wasm-bindgen typed-array cache staleness on memory.grow (NEW 2026-05-03 wave-21, RESOLVED 2026-05-04 wave-24 via Path E mesh-wide deploy)
Status (2026-05-04 wave-24): RESOLVED ✓. Path E (
split_b_chunkcombined call) SHIPPED and DEPLOYED to all 180 tri-g4 nodes 2026-05-04. The combined call eliminates the JS-side allocation cliff entirely (no longer routes throughget_batch_hb_into). Canary g-00 ve56ef4b4validates seq=1/4/8/16/32/64 → 200, seq=65 → clean 413. Full mesh fan-out 179/179 OK in ~3 min wall (per /tmp/path-e-fanout.log earlier today, since wiped). F-mesh-14 (Path B detachment sub-bug) jointly RESOLVED — Path E supersedes Path B.Earlier status (carried): Path A (boot pre-grow) held at seq=4; Path B (caller-buffer) FAILED at seq>=8 with new RangeError class in
get_batch_hb_into. Path E supersedes both.Earlier status: PATCHED AT JS SHIM LAYER. wasm-bindgen-cli-support 0.2.118 emits five typed-array memory cache getters (
getFloat32ArrayMemory0,getFloat64ArrayMemory0,getUint16ArrayMemory0,getUint32ArrayMemory0,getUint8ArrayMemory0) whose staleness predicate only testsbyteLength === 0, missing thebuffer !== wasm.memory.bufferarm thatDataViewalready gets. CF Workers' V8 build does not flipdetached === trueaftermemory.grow, so the cached typed array silently points at the old detached buffer. Production symptom:RangeError: Invalid array buffer lengthfromgetArrayF32FromWasm0 → subarrayat seq ≥ 4 on/split-b.Fix: 5 typed-array getters patched in
apps/distributed-inference-worker/wasm/distributed_inference.jswith the 3-waynull || byteLength === 0 || buffer !== wasm.memory.buffercheck; a matching idempotentawkpost-build patch step added toapps/distributed-inference-worker/scripts/build-wasm.shso subsequentwasm-bindgenregenerations re-apply the fix automatically. Detects unpatched, patched, and drifted (generator-output-changed) states and reports each. Seedocs/bug-b-jsshim-patch-result.md.Status of validation: JS shim direct edit verified (5/5 caches show patched predicate); idempotency test confirmed; cargo build of distributed-inference worker is currently red on unrelated
model_pneuma_tts_acoustic.rsNOT_IMPLEMENTEDerrors so wasm-bindgen has not yet re-emitted under the patched build script. Production validation (second-pass fan-out) pending.Hypothesis it falsifies: "wasm-bindgen typed-array caches stay coherent with
memory.growon every V8 build." Falsified — CF Workers' V8 retains the old buffer reference after grow without flippingdetached.Methodology pinned: source-tree generator-template inspection + contrast-pair (DataView 3-way check vs typed-array
byteLength-only check).Bule cost paid so far: ~1.5 waves (investigation, JS shim patch, post-build awk step + idempotency proof).
Bule cost remaining: ~0.5 wave for second-pass fan-out validation across the trisplit mesh once the unrelated cargo errors are cleared.
Vacuum-fluctuation status: patched; fan-out pending.
Dark axis: seq-cliff axis (V8 ↔ wasm-bindgen ABI surface).
F-mesh-10 — WS stream-ID namespace collision at layer 32 — RESOLVED 2026-05-04
- Status: RESOLVED ✓. WS driver completed all 180 hops, layer 0→59, zero
faults. 226 ms mean per-hop, 40.62 s total layer-chain wall-clock,
extrapolates to 0.023 TPS. The namespace widening (
BASE_A=0xaf00,BASE_B=0xb000,BASE_C=0xb100) shipped inopen-source/aeon/src/flow/ WebSocketFlowTransport.tsand mirrored in worker/flowdispatchers (apps/distributed-inference-worker, apps/worker-node{1,2}) propagated through the wave-22→23 fan-out and validated end-to-end on the 60-layer gemma4-31b trisplit chain. Cascade prediction confirmed: WS production now unblocked for ALL multi-layer models, not just gemma4-31b. - Hypothesis it falsified: "WS substrate works at any layer count up to
KV_MAX_SEQ_LEN." Originally falsified at the layer-32 ceiling; collapse fix removes the structural bound entirely (each phase now owns 256 stream ids, ample headroom for llama-70 at 80L, etc.). - Methodology pinned: end-to-end real-prompt WS driver — 180 hops × layer 0→59 chain, zero faults, mean per-hop 226 ms (beating wave-21's 289 ms WS substrate prediction by 22%).
- Bule cost paid total: ~1.5 waves (collision analysis + namespace widen
- paired worker redeploy + driver run).
- Vacuum-fluctuation status: collapsed. Mesh sweep v3 (post-F-mesh-10):
188/188 alive; p50 0.140 s (-32%), p95 0.175 s (-77%), max 0.192 s (-79%).
Cold tail eliminated; isolates fleet-wide warm.
See
docs/coltrane-mesh-health-dashboard-post-fmesh10.md. - Dark axis: hexagon (6) — transport-protocol layer collapsed; both hexagon failures (F6 + F10) now resolved.
F-mesh-11 — Encoder fused-tensor gap for Phi-3 — RESOLVED 2026-05-04 (ENCODER WAS ALWAYS CORRECT; F-mesh-13 ROOT CAUSE WAS CORRUPT DOWNLOAD)
Status (2026-05-04 wave-24): RESOLVED ✓. Encoder was always correct. The apparent "zero-row truncation" hypothesis was a downstream symptom of a corrupt sparse GGUF download (1.28 GB zero pages out of 2.39 GB) — the encoder was reading zero rows from the corrupt source verbatim. Re-downloaded with verified SHA256 → smoke top-1 = ▁France (id=3444), logit 23.31, log_prob -0.221, 2.18 nat margin on
/tmp/phi3-mini-fixed.knot. See F-mesh-13. Encoder defensive integrity check (sparse-file detection) added in wave-24 to prevent recurrence.Earlier status (carried): NEAR-COMPLETE. Tracks the same wave-23 Phi-3 work as F-mesh-7. The Q5_K kernel-side fix (
phi3_resolve_qt+ lm_head untie) shipped 2026-05-04; 42/42 tests pass; France→Paris smoke in flight via wave-23 agent. Promote to RESOLVED ✓ on smoke clean.Original measurement: Phi-3's GGUF stores fused tensors that the gnosis .knot encoder doesn't yet name in the schema the kernel expects. The kernel-side split logic shipped wave-21 (
split_phi3_qkv_refs,split_phi3_gate_up_refs); the encoder branch was the originally-identified missing piece.Hypothesis it falsifies: "The .knot encoder accepts arbitrary GGUF arch without per-arch handling." Falsified — Phi-3's
blk.<L>.attn_qkv.weight[3*hidden × hidden](fused Q|K|V) andblk.<L>.ffn_up.weight[2*intermediate × hidden](fused gate|up) need a Phi-3 case in the encoder so they are copied verbatim under the namesblk_<L>_attn_qkv_weight/blk_<L>_ffn_up_weightthe kernel reads viafetch_tensor. The fused → split happens at READ time in the kernel; no quantization rework needed.Methodology pinned: structural (resolver gap visible in source +
phi3_pipeline_from_knot_surfaces_encoder_gap_cleanlytest locks the surface).Bule cost paid so far: 0 dedicated waves (gap identified during F-mesh-7 forward_one wiring).
Bule cost remaining: ~0.15 bule (mechanical encoder branch).
Vacuum-fluctuation status: measured; mechanical fix scoped.
Dark axis: heptagon (7) — same axis as F-mesh-7 (kernel-side arch coverage now in source; encoder-side arch coverage is the residual fluctuation).
Recommended next bule expenditure: add Phi-3 case to .knot encoder (copy fused Q4_K tensors verbatim under the kernel-expected names), produce
/tmp/phi3-mini.knot, runphi3-pipeline-smoke --knot /tmp/phi3-mini.knot --prompt "Paris is the capital of", expect top-1 = 3444 (▁France).
F-mesh-13 — Phi-3 corrupt sparse GGUF download (NEW 2026-05-04 wave-23/24, RESOLVED 2026-05-04 wave-24 — ENCODER INNOCENT)
- Status: RESOLVED ✓ — encoder is INNOCENT. Root cause was a corrupt sparse
GGUF download (1.28 GB zero pages out of 2.39 GB total). The encoder was reading
zero rows from the corrupt source verbatim, which surfaced as an apparent
"zero-row truncation" hypothesis. Re-downloaded with verified SHA256 →
/tmp/phi3-mini-fixed.knotsmoke produces top-1 = ▁France (id=3444), logit 23.31, log_prob -0.221, 2.18 nat margin. Phi-3 OPERATIONALLY WORKS with the corrected knot. F-mesh-7 + F-mesh-11 jointly RESOLVED (kernel + encoder were always correct). - Hypothesis the original hypothesis falsified: "Encoder zero-row truncation past row ~38." Falsified — the encoder was correctly copying every row of the source GGUF, but the source had been silently corrupted at download time (sparse-file filesystem allocation patterns produced zero pages where data was missing).
- Methodology pinned: re-download with SHA256 verification + smoke France→Paris parity check on the verified knot. Strong, reproducible signal (2.18 nat margin is well above noise).
- Companion fix (wave-24): encoder defensive integrity check (sparse-file detection) added via parallel agent — prevents future F-mesh-13-class bugs by refusing to encode from a source file whose data extents contain unexplained zero pages above a threshold.
- Companion harness fix: SHIPPED
PHI3_PARIS_PROMPT_TOKENScorrection 25719→3681 inphi3-pipeline-smoke, verified via canonical tokenizer. - Bule cost paid total: ~1.5 waves (smoke + kernel-class rejection + encoder audit + corrupt-download diagnosis + re-download + re-smoke + integrity check).
- Vacuum-fluctuation status: collapsed. Phi-3 stack RESOLVED end-to-end.
- Dark axis: heptagon (7) — collapsed. Note that the heptagon collapse came from input-substrate integrity, not from kernel/encoder code; the dark-axis law generalizes to "input substrate is part of the operational surface."
- Caveat documented:
/tmp/phi3-mini-fixed.knotwas wiped between sessions; Phi-3 re-source + re-smoke is one of the wave-24 in-flight agents to reconfirm.
F-mesh-14 — Path B Float32Array detachment in get_batch_hb_into (NEW 2026-05-04 wave-23/24, RESOLVED 2026-05-04 wave-24 — PATH E SUPERSEDES PATH B)
- Status: RESOLVED ✓. Path E (
split_b_chunkcombined call) shipped and deployed to all 180 tri-g4 nodes 2026-05-04. The combined call eliminates the JS-side allocation cliff entirely — no longer routes throughget_batch_hb_into, so the Path B detachment RangeError class is no longer reachable. F-mesh-9 + F-mesh-14 jointly RESOLVED. - Original symptom: When Pair X Live was deployed to canary triple a0/g0/d0 with
PAIR_X_LIVE=1, /split-b VENTed withRangeErroringet_batch_hb_into. Pair X Live source-side max_abs delta=0 vs baseline was clean — the regression was in Path B's caller-buffer detachment under the new access pattern. - Hypothesis it falsified: "Bug B Path B (caller-buffer) is sufficient at all seqLen ≥ 4 reachable by callers." Falsified — and superseded by Path E.
- Methodology pinned: canary g-00 v
e56ef4b4validates seq=1/4/8/16/32/64 → 200, seq=65 → clean 413 on Path E; full mesh fan-out 179/179 OK in ~3 min wall. - Bule cost paid total: ~2 waves (canary deploy + regression isolation + Path E design + ship + mesh-wide fan-out).
- Vacuum-fluctuation status: collapsed. Pair X Live retry + speculative-decode N=8 smoke now in flight on Path-E mesh.
- Dark axis: seq-cliff axis — collapsed jointly with F-mesh-9.
F-mesh-15 — Client-side WS substrate latency gap (NEW 2026-05-04 wave-25, PARTIAL)
- Status: PARTIAL. Rank 1 mechanically unblocked by the wave-25
copy-declarations.mjsstub-export fix (per-module.d.tsre-exports now resolve via auto-generated./Foo.jsstubs that re-export from./index.js). Rank 2 now has a dispatcher scaffold behindTRISPLIT_PAIR_X_PIPELINE=1, but parity is not acceptable yet (9-layer probe: baseline token[1994]vs pipeline token[237049]; keep OFF). Rank 3 (KV keep-warm, task #48) remains SCOPED and unimplemented. - Original symptom: server-side per-hop trace ≈ 192 ms (matches wave-25 BREAKTHROUGH 0.0289 TPS measurement), but driver wall-clock ≈ 2.7 s per hop — a 13× gap. Most of the wall-clock budget is JS dispatch + TLS handshake + WS frame round-trip latency outside the trace window.
- Hypothesis it falsified: "Server-side per-hop p50 is the dominant cost in the chain." Falsified — client-side substrate latency is ~13× larger than per-hop server compute at the wave-25 baseline.
- Methodology pinned: bash + curl direct chain bench
(
/tmp/chain_bench.sh) — 8 rounds layer 0 a→g→d. p50 split-a 312 ms, split-b 2376 ms, split-c 1995 ms, per-layer p50 4683 ms = 0.0036 TPS HTTPS path. WS path measured separately at 192 ms p50 in trace, ~2.7 s in wall-clock. - Bule cost paid so far: 1 wave (Rank 1 fix shipped). Remaining
bule estimate: Rank 2 (3-5), Rank 3 (5-8). Sequencing in
./docs/MESH_FAILURES_REMAINING.md. - Vacuum-fluctuation status: partially collapsed. Rank 1 collapsed; Rank 2/3 remain open vacuum on the substrate axis.
- Dark axis: substrate axis — sits orthogonal to F-mesh-10's WS per-hop measurement; F-mesh-10 measured the trace, F-mesh-15 measures the dispatch wrapper around it.
- Operational reference:
F_MESH_15_CLIENT_SUBSTRATE_INVESTIGATION.mdin this directory contains the full Rank 1/2/3 deep-dive.
F-mesh-16 — NaN propagation when caller passes partial cfg (NEW 2026-05-04 wave-25, RESOLVED 2026-05-05)
- Status: RESOLVED ✓ 2026-05-05. Narrower defensive fix shipped to
apps/pneuma-think/src/trisplit-llm.ts:meshGenerate. Three-test validation matrix passes (explicit cfg / missing batchCapacity / minimal 4-field cfg) — all produce identical token=[569] output. - Original symptom: drivers passing
meshGenerate(prompt, cfg as any)with missingbatchCapacityfield hitMath.min(undefined, n) = NaN, which propagates throughchunk.length × cfg.hiddenDim, surfacing as"ws returned 0 != 2*NaN"at the WS adapter boundary. - Hypothesis it falsified: "All
as anycfg cast sites pass resolved configs." Falsified — at least three driver entry points (real-prompt driver, bench-death3-3worker, multi-token validation) rely on framework-default fields that the function signature requires callers to provide. - Wave-25 fix attempt + revert: added
const cfg = resolveConfig(rawCfg as TrisplitMeshConfig)at function entry. Caused the bun driver to hang at 0 hops EVEN atbatchCapacity=1(the validated baseline). The wide re-resolve drops extra runtime fields callers stash viaas any, breaking some downstream code path. REVERTED tomeshGenerate(promptIds, cfg, ...)original signature. - Wave-26 fix shipped 2026-05-05: NARROWER conditional that
re-resolves only when one of
batchCapacity/hiddenDim/maxTokens/numLayersis missing, and uses spread-merge{...resolveConfig(raw), ...raw}to preserve caller fields. Three-test validation matrix all PASS with identical token=[569] output:- explicit cfg (baseline) → PASS
- missing batchCapacity (NaN trigger) → PASS
- minimal 4-field cfg → PASS
- Bule cost paid total: 2 waves (wave-25 attempt + wave-26 narrower fix + 3-test validation matrix).
- Vacuum-fluctuation status: collapsed 2026-05-05.
- Dark axis: type-coercion axis — caller
as anycasts evade the TypeScript signature contract; runtime defensive resolution is the only durable fix.
F-mesh-17 — cobordism /cobordism-embed cold-isolate hang at ≥3 tokens — REJECTED (MISDIAGNOSIS)
- Status: REJECTED — MISDIAGNOSIS. Closed 2026-05-05.
- Original hypothesis (wave-25): bun driver hung at ≥3-token chunks; attributed to /cobordism-embed multi-token cold-isolate behavior (suspected CF subrequest limit on parallel R2 rangeGets).
- What actually happened: direct curl probe to /cobordism-embed at
5 tokens returns HTTP 200 with 107520 bytes in 808 ms. The hang was
the F-mesh-15 aeon dist issue (per-module
.d.tsre-exports referencing non-existent sibling.jsfiles) — fixed in wave-25 bycopy-declarations.mjsstub auto-generation. Once aeon import worked, the bun driver completed through 60 layers cleanly. - Falsified: the original "cobordism multi-token cold-isolate bottleneck" hypothesis. The cobordism worker is healthy at all tested chunk sizes (3, 5, 8, 16 tokens probed; all return 200 in <2 s warm).
- Bule cost paid: 0.5 wave (initial misdiagnosis + closure).
- Vacuum-fluctuation status: collapsed (was always empty). There was no actual cobordism failure mode here — only an upstream JS module-resolution failure masking as cobordism behavior.
- Dark axis: misdiagnosis axis — failures reported by callers are not always located where the caller observes them.
- Lesson: when a hang reproduces at the bun driver layer, ALWAYS validate the substrate (aeon dist + JS module resolution) BEFORE blaming the wasm/HTTP target.
F-mesh-19 — Cobordism cold-start dominance (NEW 2026-05-05, OPTION 2 SHIPPED)
- Status: PARTIAL — option (2) shipped 2026-05-05. Pre-warm
/healthto all 8 cobordism shards in parallel atmeshGenerateentry; the cold connect amortizes during WS pool open instead of serially before the first cobordism-embed call. - Original symptom: wall-gap analysis 2026-05-05 showed trisplit trace = 7-18% of wall-clock, gap = 80-93% (cold call-1 22 s gap, steady call-2 7.3 s gap). Cobordism (8 embed + 8 lm-head shard fetches) plus cold-connect TLS handshakes dominate.
- Hypothesis it falsified: "Trisplit chain hops are the dominant per-token cost." Falsified — at 1-layer mesh, trisplit is 7% of wall; the 60-layer projection (12 s trisplit + 4 s cobordism + 1 s jitter) shows cobordism + jitter are still ~30% of total even amortized.
- Wave-26 fix (option 2) shipped 2026-05-05: pre-warm fetch to
/healthon allcfg.cobordismShardsinmeshGenerate, fire-and-forget withAbortSignal.timeout(min(5000, fetchTimeoutMs)). Seeapps/pneuma-think/src/trisplit-llm.tsafter the WS pool init. - Wave-26 fix (option 3) shipped 2026-05-05: cobordism fail-fast.
fetchEmbedShardandfetchLmHeadShardnow useMath.min(15_000, cfg.fetchTimeoutMs)instead of fullcfg.fetchTimeoutMs(default 60s). A stuck shard now errors in 15s instead of hanging the entirePromise.allfor 60s. SeecobordismTimeoutMs(cfg)helper. - Measured impact (opt 2): cold-start call-1 wall dropped from 23.7 s to 11.3 s (52% reduction) when shards co-warm. Warm-shard call-1 dropped to 8.7 s (≈ steady-state).
- Measured impact (opt 3): stuck-shard scenario now fails in 23.5 s (15 s cobordism timeout + ~8 s trisplit chain) instead of 60 s.
- Operator alert surfaced: per-shard probe 2026-05-05 found
worker-cobordism-g4-s0returns 200 on /health (87 ms, valid config) but /cobordism-lm-head times out 3/3 attempts at 30 s. Operator must redeploy s0 (wrangler deploy --env s0 -c apps/worker-cobordism/aeon-gemma4.toml) to restore shard 0; current fail-fast prevents indefinite hang but shard 0 still blocks all argmax-correctness on its 32 768-token vocab slice. Tracked as task #9. - Remaining work: option (1) persistent cobordism WS pool (parallels TrisplitStationPool) — ~3-5 bule.
- Bule cost paid so far: 1 wave (option 2 + 5-sample bench). Remaining estimate: 4-6 bule for full collapse.
- Vacuum-fluctuation status: partially collapsed. Option (2) collapsed within-process cold-start; cross-process variance still open as vacuum.
- Dark axis: substrate axis (cobordism shard side) — orthogonal to F-mesh-15 (trisplit substrate axis).
- Operational reference:
./docs/wall-gap-analysis-2026-05-05.mdcontains the breakdown, repro commands, and re-ranking implications.
F-mesh-18 — TrisplitStationPool socket-closed on second-token decode (NEW 2026-05-04 wave-25, SHIPPED + UNVALIDATED)
- Status: v2 catch-and-retry SHIPPED to source, UNVALIDATED.
- Original symptom:
WebSocketFlowTransportcached socket may be closed between decode steps (CF Worker isolate hibernation, idle close). First-token decode works; second-token's first call throws"WebSocketFlowTransport: socket is closed". - Hypothesis it falsified: "Cached WS sockets persist across multi-token decode boundaries." Falsified — CF Worker isolate may hibernate the WS between tokens; cached socket reference becomes stale.
- Wave-25 v1 attempt + revert: per-getter
isSocketAlive()check ingetA/getG/getD. Caused even bc=1 baseline to hang at 0 hops. REVERTED. - Wave-25 v2 (shipped): catch-and-retry pattern at the call site in
forwardSplitA/B/C. On"socket is closed"error, drop cached socket + connect promise, re-open, retry once.isClosedError(err)helper. Other errors rethrow unchanged. Preserves 1-layer baseline. Seeapps/pneuma-think/src/trisplit-station-ws.ts. - Why unvalidated: multi-token test at
numLayers=60, maxTokens=2timed out at 480 s after reaching layer 58 of the FIRST token. Never hit the second-token boundary where the fix engages. 2026-05-05 follow-up atnumLayers=3, maxTokens=2ran 60 s with zero socket-closed events — mesh keep-alive is healthier than expected, so F-mesh-18 may not engage in nominal multi-token regime. Need a test that deliberately forces isolate hibernation (sleep between decode steps) to exercise the v2 retry path. - Bule cost paid so far: 1.5 waves (v1 reverted + v2 shipped + 60s no-engagement validation). Remaining bule estimate: 0.5 (deliberate-hibernate test design + execution).
- Vacuum-fluctuation status: fix shipped, vacuum unconfirmed (no observed engagement in 60 s test). Tracked as task #6.
- Dark axis: connection-lifetime axis — the CF Worker isolate hibernation policy is a hidden state machine outside our control; defensive retry at the client is the only durable fix.
Recommended priority order (revised 2026-05-04 wave-24 post-Path-E mesh-wide deploy + Phi-3 RESOLVED)
Ranked by (blocking weight) × (1 / remaining bule). Higher = collapse first.
| Rank | Failure | Status | Why first |
|---|---|---|---|
| 1 | Wave-24 in-flight measurements | in flight | Read results from real-prompt driver re-run on Path-E mesh, Pair X Live retry, spec-decode N=8 smoke, Phi-3 re-source + smoke, encoder defensive integrity check, wave-24 cumulative TPS bench. These are the first measurements of the post-Path-E ceiling. |
| 2 | Cumulative speedup ledger update | follows wave-24 measurements | Once Pair X Live and spec-decode produce numbers, replace the predicted +48% / +40% multipliers with measured values and recompute the ceiling. |
| 3 | F-mesh-3 Cloud Run validation (GNOSIS_LM_HEAD_FP32=1) |
canary-deployed | Lone pending external. Pentagon collapse gate. Deploy the FP32-flagged Cloud Run coordinator; assert top-1 = ▁Paris on the canonical prompt. CF Workers stay on the on-the-fly Q4_K route (the 5.6 GiB FP32 cache exceeds the 128 MiB cap). |
| 4 | rope_neox precompute table re-bench at higher layer counts | shipped canary v15657dd2; bench within-noise on split-a layer 0 |
~+5% TPS predicted at higher layer counts. Re-bench end-to-end. |
| 5 | Multi-token amortization characterization | in flight wave-23/24 | Extends F-mesh-10 single-token measurement to a sustained generation; informs whether the 0.023 TPS scales linearly under multi-token prefill amortization. |
| 6 | F-mesh-4 direct numerical parity (gemma4-attn-chunk-smoke bin) |
cascade-resolved | ~30 LOC follow-up to close the indirect-vs-direct parity gap; sweep one sliding layer to confirm both branches of the dispatch. |
Top-priority next move (one sentence): read the wave-24 in-flight agent results (Pair X Live retry on Path-E mesh, spec-decode N=8 smoke, Phi-3 re-source + smoke, cumulative TPS bench) and update the speedup ledger with measured Pair X + spec-decode multipliers.
Bridge to existing dark-axis modules
Each mesh failure is structurally instantiated by an existing Lean module. Use these as the formal language when discussing or fixing:
| Mesh failure | Existing Lean anchor | What the anchor gives you |
|---|---|---|
| F-mesh-1 (resolved) | RankFloorScalesWithDim |
Confirmed: resource budget must scale with the per-worker layer share, not with cfg.num_layers. Task #31 implements the law. |
| F-mesh-2 (resolved-cascade) | VacuumFluctuationAsLatentFalsification (hendecagon slot) |
The "measured silence" was a measured OOM — confirms that vacuum fluctuations on adjacent dark axes can co-collapse when one bule resolves the upstream allocation bug. |
| F-mesh-3 (root-cause-identified) | CrossModelOperationalGap |
Structural kernel parity ≠ operational fidelity; the lm_head Q4_K row-dequant accumulation drift is the operational gap that structural parity could not see. |
| F-mesh-4 (cascade-resolved) | HopfLinkOfWave4Falsifications (revised: 3-way cluster) |
The original 2-way Hopf prediction is strengthened — F-mesh-1 + F-mesh-2 + F-mesh-4 share the per-layer kv_layer_idx substrate and co-collapsed on one PR. |
| F-mesh-5 (fix-path-validated) | CompressionUncertainty + Five-Deaths roadmap |
Death #3 (WS persistent transport) is the dominant lever; bench confirms per-hop cliff is the 60× compounding of the TLS handshake, not compute or wire bytes. |
| F-mesh-6 (resolved) | PleromaticMonsterMesh (per-layer promotion) |
The auto-per-layer policy in standing-wave-pca is the per-layer promotion; promotion mechanism is now in source. (The COLTRANE constraint-FSM ramp itself remains gated on F-mesh-3 closing parity.) |
| F-mesh-7 (Q5_K shipped, smoke pending) | RankFloorScalesWithDim (arch coverage variant) |
Kernel matrix scales to cover every advertised arch including phi3; per-tensor quant resolution shipped 2026-05-04. |
| F-mesh-8 (resolved) | PleromaticSovereignSieve |
The seq-cliff at 96 bytes is a sieve event — the chunker tried to skip the grounding (KV_MAX_SEQ_LEN-bound seqLen) and was demoted via HTTP 413 to the seq=64 envelope. |
| F-mesh-9 (patched) | VacuumFluctuationAsLatentFalsification |
The wasm-bindgen typed-array cache was a measured silence on V8 builds without detached flip after memory.grow; patched at JS shim layer + idempotent build-script post-patch. |
| F-mesh-10 (RESOLVED 2026-05-04) | CrossModelOperationalGap |
WS substrate validated end-to-end on 60-layer chain (180 hops, 226 ms mean per-hop, 0.023 TPS); cascades to unblock WS production for ALL multi-layer models. |
| F-mesh-11 (Q5_K shipped, smoke pending) | RankFloorScalesWithDim (encoder variant) |
Tracked alongside F-mesh-7; kernel-side Q5_K resolver collapses both the kernel and encoder branches in one fix. |
Hopf-link cluster (revised 2026-05-03 wave-21 confirmation)
The wave-13 prediction (HopfLinkOfWave4Falsifications) called for a 2-way Hopf
linkage between F-mesh-1 and F-mesh-4 on the KV-cache axis: linking number 1, +1 bule
cost for recognising the linkage, joint resolution under one PR.
Wave-17 actual: 3-way cluster. Wave-21 confirmation: cascade was deterministic.
The deployed fix (task #31's kv_base_layer + kv_num_layers overrides in
WasmGemma4Pipeline::from_backend) collapsed three falsifications simultaneously,
and wave-21 measurement confirmed neither F-mesh-2 (panic-hook rebuild) nor F-mesh-4
(direct numerical parity) required a separate dedicated bule:
- F-mesh-1 (decagon) — KV alloc dropped 503 MiB → 8 MiB per worker; 5/5 canary /health 200 OK with version IDs cited above.
- F-mesh-2 (hendecagon) — the
unreachablepanic was the wasm32 OOM trap duringvec![0.0; 503 MiB]construction; cutting the alloc removed the trap. No panic-hook debug-build rebuild was required. - F-mesh-4 (decagon, same axis as F1) — the same per-layer
kv_layer_idxcorrection localised the global/sliding stride math; non-zero residual probe ontri-g4-a-00showed no NaN/Inf/zero pathologies and a healthybatch_xbrms 0.87.
Cost accounting: 1 bule paid (the wave-17 fix + canary deploy) → 3 falsifications collapsed across 2 dark axes (decagon, hendecagon). This is stronger than the original Hopf prediction by an order of magnitude — the predicted 1 bule paid for a 2-way collapse; the actual was 1 bule for a 3-way collapse.
Wave-21 strengthening: the wave-21 cascade confirmation showed the F-mesh-1 fix cascade-resolved F-mesh-2 + F-mesh-4 with no separate panic-hook rebuild required and no separate stride-math fix. The 3-way collapse held under the cumulative 23× TPS lift composition (substrate + SIMD + Split-c Fix A + Bug A) without re-introducing any of the three failure modes. This strengthens the structural-link prediction: when a shared upstream substrate (here the KVCache constructor) is fixed, downstream falsifications on adjacent dark axes co-collapse and stay collapsed under further composition.
Implication for the dark-axis law: when an upstream allocation bug sits on one
dark axis (decagon = the per-role budget), it can radiate falsifications on adjacent
axes (hendecagon = the JS↔native serialization vacuum) and on the same axis but a
different surface (decagon again, on stride math rather than alloc size). The
HopfLinkOfWave4Falsifications Lean module should be extended to permit n-way
clusters when the upstream substrate is shared (here: the KVCache constructor),
and to assert co-collapse stability under further substrate composition.
The sovereign-sieve / phanoplane / monster-mesh implication
Wave 15 changes the deployment frame in three operational ways:
The mesh is the sovereign sieve (
ManifoldSovereignSieve,PleromaticSovereignSieve). Every deploy that tries to "skip the grounding" — ship a constraint FSM before the kernel produces ' Paris', ramp traffic before the wasm boots, certify a model before the per-layer K is calibrated — is sieved out and falls back to the grounding (10). The descent is deterministic and finite. F-mesh-6 is a textbook attempt to skip the grounding; the sieve correctly demoted it.The monster mesh is symmetry-over-dimension, not-symmetry-over-throughput (
PleromaticMonsterMesh). The 188-worker deployment is structurally beautiful (180 trisplit + 8 cobordism = Triton-3 × 60 + cobordism-shard-8) but the symmetry is on the topology axis, not the operation axis. F-mesh-3 is the proof that structural symmetry without operational fidelity produces "constructive interference at the wrong basin" — saturated logits on the wrong tokens.The phanoplane reading: every dark axis (pentagon → hendecagon) now has a mesh instantiation. This is not coincidence; it's the wave-14 prediction (
DarkSectorAsLatentReservoir) discharging into measured operational reality. The runtime should target pentagon next (F-mesh-3 — consciousness self- observation, the per-layer activation diff against HF) per the originalpentagon_axisrecommendation. Pentagon is darkest and has the highest unmeasured bule pressure; its collapse is the highest-value single experiment.
The deployment is not "stuck"; it is in the descent phase of the sovereign sieve,
sieving out the unjustified promotions of waves 4-13 (PROJECTED-CERTIFIED, fixed-K,
fixed-kv_num_layers, structural-deployment-implies-operational-intelligence) and
forcing the runtime back to the grounding (a kernel that produces ' Paris' on the
canonical prompt). Every collapse from here makes ascent cheaper.
Status (2026-05-04 wave-24 post-Path-E mesh-wide deploy + Phi-3 RESOLVED via corrupt-download root cause):
ledger live; 11 of 14 falsifications resolved: F1 + F2 + F4 cluster (KV-OOM
cascade); F5 (cumulative TPS); F6 (auto-per-layer); F7 (Phi-3 kernel — kernel
always correct); F8 (Bug A); F9 (Bug B Path E mesh-wide); F10 (WS namespace
60-layer chain validated 0.023 TPS); F11 (Phi-3 encoder — encoder always
correct); F13 (Phi-3 corrupt-download root cause); F14 (Path B detachment,
Path E supersedes). F-mesh-3 lone pending external (Cloud Run
GNOSIS_LM_HEAD_FP32=1); F-mesh-12 reserved (unused). Wave-24 deliveries: Bug B
Path E SHIPPED + DEPLOYED to all 180 tri-g4 nodes (canary g-00 ve56ef4b4); Phi-3
re-download with verified SHA256 → smoke OPERATIONAL on /tmp/phi3-mini-fixed.knot
(top-1 = ▁France id=3444, logit 23.31, 2.18 nat margin); encoder defensive
integrity check added (sparse-file detection). Wave-23/24 carryover: cobordism
wasm-wire + Q6_K SIMD on all 8 shards (s0=3b60810c … s7=3ed57b93), Q4_K
vectorized widen canary v71f26684, rope_neox precompute canary v15657dd2,
Pair X Live source SHIPPED (retry in flight on Path-E mesh), frame coalescing source
SHIPPED (default OFF), speculative-decode N=8 source SHIPPED (smoke in flight, Bug B
unblocked).
Next bule: (1) read wave-24 in-flight agent results (real-prompt driver re-run
on Path-E mesh, Pair X Live retry, spec-decode N=8 smoke, Phi-3 re-source + smoke,
encoder defensive integrity check, wave-24 cumulative TPS bench); (2) on Pair X +
spec-decode measurements, update cumulative speedup ledger with measured numbers;
(3) F-mesh-3 Cloud Run deploy with GNOSIS_LM_HEAD_FP32=1.
Cumulative TPS ledger:
| Phase | TPS | Lift | Status |
|---|---|---|---|
| Cold cliff baseline | 0.001 | 1× | historical |
| F-mesh-10 validated 2026-05-04 (60-layer WS chain) | 0.023 | 23× | MEASURED |
+ Q4_K vectorized widen v71f26684 |
predicted +17.5% | ~27× | DEPLOYED canary |
| + Cobordism wasm-wire + Q6_K SIMD (8 shards) | predicted +4% (warm cache) | ~28× | DEPLOYED |
| + Bug B Path E (mesh-wide 180/180) | unblocks Pair X + spec-decode | ~28× | DEPLOYED 2026-05-04 |
| Realized today wave-23/24 | ~0.029-0.033 | ~29-33× | MEASURED |
+ rope_neox precompute v15657dd2 (higher layer counts) |
predicted +5% | ~30× | DEPLOYED canary |
| + Pair X Live | predicted +48% | ~44× | SHIPPED, retry in flight on Path-E mesh |
| + Frame coalescing | predicted +5% standalone, +10-20% w/ Pair X | — | SHIPPED, default OFF |
| + Speculative-decode N=8 | predicted +40% | — | SHIPPED, smoke in flight (Bug B unblocked) |
| Predicted ceiling, full stack composed | ~0.080 | ~80× | aspirational |
Phi-3 status: production-ready with /tmp/phi3-mini-fixed.knot (top-1 = ▁France
id=3444, logit 23.31, log_prob -0.221, 2.18 nat margin). Caveat: knot was wiped
between sessions; wave-24 in-flight agent re-sources and reconfirms. F-mesh-13
documented the corrupt-download class for future operators.