System Invariants (Stable Theory Contract)
This document defines the non-negotiable behavioral invariants for the DEXBot2 system. It is a contract for code review and release safety, not a design tutorial.
Scope
- COW Pipeline: planning, projection, reconciliation, commit, and fund accounting flows.
- Sync Engine: fill history sync, open-order sync, orphan detection, cache vs chain consistency.
- Maintenance Runtime: pipeline gating, illegal-state handling, cooldown semantics, dust ordering.
- Grid & Reconcile: price-level uniqueness, dust detection scope, reconcile cancel semantics.
- Batch & Pipeline: stale handling, retry gating, stale-flag cleanup.
- Fund Registry: cross-bot allocation invariants.
- Subscriptions: health watchdog, silent-death detection.
- Broadcast: uncertain-broadcast recovery, retry, deadlock-free reconcile.
Invariant Prefixes
| Prefix | Subsystem |
|---|---|
INV-COW |
COW pipeline |
INV-PROJ |
Projection |
INV-ID |
Order identity |
INV-ACC |
Accounting / fund tracking |
INV-SYNC |
Sync engine |
INV-MAINT |
Maintenance runtime |
INV-GRID |
Grid structure |
INV-RECON |
Reconcile |
INV-BATCH |
Batch / pipeline |
INV-STATE |
State / lifecycle |
INV-REG |
Fund registry |
INV-SUB |
Subscriptions |
INV-BROADCAST |
Broadcast |
COW Pipeline
INV-COW-001Master immutability until commit- The master grid must not be mutated during planning/execution prep.
- All intermediate mutations happen in
WorkingGrid. - Master updates occur only during commit after guard checks pass.
INV-COW-002Commit atomicity- Commit swaps working state to master atomically.
- On failed/aborted execution, working state is discarded and master remains unchanged.
INV-COW-003Pre-broadcast staleness guard + bounded re-plan- A plan whose working grid went stale before broadcast (master mutated mid-planning: fills, syncs) must never be broadcast as-is.
- Re-plan ONCE from fresh master with the same fills (
replanStaleBatch, bounded bySTALE_PLAN_REPLAN_LIMITindexbot_cow_runtime.ts); restore the boundary-shift budget consumed by the abandoned plan; clear only the abandoned batch's OWN pending-broadcast entries (earlier unresolved batches' entries are kept — the recursion guard then aborts + reconciles instead of re-creating possibly-landed slots). - Still stale or no fill context → proceed + structural resync (never a silent abort that drops the fill set); commit-time guard + chain adoption close residual divergence.
INV-COW-004Working-grid stack exactly-once push/pop- Pushes go through the single
_pushWorkingGridRef(sets_workingGridPushedmarker); releases go through_popWorkingGridRef(marker-guarded, early-return sites) or_releaseWorkingGridRef(identity-checked, commit sites). - A result that was never pushed must never pop — that would underflow or steal a nested grid's stack entry.
_commitWorkingGridreleases the entry exactly once on every settle path (return or throw).
- Pushes go through the single
INV-COW-005Grid regeneration bumps_gridVersion_clearOrderCachesLogicmust bump_gridVersionso an in-flight COW plan (baseVersion from the pre-swap grid) is refused at commit instead of committing over a regenerated zero-slot grid.
INV-COW-006Boundary moves only on fills and same-batch spread promotion- Fund changes never move the boundary. Only boundary crawl (
deriveTargetBoundary,order/utils/order.ts) and spread promotion move it, and promotion shifts it only onto orders placed in the same atomic batch (prepareSpreadCorrectionOrders). - Funds drive order sizing/budget allocation only;
syncBoundaryToFunds/calculateFundDrivenBoundaryare deleted (1.5.3).
- Fund changes never move the boundary. Only boundary crawl (
INV-COW-007Boundary hold on guard-skipped refills- When a slot on the plan's refill wire (
collectRefillSlotIds: the plan's CREATE ids minus reserve-ladder edge ids) was guard-skipped at broadcast (skippedUpdateSlotIds∪clampedUpdateSlotIds∪skippedCreateSlotIds),resolveRefillBoundaryHoldkeeps the COMMITTED boundary instead of committing the plan's shifted one over the stranded rail hole; the grid still commits. - Reserve-ladder CREATEs are excluded because reserves are static edge insurance whose fills never crawl — a skipped reserve must not pin geometry.
- The hold rides to the commit as
options.boundaryHeld, which suppresses the pending-crawl clear. trackBoundaryHoldcounts consecutive hold batches (manager._consecutiveBoundaryHolds, cleared on the first clean batch) with a last-hold snapshot (manager._lastBoundaryHoldInfo); runs of 3+ escalate visibly. It also records a hold signature (manager._lastHeldPlanSignature: committed boundary + guard pivot +_lastFilledAt+ held wire) so a fill-less re-plan of an unchanged state defers instead of re-broadcasting the identical vetoed plan.- Fill-driven rebalances never sleep on
_broadcastingFlaginside_fillProcessingLock:performSafeRebalance({deferIfBroadcasting:true})returns adeferredresult and the region-end hook schedules one no-fill rebalance (schedulePostRecoveryRebalance). This prevents the 20s lock-acquisition-timeout cascade when a broadcast region outlives the 30s_awaitBroadcastIdlewait. The stale-flag watchdog (_clearStaleBroadcastFlag, 120s) fires the same region-end hook when it hard-clears a leaked flag, so deferred fills are woken even when the region never ends throughstopBroadcasting(). - A hold run carrying fresh fills (>=
TIMING.BOUNDARY_HOLD_RESYNC_THRESHOLD) requests a guard-aware structural re-center (requestStructuralGridResync, cooldownTIMING.BOUNDARY_HOLD_RESYNC_COOLDOWN_MS) instead of holding forever — the heal path when the grid is trailing the market.
- When a slot on the plan's refill wire (
INV-COW-008Owed fill crawls survive unapplied commitsorder/strategy.tsrecords a crawl per shift-eligible fill at intake (slot-level dedupe, capped at 500);deriveTargetBoundaryfolds still-owed_pendingFillCrawlsinto the next derivation, excluding the current batch's slots, reserve slots, and window members — a window that reaches the grid edge is not a reserve, using the same window exclusion every placement picker andcountLiveReserveOrdersapply._commitWorkingGridclears them only when the plan's boundary was actually applied — a held boundary (boundaryHeld), a gate-rejected boundary, and a null boundary all leave them owed.consumePendingFillCrawlsapplies them onto a restored finite boundary;applyPersistedPendingCrawlsis the shared startup/recovery wrapper used before sync/reconcile;_clearPendingFillCrawlsdrops them when the boundary is re-anchored (grid rebuild viainitializeGrid, rejected snapshot viarejectCorruptedGridSnapshot, persisted snapshot wipe viaAccountOrders.clearGrid).
INV-COW-009LAST-FILL-GUARD pivot survives restarts with the snapshot, and only as a fill fact- Single source of truth: the manager's in-memory pivot is authoritative at runtime; the disk row is a mirror that rides the grid snapshot (
persistGridSnapshot→storeMasterGrid10th param →lastFillPivot), so it invalidates in lockstep with the boundary/genesis instead of forming a second ledger (the_pendingFillCrawlslockstep contract). - Provenance gate: only a pivot written through the shared
setLastFillPivot(type, price, 'fill')writer is persist-eligible. A book-seeded pivot (seedLastFilledPricesFromBook, provenance'book') is a heuristic — max resting buy / min resting sell is not market truth — and is never fossilized into the snapshot. The fill-batch and queued-fill refresh sites funnel through one implementation (free function inutils/system.ts;OrderManager._setLastFillPivotdelegates to it), so value, timestamp, and provenance cannot drift apart; the queued-fill refresh and the restore path call the same free function directly, with the legacy-stub fallback built into the shared writer rather than copy-pasted. A single-sided book arms through the same writer with provenance'book'; the two-sided midpoint and per-side mirror seeds stay direct writes by design, because they carry no pivot provenance and must not be persist-eligible. - The row carries
{price, type, fillsAt, genesisHash}validated by one shared gate (normalizeLastFillPivot) used by BOTH thestoreMasterGridsanitizer and the loader — the two shape checks cannot drift.nullfrom the payload builder always clears the stored row (consumed pivot / cold manager / no live genesis),undefinedstays a legacy no-op for old callers, and a malformed shape also clears. - Restore ordering:
grid.loadGridrestores the pivot right after_restoreBoundary— with the persisted genesis applied and the grid re-typed, BEFORE the first reconcile/broadcast — so a restart cannot re-open the cold window where evacuation stamps get applied against a boundary still being rebuilt. AllloadGridcallers (startup resume, price-match resume, recovery reload) inherit it; there is no second restore call site to drift. - Validation chain (
restoreLastFillPivot): TTL expiry (GRID_LIMITS.LAST_FILL_PIVOT_TTL_MS, 24h; the pivot is a "latest fill" fact, not a permanent ratchet; expired pivots hand over to the startup book seed) and genesis binding (mismatched hash = dead generation) erase the row through one shared drop path (clearPersistedLastFillPivotunder the bot's persistence lock, best-effort). Off-grid refusal (resolveOnGridPivot— the runtime's own ladder validator, imported lazily from the restore site so ONE implementation serves both the per-probe live guard and the restore; the restorer can never accept a value the runtime would refuse per probe) is a no-op that falls back to the book seed and leaves the row for the next snapshot flush to clear (the cold in-memory pivot makes the payload builder return null). The snapped ladder level is restored, never the raw float, and the ORIGINALfillsAtis preserved throughsetLastFillPivot'satMsparam so the TTL keeps meaning "age of the last fill", not "time since this restart". - Generation invalidation mirrors the pending-crawl ledger:
initializeGridandrejectCorruptedGridSnapshotclear the in-memory pivot viaresetLastFillPivot(exposed asOrderManager._resetLastFillPivot), which clears the FULL scalar family — per-side mirrors included — soseedLastFilledPricesFromBook's cold gate cannot be silently suppressed by stale per-side values after a grid rebuild, andAccountOrders.clearGridwipes the persisted row with the snapshot. The guard re-arms on the first real fill of the new generation.
- Single source of truth: the manager's in-memory pivot is authoritative at runtime; the disk row is a mirror that rides the grid snapshot (
INV-PROJ-001New projected orders remain virtual- Orders projected into empty slots must be
VIRTUALwith noorderIduntil chain confirmation.
- Orders projected into empty slots must be
INV-PROJ-002Preserve on-chain PARTIAL size in projection- If identity is retained (
keepOrderId=true) and current state isPARTIAL, projected size must preserve current on-chain remaining size. - It must not be overwritten by ideal geometric
targetSize. - Exception: a
PARTIALwith a rotation/size-update action targeting itsorderIddoes usetargetSize(the explicit-UPDATE path atmodules/order/utils/validate.ts:1129). - Preserve-path size must be normalized to finite, non-negative value.
- If identity is retained (
INV-PROJ-003ACTIVE on-chain projection preserves current size (same as PARTIAL)- If identity is retained and state is
ACTIVE, projection preserves current on-chain size via the sameshouldPreserveSizepath asPARTIAL(validate.ts:1129). - An explicit UPDATE action targeting the
orderIdis required to applytargetSize.
- If identity is retained and state is
INV-ID-001Order identity retention ruleorderIdand on-chain state are retained only when order is on-chain and side/type is unchanged.- Otherwise projected order becomes
VIRTUALwithorderId=null.
INV-ACC-001Committed accounting source of truth- Committed chain/grid totals derive only from on-chain orders (
ACTIVE/PARTIALwithorderId) and their projected sizes. - Virtual orders contribute only to virtual pools, not committed chain totals.
- Committed chain/grid totals derive only from on-chain orders (
INV-ACC-002Fund invariant consistency (INVARIANT 1)- Tracked totals must remain consistent with blockchain totals within
FUND_INVARIANT_PERCENT_TOLERANCE. Total = Free + Committedmust hold per side.- False violations due to ideal-size projection overstatement are prohibited.
- Tracked totals must remain consistent with blockchain totals within
INV-ACC-003Cross-bot fund registry invariant (INVARIANT 3)- Shared-account per-bot commitment must not exceed the bot's proportional share of chain balance.
- Checked with widened tolerance
max(PERCENT_TOLERANCE * 3, 0.15). - Registry failure logs an error (
order/accounting.ts:574-590, with a "CRITICAL FIX: Log as ERROR instead of WARN" comment), not a silent skip.
Sync Engine
INV-SYNC-001Drift-refetch for partial fill correctnessisEffectivelyFullmust not unconditionally trust cachedrawOnChain.for_sale.- When cached
for_saleis smaller than the grid's own size (drift signal), refetch from chain viareadSingleOrder. isEffectivelyFullrequires either: chain-confirmed empty (refetch returned null), grid also at 0, or the other side rounding to 0.- The legacy
newSizeInt <= 0fast-path is forbidden.
INV-SYNC-002No TTL-based chain refetch- Chain is only consulted on a drift signal (cache < grid size).
- Time-based TTL refetches are prohibited — a stale-but-consistent cache is still correct.
INV-SYNC-003Drift refetch null = chain-confirmed empty- When drift refetch returns null from
readSingleOrder, the order is authoritatively gone. - The sync engine must treat this as a chain-confirmed empty, not a connection failure.
- When drift refetch returns null from
INV-SYNC-004Orphan adoption rejects duplicate price levels- Pass-2 fallback adoption must check whether an active grid order of the same type already exists at the same price (within
calculatePriceToleranceusingMath.max(size)). - If a duplicate exists, the orphan is NOT adopted; it is pushed to
unmatchedChainOrdersfor chain cancel. - Size is irrelevant — any duplicate violates the one-order-per-price-level invariant.
- Pass-2 fallback adoption must check whether an active grid order of the same type already exists at the same price (within
INV-SYNC-005Fill queue back-pressure- Subscription callbacks must not push fills beyond
MAX_INCOMING_FILL_QUEUE. - Rejected fills must leave subscription cursors retryable; the queue must remain unchanged.
- Back-pressure must not start a consumer for rejected fills.
- Subscription callbacks must not push fills beyond
INV-SYNC-006syncFromOpenOrders acquires fill-processing locksyncFromOpenOrdersacquires_fillProcessingLockby default.AsyncLockis re-entrant — callers already inside the lock rely on the intrinsicisReentrant()check instead of afillLockAlreadyHeldparameter.- Direct call sites without the lock contract are prohibited.
INV-SYNC-007Authoritative sync preserves fetched free balances- Free-balance values fetched during a sync must not be silently discarded; they seed the optimistic balance model.
- If fetched balances are unavailable (failure/error), the optimistic model remains untouched rather than zeroed.
synchronizeWithChainmust not double-deduct already-locked funds from fetchedbuyFree/sellFree.- After authoritative open-order sync,
checkFundDriftAfterFillsis expected to returnisValid=trueas a design consequence (no enforcement code exists — this sub-clause expresses the intended post-condition, not an assertion).
INV-SYNC-008SPREAD-typed fills resolve the real side- A SPREAD slot can carry an on-chain order (spread-correction activation). Fill processing must resolve the real BUY/SELL side (chain order type, pays asset, or price-vs-startPrice convention) before pushing the fill or computing the transition.
- A SPREAD-typed fill would silently drop the
deriveTargetBoundaryshift and produce an illegal SPREAD+on-chain state rejected byvalidateOrder.
Maintenance Runtime
INV-MAINT-001Pipeline in-flight defers maintenance- When
isPipelineEmptyreturnsisEmpty=false(batch in-flight, recovery-in-flight, or broadcasting active),checkSpreadConditionand divergence corrections must be skipped. - Dust cancels broadcast directly to chain and are safe to call regardless of pipeline state.
- Pipeline signals (
batchInFlight,recoveryInFlight,broadcasting) must be passed toisPipelineEmpty.
- When
INV-MAINT-002Illegal state abort- When
consumeIllegalStateSignalreturns a non-null signal, the maintenance cycle must:- Trigger immediate recovery sync (
_triggerStateRecoverySync). - Skip persistence (
_persistAndRecoverIfNeeded) in the same cycle. - Skip remaining maintenance steps (spread checks).
- Arm one maintenance cooldown cycle (
_maintenanceCooldownCycles = 1).
- Trigger immediate recovery sync (
- When
INV-MAINT-003Maintenance cooldown- After a hard-abort (illegal state), the next maintenance cycle must skip all checks (spread, health, etc.).
- Cooldown is exactly one cycle; normal checks resume on the subsequent cycle.
- Cooldown must not stack or compound.
INV-MAINT-004Maintenance idle gate- Maintenance must wait for a quiet window with no fill queue, sync, or batch activity (
BLOCKCHAIN_SETTLE_DELAY_MS). - Recent activity tracking covers: fill queueing, fill processing completion, COW batch start/end, open-order sync, periodic fetches.
- Maintenance must wait for a quiet window with no fill queue, sync, or batch activity (
INV-MAINT-005Dust-first ordering- Dust partials are cancelled immediately on detection — no delay, no timer.
- Grid resync and structural maintenance wait only for the idle settle delay.
Grid Structure
INV-GRID-001One-to-one order mapping- One grid slot = at most one on-chain order. No two chain orders may map to the same grid slot.
- Sync engine tracks
matchedGridOrderIdsthrough both sync passes and skips already-matched slots. - Surplus orders (matched count above the per-side
activeOrders+reserveOrderstargets) are flagged for cancellation.
INV-GRID-002One order per price level- The active grid must have at most one on-chain order per (type, price) pair.
- Duplicate price levels are a structural violation — the sync engine must reject orphan adoption at a duplicate price, and the reconcile layer must cancel offenders on chain.
- Size is irrelevant; any duplicate violates this invariant.
INV-GRID-003Interior dust detection at duplicate price levels- Interior partials (not top-of-window) are eligible for dust detection if they share a price level with an ACTIVE sibling.
- Top-of-window partials remain always eligible.
- Two PARTIALs sharing a price with no active sibling do not qualify (left to rebalancer).
INV-GRID-004Slot price equals its genesis level (GRID_PRICE_INVARIANT.md)order.pricefor a slot-idxorder must equalpriceForSlot(idx, genesis); the genesis ladder is the only authoritative price for a slot.- Enforced at all six emission sites (CREATE / UPDATE / CREATE-FALLBACK, RECONCILE-CREATE / RECONCILE-UPDATE, STARTUP-CREATE): an off-grid emission is blocked, never broadcast.
- Range guards (
isChainPriceOutOfGrid) are bounds checks, not membership checks — they cannot substitute for this invariant. - Adoption keeps the slot's own level (a fill/chain price is metadata, not the slot's price);
loadGridrepairs a pre-existing off-grid slot price at load. - A snapshot carrying orders but no usable ladder is refused before any mutation (
loadGrid, before the startup resume decision, and at the persisted-row schema), and the sync entry refuses to reconcile a ladder-less grid at all — without a ladder no consumer has authority for a slot's price, which is what this invariant exists to prevent. The tolerance matcher that used to cover that state has been removed.config.gridLimits.MISSING_GENESIS_POLICYpicks the response:'rebuild'(default, structural resync) or'halt'(manual reset). One decision, one module:modules/order/genesis_policy.ts. - Tests: GPI-001..015 (
tests/test_grid_price_invariant_guard.ts), GPI-WIRE-001..009 (tests/test_grid_price_invariant_wiring.ts), LEGACY-ADOPT/MATERIALIZE/ADOPT-NAME (tests/test_sync_out_of_grid_defer.ts), GEN-01..21 (tests/test_missing_genesis_policy.ts).
Reconcile (GRID_RECONCILE.md)
INV-RECON-001Rotation-only size updates in reconcilereconcileGriddoes not emit generic in-place size UPDATEs for active slot diffs.- Size-changing UPDATE actions are rotation updates (
newGridIdpath). - Non-rotation size correction is handled by dedicated maintenance flows.
INV-RECON-002Dust health gating parity- Dust health thresholding applies consistently to both CREATE and rotation destination holes.
INV-RECON-003Reconcile cancels duplicate chain orders unconditionally- An unmatched chain order whose price equals an active same-type grid slot's price (exact slot-price equality via
priceSlotEqualat the asset precision) is a suspected duplicate and must be cancelled on chain via_cancelChainOrderwithreleaseUntrackedFunds: true. - Cancelled IDs are filtered out of
unmatchedParsedto prevent reprocessing. - No size guard — any duplicate at the same price is a violation.
- The earlier fuzzy
SUSPECTED_DUPLICATE_TOLERANCE_MULTIPLIER(5×calculatePriceTolerance) andSUSPECTED_DUPLICATE_TOLERANCE_FLOORare removed; only exact price-level equality triggers a reconcile cancel.
- An unmatched chain order whose price equals an active same-type grid slot's price (exact slot-price equality via
INV-RECON-004Rebalance must not convert on-chain slots to SPREAD via CREATEperformSafeRebalancemust not emitCREATEactions that convert existing on-chain slots into SPREAD orders.- On-chain mid-slot must keep its BUY/SELL type before commit.
INV-RECON-005Extreme placement ordering- BUY placements must use nearest available free slots first (descending price, so the nearest-to-center slots fill first —
order/utils/order.tsbuildOutsideInPairGroups). - SELL placements must use nearest available free slots first (ascending price).
- BUY placements must use nearest available free slots first (descending price, so the nearest-to-center slots fill first —
Batch / Pipeline
INV-BATCH-001Illegal state batch abortexecuteBatchthrowsILLEGAL_SPREAD_STATEon an illegal grid layout (emitted atmodules/order/utils/validate.ts, propagated viamodules/order/manager.ts_throwOnIllegalState).- The
_handleBatchHardAbortcatch forILLEGAL_ORDER_STATE(dexbot_class.ts:455) is a test-only dead branch — production never emits that code; only a test stub uses it. - In production, recovery + cooldown are armed on the next maintenance tick via
_abortFlowIfIllegalState(theINV-MAINT-002path), returningabortedForIllegalState: trueto the caller. The caller does not need to return immediately; the maintenance tick handles recovery. - Hard abort triggers one immediate recovery sync (
_triggerStateRecoverySync) plus arms one maintenance cooldown cycle (_maintenanceCooldownCycles = Math.max(current, 1)).
INV-BATCH-002Stale-only cancel fast path- When a cancel-only batch fails with "order does not exist" (stale), the handler must:
- Virtualize the slot (
state = VIRTUAL,orderId = null,size = 0,rawOnChain = null). - Return
stale: true. - NOT trigger a recovery sync.
- Track the stale order id in
_staleCleanedOrderIdsto prevent double-credit. - Preserve manager index validity.
- Virtualize the slot (
- When a cancel-only batch fails with "order does not exist" (stale), the handler must:
INV-BATCH-003"Cannot deduct" → recovery sync- When
executeBatchthrows "Cannot deduct all or more from order than order contains", the handler must:- Return
recoveredBySync: true,reason: 'ORDER_SIZE_DRIFT'. - Trigger one recovery sync.
- NOT virtualize the slot.
- Preserve
orderIduntil sync reconciles it. - NOT mark the order as stale-cleaned.
- Return
- Fast path: if the batch result indicates
ORDER_SIZE_DRIFT_TARGETED(dexbot_state_recovery.ts:269), a targeted repair applies the correction directly and skips_triggerStateRecoverySync.
- When
State / Lifecycle
INV-STATE-001Bootstrap suppresses invariant checks- While
isBootstrapping()is true,_verifyFundInvariantsmust be suppressed. - Suppression covers
recalculateFundsand_updateOrdertrigger paths.
- While
INV-STATE-002Recovery validation not masked by bootstrap_performStateRecoverymust detect drift even whenisBootstrapping()is true.- Bootstrap suppression of invariant checks must not also suppress recovery validation.
INV-STATE-003Grid resize respects budget after capping- BUY allocation must not exceed calculated budget after
_recalculateGridOrderSizesFromBlockchaincapping.
- BUY allocation must not exceed calculated budget after
Fund Registry
INV-REG-001Cross-bot allocation ≤ proportional share- Per-bot committed amounts (sum of on-chain orders) must not exceed
totalChainBalance × allocatedPercent. - Violation triggers an error-level log entry (not silent), with tolerance
max(PERCENT_TOLERANCE * 3, 0.15). - Registry registration is pre-flight + atomic; only shared-account bots register (
dexbot.ts:611filtersaccountGroups[a].length > 1), and registration completes before any shared-account bot starts. - Release happens in
DEXBot.shutdown.
- Per-bot committed amounts (sum of on-chain orders) must not exceed
INV-REG-002Async-locked registry writes- All
fund_registrymutations (register, update, release) must hold an async lock. - Concurrent writes from multiple bots sharing the same account must be serialized.
- All
Subscriptions
INV-SUB-001Fill discovery via periodic pollingstartFillPollinginvokesprocessObjectsper active subscription everyFILL_POLL_INTERVAL_MS(default 60s).entry.activeandentry.reconnectingflags gate per-subscription work.fillPollInProgressflag prevents overlapping poll ticks.
Broadcast
INV-BROADCAST-001BROADCAST_DEADLINE graceful recovery- On
BroadcastUncertainError, the bot must:- Retry once with a fresh deadline window (
_executeWithRetryOnUncertain). - Skip retry when
err.partialOnChainStateis true (pair-mode grouped execution). - After retry expiry, reconcile —
AsyncLockis re-entrant so no specialfillLockAlreadyHeldparameter is needed.
- Retry once with a fresh deadline window (
- Daemon broadcast (
broadcastWithDeadline) pins allCREDENTIAL_DAEMON_BROADCAST_RETRIESattempts to one node; only when they all fail with provably untransmitted failures (pre-send errors only) does it report the node to the health ledger (threshold → blacklist + shared health-cache exclusion, logged) and rotate to the next best node. Re-signing an uncertain broadcast would duplicate a landed transaction on chain (new tx ID per signature), so BROADCAST_DEADLINE is reported to the bot and retries happen only after chain verification.
- On
INV-BROADCAST-002Deadlock-free reconcile after uncertain broadcast_reconcileAfterUncertainBroadcastdoes not need afillLockAlreadyHeldflag becauseAsyncLockis re-entrant — a secondacquire()from within the same execution context runs the callback directly instead of queueing.- The lock hierarchy requires
_syncLock(2)to be acquired before_gridLock(3)in all paths; nogridLockAlreadyHeldbypass flag exists.
INV-BROADCAST-003Verify-before-retry; never re-sign an uncertain broadcast- An uncertain outcome (RPC timeout, connection dropped with a response pending, unknown code) must never be re-signed — a re-sign would land a duplicate transaction on chain (new tx ID per signature).
- Retries are limited to provably-untransmitted failures (pre-send connection/frame errors) and re-broadcast happens only after an AUTHORITATIVE absence read (non-empty, non-truncated chain read containing none of the batch's CREATEs).
- Direct-key and claw broadcast paths classify failures via
modules/broadcast_failure.tsand surface typedBroadcastUncertainErrorso the COW machinery engages instead of a blind error.
INV-BROADCAST-004Truncated/empty reads are ambiguous, not authoritative absence- A
get_full_accountswindow capped at the order limit omits the freshest orders (by_account index = seller,id ascending — fresh creates sort last), so absence on a truncated/empty read can never free slots/capital, confirm a cancel, discard a batch, or clear pending-broadcast protection. - All absence decisions use
readOpenOrdersWithMeta; truncated/empty snapshots defer (keep protection + request structural resync) instead of acting.
- A
Change Policy
- Any intentional invariant change must:
- Update this document in the same PR/commit.
- Include explicit rationale and risk note.
- Add or update regression tests in the relevant test files.