
A market-making pipeline built to stream a 44 GB day through a 16 GB machine.
One day of order book deltas on a liquid crypto perpetual is 73.3 million events. Held as Python objects that is roughly 44 GB, and the machine this was built on has 16 GB. Everything in Kestrel falls out of that number: a day can never be loaded, only streamed, and it has to stream identically every time, because the whole point is to replay it twice and get the same answer to the byte.
The trading question comes second, and it is narrow. Does a simple inventory-aware quoter make money here, and how much of the backtest do I believe? Every naive market-making backtest looks profitable, because queue position cannot be simulated: a replay fills orders a real exchange would have left behind hundreds of others. The edge worth chasing is pricing and inventory risk, not latency. Backtest and testnet only.
Nautilus ships a Rust CSV streamer that returns two million rows at a time and keeps its reader alive across calls, so ingest feeds the verification book chunk by chunk and writes each chunk before loading the next: peak memory is bounded by the chunk size, not the day. Chunks land in a fresh staging catalog and move into the real one only after the whole day verifies, because once bad data is in the catalog every downstream result inherits the defect. Replay reads it back in 15-minute windows, batched per venue timestamp group, so a fill can never occur against a half-built book. Trades are part of the spine, not an extra: with only deltas, a resting limit order can never be traded through and the backtest is structurally fill-less.
Determinism is proven by a streaming SHA-256 digest of every observation, run twice, never by holding two observation lists and comparing them. The lists are the thing that does not fit.
Verification is a gate, not a report, because a plausible wrong book is invisible. Crossed books are judged at venue batch boundaries rather than per delta: 2.39 million deltas on the real day, 3.3% of them, were seen crossed, and every one was a transient inside a single venue message. Zero of 370,988 sampled completed batches were crossed. The per-delta model also burned about 85% of ingest wall clock, which fell from 75 minutes to 9.
Sequence semantics belong to the source, not the verifier. The first live Bybit capture flagged 10,912 gaps in 10,913 perfectly coherent batches, because Bybit's sequence is a venue-wide counter that jumps between messages for one symbol. Only a decrease proves lost ordering there, so callers declare the semantics; loosening the default was rejected, since on a per-topic feed a jump really is data loss.
Three traps are handled where they bite. The Rust CSV reader parses positionally, so a deltas file fed to the trades loader yields garbage without erroring: the header is validated before every ingest. The streamer cannot know a group ended at a chunk boundary, so the last-in-group flag is recomputed at write time, making catalog content chunk-size invariant. And the catalog's write silently skips a filename that already exists, so ingest stages fresh and refuses a re-ingest loudly rather than reporting a success it did not perform.
Spread capture per fill and markout at 1s, 10s and 60s after it are deliberately decoupled, so adverse selection shows up independent of spread earned. P&L decomposes into spread plus inventory minus fees as an exact identity, tested against independently computed mark-to-market. The first probe strategy, one with no signal at all, produced markouts of -0.42, -0.81 and -1.05. Textbook adverse selection, which is the metrics working.
An inventory-skewed quoter posts two-sided quotes around the microprice, the centre pushed against inventory so the book pulls the position back toward flat. Risk lives in a shared guard, never inside the strategy: a position bound checked against projected inventory, so it binds on what can execute rather than what already has, and a latching kill switch, because a strategy that trades through its own drawdown limit is one that will not come back one day.
The fill assumption is recorded on every run, as one of two named presets. PESSIMISTIC double-gates a touch through both the queue model and a fill probability; QUEUE_GATED makes the queue model the only gate. On the real day they produced 2,731 and 2,747 fills, 0.6% apart, because this quoter's fills are pick-offs, not queue-front touches. Latency is measured, not assumed: p50 34 µs and p99 61 µs across 1.96 million acting callbacks, three orders of magnitude under the 1.5 ms simulated venue round trip, so a Rust port stays parked until a loop is provably latency bound.
The number the whole phase was built to produce is fill inflation: backtest fills divided by live fills, measured by replaying a keyed testnet session's own captured window through the same engine with that session's quoter config and venue-true fees. Until it exists, backtest P&L is an upper bound, never a forecast.
Every run persists as a ULID directory: metadata carrying the assumptions, instrument, data range and git SHA, Parquet for fills, mid series, inventory and markouts, one SQLite row for the list and compare screens. A read-only FastAPI layer sits over it and a Next.js dashboard over the API: equity, inventory and a markout scatter beside a provenance panel, plus side-by-side comparison. Runs launch from the CLI, so there is no write surface at all.
Kestrel is not deployed. What exists is a recipe: two systemd units, a Dockerfile, and a runbook that targets one small VM.
The runbook targets a GCE e2-micro on the free tier, or an e2-small at about $13 a month if capture, trade and the API run together. Vercel and Cloud Run were rejected on shape, not preference: the engine is one long-lived process holding venue websockets open around the clock, request-scoped functions cannot hold a socket, and the catalog needs a disk that survives a restart.
The split between the units is the interesting part. Capture is keyless and restarts forever, because nothing it can get wrong is unfixable by waiting. Trade holds keys, and refuses to start unless the venue is flat: inventory is the strategy's own accumulator starting at zero, so an inherited position would skew, size and calibrate against a number that is not real. That refusal exits 2, and the unit sets RestartPreventExitStatus=2, so a non-flat venue stops the retry loop dead instead of refusing every thirty seconds forever. Tailscale means there is no auth code to write, and the read-only API could not trade even if reached.
Python 3.13 on Nautilus Trader 1.231.0, pinned over the v2 release candidate on purpose: the matching engine is the one place a subtle bug is indistinguishable from a strategy result. FastAPI, Next.js and React, Parquet and SQLite, uv, ruff and strict mypy. Load-bearing tests are mutation-verified: the code is broken on purpose and the test watched to fail before it is believed.