TFL — Token Flow Laboratory
An interactive, deterministic simulation of how a language-model
serving system carries one request from prompt to next token.
Everything runs in your browser as a teaching model.
Scope: timings, capacities and token text are
synthetic. No model, neural network or inference engine runs here, no
hardware benchmark is executed, and tokenization is illustrative.
Time-to-first-token is calculated from simulated events, not measured
latency.
What it covers
-
Queueing, prefill, KV-cache reservations, decode steps and token
delivery.
-
Nine experiments: one request, load storm, steady arrivals, 32K
context, code generation, many short users, KV wall, overload and
batching trade-off.
- Ten guided checkpoints in Token Flow 101.
-
Assumptions and evidence shown beside every result, so a simulated
number is never read as a measurement.