# Contextual speculative decoding measurement

Intel Core i9-14900KF, CPU only; Qwen3.5-0.8B Q8_0; identical200-token prompt IDs;64 emitted predictions;63 decode steps timed after prefill and first prediction. Each process warms one identical workload before a fresh-context measurement. Five rotated rounds, CPU busy fraction below0.05 before each sample. No EOS early stop. No sample is discarded.

llama.cpp headers and libraries bind to build10152, revision0324696b8; the source checkout is newer. The comparator uses lowest-index greedy argmax, f16 KV, flash attention auto, context268, batch/ubatch512, no GPU or operation offload. Threads8/16/24/32 are measured in every round. Eight threads wins every round on both cases. Apeiron uses its existing resident worker pool; this is tuned throughput, not an equal-worker-budget or energy comparison. Apeiron lookup speculation is compared with llama.cpp greedy; llama.cpp speculation is not measured.

| Context / engine | Median tok/s | Median paired ratio to fastest llama arm | Median paired ratio to own greedy |
|---|---:|---:|---:|
| README / llama greedy | 51.3103 | 1 | 1 |
| README / Q32 greedy | 75.5830 | 1.42308 | 1 |
| README / Q32 lookup4 | 92.1430 | 1.79580 | 1.26191 |
| README / f32 greedy | 90.9610 | 1.68849 | 1 |
| README / f32 lookup2 | 94.7677 | 1.90823 | 1.06794 |
| Code / llama greedy | 51.0165 | 1 | 1 |
| Code / Q32 greedy | 74.2637 | 1.45568 | 1 |
| Code / Q32 lookup4 | 110.7579 | 2.14408 | 1.48519 |
| Code / f32 greedy | 79.8245 | 1.47952 | 1 |
| Code / f32 lookup2 | 98.9212 | 1.84617 | 1.13540 |

Ratios are medians of per-round ratios, not ratios of the displayed medians. Code Q32 lookup4 spans2.01112–2.22282x versus the fastest llama arm over the five rounds. README Q32 lookup4 spans1.60709–1.91751x. Code f32 lookup2 spans1.12149–2.00648x; that variability prevents a dependable2x claim.

Every lookup arm emits the same64 IDs as its own greedy arm in every measured round. README emits the same64 IDs across all three engines. Code Q32 first diverges from llama at zero-based token27;28/64 IDs match. Code f32 first diverges at token58;58/64 IDs match. Numeric-carrier quality and longer output equivalence need separate acceptance. All llama thread settings emit identical IDs within each round.
