Advanced Features

The systems that make TensorSharp fast and scalable: continuous batching with a paged KV cache, speculative decoding, text diffusion, and the kernel- and memory-level optimizations underneath.

Continuous batching & paged KV cache

The server's InferenceEngine is a vLLM-style continuous-batching engine, on by default. Instead of running one request at a time, it interleaves many at the granularity of a single decode step.

Models that have not implemented the batched path still run on the engine's isolated per-sequence KV-swap fallback. Tune it with the TS_SCHED_* variables, or disable it entirely with --no-continuous-batching.

Native paged attention

The native kernel TSGgml_PagedAttentionForward (and a WithSinks variant for GPT OSS) gathers K/V from the paged buffer in C++, builds a small GGML graph per sequence, and dispatches ggml_flash_attn_ext — the same fused Metal/CUDA flash-attention kernel the single-sequence path uses. On a long-context Ministral-3-14B workload (4×~800 tokens) it runs ~21% faster than the per-sequence GGML path.

MLA / sparse-attention architectures (DeepSeek V4, GLM 5.x) take a third path. A single compressed cache row per token has no paged layout to page, so the native executor gives each request its own slot — a full set of per-layer MLA and indexer caches plus its own n_past — and binding a request switches the active slot without moving KV bytes. GLM's batched fused decode is enabled by default: one graph takes one token from each of N sequences, so the weights are read once between them — 1.81× aggregate decode at four concurrent requests. Set TS_BATCHED_FUSED_DECODE=0 to use serial fused decode for an exact path comparison. Batching changes every GEMM shape, and 75 of the 78 layers routing top-8-of-256 can turn a last-bit difference into a different expert pick on a 2-bit checkpoint.

Batched forward passes (Mistral 3, Gemma 4, GPT OSS, Hunyuan Dense, Qwen 3.5/3.6 with a GatedDeltaNet recurrent-state pool, and Nemotron-H with a Mamba2 recurrent-state pool) pack N sequences into one ForwardBatch call with one batched linear-projection matmul per layer and paged K/V scatter. Gemma 4 reaches ~1.5–1.6× single-sequence throughput; Nemotron-H Mamba2 batched reaches ~3.95× at batch=3 on an Apple M4 Pro.

Shared-prefix checkpoints

Reusing cached KV pages covers a plain attention model. It is not enough at the end of the prompt every conversation begins with — the system prompt, the tool declarations, the selected skills — for a model whose state lives in a recurrent or fused holder, which each new chat would otherwise prefill from scratch. On Gemma 4 and the Qwen 3.5/3.6/3.8 family over the GGML backends, and on Qwen 3.8 Flash Next (including its layer split), the engine checkpoints the complete model state at the end of that shared prefix and starts each new chat from a clone of it, so a new chat re-prefills only its own message; neither Qwen family takes checkpoints under tensor parallelism. The host declares shared prefixes of at least 64 tokens. TS_PREFIX_CHECKPOINTS_MAX (default 4) bounds how many distinct prefixes stay checkpointed; 0 turns it off.

A host can make a checkpoint outlive the process by attaching an IPrefixCheckpointStore to the engine — the executor reads a saved checkpoint at admission and writes one when it takes a checkpoint, with the model family owning and validating the byte format (Gemma 4 and the Qwen 3.5 family export theirs). That is how TensorAgent makes the first message of a launch cost a restore rather than a full prefill: measured on iPhone 17 Pro Max with Qwen3.5 9B, a 54 s cold first message becomes a 1.2 s warm-up and a ~0.6 s first message. The server does the same by default: it forwards the shared prompt of its startup model before opening its port and stores the checkpoint beside the binary in prefix-cache/, one directory per model (TENSORSHARP_PREFIX_CACHE_DIR moves it) — measured 21.8 s → 0.7 s for the first message on an agent configuration. --no-prefix-cache turns off reuse, that startup warm-up and persistence together. The CLI keeps its cache in memory for the life of the process. PrefixCheckpointExactnessTests proves a chat started from a restored copy produces the same tokens as a cold prefill.

Alongside it, retention (on by default; TS_RETAINED_FUSED_CACHE_MAX caps how many are kept and defaults to the running-request limit, so every request that runs in parallel keeps its own, and 0 turns it off) keeps a finished request's fused holder so an exact-prefix continuation skips re-prefilling it — on Gemma 4, the Qwen 3.5 family and Qwen 3.8 Flash Next. DeepSeek V4.1 and GLM 5.x keep each finished conversation's native slot for its next turn (V4.1 budget TS_DSV41_RETAINED_CACHE_MB, default 2048; not while a DSpark drafter is loaded). Three K/V sizing knobs — TS_KV_INITIAL_TOKENS, TS_KV_GENERATION_RESERVE_MAX, TS_KV_HOLDER_POOL_MAX — let a memory-constrained host start caches small and grow them on demand instead of reserving a whole context window per request.

📖

Full deep dive: docs/PAGED_ATTENTION_AND_CONTINUOUS_BATCHING.md in the repository.

MTP / NextN speculative decoding

Some architectures ship a multi-token-prediction (MTP / NextN) draft head that lets the server run lossless speculative decoding for solo (non-concurrent) sequences. The draft proposes several future tokens cheaply, the trunk verifies all of them in one batched forward, and accepted tokens are committed in a single step.

🎯

Because the request's own sampler — temperature, top-k/p, and all penalties — drives both the draft and the verify, the output is identical to standard decode. Speculation only changes how many forward passes it takes to produce the same tokens.

It is off by default. For a draft head embedded in the trunk GGUF, enable it on the CLI or the server with --spec (env TS_SPEC=1); a draft head that ships as its own GGUF is enabled simply by naming it with --draft-model (an explicit --no-spec still vetoes it):

# Qwen 3.6 — the NextN block is embedded in the trunk GGUF, no extra file needed
# (use a GGUF from unsloth/Qwen3.6-35B-A3B-MTP-GGUF; base-repo exports strip the NextN block)
dotnet run --project TensorSharp.Server.Host -c Release --no-build -- --model Qwen3.6-35B-A3B-UD-IQ2_XXS.gguf --backend ggml_cuda \
    --spec --spec-pmin 0.75

# Gemma 4 — load the separate gemma4-assistant draft GGUF that matches the target
dotnet run --project TensorSharp.Server.Host -c Release --no-build -- --model gemma-4-12B-it-qat-UD-Q4_K_XL.gguf --backend ggml_cuda \
    --draft-model mtp-gemma-4-12B-it.gguf

Draft-head shapes

Speculation and the radix prefix cache. --spec no longer disables prefix reuse. The Qwen 3.5/3.6/3.8 NextN head restarts only its private attention cache at the end of a reused prefix — a radix hit, a startup-warmed or disk-restored shared prefix, or a retained chat turn — while the trunk keeps its full cached state, so prefill stays on the normal fast path; initial acceptance can be lower until the head rebuilds its history. This covers the default fused-verify route for a solo request; the paged speculative route selected by disabling fused verify still needs a position-0 prefill. Gemma 4's assistant head keeps no state and drafts from any position, and block drafters (DFlash, DSpark) and the n-gram speculator also arm after a reused prefix; the Qwen 3.8 Flash Next head does not.

GLM-5.3 — the non-Flash release, GGUF arch glm-dsa — ships the embedded shape too, carrying one complete NextN block at blk.78 of the trunk GGUF, the same place GLM-5.2 carries its own. --spec has to be on the command line before load, because the loader decides whether to page that block in while it reads the trunk (TS_SPEC). The block ships no nextn.shared_head_head.weight and so borrows the trunk LM head, which is why speculation engages on single-device or explicit layer-split placement (--layer-split N, no active tensor parallelism): under --tp N>1 that head is column-parallel and the loader refuses to draft from one rank's strip of the vocabulary. That constraint is read off the native and managed loaders rather than covered by a test, and it is not GLM-5.3-Flash's limitation — glm5next's NextN/MTP is not implemented at all, which is a different thing from a mode restriction. The matrix below scores Qwen 3.6 and Gemma 4 only; GLM-5.3's embedded NextN has no acceptance or speedup figures of its own. Full mechanism: the repository's docs/models/glm.md card.

Where it's profitable

BackendQwen 3.6Gemma 4
GGML CUDA / GGML Metal✅ fused verify + draft kernels✅ fused verify + draft kernels
Direct CUDA (cuda, pure C#)✅ GPU-resident per-op verify/draft✅ GPU-resident per-op verify/draft
GGML CPU / GGML Vulkan✅ speculates — measured 1.21× on Qwen3.6-35B-A3B ggml_cpu at an 8-token window (86% acceptance), before this family's default narrowed to 3 tokens; no figure is committed for the current default✅ speculates — Gemma 4's gate is any GGML backend
CPU / MLXnot gated off by backend — --spec is honoured wherever the head loads; no speedup measuredstandard decode

Tuning: --spec-type auto|draft-head|block|ngram picks the algorithm — auto (the default) uses whatever drafter the checkpoint carries, and ngram is a weight-free suffix matcher over the sequence's own tokens that needs no draft file. N-gram still needs a speculative target, so it is not universal: it runs on the Qwen 3.5/3.6/3.8 family, Gemma 4 (GGML backends or cuda), GLM 5.x (GLM-5.3-Flash only where its rollback is available) and Qwen 3.8 Flash Next (GGML), but not on GPT OSS, Mistral 3 or Hunyuan Dense; DeepSeek V4/V4.1 and Muse-Glimmer speculate only with their own block drafter attached, and Nemotron-H refuses every speculator. --spec-draft (default 8) bounds tokens drafted per step; --spec-pmin is the minimum draft confidence to keep a token — its default depends on the drafter kind (0.15 for a per-token draft head, 0.35 for a block drafter, where the gate is the cumulative prefix probability and so the same number is far stricter). On Qwen 3.5 under the GGML backends, keep a typed --spec-draft at 7 or lower: a verify batch of nine rows diverges from plain greedy, which is a correctness bug rather than a speed one. The default is unaffected, since that model asks for a 3-token window.

DSpark block speculative decoding (DeepSeek V4)

DeepSeek V4 ships DSpark ("Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation") as a support module in the checkpoint: three DSV4 blocks that read the trunk's hidden states and propose a whole block of tokens per step instead of one, a Markov head that conditions each block position on the token before it, and a confidence head that predicts each position's acceptance probability. The trunk then verifies the block in one batched forward and keeps only the prefix its own sampler would have produced.

The drafter is a separate GGUF loaded with --draft-model — every GGUF conversion of the trunk drops the mtp.* tensors. Pre-built drafters are listed in the repository's MODEL_DOWNLOADS.md, or you can convert one from the upstream safetensors checkpoint with eng/dsv4-dspark-to-gguf.py (only the three shards holding mtp.* are downloaded).

# CLI — every single-sequence path (--input, --multi-turn-jsonl, --interactive) uses it
dotnet TensorSharp.Cli/bin/TensorSharp.Cli.dll --model DeepSeek-V4-Flash-...-00001-of-00005.gguf \
    --backend ggml_cuda --layer-split 4 --draft-model DSpark-drafter-Q2K-Q8-0731.gguf \
    --input prompt.txt --max-tokens 200 --temperature 0

# Server — the same drafter flag; naming the file turns speculation on
dotnet TensorSharp.Server.Host/bin/TensorSharp.Server.Host.dll --model DeepSeek-V4-Flash-...-00001-of-00005.gguf \
    --backend ggml_cuda --layer-split 4 --draft-model DSpark-drafter-Q2K-Q8-0731.gguf

It runs on both GPU engines — --backend cuda (direct CUDA) and --backend ggml_cuda (native ggml); ggml_vulkan, ggml_metal, ggml_cpu and cpu have no speculative path for this architecture and log a warning if a drafter is configured. On ggml the drafter is three extra graph layers whose key ring the trunk graph commits itself, so speculation costs no host round-trips. The CLI and the server both run it through the shared engine, where every verify row is drawn with the request's own sampler, so it composes with any sampling settings. Speculation serves solo sequences only — as soon as a second request is in flight, DSV4's per-sequence slots serve the batch at normal decode speed.

Measured (4×A40, DeepSeek-V4-Flash-0731 UD-Q8_K_XL, greedy)

Metriccuda baselinecuda + DSparkggml_cuda baselineggml_cuda + DSpark
Decode (200-token generation)26.0 tok/s34.0 tok/s (1.31×)26.4 tok/s37.1 tok/s (1.41×)
Prefill (15K prompt)962 tok/s955 tok/s952 tok/s954 tok/s
Acceptance—69%—69%

Multi-turn chat benefits most — a turn that continues an established context is exactly where the drafter is confident. On a 5-turn interactive session the same box measured 1.50× to 2.02× per turn (66–87% acceptance), with prefill at parity. Greedy output was byte-identical to the non-speculative baseline on both the 200-token and the 15K-context runs.

Why 1.3× and not more: the trunk is 6-of-256 sparse, so each extra token in the verify batch pulls its own set of routed experts through VRAM — a verify row costs roughly a quarter of a full decode step no matter how cheap the draft was. --spec-pmin (default 0.35 for a block drafter, the minimum cumulative acceptance probability for a drafted position) is the knob that keeps that trade positive; --spec-draft caps tokens per block.

None of these measurements carries to DeepSeek V4.1, whose DSpark support is experimental. The loader accepts a deepseek41-dspark drafter through --draft-model on ggml_cuda and ggml_cpu only, and refuses a V4 drafter or any other draft architecture. On every other backend — including --backend cpu, the 100% pure-C# V4.1 executor, and the direct-CUDA engine — a configured drafter is a hard refusal ("DeepSeek V4.1 DSpark requires the TensorSharp ggml executor (--backend ggml_cuda or ggml_cpu) and a matching deepseek41-dspark drafter. Other executors do not implement the V4.1 draft graph."), not V4's warn-and-continue. Initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified.

DFlash / DFlash2 block speculative decoding (Muse-Glimmer, Qwen 3.5 family)

Muse-Glimmer has its own block drafter, DFlash: a separate 5-layer GGUF (general.architecture = dflash) that proposes the whole speculative window in a single forward. It borrows the target's token embedding and LM head, keeps its own sliding-window KV ring, and is driven in three passes — encode the trunk's per-layer input residuals at dflash.target_layers into one wide row, inject that row as the K/V of every draft layer, then draft [anchor, MASK × (block−1)] through the five blocks and score it with the target's LM head. The trunk then verifies the block in one batched forward.

dotnet TensorSharp.Cli/bin/TensorSharp.Cli.dll --model Muse-Glimmer-30B-UD-IQ2_XXS.gguf \
    --backend ggml_cuda --draft-model dflash-kquant.gguf --spec-draft 15 \
    --input prompt.txt --max-tokens 256

Both halves are fused native graphs that CUDA-graph-capture and replay: TSGgml_DFlashInject and TSGgml_DFlashDraftBlock, the latter finishing with an on-device argmax so the 202048-wide probability block never crosses PCIe. Every verify row is drawn with the request's own sampler, so a greedy run emits the plain-greedy stream. A runtime cost governor measures speculation against plain decoding and parks the drafter while it is measurably slower, which means speculation can only help — but it also needs a few hundred generated tokens to settle, so very short answers may not see the win. Measured against llama.cpp's own DFlash implementation on one RTX PRO 6000 Blackwell (Q8_0, greedy, 60-token prompt): 50.9 tok/s against llama.cpp's 45.5 and 35.0 for plain decode. Full ladder, including where llama.cpp is ahead: the repository's docs/models/muse-glimmer.md card.

DFlash2 is the same backbone plus two additions, both keyed off the GGUF so one code path serves either generation: a grouped dynamic depthwise convolution around every attention and FFN sublayer, which gives a block-diffusion draft a local left-to-right signal without a second forward; and a candidate selector that scores adjacent positions' top-K candidates pairwise through two low-rank [vocab, r] codebooks and reads the block off as a walk through that lattice, so position i+1 is no longer chosen without knowing what i chose. Attach either with --draft-model — the file says which it is. Both also target the Qwen 3.5 / 3.8 trunk; a Qwen 3.8 checkpoint can carry an embedded NextN block as well, and an attached DFlash drafter takes precedence over it.

Nemotron 3.5 Lightning's DSpark module is not attached. NVIDIA ships a 6-layer DSpark drafter for it (SWA-1024, per-head attention-sink biases, an encoder reading the trunk residuals at layers 2/6/20/30/42/52, and a rank-512 Markov head), and eng/nemotron-dspark-to-gguf.py still exports it to a dflash-architecture GGUF. But Nemotron-H refuses every speculator, n-gram included: its multi-token verify and its single-token decode run different attention and MoE kernels whose logits differ, so a speculative stream would not match plain greedy decoding, and a bit-exact verify costs as much as decoding the rows one at a time. A drafter named with --draft-model is not attached: the CLI reports it and decodes plainly, and the server refuses to start until the flag is dropped.

DiffusionGemma text diffusion

DiffusionGemma is fundamentally different from autoregressive models: it does not call Forward() one token at a time. Instead it uses block-wise EntropyBound denoising over fixed-length canvases on a Gemma-4-derived MoE backbone — the whole answer is refined iteratively rather than written left to right.

Performance optimizations

A cross-architecture summary; each per-model card in docs/models/ walks through the same kernels with the exact GGML graph dispatched.

⚡

Verified Gemma 4 E4B Q8_0 fast path: repository benchmarks verify the native-GGML E4B Q8_0 family and execution path on GPU backends. Use ggml-org/gemma-4-E4B-it-GGUF as the recommended public artifact.

Memory optimizations