Features
A complete catalog of what TensorSharp does today. Each item links to the page where you can use it.
Sentence encoders: Snowflake Arctic Embed L v2.0 / all-MiniLM-L6-v2, with pure C# CPU and native GGML CPU/Metal/CUDA execution, served through --embeddings and OpenAI/Ollama embedding APIs. Embedding models, supported backends, and validation โ
Highlights
Multi-architecture
DeepSeek V4 Flash, DeepSeek V4.1 Flash, GLM 5.2, GLM-5.3 & GLM-5.3-Flash, Gemma 4, Qwen 3.5 / 3.6 / 3.8 Flash Next, Bonsai2, GPT OSS, Nemotron-H, Mistral 3, Hunyuan Dense, Muse-Glimmer, DiffusionGemma, Qwen-Image-2.1, MiniMax-H3 audio+video, Wan video.
Multimodal
Image, video, and audio inputs (Gemma 4); image input for several others.
PDF documents
Upload PDFs in the Web UI or pass --pdf on the CLI โ text PDFs are inlined, scanned pages go to vision models.
Image generation & editing
Qwen-Image-2.1 generates images, edits photos with references, and preserves pixels outside a painted selection. The Web UI and TensorAgent share the selection editor; LoRA plug-ins, prefix K/V caching and supported multi-GPU execution are available.
Video generation
MiniMax-H3 generates video and native 32 kHz stereo audio together in one packed latent, CFG-free at 4โ8 steps. Wan 2.1 / 2.2 generate video alone, where a step-distilled checkpoint cuts a 5-second 720p clip from hours to minutes.
Thinking mode
Structured chain-of-thought, separated from the visible answer.
Tool calling
Multi-turn function calling across all three API styles.
Agent Skills
Folders of model-facing instructions that load only when a task needs them โ selected with --skill or "skills": ["pdf"], fetched through built-in tools TensorSharp answers itself.
Agentic work
A bounded model/tool loop with four code tools, isolated workspaces, sandboxed commands, dependency installation, bounded sub-agents on the server, live progress, and downloadable artifacts.
Native quantized compute
Q4_K_M, Q8_0, MXFP4, IQ2_XXS and more run in matmul without dequantizing to FP32.
Continuous batching
vLLM-style paged KV cache with cross-request prefix sharing.
Speculative decoding
MTP / NextN draft heads, DeepSeek V4's DSpark and the DFlash block drafters (Muse-Glimmer, Qwen 3.5 family), and n-gram lookup on supported models accelerate solo decode.
Multi-GPU & multi-node
Tensor parallelism shards every layer across GPUs โ CUDA and GGML alike โ and across machines over a TCP mesh; the whole-model executors take a layer split instead.
Ollama & OpenAI APIs
Drop-in endpoints for existing tooling, plus a browser chat UI.
TensorAgent: local text, images, video, files, code, skills and desktop browser workflows in one app for iPhone, iPad, macOS and Windows. Eight interface languages; device memory and platform policy determine which workflows are available. Capabilities and validation limits โ
Models & modalities
- Multi-architecture support โ DeepSeek V4 Flash, DeepSeek V4.1 Flash, GLM 5.2, GLM-5.3, GLM-5.3-Flash, Gemma 4, DiffusionGemma, Qwen 3.5/3.6-family (including Bonsai2), Qwen 3.8 Flash Next, GPT OSS, Nemotron-H, Mistral 3, Hunyuan Dense, Muse-Glimmer, Qwen-Image-2.1, MiniMax-H3 audio+video, Wan 2.1 / 2.2 video. โ Supported models
- DeepSeek V4 Flash (284B MoE) โ a compressed-sparse-attention, 1M-context architecture with three dedicated whole-model executors: a direct-CUDA engine (
--backend cuda), the native ggml executor (ggml_cuda/ggml_vulkan), and a 100% pure-C# CPU executor (--backend cpu, no native dependencies). Weights use whole-layer placement with--layer-split N, and the server hosts it with native per-sequence slots and continuous batching. โ DeepSeek V4 - DeepSeek V4.1 Flash (384 routed experts) โ a dedicated native V4.1 graph: four residual streams with delayed hyper-connection mixing, a lightning indexer with candidate block filtering, and Engram n-gram features injected at layers 1 and 14, with two Engram tables totaling 60.1 GiB at Q2_K, placed on GPUs and gathered in-graph by default when they fit.
ggml_cudais the serving backend;ggml_cpuis a scalar correctness path,cpuis a 100% pure-C# V4.1 executor with no ggml, no native dependency and no GPU, held to the independent PyTorch oracleeng/dsv41-reference.pyat atol=rtol=2e-5 plus exact greedy argmax on a five-layer F32 fixture (architectural agreement with the oracle, not parity on the real quantized weights) โ a correctness and portability path rather than a serving one, and text-only there, since the vision companion is a native ggml component โ andcudaruns V4.1 through the direct-CUDA engine's own kernels with targeted synthetic kernel and sequence-state tests rather than a full-checkpoint numerical gate.mlxrefuses the checkpoint. Images and video need a separately prepared vision companion, audio is refused, and the current Q2_K and Q4_K_M GGUF releases embed the Engram constants and weights for direct loading. โ DeepSeek V4.1 - GLM 5.2 (744B-A40B MoE) โ Multi-head Latent Attention with weight absorption plus DeepSeek Sparse Attention's lightning indexer, 256 routed experts at top-8 with one shared expert, and a 1M advertised context, loaded straight from a 6-shard split GGUF. Two implementations: a native whole-model ggml executor (
ggml_cuda/ggml_vulkan/ggml_cpu/ggml_metal) that accepts--layer-split Non CUDA or Vulkan to place 226 GiB across N local GPUs (default: one device), and a managed per-op path used by--backend cpu(100% pure C#, no native dependencies) and--backend cuda. โ GLM 5.x - GLM-5.3 (256 routed experts, text only) โ the non-Flash 5.3 release is the same
glm-dsablock shape as GLM-5.2 โ 79 blocks (78 trunk + one NextN), 256 routed experts at top-8 plus one shared expert, MLA with the lightning indexer, rope base 8e6 (n_rot64) โ so it loads on the GLM-5.2 path with no new code and no new flag: one architecture descriptor serves GLM-5.2, GLM-5.3 and GLM-5.3-Flash, and Flash vs non-Flash is a single boolean ongeneral.architecture. Download unsloth/GLM-5.3-GGUF, one subdirectory per quant and each a multi-shard set โ point--modelat the-00001-of-shard; UD-Q2_K_XL is seven shards, 236.4 GiB. It is text only twice over: the repo publishes no mmproj at any quant, andLoadVisionEncoderwarns about and ignores an--mmprojonglm-dsa, so a vision flag yields a text-only run rather than a failure. The advertised 1M context is a ceiling the loader caps to device capacity unlessMAX_CONTEXTnames the number. The native whole-model executor runs onggml_cuda/ggml_vulkan/ggml_metal/ggml_cpu, and the managed per-op path on--backend cpu/cudaor on a GGML backend withTS_GLM_NATIVE=0.--spec(on the command line before load) drafts from the complete NextN block atblk.78, which ships nonextn.shared_head_head.weightand so borrows the trunk LM head โ speculation engages on single-device or explicit--layer-split Nplacement, with no active tensor parallelism.--tp Nis accepted but local / single-process only and unmeasured above one rank, with the MLA and indexer caches replicated per rank, while--tp-node-id/--tp-peersare refused for the whole GLM family before the model is built. Measured on 8ร A40 46 GB with no NVLink (UD-Q2_K_XL, a 10,531-token prompt, 300 decode tokens, median of 3, whole-layer placement) against llama.cpp, which is a valid reference engine for this release unlike Flash: 251.6 t/s prefill and 20.48 t/s decode after a 264 s load, against llama.cpp's 20.28 t/s decode and 753 s load โ a decode tie with a 2.9ร faster load of a 236.4 GiB checkpoint, and one honest gap, TTFT 41.9 s vs 29.0 s (~1.4ร slower). โ GLM 5.x - GLM-5.3-Flash (320B MoE, text + image) โ the hybrid successor loads through the same native executor and the same
GlmDsaModelas GLM-5.2, under the GGUF arch idglm5next, with four architectural changes layered on: KDA linear attention on 34 of the 45 trunk layers, NoPE MLA + DSA on the other 11, a pooled indexer (4-cell pools, top-k 2048 over pools, then expanded to their members), Sinkhorn hyper-connections over ร4 streams, and a SwiGLU clamp at 10; 288 routed experts at top-8 with one shared and a ร2.5 routed scale, 46 blocks = 45 trunk + 1 NextN. It places whole layers on N local GPUs with--layer-split N, takes--cpu-moe/--n-cpu-moe, and serves through per-sequence native slots. On GGML GPU backends,--tp Nselects native local/single-process tensor parallelism, as it does for GLM-5.2: KDA/MLA heads and routed-expert hidden rows are sharded, and KDA recurrent state is per-rank. With full sharding on one rank per local GPU, no CPU MoE or tracing, and native hyper-connection kernels, segmented rank-local graphs submit concurrently: routed-MoE partials reduce first, then every rank computes and adds its replicated shared expert locally before Sinkhorn hyper-connections; hyper-connections, pooled indexer, router, norms, dense layers and embedding remain unsharded and execute per rank, while output norm / LM head stay on rank 0. CPU MoE, tracing, partial sharding, oversubscription, or missing native hyper-connection kernels select the combined scheduler fallback, where the shared expert runs once on rank 0. NextN/MTP speculation remains unsupported. Vision comes frommmproj-BF16.gguf(the GLM-OCR ViT):--image, multi-image, and multi-turn image sessions. Measured on 2ร RTX PRO 6000 Blackwell (96 GB) against llama.cpp, UD-Q2_K_XL (101 GiB), both engines layer-split atn_ubatch2048, back to back: tg64 73.5 vs 36.6 t/s โ decode at 2.0ร llama.cpp, with prefill within a few percent either way (pp2048 2014 vs 2070, pp16384 1692 vs 1690, pp32768 1446 vs 1483). It also runs on the 100% pure-C#--backend cpupath for text, as a reference implementation to A/B against rather than a fast one: ~5.7ร offggml_cpuon prefill and ~2.4ร on decode, with a prefill-logit cosine of 0.9567 againstggml_cpuโ close, but not established as bit-parity, so greedy text can diverge on a near-tie. โ GLM 5.x - Qwen 3.8 Flash Next (hybrid MoE, text + image + video) โ GatedDeltaNet recurrent layers on 36 of its 48 layers, interleaved with full-attention layers (some behind Qwen Sparse Attention's indexer), a PLE n-gram embedding block on layer 1, ร4 hyper-connection streams and a 512-expert MoE at top-10 (hidden 2560, 24 query / 2 KV heads, head_dim 256, vocab 248320). The GGUF arch id is
qwen4exp. On the GGML backends the whole token runs as (almost) one graph โ embedding, in-graph PLE, all 48 layers, the final mixer and the LM head โ replayed from a shape-keyed cache of captured graphs. Vision rides the Qwen3.5-VL tower with (T,H,W) IMRoPE, so multi-image and multi-turn image sessions work with KV reuse across turns (extend-only: the GDN recurrence cannot rewind), and avideo_urlpart is sampled into timed frames. Switchable thinking, the server, and continuous batching through per-sequence state holders are supported, and--layer-split Nruns a multi-GPU layer split. Tool calls parse through the Qwen 3.5<tool_call>parser (XML or JSON bodies), so skills, the code tools, and sub-agents are offered to it; a separate shared MTP head attaches with--draft-modelon the GGML backends. โ Qwen 3.8 Flash Next - Multimodal inference โ image, video, and audio inputs for Gemma 4; images for Qwen 3.5-family, Qwen 3.8 Flash Next, GLM-5.3-Flash, Mistral 3, Nemotron-H Omni, Muse-Glimmer, and DiffusionGemma; video for Qwen 3.8 Flash Next too; image and video for DeepSeek V4.1 once its separately prepared vision companion is attached. โ Multimodal
- PDF document input โ born-digital PDFs have their complete text layer extracted and inlined into the prompt; scanned PDFs fall back to page images for vision-capable models. Available as a Web UI upload and via the CLI's one-shot
--pdfflag; cap the pages read withTS_PDF_MAX_PAGES(default: all). โ Web UI - Mixture of Experts (MoE) โ Gemma 4 MoE (e.g. 26B-A4B), GPT OSS MoE (gpt-oss-20b), Qwen 3.5/3.6 MoE (35B-A3B), and Nemotron-H MoE FFN layers, with a fused batched GPU MoE dispatch.
- Hybrid SSM-Transformer โ Nemotron-H mixes Mamba2 SSM layers, attention layers, and MoE FFN in one model.
- Hybrid Attention-Recurrent โ Qwen 3.5/3.6-family mix full-attention layers with GatedDeltaNet recurrent layers; Qwen 3.8 Flash Next does the same on 36 of 48 layers and adds a PLE n-gram block, ร4 hyper-connection streams and a 512-expert MoE, while GLM-5.3-Flash pairs KDA linear attention (34 of 45 trunk layers) with NoPE MLA and a pooled sparse indexer on the rest.
- Video generation with audio (MiniMax-H3) โ one 50-block diffusion transformer denoises video and a native 32 kHz stereo soundtrack as a single packed latent, so the audio is model output rather than something added afterwards. Four conditioning modes (
--video-mode t2v/i2v/fl2v/ref): a prompt alone, a photo animated as the first frame (--image), a first-and-last keyframe pair (--end-image), or up to nine identity references โ stills, clips and soundtracks โ for a brand-new scene (--ref-image/--ref-video/--ref-audio). The two denoisers are separate checkpoints, not settings: keyframes need the FL2VA file, references the Ref2VA one. It is CFG-distilled, so--cfg 1.0is required and 4โ8 steps is the fast operating point against a 20-step default. Drives from the CLI,/api/video-generate,/v1/videos/generations, or the Web UI; the MP4 arrives with a sidecar.wav. All four conditioning modes also run on the 100% pure-C#--backend cpupath, with no native dependencies. โ MiniMax-H3 - Video generation, video-only (Wan 2.1 / 2.2) โ Wan 2.1 (text โ video) and Wan 2.2 TI2V-5B / A14B (text โ video and image โ video, where the uploaded image becomes the first frame) render an H.264 MP4 from the CLI,
/v1/videos/generations, or the Web UI. Point--modelat a step-distilled checkpoint (Turbo / Lightning / FastWan) and the pipeline auto-detects it, dropping the denoise loop from 100 DiT passes to 4. โ Video generation - Text-diffusion generation โ DiffusionGemma uses an iterative EntropyBound denoising sampler instead of autoregressive decode. It reads images through the Gemma 4 vision tower, loaded from an mmproj GGUF or the upstream
model-00011-of-00011.safetensorsshard (audio and video are refused, and so are tools). The same model serves a Jev-compatiblePOST /v1/systemoneendpoint that answers typed questions โ yes/no probabilities, choices, and expected scores โ about text or up to eight images (Jev card). โ DiffusionGemma - Image generation & editing (Qwen-Image-2.1) โ a prompt produces an image, and a prompt plus one or more reference images produces an edited image, via the Qwen-Image-2.1 diffusion transformer, its dedicated 2.1 VAE and the Qwen3-VL-8B text encoder (FlowMatch-Euler on the official exponential dynamic-shift schedule, run as one complete GGML graph with resident quantized weights on the GGML backends, or as managed code on the pure-C#
cpubackend). Defaults are 2048ร2048, 40 steps and CFG 1.0, and RGBA input and PNG output keep transparency.--lora/--lora-scale/--lora-configapply LoRA plug-ins unmerged on the quantized weights;config/lora/has twelve presets, four of them step-distillation recipes that run 4โ8 steps. A prefix K/V cache for the text and reference tokens is on by default (TS_QWEN21_PREFIX_CACHE=0disables it) and mostly helps edits: 1.7โ2.9ร faster per step on 1024ยฒ edits on an M5 Pro, with generation about unchanged.--tp Nshards the diffusion transformer onggml_cuda/ggml_vulkan. โ Image generation
Generation & control
- Thinking / reasoning mode โ structured chain-of-thought with
<think>/<|channel>tags (Qwen 3.5/3.6, Qwen 3.8 Flash Next, Gemma 4, GPT OSS, Nemotron-H, DeepSeek V4, DeepSeek V4.1, GLM 5.x), or on Muse-Glimmer'sto=selfchannel. โ Thinking mode - Tool calling / function calling โ architecture-agnostic output parsing turns raw model output into structured
tool_calls, whether the model emits JSON inside a<tool_call>block (Nemotron-H), XML inside that block (Qwen 3.5/3.6, Qwen 3.8 Flash Next), Gemma 4's<|tool_call>tokens, GLM 5.x's<arg_key>/<arg_value>pairs, Harmony commentary (GPT OSS), ATEM markup (Muse-Glimmer), or DSML markup (DeepSeek V4, and its spaced, grammar-constrained variant on DeepSeek V4.1). โ Tool calling - Configurable sampling โ temperature, top-k, top-p, min-p, repetition / presence / frequency penalties, seed, and stop sequences. โ Sampling
- Structured outputs โ OpenAI
response_formatwithtext,json_object, and validatedjson_schema. โ Structured outputs - Chat templates โ auto-loaded from GGUF metadata (Jinja2), with hardcoded fallbacks per architecture.
- Streaming โ token-by-token output via SSE (web) or stdout (console), with abort/stop support for in-flight generations.
Agent Skills
- What a skill is โ a folder holding a
SKILL.md(YAML frontmatter plus Markdown instructions written for the model) together with the scripts, reference documents and assets those instructions refer to. TensorSharp scans one or more skill directories (--skills-dir; without it, every existing.agents/skillsfrom the working directory up to the Git root, plus askillsfolder next to the binary, which the server always scans first as its upload directory), advertises each skill's one-line description to the model, and loads the rest only when the model asks for it. โ CLI flags ยท HTTP fields & endpoints ยท C# API - Two built-in tools (three with
--skills-allow-exec) โskills_list()returns every reachable skill and its bundled paths;skills_read(skill, path, offset)pages a file, includingSKILL.md.skills_run(skill, path, args)is off by default. When enabled, scripts use an interpreter allow-list, a scrubbed environment, time/output limits and an OS sandbox that isrequiredby default. - TensorSharp answers those calls itself โ in process, next to the weights. That is what makes the feature work for clients that know nothing about skills: an ordinary OpenAI client sends
"skills": ["pdf"]and gets back a finished completion, never a tool call it has no implementation for. The caller's own tools are never executed โ they are returned to the caller as usual. - Progressive disclosure โ for a tool-capable family, startup carries names and descriptions only, including explicitly selected skills; the metadata tier uses about 2% of context (1024โ10000 approximate tokens). The model activates a matching skill by reading its
SKILL.md, and bundled files stay on demand in 48 KB pages. A family that cannot complete a tool round trip receives selected bodies inline instead. - Prompt shape โ the block is merged into the leading
system/developermessage rather than appended as a second one, which is the only injection point every chat template in the repository handles. Its bytes are a pure function of the sorted skill selection โ no timestamps, paths or counters โ so a conversation re-hashes identically turn to turn and the KV prefix cache keeps matching from block 0. - Confinement โ every path the model names is resolved through
SkillPathGuard, which closes lexical (.., absolute,~, UNC, drive-qualified), canonical and symlink escapes, and confines each skill to its own directory. ZIP installs run every entry through the same guard (zip-slip), enforce size on the decompressed stream, and cap per-entry (64 MB), per-archive (256 MB), entry-count (4096) and compression ratio (200ร). - Model families โ tools are offered only when both the chat renderer and output parser support them. Mistral 3, Hunyuan Dense, and DiffusionGemma carry no declarations, so they inline selected skill bodies and withhold skill/code tools rather than emitting calls nobody can service.
qwen4exp(Qwen 3.8 Flash Next) completes the round trip. - Agentic code execution โ
--code-execaddsread_file,write_file(new files only),shell, andapply_patch. TensorSharp executes those built-ins inside the same bounded loop; caller-owned tools are never run and still return to the caller. The normal eight-round skill budget rises to 24 when code execution is offered unless the operator set it explicitly. โ Architecture, permissions and platform limits - Sub-agents โ on the server's chat paths (OpenAI chat and Responses, Ollama chat, the Web UI), a tool-capable model can delegate independent work with
spawn_agent,wait_agent,send_input,close_agent, andlist_agents. It is on by default and needs neither skills nor--code-exec; children run on the same loaded model, are read-only by default (onlyworkerchildren can use the parent's mutable tools, and only with--agents-allow-worker-tools), and are bounded by the--agents-*limits.--no-multi-agentor a request's"multi_agent": falseturns it off; TensorAgent delegates by default with no setting to turn it off, and the CLI does not delegate. No latency or quality numbers are published. โ Sub-agents - Browser automation โ the
playwrightskill inTensorAgent/skillsdrives a real browser throughskills_runand a pinned@playwright/cli; there is no screen or mouse tool. It needs script execution, networking, and Node.js on the host, and has recorded model workflows on macOS arm64 and separate bounded Windows synthetic skill-runner/session probes. โ Playwright skill - Workspaces and results โ Web/CLI chats keep one workspace across turns; each OpenAI/Ollama request gets a private workspace across its internal rounds and deletes it afterward. Skill scripts share it with code tools. The Web UI shows transient
tool_progressactivity and retains produced-file links; server artifacts are downloadable below/api/code/artifacts/. - Safety boundary โ generated commands are offline by default. macOS Seatbelt and Linux bubblewrap 0.12+ confine writes and home reads; macOS cannot guarantee cleanup of a deliberately detached child. Windows job objects bound the process tree but not files or network, so required mode refuses code and an operator must explicitly choose
--code-exec-unconfined(or--skills-sandbox preferredfor skill scripts). TensorSharp has no per-command approval UI: startup flags are the operator's decision. - Structured output โ a
response_formatrequest can still select skills, but their bodies are delivered inline and built-in skill/code tools are suppressed because schema-constrained output cannot emit the tool-call markup. - Selecting skills โ
--skill pdfon the CLI, or"skills": ["pdf"]on every chat API (/v1/chat/completions,/v1/responses,/api/chat/ollama, and the Web UI's/api/chat), with"skills_discovery": falseto restrict a request to exactly the skills it named. The server also exposes the registry itself โGET /v1/skills,GET /api/skills,POST /api/skills(upload a.zip),DELETE /api/skills/{name}โ and/api/modelsreports askillsblock (enabled,installable,count) so a UI knows whether to offer the control. Open-source skills to start from: github.com/anthropics/skills.
Performance & scale
- Measured against llama.cpp โ comparisons use identical GGUF files and hardware; results depend on the model, backend, and workload. In the current comparison run (short/long/multi-turn text on both engines' GGML CUDA and Vulkan builds, reproducible via
benchmarks/engine_comparison): Gemma 4 E4B and the 2-bit Qwen 3.6 35B-A3B MoE prefill 1.28ร faster on CUDA with first tokens 1.27ร sooner (multi-turn prompts up to 1.49ร), Gemma 4 12B decodes 1.21ร faster on Vulkan (up to 1.32ร on long context), and CUDA decode holds parity or better on three of four models (up to 1.07ร geomean on Qwen 3.6 27B). โ Head-to-head benchmarks - GPU-accelerated โ GGML Metal (macOS), GGML CUDA (Windows/Linux + NVIDIA), GGML Vulkan (Windows/Linux + AMD/Intel/NVIDIA), a direct CUDA/cuBLAS backend, and an MLX backend for Apple Silicon โ all with CPU fallbacks. โ Backends
- Continuous batching & paged KV cache โ block-paged KV pool with a radix-tree prefix cache on by default, an iteration-level scheduler that admits/preempts sequences mid-batch, and a native fused paged-attention kernel. The paged KV cache is host-resident, so concurrency buys capacity and fair scheduling rather than throughput โ the batched paged route saturates at roughly 69 tok/s however many sequences are in flight. โ Deep dive
- Batched / parallel inference โ N sequences packed into a single forward pass with paged K/V scatter (Mistral 3, Gemma 4, GPT OSS, Qwen 3.5 / 3.6, Nemotron-H, Hunyuan Dense).
- Multi-GPU modes โ Use
--tp Nfor tensor parallelism and--layer-split Nfor whole-layer placement. The modes are mutually exclusive; unsupported requests fail at startup. Layer split is local to one node. Supported tensor-parallel architectures can add--tp-node-idand--tp-peersto span nodes. โ Multi-GPU & Multi-Node - MoE CPU offload โ
--cpu-moe/--n-cpu-moe Nkeeps the routed experts of the first N layers in system RAM, served straight from the GGUF mapping with no private copy, and composes with--tp. It is how a checkpoint that does not fit runs at all rather than a speed knob: GLM 5.2 at--n-cpu-moe 30measures pp2048 94.7 / tg64 16.4 against 915.9 / 43.9 fully resident, and frees enough VRAM to raise the sized context from 342,272 to 646,400 tokens. โ Memory - MTP / NextN speculative decoding โ multi-token-prediction draft heads accelerate solo decode. It samples from the same distribution as plain decoding: every emitted token is drawn by the request's own sampler from a verified trunk row, and a draft only decides how many rows one forward pass computes. Greedy output matches plain decoding except at logit near-ties, where the multi-row verify kernel, which accumulates in a different order from one-row decode, can pick the other token. โ Speculative decoding
- DSpark block speculative decoding โ DeepSeek V4's drafter proposes a whole block of tokens per step (a Markov head conditions each block position on the one before it, a confidence head gates how far to draft) and the trunk verifies the block in one batched forward. Loaded as a separate GGUF with
--draft-model; measured 1.3โ1.4ร decode on 4รA40, up to 2.0ร on multi-turn chat, with greedy output byte-identical to the baseline. DeepSeek V4.1 accepts adeepseek41-dsparkdrafter experimentally, onggml_cuda/ggml_cpuonly; initial text/image HTTP probes with trained weights passed using two-GPU layer split onggml_cuda; broad quality and throughput remain unqualified. โ DSpark - Whole-model fused decode graphs โ Gemma 4, Qwen 3.5/3.6, and GPT OSS submit an entire decode token as one GGML graph instead of one dispatch per layer, so the GPU never waits on the host between layers. GPT OSS decode: 24 โ 154 tok/s on an A40, and flat in context length. โ Performance optimizations
- Diffusion fast paths (video & image) โ MiniMax-H3 ships CFG-distilled, so there is no fast checkpoint to hunt for:
--cfg 1.0at 4โ8 steps is the operating point against its 20-step default, and on an M5 Pro (ggml_metal, 22 frames, 8 steps, identical seed) it finishes 2.4ร faster thanstable-diffusion.cppat 256ร256 (49.3 s โ 20.9 s) and 1.7ร at 640ร384 (108.5 s โ 63.1 s). On Wan, a step-distilled checkpoint (Turbo,distill,Lightning,lightx2v,FastWan,-dmd, or an explicitโฆ-4steps-โฆin the file name) is detected at load and runs 4 guidance-free DiT passes instead of the base recipe's 100 โ no flag, just a different--model. Measured on an M5 Pro (ggml_metal, Wan2.2-TI2V-5B Q8_0, 1088ร832ร121f image-to-video): โ3 h 30 m on the base checkpoint vs 17 m 30 s on the Turbo one. Independently of that, F16 attention K/V is 2.02ร on a 27,404-token self-attention, the Metal MPSGraph VAE convolutions take a 736ร544ร81f decode from 159 s to 80 s, and--cfg-cache-stride 2/3buys 1.30ร / 1.43ร on base checkpoints. For images, Qwen-Image-2.1's CFG 1 default already runs one transformer prediction per step, so the levers are step count โ a step-distillation LoRA plug-in fromconfig/lora/cuts the 40-step default to 4โ8 transformer passes โ and resolution:--width 1024 --height 1024quarters the latent image tokens of the 2048ร2048 default for drafts. โ Video generation - Native quantized compute โ quantized weights are used directly in matmul without expanding to FP32, saving memory and bandwidth. On
--backend cputhat now covers the Wan and MiniMax-H3 diffusion linears too, which used to be expanded to F32 at load: a small Wan render measured 80.9 s against 121.4 s at 4ร less weight memory. The i-quant coverage is wider too:IQ2_XSandIQ4_XSstay quantized at load instead of being expanded, andIQ2_XS/IQ3_XXSagainstQ8_Khave direct dot kernels with AVX2 paths instead of dequantizing each row into scratch. - Optimized pure C# CPU backend โ managed GEMM fast paths plus fused SIMD kernels for RMSNorm, RoPE, softmax, and fused activations, now spread across a persistent multi-core worker pool instead of a per-matmul
Parallel.For(measured on gemma-4-E4B-it-Q8_0 with a 122-CPU allocation: prefill โ+15%, decode โ2.8ร; the pool defaults to half the usable cores on purpose, since spinning workers otherwise starve the rest of the CPU path). MiniMax-H3, the last video family without one, now has a pure-C# CPU path too. Qwen-Image-2.1 now runs its whole pipeline oncputoo, including LoRA and masked editing. Bonsai2 still requires a single-device GGML backend. Qwen 3.8 Flash Next has GGML and direct-CUDA execution; it has no managed CPU model path. See each model card for tested devices and numerical scope. TheModelBaseloader binds quantized weights zero-copy from the GGUF mapping instead of copying them into fresh anonymous memory: GLM-5.3-Flash UD-Q2_K_XL went from a load that never completed to ~48 s. Embedding models use a separate compact-weight loader and a model-owned worker pool; see embeddings. โ CPU backends
Interfaces & integration
- Ollama & OpenAI API compatibility โ drop-in replacement endpoints for existing tooling. โ HTTP API
- Browser chat UI โ multi-turn chat, file uploads up to 500 MB (images, video, audio, PDF, text/code), thinking toggle, tool calling, live sub-agent cards, message editing, and live streaming. โ Web UI
- Interactive REPL โ a turn-by-turn console chat with slash commands, hot-swappable model/backend/projector, and live sampling tuning. โ REPL
- Batch processing โ JSONL input in the console application, plus a built-in prefill/decode benchmark.
- Source-project embedding โ reference only the layers you need in your .NET app. The release workflow packages thirteen projects; use current source references for features newer than the release tag. โ Package boundaries
- Per-turn observability โ structured bounded input summaries, raw output logs, and KV-cache hit ratios surfaced through every API (
prompt_cache_hit_*,cached_tokens,kvReused*). Uploaded document bodies are omitted from logs while their metadata and user instruction remain visible.