Compute Backends
A backend is selected with --backend (CLI and server) and decides which hardware runs the math. Supported backends depend on the model architecture; check the model guide for implementation and validation limits. TensorAgent adds an iOS/iPadOS application target that uses the GGML Metal backend on physical Apple devices.
Sentence encoders: Snowflake Arctic Embed L v2.0 / all-MiniLM-L6-v2, with pure C# CPU and native GGML CPU/Metal/CUDA execution, served through --embeddings and OpenAI/Ollama embedding APIs. Embedding models, supported backends, and validation →
Which backend should I use?
| Your hardware | Recommended | Flag | Notes |
|---|---|---|---|
| Apple Silicon (Mac) | GGML Metal | ggml_metal | The server's default on macOS (the CLI defaults to ggml_cpu on every OS, so pass --backend ggml_metal there). mlx is an alternative Apple-Silicon GPU path. |
| iPhone / iPad | GGML Metal through TensorAgent | ggml_metal | Native iOS target; build the TensorAgent app with TensorSharpAppleTargets=true. iOS 17.0 or later, arm64 iPhone and iPad (the app, its share extension and the native .xcframework all target 17.0); the app declares the single-window scene lifecycle that linking against the iOS 27 SDK requires. No ASP.NET Core host or child processes. |
| Windows / Linux + NVIDIA | GGML CUDA | ggml_cuda | Most-tested NVIDIA path. cuda is the direct PTX/cuBLAS backend for experimentation. Multi-GPU tensor parallelism runs on ggml_cuda, ggml_vulkan and cuda. |
| Windows / Linux + AMD / Intel / NVIDIA | GGML Vulkan | ggml_vulkan | Vendor-neutral GPU path via ggml-vulkan (cooperative-matrix accelerated where the driver supports it). Built automatically when the machine has a Vulkan runtime; opt out with --no-vulkan. |
| No GPU / portability / debugging | Pure C# CPU | cpu | No native dependencies, and managed matmuls now run on a multi-core worker pool. Compare ggml_cpu (native kernels) on the intended workload. |
All backends in detail
| Backend | Flag | Best fit |
|---|---|---|
| Direct CUDA / cuBLAS | cuda | NVIDIA inference & experimentation |
| MLX Metal | mlx | Apple Silicon (alternative to GGML Metal) |
| GGML Metal | ggml_metal | Apple Silicon (server default on macOS) |
| GGML CUDA | ggml_cuda | NVIDIA inference through ggml |
| GGML Vulkan | ggml_vulkan | Vendor-neutral GPU inference through ggml (AMD / Intel / NVIDIA) |
| GGML CPU | ggml_cpu | Native CPU kernels |
| Pure C# CPU | cpu | Portability & debugging |
| iOS GGML Metal | TensorAgent app | On-device iPhone/iPad inference through the statically linked iOS .xcframework |
GGML Metal, GGML CUDA & GGML Vulkan
The GPU paths built on a native C++ bridge that links ggml.
- GGML Metal (
ggml_metal, macOS) — quantized weights are mapped zero-copy from the GGUF file into Metal command buffers via host-pointer buffers, so the resident set stays close to the on-disk model size. - GGML CUDA (
ggml_cuda, Windows/Linux + NVIDIA) — quantized weights are uploaded to device memory once at load time and the host copy is released afterwards. - GGML Vulkan (
ggml_vulkan, Windows/Linux + AMD/Intel/NVIDIA) — vendor-neutral GPU path for any GPU with a Vulkan 1.3 driver, using cooperative-matrix shaders (KHR coopmat / NV coopmat2) where the driver supports them. Weights are device-resident like GGML CUDA and the same fused whole-model decode/prefill graphs are used. On multi-GPU hosts (e.g. an integrated Intel GPU next to a discrete NVIDIA one), pick the device with--gpu-device N(or theTS_GGML_VULKAN_DEVICEenv var) and list the visible devices with--list-gpus.
All three run native quantized matmul (Q4_K_M, Q8_0, …) without dequantizing to FP32, plus a native paged-attention kernel that drives ggml_flash_attn_ext.
Multi-GPU. GGML CUDA and GGML Vulkan support tensor parallelism through --tp N on compatible architectures. Whole-layer placement uses --layer-split N, including Qwen 3.8 Flash Next, DeepSeek V4 and GLM native executors. The modes are distinct and mutually exclusive; the default with no placement settings is one device. Unsupported requests fail at startup. → Multi-GPU modes
Verified E4B fast lane: repository benchmarks exercise the Gemma 4 E4B Q8_0 family on these native GGML GPU backends. E4B's PLE and shared-KV layout stays on the fused whole-model prefill/verify and single-graph decode paths, with automatic fused N=1 server routing. Start with the Gemma 4 E4B guide.
MLX Metal
--backend mlx is a GPU-accelerated Apple-Silicon path built on mlx-c. It implements quantized ops (Q4_K_M, Q8_0, Q5_K, Q6_K, IQ2_XXS, IQ4_XS, IQ4_NL, MXFP4, …) without dequantizing to FP32, fused decode/prefill Metal kernels, compiled-graph kernels, async worker dispatch, batched MoE decode, and MoE expert offload. It pins the GGUF mmap in physical RAM via mlock(2) and derives allocator caps from the host's unified-memory capacity. Requires libmlxc (built locally or located via TENSORSHARP_MLX_LIBRARY / TENSORSHARP_MLX_LIBRARY_DIR).
Direct CUDA
--backend cuda is a pure-C# path using the CUDA Driver API, cuBLAS GEMM, and PTX kernels for common float32 ops (fill, unary/binary/ternary, activations, RMSNorm, softmax, RoPE/RoPEEx, SDPA, GQA prefill/decode, causal mask, gather/concat) plus native quantized matmul/get-rows for supported quant types. Unsupported ops route through CPU fallbacks while preserving tensor semantics. It is also the pure-C# backend where MTP speculative decoding is profitable.
Tensor parallelism uses --tp N on supported Direct CUDA and GGML CUDA / Vulkan architectures. Whole-layer placement uses --layer-split N on architectures that implement it, including the Direct CUDA DeepSeek V4 executor. Unsupported modes and CPU/MLX multi-GPU requests fail at startup. Only supported tensor-parallel architectures can add --tp-node-id / --tp-peers; layer split remains single-node.
CPU backends
- Pure C# CPU (
cpu) — portable inference with no native dependencies, using managed GEMM fast paths and fused SIMD kernels. Ideal for debugging and maximum portability. - GGML CPU (
ggml_cpu) — native GGML CPU kernels. Weight storage and performance depend on the model executor; benchmark the intended workload againstcpu.
The pure-C# path is multi-core. Managed matmuls run on a persistent spin-then-park worker pool (TensorSharp.Models/CpuWorkerPool.cs) instead of a per-matmul Parallel.For, and work items are sized from the work rather than from the thread count. Measured on gemma-4-E4B-it-Q8_0 with a 122-CPU allocation, A/B-ed inside one binary via TS_CPU_POOL: prefill ≈+15% and decode ≈2.8× against the old path. The pool deliberately does not take every core on a large host — its default width is every CPU up to eight and half of them above that (never fewer than eight), because spinning workers otherwise starve the rest of the CPU path, which still uses the ThreadPool (at full width the same model went backwards). Tune it with TS_CPU_THREADS, TS_CPU_POOL, TS_CPU_SPIN, TS_CPU_TASK_BYTES and TS_CPU_TASKS_PER_WORKER.
GEMM and SIMD kernels. Quantized weights (Q4_K, Q5_K, Q6_K, Q4_0, Q5_0, Q8_0) run through a multi-row int8 GEMM that decodes each pair of weight rows once for all activation rows, with AVX-512BW and AVX2 register tiles; F16, BF16, F32 and dequantize-only types through a float-panel GEMM. F32 matmuls behind Ops.Addmm use a packed, cache-blocked SGEMM (AVX-512 8×32, AVX2 6×16, or a portable kernel), and the elementwise, norm, softmax and RoPE ops use Vector512 / Vector256 kernels. TensorSharp.Models binds Core's parallel loops to the same worker pool, so all of them share one set of threads. Measured on an i7-11800H (8 cores / 16 threads, AVX-512, 32 GB): a DiffusionGemma-26B-A4B Jev read of a new 54-token prompt takes 0.92 s, against 9.5 s before this work and 1.57 s on ggml_cpu, and 0.22–0.25 s for further reads of the same prompt; Qwen-Image-2.1 at 512×512 with the Pruna 5-step LoRA takes 118 s against 295.7 s on ggml_cpu. Every kernel takes its instruction set from one decision (AVX-512 needs F/BW/DQ with Vector512 accelerated by the runtime), and TS_CPU_POOL=0 moves all of them to Parallel.For. TS_CPU_DISABLE_AVX512=1 runs the AVX2 kernels on an AVX-512 host; to emulate an AVX2-only host, start the process with DOTNET_EnableAVX512=0 (.NET 10 ignores the older DOTNET_EnableAVX512F=0). The AVX2 kernels have only been run that way, not on AVX2-only hardware.
The ModelBase loader maps quantized weights zero-copy. cpu was the last backend that copied every quantized tensor into fresh anonymous memory at load instead of binding it from the GGUF mapping the way the GGML backends always had; the loader now prints the split it got, e.g. Quantized: 103255 MB (103255 MB file-backed), F32: 983 MB. Any model whose weights used to be copied loads faster, and the effect is largest on big quantized checkpoints: GLM-5.3-Flash UD-Q2_K_XL went from a load that never completed (resident set 412 GB and still climbing) to ~48 s, most of it the page-cache prefault.
Managed i-quant coverage is wider. IQ2_XS and IQ4_XS have managed dequantizers (checked against ggml's own dequantize_row_*) and a place in the CPU quantized-storage matrix, so they stay quantized instead of being expanded to F32 at load — that expansion is what made GLM-5.3-Flash's 2-bit checkpoint ask for 765 GB. IQ2_XS × Q8_K and IQ3_XXS × Q8_K also have direct dot kernels with AVX2 paths (VecDotIq2XsQ8KAvx2, VecDotIq3XxsQ8KAvx2) rather than falling back to the generic dequantize-row-into-scratch path. Both changes apply to any model on --backend cpu. If you add another such kernel, note that ggml folds a constant into the result of some i-quant dots (0.125 for IQ2_XS, 0.25 for IQ3_XXS, 1.0 for IQ3_S) rather than into each per-block scale: omitting it is an 8× error that produces fluent-looking garbage rather than a crash.
Managed CPU coverage. Qwen-Image-2.1 runs its diffusion transformer, Qwen3-VL text and vision encoders, VAE, LoRA plug-ins and prefix KV cache in managed C#, without GGML. The default transformer uses 8-bit activations with multi-row integer GEMM; rounding differs from native GGML, so images need not be bit-identical. Its automatic output area is 1024×1024, and estimated memory is checked before inference. Current measurements and memory limits →
MiniMax-H3 also has managed text-to-video, keyframe and reference-conditioning paths; its vision route still has an unexplained numerical residual and throughput depends on the mode. Measured scope → Bonsai2 requires a single-device GGML backend. Qwen 3.8 Flash Next supports GGML and a dedicated direct-CUDA engine, but has no managed CPU model path; MLX is refused. A backend's presence does not establish support for every model, quantization or placement mode.
The Wan and MiniMax-H3 non-ggml paths share primitives: Direct{Context,Linear,Ops} (TensorSharp.Models/Direct/DirectOps.cs) back the Wan video networks and MiniMax-H3 on both cpu and cuda, and their row loops go through the same worker pool. On CPU DirectLinear no longer expands a quantized weight to F32 at load; it keeps the GGUF storage type and multiplies against it directly, which measured 80.9 s against 121.4 s on a small Wan render (256×160, 5 frames, 1 step, --backend cpu) at 4× less weight memory — and marginally closer to the native ggml_cpu render, not further (43.51 dB vs 43.39 dB). F16/BF16/F32 weights keep the plain GEMM.
DeepSeek V4.1 has a managed whole-model executor too. deepseek41 on --backend cpu runs the pure-C# DeepSeek4CpuExecutor — the ratio-1 and ratio-2 block compressors, the shared compressed and indexer caches with the lightning indexer's top-k, candidate block pruning, the Engram row gathers, the delayed hyper-connection gates and the checkpoint's trained cache quantization (FP8 E4M3 raw rows, MXFP4 indexer, NVFP4 compressed) — with no ggml, no native library and no GPU. It is held to the independent PyTorch oracle eng/dsv41-reference.py at atol=rtol=2e-5 plus greedy-argmax agreement, on a five-layer F32 fixture rather than the real weights, and fixture-dependent checks require TS_DSV41_FIXTURE_DIR; unavailable fixture or GPU scenarios are explicit skips. It reads Engram directly from the GGUF and has no vision — the companion is a native ggml component, so LoadVisionEncoder throws on this backend — and TS_DSV4_THREADS defaults to ProcessorCount here instead of the min(cores, 32) used elsewhere. Like cuda, it is a correctness and portability path, not a serving one: no throughput, load time or resident footprint has been measured for a full V4.1 checkpoint on it, and the TS_CPU_THREADS / TS_CPU_SPIN figures above are the generic worker pool, not V4.1 numbers. → deepseek41 backends
DeepSeek V4, V4.1 & GLM 5.x: dedicated whole-model executors
Three architectures skip the generic per-op forward. deepseek4's 284B compressed-sparse-attention MoE stack runs through one of three dedicated whole-model executors:
--backend cuda— a direct-CUDA engine independent of ggml. Quantized weights stream from the GGUF shards straight into per-device arenas.--backend ggml_cuda/ggml_vulkan— the native ggml executor: it loads the split GGUF itself, keeps all DSV4 KV state on-device, and runs each prefill/decode micro-batch as a single graph with a shape-signature cache so steady-state decode replays a captured CUDA graph.--backend cpu— a 100% pure-C# executor with no native dependencies, serving quantized weights straight from the memory-mapped GGUF shards.
The GPU executors use --layer-split N to place whole layers across N local GPUs, allowing models larger than one card to run. Without a placement configuration they use one device; the CPU executor streams weights from the mapped shards. Speculative decoding (DSpark) runs on --backend cuda and --backend ggml_cuda; on ggml_vulkan and the other backends a configured drafter is reported and ignored. → DeepSeek V4 downloads and commands
deepseek41 has one serving backend. DeepSeek V4.1's native graph runs under --backend ggml_cuda, its serving path — the only one with kernels for its fused ops. ggml_cpu loads the same graph on the scalar implementations those ops fall back to, which exists so the architecture can be checked without a GPU rather than served: the checkpoint reads six of 384 routed experts per layer per token out of a large split checkpoint. TS_DSV41_ALLOW_NON_CUDA_GPU=1 additionally permits ggml_vulkan and ggml_metal, where the ordinary graph runs on the GPU and only the architecture-specific ops fall to the CPU backend, at a host round trip each — opt-in, because what the refusal originally closed was a silent fallback. cpu runs a pure-C# V4.1 executor held to the PyTorch reference at 2e-5, and cuda runs V4.1 through the direct-CUDA engine's own kernels without ggml, with targeted synthetic kernel and sequence-state tests, not a full-checkpoint numerical gate; both are correctness and portability paths rather than serving ones. mlx fails before the weights are read rather than loading V4.1 weights into a graph that does not implement it. On the GPU, the architecture-specific ops are emitted as GGML_OP_CUSTOM nodes run by a TensorSharp backend that wraps its GPU's CUDA backend and takes its place in the scheduler, so ordinary nodes still go to CUDA as graph views on the same stream. An experimental V4.1 DSpark drafter (deepseek41-dspark) attaches only on ggml_cuda and ggml_cpu; see DSpark. → DeepSeek V4.1
glm-dsa's 744B-A40B MLA + DeepSeek-Sparse-Attention stack does the same across two:
--backend ggml_cuda/ggml_vulkan/ggml_cpu/ggml_metal— the native ggml executor loads the 6-shard split GGUF itself, accepts--layer-split Non CUDA or Vulkan to place 226 GiB of weights across N local GPUs (default: one device), owns the MLA and lightning-indexer caches on-device, and submits one graph per micro-batch through a shape-keyed graph cache so steady-state decode replays an allocated (and on CUDA, captured) graph.--backend cpu(100% managed, no native dependencies) and--backend cuda— the per-op path inTensorSharp.Models/Models/GlmDsa, which is also the reference the native executor is checked against;TS_GLM_NATIVE=0selects it on a GGML backend for an A/B. MLX does not run this family.
GLM-5.3 (not Flash) is the same glm-dsa block shape as GLM-5.2 — 79 blocks (78 trunk + one NextN), 256 routed experts at top-8 plus one shared expert, MLA with the lightning indexer, rope base 8e6 — so it loads on the path above with no new code and no new flag: the native whole-model executor on ggml_cuda / ggml_vulkan / ggml_cpu / ggml_metal, and the managed per-op path on cpu / cuda or on a GGML backend with TS_GLM_NATIVE=0; MLX does not run it either. It is text only twice over: unsloth/GLM-5.3-GGUF publishes no mmproj at any quant, and LoadVisionEncoder warns and ignores an --mmproj on glm-dsa, so a stray flag gets a text-only run rather than a failure. --tp N is accepted here as the same native local/single-process tensor parallelism, with the MLA and indexer caches replicated per rank — --tp-node-id / --tp-peers are refused for the whole GLM family before the model is built — and no GLM-5.3 run above one rank has been measured. --spec is engaged on single-device or explicit --layer-split N placement (no active tensor parallelism), drafting from the complete NextN block at blk.78: that block ships no nextn.shared_head_head.weight, so it borrows the trunk LM head, which is column-parallel under --tp, and the loader refuses to draft from one rank's strip of the vocabulary. Measured on 8× A40 46 GB without NVLink (UD-Q2_K_XL, seven shards, 236.4 GiB, a 10,531-token prompt, 300 decode tokens, median of 3, whole-layer placement): 251.6 t/s prefill and 20.48 t/s decode against llama.cpp's 20.28 t/s decode, and a 264 s load against 753 s — a decode tie with a 2.9× faster load, with TTFT the honest gap at 41.9 s against 29.0 s (~1.4× slower).
GLM-5.3-Flash (glm5next) loads through that same native executor and the same GlmDsaModel, and uses the same layer-split mode when --layer-split N is requested. On GGML GPU backends, --tp N selects native local/single-process tensor parallelism. Its eligible full-sharding configuration uses concurrent segmented rank-local graphs; CPU MoE, tracing, partial sharding, oversubscription, or missing native hyper-connection kernels select the combined scheduler fallback. NextN/MTP speculation is not implemented for it yet.
GLM-5.3-Flash also runs on --backend cpu — the same managed per-op path, with no native dependencies. It works for text, and it is a reference implementation to A/B against rather than a fast one: on the same file it is ~5.7× off ggml_cpu on prefill and ~2.4× on decode, and its prefill logits track ggml_cpu at a cosine of 0.9567 over the 154880-wide vocabulary — close, but not established as bit-parity, and enough for the greedy text to diverge where the top two logits are a near-tie. TS_DUMP_LOGITS writes the first real forward's logits to a file for exactly that comparison. Detail: the repository's docs/models/glm.md card.
Because 1M tokens of MLA cache is ~93 GiB, the advertised context is treated as a ceiling: after the weights land the loader measures the VRAM actually free and sizes the context to fit, logging its pick (342,272 tokens on a 3-GPU layer split). MAX_CONTEXT turns a specific length into a hard requirement instead. See GLM 5.x.
The server reports which backends are actually available on the host in GET /api/models (supportedBackends). If a CUDA or MLX backend is missing, the host did not detect a usable driver/runtime at startup. If ggml_vulkan is missing, the native bridge was not built with Vulkan enabled or no Vulkan 1.3 device/driver was found.
GLM-5.3-Flash on direct CUDA. The cuda engine has its own Glm5NextCudaEngine. Its synthetic four-layer fixture compares a 40-token prefill and six forced decode rows against ggml_cuda, requiring equal argmax and relative L2 error at most 2e-2; separate fixtures cover split placement, slots, capture and rollback. These contracts do not establish full-checkpoint quality or throughput, and this documentation refresh did not run GPU tests. Model card →
Building the native libraries
First install and verify the .NET 10 SDK for your platform. The native GGML library is built automatically on the first dotnet build. To build it manually or with CUDA:
cd TensorSharp.GGML.Native
# macOS (Metal)
bash build-macos.sh
# Linux — CPU only, force CUDA on, force Vulkan off
bash build-linux.sh
bash build-linux.sh --cuda
bash build-linux.sh --no-vulkan
# Windows — CPU only, force CUDA on, force Vulkan off
.\build-windows.ps1 --no-cuda
.\build-windows.ps1 --cuda
.\build-windows.ps1 --no-vulkan
The GGML Vulkan backend is enabled automatically when the machine has a Vulkan runtime (a loader such as vulkan-1.dll on Windows or libvulkan.so.1 on Linux, shipped by every recent GPU driver); opt out with --no-vulkan (or TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=OFF), and an explicit choice sticks across rebuilds via the CMake cache. On Windows, a LunarG Vulkan SDK is used when installed; without one the build auto-provisions a portable toolchain (Vulkan-Headers, a vulkan-1 import library generated from the system loader, glslc, SPIRV-Headers) into ExternalProjects/vulkan-toolchain/ via eng/fetch-vulkan-toolchain.ps1. On Linux, distro dev packages are used when present (apt install libvulkan-dev glslc spirv-headers); otherwise the build auto-provisions the missing pieces (Vulkan-Headers, glslc from the shaderc CI prebuilts, SPIRV-Headers) via eng/fetch-vulkan-toolchain.sh. A GPU driver with Vulkan 1.3 support is required at run time.
On Windows/Linux the script auto-detects the visible NVIDIA GPU compute capability and passes a narrow CMAKE_CUDA_ARCHITECTURES (e.g. 86-real on an RTX 3080), which cuts CUDA build time substantially. Override it explicitly:
TENSORSHARP_GGML_NATIVE_CUDA_ARCHITECTURES='86-real;89-real' bash build-linux.sh --cuda
bash build-linux.sh --cuda --cuda-arch='86-real;89-real'
You can also request CUDA from dotnet build directly:
TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON dotnet build TensorSharp.Cli/TensorSharp.Cli.csproj -c Release
MLX native library (macOS only)
The MLX backend depends on libmlxc. A helper script fetches and builds it:
bash TensorSharp.Backends.MLX/build-native-macos.sh
It writes the libraries into TensorSharp.Backends.MLX/Native/dist/. At run time the backend probes the application directory first; point it elsewhere with TENSORSHARP_MLX_LIBRARY or TENSORSHARP_MLX_LIBRARY_DIR.
With the macOS 27 SDK, which compiles Metal shaders as Metal 4.1 and rejects the pinned MLX kernels, the MLX build sets CMAKE_OSX_DEPLOYMENT_TARGET=26.2 to select Metal 4.0. An explicit CMAKE_OSX_DEPLOYMENT_TARGET or MACOSX_DEPLOYMENT_TARGET takes precedence, and older SDKs keep their defaults.
Platform binary release status
Each application is published for the following platform/backend matrix:
| Archive | Native backend(s) bundled | Format |
|---|---|---|
win-x64-cpu | GGML CPU | .zip |
win-x64-cuda | GGML CUDA + pure-C# CUDA (PTX) + CUDA 12.x runtime | .zip |
linux-x64-cpu | GGML CPU | .tar.gz |
linux-x64-cuda | GGML CUDA + pure-C# CUDA (PTX) + CUDA 12.x runtime | .tar.gz |
osx-arm64 | GGML Metal + MLX | .tar.gz |
linux-arm64-cuda13-GB10 (experimental) | GGML CUDA for SM121a + pure-C# CUDA (compute_121 PTX) + CUDA 13 runtime, for NVIDIA GB10 / DGX Spark | .tar.gz |
Each release contains both a CLI and a Server file for every suffix above. The -cuda archives still require an NVIDIA GPU and compatible driver; the macOS archives require Apple Silicon. Source builds remain available.
The GB10 archives are built by eng/Dockerfile.gb10 on a hosted Linux ARM64 runner with no GPU: CUDA 13.0.2, .NET SDK 10.0.401 and an unmodified, pinned ggml revision. The release job checks dependency loading and startup in a clean Ubuntu 24.04 image, not model inference, and publishes SHA256SUMS plus a validation record beside the two archives. The recorded GB10 hardware smoke test predates the upstream ggml reintegration and does not certify the current code. Older releases may not contain these archives; check the assets of the specific release. Building the image locally is covered in the repository's DEVELOPMENT.md.