Overview & Architecture
TensorSharp is a native .NET LLM inference engine for GGUF models — including autoregressive LLMs and DiffusionGemma-style text-diffusion models. It ships a console application, a web-based chatbot, and Ollama/OpenAI-compatible HTTP APIs.
What it is — in plain terms
A large language model (LLM) is a neural network that predicts text. To use one you need an inference engine: software that loads the model's weights and runs the math to turn your prompt into a reply. TensorSharp is that engine, written in modern C# / .NET 10, focused on the GGUF model format used widely in the local-LLM ecosystem.
It gives you three ways to use a model, all from one binary set:
- Command line — run a prompt, an image, a batch of questions, or a benchmark. → CLI
- Server — a browser chatbot plus REST endpoints that mimic Ollama and OpenAI. → Server
- Library — reference the source projects and call the engine from your own .NET code. → C# Library
Business value
For decision-makers evaluating local inference, the core trade is control and cost in exchange for running your own hardware.
Data privacy & compliance
Prompts and documents never leave your infrastructure — a fit for regulated, on-prem, or air-gapped environments.
Predictable cost
No per-token API billing. Capacity is bounded by hardware you already budget for.
No vendor lock-in
Open GGUF models from Hugging Face, and an OpenAI/Ollama-compatible surface that existing tools already speak.
.NET-native integration
Embed inference inside existing C# services instead of bridging to an external runtime.
Scales on one box
Continuous batching serves many concurrent users from a single hosted model.
Migrate easily
Point existing OpenAI SDK code at http://localhost:5000/v1 and keep your application unchanged.
Architecture
TensorSharp is a layered system. Each layer is an independently packable project, so source consumers can reference only what they need. The publish set is thirteen packages, verified by eng/verify-packages.ps1. Package and release versions follow their release tag; features merged after v2026.09.01 require current source builds. Package roles and references →
| Layer | Responsibility |
|---|---|
| TensorSharp.Core | The core Tensor type, storage abstraction, device abstraction, and the extensible operation registry (Ops). CPU implementations use System.Numerics.Vectors for SIMD. |
| TensorSharp.Runtime | GGUF parsing, tokenizers (SentencePiece / BPE), chat-template rendering, sampling, output parsing, the paged KV cache, and the continuous-batching scheduler / engine. Prompt framing and reply parsing are declared per family as a ChatProtocol in ChatProtocolRegistry rather than keyed off architecture names. |
| TensorSharp.AgentHost | The optional agentic layer over Runtime: Agent Skills, the internal model/tool loop, bounded sub-agent delegation, sandboxed script and code execution, persistent workspaces, file editing, package installation, and artifact capture. Runtime never references back into it, so plain inference hosts do not carry agentic machinery. |
| TensorSharp.Models | ModelBase plus the concrete architectures and multimodal encoders. Generative families use plug-ins — a ModelArchitectureDescriptor registered in BuiltInArchitectures — and ModelBase.Create() auto-detects the architecture from GGUF metadata, then hands off to the matching descriptor’s factory. Sentence encoders use EmbeddingModel.Load(); see embeddings. |
| TensorSharp.Backends.GGML | Accelerated ops via a native C++ bridge (libGgmlOps) linking ggml — Metal on macOS, CUDA and Vulkan on Windows/Linux, and native CPU. |
| TensorSharp.Backends.Cuda | The direct CUDA path: CUDA Driver API, cuBLAS GEMM, and PTX kernels for hot ops, with CPU fallbacks. |
| TensorSharp.Backends.MLX | The Apple-Silicon MLX path, wrapping mlx-c with quantized, fused, and compiled kernels. |
| TensorSharp.Distributed | Peer-to-peer TCP coordination for multi-node tensor parallelism, kept separate from the model/runtime layers. |
| TensorSharp.Chat | The host-neutral chat pipeline behind the server and TensorAgent: model service, sessions, generation, the per-model continuous-batching engine host, the skills and sub-agent loop, and the Web UI request/stream contract, with no ASP.NET Core dependency. The CLI project references the library but does not run on its pipeline: it drives TensorSharp.Runtime's inference engine with its own loop. |
| TensorSharp.Server | The ASP.NET Core HTTP / application layer over TensorSharp.Chat: Ollama- and OpenAI-compatible REST APIs, the browser chat UI's routes, upload handling, and option parsing. TensorSharp.Server.Host is the executable that serves the page. |
| TensorSharp.Cli | The console host for local prompts, multimodal experiments, prompt inspection, JSONL batch workflows, the interactive REPL, and benchmarks. |
Architectures are plug-ins, not switch statements. A family declares itself once and nothing else has to learn its name: a ModelArchitectureDescriptor (id, aliases, factory, display name, tensor-table detection for GGUFs that carry no metadata, multi-GPU mode, native tunables, projector file hints) in BuiltInArchitectures, plus a ChatProtocol for its text format. Everything downstream asks capability interfaces instead of testing an architecture string — IVisionCapableModel, IAudioCapableModel, IMRoPEPositionSink, IBatchedPagedModel, ISpeculativeTarget, ISpeculator — so adding a model, a modality or a chat format touches its own descriptor and its own folder, and ModelBase itself is split by concern into partials (ModelBase.TensorParallel.cs, .CpuAttention.cs, .WeightLoading.cs, .KvCache.cs, .Warmup.cs). The procedure is written up in DEVELOPMENT.md and the model-card README.
The Wan and MiniMax-H3 non-ggml paths share execution primitives: Direct{Context, Linear, Ops} in TensorSharp.Models/Direct/DirectOps.cs, used by the Wan video networks and MiniMax-H3 alike on BackendType.Cuda and BackendType.Cpu. On CPU a DirectLinear keeps the weight in its GGUF storage type instead of expanding it to F32 at load: a Wan 256×160 5-frame single-step render on --backend cpu went from 121.4 s to 80.9 s at a quarter of the weight memory, and landed marginally closer to the native ggml_cpu render (43.51 dB against the old path’s 43.39 dB). F16/BF16/F32 weights keep the plain GEMM.
The ModelBase loader on the managed cpu backend now loads the way its GGML backends do: quantized weights are bound zero-copy from the GGUF mapping instead of being copied into fresh anonymous memory, and the loader prints the split it got — Quantized: 103255 MB (103255 MB file-backed), F32: 983 MB for GLM-5.3-Flash UD-Q2_K_XL, whose load went from never completing (resident set 412 GB and still climbing) to ~48 s, most of that the page-cache prefault. Any model whose weights were previously copied benefits; the effect is largest on big quantized checkpoints. ManagedQuantizedOps also gained IQ2_XS and IQ4_XS — managed dequantizers verified against ggml's own dequantize_row_*, plus entry into the CPU quantized-storage matrix, so they are kept quantized rather than expanded to F32 at load — and direct IQ2_XS × Q8_K and IQ3_XXS × Q8_K dot kernels with AVX2 paths (VecDotIq2XsQ8KAvx2, VecDotIq3XxsQ8KAvx2); both types previously fell back to the generic dequantize-row-into-scratch path.
Generic operations can fall back to CPU, but model-specific executors have explicit backend and placement requirements. Unsupported combinations fail at load; numerical and performance coverage remains specific to each model, checkpoint and device. See the backend guide and model cards.
How a request flows
- Load —
ModelBase.Create(path, backend)reads GGUF metadata, picks the architecture, and maps the quantized weights into the chosen backend. - Render — the prompt (and any system message, images, audio, tools) is turned into tokens via the architecture's chat template and tokenizer.
- Prefill — the prompt tokens are processed in a batched forward pass that populates the KV cache.
- Decode — tokens are generated one at a time (optionally several at once via speculative decoding), sampled with your settings, and streamed back.
- Serve — in the server, the continuous-batching engine interleaves many requests against one model, sharing KV-cache prefixes across them.
Project structure
The repository is organized by the layers above. The most useful entry points:
| Path | Contents |
|---|---|
TensorSharp.Core/ | Tensor library, ops, memory, device abstraction, CPU SIMD/quantized kernels. |
TensorSharp.Runtime/ | GGUF, tokenizers, templates, sampling; Paged/ KV primitives, Scheduling/ the inference engine, scheduler and radix prefix cache, and Speculative/ the shared draft/verify core (MTP/NextN draft heads, DSpark/DFlash block drafting, n-gram). |
TensorSharp.AgentHost/ | Skills, progressive disclosure, the bounded tool loop, bounded sub-agents (Agents/), OS sandboxes, session/request workspaces, four code tools, package installs, and generated artifacts. |
TensorSharp.Models/Architecture/ | The architecture plug-in table: ModelArchitectureDescriptor, ModelArchitectureRegistry, BuiltInArchitectures, and the multimodal capability interfaces. |
TensorSharp.Models/Models/<Family>/ | One folder per architecture (DeepSeek4 — V4 and V4.1, GlmDsa, Gemma4, Qwen35, Qwen4Exp, GptOss, Nemotron, Mistral3, HunyuanDense, MuseGlimmer, DiffusionGemma, QwenImage, MiniMaxH3, WanVideo), plus Video/ for the contracts the video families share. The autoregressive text families carry a single-sequence and a batched forward; DeepSeek4 and GlmDsa use whole-model executors instead, and the media families (DiffusionGemma, QwenImage, MiniMaxH3, WanVideo) run diffusion pipelines rather than a decode loop. |
TensorSharp.Models/Direct/ | Direct{Context, Linear, Ops} — the non-ggml execution primitives shared by the Wan video networks and MiniMax-H3 on the cuda and cpu backends. |
TensorSharp.GGML.Native/ | The native C++ bridge to ggml (matmul, fused transformer kernels, paged attention, MoE, Mamba2, GatedDeltaNet, diffusion). |
TensorSharp.Distributed/ | The TCP mesh and collectives used by multi-node tensor parallelism. |
TensorSharp.Chat/ | Host-neutral model service, sessions, chat pipeline, inference-engine host, skills and sub-agent loop, telemetry. |
TensorSharp.Server/, TensorSharp.Server.Host/ | ASP.NET Core endpoints, protocol adapters and option parsing; the server executable, its startup, and the Web UI page. |
docs/ | Per-model architecture cards, paged-attention deep dive, env-var matrix, benchmark matrix. |
Current status & capabilities
| Area | Status |
|---|---|
| Model families | Fifteen architecture plug-ins, dispatched by ModelBase.Create() from the GGUF's general.architecture: DeepSeek V4 Flash (deepseek4), DeepSeek V4.1 Flash (deepseek41), GLM 5.x (glm-dsa, glm_dsa, glm5next), Gemma 4 (gemma4), DiffusionGemma (diffusion-gemma, diffusion_gemma), Qwen 3.5/3.6-family (qwen35, qwen35moe, qwen3next; Bonsai2's low-bit checkpoints also declare qwen35 and are recognised by their PRISM Hadamard metadata or tensor types), Qwen 3.8 Flash Next (qwen4exp), GPT OSS (gptoss, gpt-oss), Nemotron-H incl. Nemotron 3 Nano Omni (nemotron_h, nemotron_h_moe, nemotron_h_omni), Mistral 3 (mistral3), Hunyuan Dense (hunyuan-dense), Muse-Glimmer (muse-glimmer, muse_glimmer), Qwen-Image-2.1 text-to-image and image editing (qwen_image, qwen-image), MiniMax-H3 joint audio-video generation (minimax-h3, minimax_h3), and Wan 2.1 / 2.2 video generation (wan, wan2.1, wan2.2). Two families do not need that metadata: MiniMax-H3's published GGUFs carry none at all, and neither do some community Qwen-Image-2.1 GGUFs (such as Unsloth's Q8_0), so ModelBase.Create() recognises both by their tensor table instead. A qwen_image file that is not 2.1, such as Qwen-Image-Edit-2511, is refused at load. → Supported architectures |
| Inference hosts | CLI, interactive REPL, ASP.NET Core web UI, Ollama-style API, OpenAI Chat Completions- and Responses-style APIs, a Jev-compatible POST /v1/systemone typed-decision endpoint served by DiffusionGemma, and TensorAgent on iOS/iPadOS, macOS and Windows (source builds). |
| Agentic work | TensorSharp.AgentHost runs a bounded tool loop for Agent Skills and, when the operator enables --code-exec, read_file, write_file (new files only), shell, and apply_patch. TensorSharp executes only its own built-ins; caller-defined tools are returned to the caller. Workspaces persist for a Web/CLI chat or one isolated OpenAI/Ollama request, and OS confinement is required by default. On the server's chat paths a tool-capable model can also delegate independent work to bounded sub-agents — on by default, read-only by default (only worker children can gain the parent's mutable tools, and only with --agents-allow-worker-tools), and turned off with --no-multi-agent; TensorAgent runs the same delegation and can turn it off in Settings → Sandbox → Sub-agents; the CLI runs a single agent. There is no interactive approval service. → Agentic Work |
| Backends | Pure C# CPU, direct CUDA/cuBLAS (cuda), MLX Metal (mlx), GGML CPU, GGML Metal, GGML CUDA, GGML Vulkan, plus the iOS/iPadOS TensorAgent target using ggml_metal through a statically linked .xcframework. DeepSeek V4 additionally has three whole-model executors of its own — direct CUDA, native ggml, and a pure-C# CPU one — with GPU executors placing whole layers through --layer-split N. DeepSeek V4.1 Flash has its own set too: ggml_cuda is its serving backend, ggml_cpu runs the same native graph on the scalar kernels those fused ops fall back to, --backend cpu runs the pure-C# DeepSeek4CpuExecutor — no ggml, no native library, no GPU — and cuda runs V4.1 through the direct-CUDA engine's own kernels; the last two are correctness and portability paths rather than serving ones. GLM 5.x — GLM-5.2, GLM-5.3 and GLM-5.3-Flash alike — has two: a native ggml executor on ggml_cuda / ggml_vulkan / ggml_cpu / ggml_metal (whole-layer placement with --layer-split N; on GGML GPU backends --tp N selects native local/single-process TP for GLM-5.2, GLM-5.3 and GLM-5.3-Flash alike) and the managed per-op path on cpu and cuda, which a GGML backend can also be forced onto with TS_GLM_NATIVE=0; MLX does not run it. |
| Speed levers | The choice that matters most is usually the checkpoint, not the flag. MiniMax-H3 is the one family where it is not: it ships CFG-distilled, so there is no faster file to fetch — --cfg 1.0 at 4–8 steps is the operating point against its 20-step default, and --width / --height is the dominant quality-versus-cost lever (640×384 is the documented starting point). A step-distilled Wan DiT (Turbo / Lightning / lightx2v / FastWan / …-4steps-… in the file name) is detected at load and runs 4 guidance-free DiT passes instead of the base recipe's 100: the same 1088×832×121-frame image-to-video request measured ≈3 h 30 m on the base Wan2.2-TI2V-5B and 17 m 30 s on the Turbo one (M5 Pro, ggml_metal). Qwen-Image-2.1's CFG 1 default already runs one transformer prediction per step, so its levers are resolution — --width 1024 --height 1024 quarters the latent image tokens of the 2048×2048 default — and step count: a step-distillation LoRA plug-in from config/lora/ (Viggle Turbo, Pruna 8- or 5-step, PAI Fun-Acc 4-step) cuts the 40-step default to 4–8 transformer passes. For text models the levers are the backend (ggml_cuda on NVIDIA, ggml_metal on Apple Silicon, ggml_cpu rather than the pure-managed cpu), whole-model fused decode graphs (GPT OSS 24 → 154 tok/s on an A40), speculative decoding, --tp N, and --n-cpu-moe for a model that would otherwise not fit. → Models · Benchmarks |
| Multimodal | Gemma 4 image/video/audio; Qwen 3.5-family, Qwen 3.8 Flash Next, GLM-5.3-Flash, Mistral 3, Nemotron-H Omni, and Muse-Glimmer image input, with video_url video on Qwen 3.8 Flash Next too; DiffusionGemma image input through the Gemma 4 vision tower, loaded from an mmproj GGUF or the upstream model-00011-of-00011.safetensors shard (audio and video are refused), which also lets Jev's /v1/systemone read up to eight images; DeepSeek V4.1 image and video once its separately prepared vision companion is attached with --mmproj (audio is refused, not ignored). PDF documents via Web UI upload and CLI --pdf (text extracted and inlined; scanned pages rendered as images for vision models). Media out: MiniMax-H3 (H.264 MP4 plus a native 32 kHz stereo .wav, denoised together in one packed latent — text → video, image → video, first/last frame, reference → video), Qwen-Image-2.1 (image), and Wan 2.1 / 2.2 (H.264 MP4 video only, text → video and image → video). → Video generation |
| Continuous batching | vLLM-style paged KV cache, a radix-tree prefix cache that reuses shared prefixes across requests by default (--no-prefix-cache turns reuse off), iteration-level scheduler (on by default; opt-out via --no-continuous-batching). DeepSeek V4 and GLM 5.x serve through their own native per-sequence slots on the same engine — a compressed MLA cache row per token has no paged layout to page — and GLM uses a default-on batched fused decode (1.81× aggregate at four concurrent requests; TS_BATCHED_FUSED_DECODE=0 disables it). Qwen 3.8 Flash Next uses per-sequence state holders: each in-flight request owns its KV, GatedDeltaNet, PLE and indexer state, so switching between them is a reference swap rather than a state download. |
| Speculative decoding | MTP / NextN draft heads for Qwen 3.6 and Qwen3.8-27B (embedded), GLM-5.2 and GLM-5.3 (glm-dsa, embedded — --spec loads the NextN block the checkpoint already carries, with nothing to download, and it engages on single-device or explicit --layer-split N placement: under --tp N>1 the draft block borrows the trunk's column-parallel LM head and the loader refuses to draft from one rank's strip of the vocabulary; GLM-5.3-Flash has no NextN implementation), Gemma 4 (separate draft GGUF), and Qwen 3.8 Flash Next (a separate shared MTP head, GGML backends only); DSpark block drafting for DeepSeek V4 (separate drafter GGUF via --draft-model on the cuda / ggml_cuda backends) and, experimentally, for DeepSeek V4.1 (a deepseek41-dspark drafter on ggml_cuda / ggml_cpu; initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified); DFlash / DFlash2 block drafters for Muse-Glimmer and the Qwen 3.5 family (separate drafter GGUF via --draft-model); and a weight-free n-gram drafter on models that accept speculation: not GPT OSS, Mistral 3, Hunyuan Dense, or DiffusionGemma; DeepSeek V4 / V4.1 and Muse-Glimmer speculate only while their own drafter is loaded; and Nemotron-H refuses every speculator. Every emitted token is drawn by the request's own sampler from a trunk row computed for the right prefix, so drafting does not change the distribution being sampled. The multi-row verify kernel accumulates in a different order from the one-row decode kernel, though, so at a logit near-tie a greedy stream can diverge from plain decoding. Off by default — opt in with --spec on the CLI or the server (env TS_SPEC); naming a --draft-model enables it by itself unless --no-spec is given. → DSpark · DFlash |
| Multi-GPU | --tp N selects tensor parallelism, sharding weights within layers. --layer-split N selects whole-layer placement on supported architectures. The flags are mutually exclusive; unsupported requests fail at startup. Layer split is local to one node. Only supported tensor-parallel architectures can use --tp-node-id / --tp-peers across nodes. Migrate old layer-placement commands from --tp N to --layer-split N (environment: TENSORSHARP_LAYER_SPLIT_DEGREE=N). → Multi-GPU & Multi-Node |
| Observability | Structured per-turn logs, queue status, and KV-cache reuse metrics across Web UI, Ollama, and OpenAI response shapes. |