⚡ C# · .NET 10 · GGUF · AgentHost · GPU-accelerated

TensorSharp

A native .NET inference engine for GGUF models — with a CLI, browser chat, Ollama- & OpenAI-compatible APIs, and an optional AgentHost for progressive-disclosure skills and sandboxed model-authored code.

Everything runs on your own hardware: your laptop, workstation, or server. Inference and agent work stay local by default, there are no per-token fees, and the same engine powers a quick command-line test, a shared internal chatbot, and a production REST endpoint. This wiki is the complete guide: pick a starting point below or use / to search.

TensorAgent: text, images, video and agentic work

One local app for different kinds of work: chat and reason with text models, read photos and files, generate and edit images with Qwen-Image 2.1, make video with sound using MiniMax-H3, and use skills and tools to create files, run code or work in a browser. Choose a model for the task; the available models depend on the device's memory.

Source builds target iPhone, iPad, macOS and Windows. The interface supports English, Simplified and Traditional Chinese, Japanese, Korean, Spanish, French and German. Explore TensorAgent's capabilities and platform limits →

Current source: Qwen-Image 2.1 and mask editing, TensorAgent desktop builds and translations, browser skills and sub-agents, and the newest native parallelism changes are newer than v2026.09.01. Build from this checkout to use them; downloaded release assets reflect their own tag.

Supported model families

Image & video generation

Embeddings · BERT / XLM-R

Backend availability, input modalities, tool calling, and validation vary by family. Compare model capabilities →

Learn with the books

Two books by Zhongkai Fu turn the TensorSharp source into guided implementation journeys.

Cover of Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgent by Zhongkai Fu

Qwen · Inference & agents

Building LLM Inference Engines and Agentic Runtimes from Scratch

Qwen Dense and MoE Models with TensorSharp and TensorAgent

Follow a C# implementation from tensors and tokenization to Qwen dense and MoE inference, then to agent tools, skills, code execution, and sandbox boundaries. TensorSharp and TensorAgent connect the design to GPU optimization, multimodal input, and desktop and mobile deployment.

Cover of From Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B by Zhongkai Fu

Gemma 4 · Multimodal inference

From Tensors to Tokens

Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B

Build a multimodal inference engine in C# and .NET using Gemma 4 E4B. Work through tensor storage, GGUF loading, quantization, and tokenization, then explore image, video, and audio inputs, numerical checks, caching, and batching.

Explore the wiki

🚀

Getting Started

Prerequisites, build, download a model, and stream your first reply.

⌨️

Command Line

Run prompts, images, audio, batches, and benchmarks from the CLI.

🌐

Server & Web UI

Host a browser chatbot and HTTP endpoints on localhost.

🛠️

Agentic Work

Load skills on demand, run sandboxed code tools, delegate to bounded sub-agents on the server, repair outputs, and download generated artifacts.

📱

TensorAgent

Local text, image editing, video, skills and tools on iPhone, iPad, Mac and Windows, within each device’s limits.

🔌

HTTP API

Call it from curl, Python, or any Ollama/OpenAI client.

🧩

C# Library

Embed the engine directly in your .NET application.

📚

API Reference

Searchable tables of flags, env vars, endpoints, and types.

🧠

Models

Supported model families, downloads, multimodal, reasoning — and which checkpoint to pick when you want it fast.

🔗

Multi-GPU & Multi-Node

Tensor parallelism across GPUs — CUDA and GGML — and across machines.

📖

Glossary & FAQ

New to LLMs? Plain-language definitions and common questions.

Quick start in ~30 seconds

After installing the .NET 10 SDK for your platform, git, and curl, this path runs the benchmark-verified Gemma 4 E4B Q8_0 model on a native GGML backend. Copying the commands takes about 30 seconds; the 7.48 GiB model download and the first restore/build take longer and depend on your connection and machine. This block is for Linux + NVIDIA — see other backends below. On Windows PowerShell, use curl.exe.

  1. Clone

    The first dotnet run below restores the projects and compiles the native GGML CUDA backend; that first build can take several minutes.

    git clone https://github.com/zhongkaifu/TensorSharp.git
    cd TensorSharp
  2. Download the model

    The recommended artifact is gemma-4-E4B-it-Q8_0.gguf (7.48 GiB); the lower-memory gemma-4-E4B-it-Q4_K_M.gguf lives in the same repository.

    curl --create-dirs --fail -L "https://huggingface.co/ggml-org/gemma-4-E4B-it-GGUF/resolve/main/gemma-4-E4B-it-Q8_0.gguf?download=true" -o models/gemma-4-E4B-it-Q8_0.gguf
  3. Run it

    The environment assignment enables the CUDA native build. E4B also supports thinking (--think) and tool calling (--tools).

    echo "What is TensorSharp? Answer briefly." > prompt.txt
    TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON dotnet run --project TensorSharp.Cli -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --input prompt.txt --max-tokens 128 --backend ggml_cuda
  4. Prefer a UI + API?

    Start the server with the same model. The browser chat is at /; the plain liveness check is at /health.

    TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON dotnet run --project TensorSharp.Server.Host -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --backend ggml_cuda --max-tokens 128
    # open http://localhost:5000/

Other backends and multimodal

On Apple Silicon, omit the CUDA environment assignment and use ggml_metal; on a supported Windows/Linux Vulkan GPU, request TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=ON instead and use ggml_vulkan; without a GPU, use ggml_cpu (native CPU kernels, no GPU toolchain). Text inference needs only the model GGUF. For image, video, or audio input, also download the matching mmproj-gemma-4-E4B-it-Q8_0.gguf and pass it with --mmproj. Windows PowerShell and full platform syntax are in the E4B fast-lane notes.

See it running

One engine, four ways to use it, each an unedited capture of a real run. Select one to see it full size.

TensorSharp.Cli in a terminal: an interactive chat with Gemma 4 E4B on Metal that reads the project README and answers two questions about it

TensorSharp.Cli

Run models, chats and benchmarks from a terminal.

Command Line →
The TensorSharp Web UI: Qwen3.8 27B compared two mortgages by writing and running a Python script, answered with a table and a recommendation, and offered the script and the CSV as downloads

Web UI chat · TensorSharp.Server.Host

A browser chat plus Ollama- and OpenAI-compatible APIs.

Server & Web UI →
TensorAgent on an iPhone: Gemma 4 E2B scaled a recipe from 4 to 10 people by running a Python script in the app's built-in Python, and answered with a table

TensorAgent on iPhone

A private agent that runs the model on the phone.

TensorAgent for iOS →
TensorAgent on a Mac: Qwen-Image 2.1 edits a TensorSharp banner to a deep blue night sky with stars while preserving the text, with Compare original and Edit again controls below the result

TensorAgent on the desktop

Qwen-Image 2.1 edits an attached image from an instruction. Text chat, files, code, browser skills and video share the same app.

TensorAgent on the Mac and Windows →

Why TensorSharp?

🔒

Private by default

Inference stays on your hardware. Command and skill-script networking is denied by default; Windows unconfined code execution is an explicit exception.

🛠️

Skills that do the work

Skills load on demand, opt-in sandboxed file and code tools let the model do real work, and on the server it can hand independent tasks to sub-agents that are read-only by default.

💸

No per-token bill

Run as much as your hardware allows — predictable cost, no metered API.

🔁

Drop-in compatible

Speaks the Ollama and OpenAI wire formats, so existing tools and SDKs just work.

🖥️

Runs anywhere

NVIDIA (CUDA), AMD / Intel / NVIDIA (Vulkan), Apple Silicon (Metal/MLX) and iOS (Metal), or pure CPU — with automatic fallbacks.

🧠

Modern model support

Twelve text families with vision, audio, PDF documents, reasoning and tools, plus three that generate images and video instead of text.

🎬

Images and video out, too

Qwen-Image-2.1 generates, edits and makes precise changes inside a painted selection, MiniMax-H3 makes video with a native stereo soundtrack, and Wan 2.1 / 2.2 make video alone, all on the same engine.

⚙️

Built in .NET

A native C# engine you can embed in your apps, not just a black-box binary.

🔗

Scales past one GPU

Spread one model over several GPUs with tensor parallelism (--tp N) or whole-layer placement (--layer-split N), where the model and backend support it.

🏁

Benchmarked vs llama.cpp

On identical GGUF files and the same GPU it trades wins with llama.cpp: up to 1.28× faster CUDA prefill and 1.21× faster Vulkan decode in the current comparison.

Who is this for?

TensorSharp serves a wide range of visitors. Here is the fastest path for each.

Beginners & students

Start with the guided book or Glossary & FAQ, then try Getting Started.

Developers

Jump to Agentic Work, the HTTP API, C# Library, and API Reference.

Senior / principal engineers

Read Advanced Features — paged KV, continuous batching, speculative decoding.

Managers, CTOs & CEOs

See the business value and capability matrix.

Sales & marketing

Use the feature catalog and benchmarks for positioning.

Researchers & professors

Explore model architectures and the head-to-head benchmarks.