Getting Started
From a source checkout to a verified CLI reply or local API. The quick start runs Gemma 4 E4B on a native GGML backend; a managed CPU path (no CMake or GPU toolchain) remains available when you cannot build native code.
Following the Gemma 4 E4B path? From Tensors to Tokens turns this quick start into a guided, from-scratch explanation of the engine behind every step. Explore the TensorSharp book β
Current source: Qwen-Image 2.1 and mask editing, TensorAgent desktop builds and translations, browser skills and sub-agents, and the newest native parallelism changes are newer than v2026.09.01. Build from this checkout to use them; downloaded release assets reflect their own tag.
1 Β· Prerequisites
Install the .NET 10 SDK
TensorSharp targets net10.0, so a source build requires the full .NET 10 SDKβinstalling only the .NET Runtime or ASP.NET Core Runtime is not enough. The SDK includes the runtimes needed to run the CLI and server. Start with Microsoft's cross-platform .NET installation guide, then use the instructions for your operating system:
| Platform | Install the SDK | Official instructions |
|---|---|---|
| Windows | Open PowerShell or Command Prompt and run winget install Microsoft.DotNet.SDK.10. | Install .NET on Windows |
| macOS | Download the .NET 10 SDK installer for your processor: Arm64 for Apple Silicon or x64 for an Intel Mac. | Install .NET on macOS |
| Linux | Follow the page for your distribution, configure its package feed if directed, and install dotnet-sdk-10.0. On Ubuntu, after completing the repository setup for your release: sudo apt-get update, then sudo apt-get install -y dotnet-sdk-10.0. | Choose a Linux distribution (for example, Ubuntu) |
Open a new terminal after installation and confirm that a 10.0.x SDK is listed:
dotnet --list-sdks
Other prerequisites
git,curl, and network access β used to clone TensorSharp and download a model. A full native build also clones ggml intoExternalProjects/ggml/; setTENSORSHARP_GGML_NO_UPDATE=1to skip later network updates.- A GGUF model file β e.g. from Hugging Face. See Model downloads.
Per-platform toolchains (only for GPU acceleration)
| Platform | Needed for | Install |
|---|---|---|
| macOS (Metal) | ggml_metal / mlx | CMake 3.20+ and Xcode command-line tools. MLX additionally builds libmlxc. |
| Windows | ggml_cuda / cuda | CMake 3.20+, Visual Studio 2022 C++ build tools, NVIDIA driver + CUDA Toolkit 12.x (with cuBLAS). |
| Linux | ggml_cuda / cuda | CMake 3.20+, NVIDIA driver + CUDA Toolkit 12.x (with cuBLAS). |
| Linux ARM64 + NVIDIA GB10 (DGX Spark, experimental) | ggml_cuda / cuda | CUDA 13. Build inside the eng/Dockerfile.gb10 container on a native Linux ARM64 Docker builder (no GPU needed to compile), or use the linux-arm64-cuda13-GB10 release archives, which need only a compatible NVIDIA driver. See platform binaries. |
| Windows / Linux (any Vulkan GPU) | ggml_vulkan | Enabled automatically when the machine has a Vulkan runtime (loader installed); opt out with --no-vulkan (or TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=OFF). Windows auto-provisions a portable Vulkan toolchain when no SDK is installed; on Linux install e.g. libvulkan-dev glslc spirv-headers. Needs a Vulkan 1.3 driver at run time. |
| Any (CPU only) | cpu | Nothing beyond the .NET SDK. ggml_cpu is faster but is a native build. |
2 Β· Clone and choose a build path
git clone https://github.com/zhongkaifu/TensorSharp.git
cd TensorSharp
Managed-CPU path (no native toolchain)
To skip every native build, pass both skip properties and use the managed cpu backend. The first run restores and builds the managed projects automatically.
-p:TensorSharpSkipGgmlNative=true -p:TensorSharpSkipMlxNative=true -- --backend cpu
Full native / GPU path
Build the whole solution when you want GGML CPU, Metal, CUDA, Vulkan, or MLX. The first build compiles the native bridge and can take several minutes; subsequent builds are incremental.
dotnet build TensorSharp.slnx -c Release
On macOS and Windows this also builds the TensorAgent app when the .NET SDK running the build has the .NET MAUI workloads, and warns when it does not; add -p:TensorSharpSkipTensorAgentApp=true to leave the app out (details).
Or build just one application:
# Console application
dotnet build TensorSharp.Cli/TensorSharp.Cli.csproj -c Release
# Web application
dotnet build TensorSharp.Server.Host/TensorSharp.Server.Host.csproj -c Release
The CLI binary lands in TensorSharp.Cli/bin/... and the server in TensorSharp.Server.Host/bin/.... For a CUDA- or Vulkan-enabled native build, manual native builds, or MLX, see Building the native libraries.
3 Β· Download a model
For the quick start, download the recommended benchmark-verified gemma-4-E4B-it-Q8_0.gguf file (7.48 GiB) from the public ggml-org/gemma-4-E4B-it-GGUF repository. The lower-memory gemma-4-E4B-it-Q4_K_M.gguf lives in the same repository, and the Models page lists other options.
curl --create-dirs --fail -L "https://huggingface.co/ggml-org/gemma-4-E4B-it-GGUF/resolve/main/gemma-4-E4B-it-Q8_0.gguf?download=true" -o models/gemma-4-E4B-it-Q8_0.gguf
For multimodal models, download the matching projector (mmproj) and pass its exact path with --mmproj. The server does not auto-detect a projector; the CLI looks beside the model only when an image, audio or video input is given, and only for each architecture's known projector file names. Explicit paths are reliable.
Generating images or video? Pick the fast path on the way in. Each of these families has one, and the difference is not subtle:
- MiniMax-H3 video + audio β the one family here that comes back with a soundtrack: the picture and native 32 kHz stereo audio are denoised together in a single packed latent, and a run writes an
.mp4plus a sidecar.wav. It is four files across two repositories β the FL2VA denoiser (10.64 GiB), the shared Qwen3-VL-32B text encoder (16.97 GiB), the video VAE (5.21 GB) and the audio VAE (0.61 GB), β35.5 GB in all β or one--config config/minimax-h3-fl2va.json, which downloads all four on first run. The checkpoint ships CFG-distilled, so--cfg 1.0is required (TensorSharp refuses anything higher) and 4β8 steps against the 20-step default is the fast operating point. The one gap auto-download cannot fill: the text-encoder GGUF carries no tokenizer, so fetchvocab.jsonandmerges.txtfrom MiniMaxAI/MiniMax-H3 yourself and drop them beside the encoder (or pointTS_VIDEO_TOKENIZERat their folder). - Wan video, video only β a base
Wan2.2-TI2V-5Bfollows the official 50-step Γ 2-CFG recipe = 100 DiT passes; a Turbo / Lightning / FastWan checkpoint is trained guidance-free and costs 4. TensorSharp detects it from the DiT file name and needs no extra flag. On an M5 Pro a 1088Γ832 Γ 121-frame image-to-video takes β3 h 30 m on the base checkpoint and 17 m 30 s on Turbo β the same command, a different--modelpath. - Qwen-Image-2.1 β its CFG 1 default already means one transformer pass per step, so the levers are step count and size. The biggest is a step-distillation LoRA plug-in from
config/lora/:--lora config/lora/qwen-image-2.1-viggle-turbo.jsondownloads its weights on first use and brings its own recipe, 6 transformer passes instead of 40 (4β8 supported); Pruna 8-step and 5-step and a Fun-Acc 4-step preset ship beside it. All four adapters are under the Qwen Research License (non-commercial). Then size: for drafts use--width 1024 --height 1024(a quarter of the 2048Γ2048 default's latent image tokens). Without a distillation LoRA,--diffusion-steps 25is the plain step cut.
β Model downloads for the exact repos and files Β· the measurements behind these numbers Β· Qwen-Image-2.1 LoRA timings
Gemma 4 E4B backend and platform notes
The first-run commands below are for Linux + NVIDIA. On Windows + NVIDIA, set $env:TENSORSHARP_GGML_NATIVE_ENABLE_CUDA='ON' in PowerShell, then run the same commands without the Bash assignment prefix. On Apple Silicon, omit the CUDA environment assignment and use ggml_metal. For Vulkan on Windows/Linux, set TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=ON (PowerShell: $env:TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN='ON') and use ggml_vulkan. Without a GPU, use ggml_cpu (native CPU kernels). Text needs no projector. For image, video, or audio, download mmproj-gemma-4-E4B-it-Q8_0.gguf from the same repository and add --mmproj models/mmproj-gemma-4-E4B-it-Q8_0.gguf.
4 Β· First run
The one-shot prompt comes from a file via --input; --prompt belongs to the generative families β the image prompt for Qwen-Image-2.1, and the video prompt for MiniMax-H3 and Wan. The CLI defaults to ggml_cpu, greedy decoding, and 100 generated tokens. The environment assignment enables the CUDA native build; the first run compiles it and can take several minutes, later runs are incremental.
Option A β one-shot generation (CLI)
echo "What is TensorSharp? Answer briefly." > prompt.txt
TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON dotnet run --project TensorSharp.Cli -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --input prompt.txt --max-tokens 128 --backend ggml_cuda
Option B β interactive chat (REPL)
dotnet run --project TensorSharp.Cli -c Release --no-build -- --model models/gemma-4-E4B-it-Q8_0.gguf -i --max-tokens 128 --backend ggml_cuda
Type messages turn-by-turn; drive the session with slash commands like /reset, /think on, or /image photo.png.
Option C β browser UI + HTTP APIs (server)
TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON dotnet run --project TensorSharp.Server.Host -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --backend ggml_cuda --max-tokens 128
# open http://localhost:5000/index.html
The health response is at http://localhost:5000/; the browser chat is at /index.html. The Ollama- and OpenAI-compatible endpoints use the same fixed port.
Skip the long command lines. Put your options in a JSON file and pass --config config/server-basic.json (or config/cli-basic.json). A config entry can even auto-download the model on first run, so a fresh machine needs no manual download step. The repository's config/ folder ships ready-to-run examples. β Configuration file
Where to go next
Pick a backend
Match --backend to your hardware.
CLI reference
All flags, the REPL, and batch workflows.
HTTP API
Call the server from curl, Python, or SDKs.
Models
What's supported and where to download.
Video and audio
MiniMax-H3 writes a clip and its 32 kHz stereo track in one pass; Wan 2.1 / 2.2 generate video alone.
Benchmarks
Measured speed per model, and the levers that move it.