Supported Models

TensorSharp loads models in GGUF format and auto-detects the architecture from the file's general.architecture metadata — except MiniMax-H3 and metadata-free Qwen-Image-2.1 GGUFs (such as Unsloth's Q8_0), which carry no metadata at all and are identified by their tensors instead. Pick a quantization that fits your hardware (Q4_K_M for low memory, Q8_0 for higher quality).

📘

Gemma 4 E4B is the example model in From Tensors to Tokens. Read the book for a guided build from tensor fundamentals through multimodal inference; use this reference for current downloads and capability details. Explore the book →

Sentence encoders: Snowflake Arctic Embed L v2.0 / all-MiniLM-L6-v2, served through --embeddings and OpenAI/Ollama embedding APIs. Embedding models, supported backends, and validation →

Browse the model reference

The full reference is split across five pages. Start with downloads if you know which model you want, or with a category page if you are still choosing.

🔎

Embedding Models & Semantic Search

Snowflake and MiniLM: text/code vectors, APIs, and retrieval quality.

⬇️

Model Downloads

Every GGUF, with the projector, VAE, encoder and drafter files each family needs.

🧠

Text & LLM Models

Twelve architectures, their run recipes, plus multimodal input, thinking and tool calling.

🖼️

Image Generation

Qwen-Image-2.1 text-to-image and image editing, its defaults and draft settings.

🎬

Video Generation

MiniMax-H3 with native stereo audio, and Wan 2.1 / 2.2 with the distilled checkpoint that turns hours into minutes.

Supported architectures

ArchitectureGGUF arch keysExample modelsMultimodalThinkingToolsMTP spec
DeepSeek V4.1 Flashdeepseek41DeepSeek-V4.1-Flash (40 layers, 384 routed experts at top-6 plus one shared expert, four residual streams with delayed hyper-connection mixing, Engram n-gram features on layers 1 and 14, 1M declared context). Served on ggml_cuda, with cuda and ggml_cpu as portability and correctness paths and cpu the 100% pure-C# executor — no ggml, no native library and no GPU, the whole V4.1 graph held to an independent PyTorch oracle at atol=rtol=2e-5 plus greedy-argmax agreement (fixture-scale — a five-layer, 256-hidden synthetic model, so architectural agreement with the oracle rather than CPU/CUDA parity on the real 246 GiB Q2_K weights), a correctness and portability path rather than a serving one; current Q2_K and Q4_K_M GGUF releases embed Engram metadata and need no Engram preparationImage and video with the separately prepared vision companion (--mmproj); audio is refused, not ignored. The companion is a native ggml component, so it is unavailable on the pure-C# cpu backend, which is text-onlyYesYes (spaced DSML, grammar-constrained)Experimental — loads a deepseek41-dspark drafter via --draft-model on ggml_cuda / ggml_cpu; initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified. V4 drafters are rejected
DeepSeek V4 Flashdeepseek4DeepSeek-V4-Flash (284B MoE, 256 experts, compressed sparse attention, 1M context)Text onlyYesYes (DSML)Yes (DSpark block drafter, separate GGUF)
GLM 5.xglm-dsa, glm5nextGLM-5.2 (744B-A40B MoE, 256 experts, MLA + DeepSeek Sparse Attention, 1M context); GLM-5.3 (the same 79-block glm-dsa shape as 5.2 — 78 trunk layers plus one NextN, 256 routed experts at top-8 with one shared expert, MLA with the lightning indexer, rope base 8e6, 1M declared context capped to device capacity unless MAX_CONTEXT names the number — so it loads on the GLM-5.2 path with no new code and no new flag; at UD-Q2_K_XL it is seven shards, 236.4 GiB, in unsloth/GLM-5.3-GGUF, with --model pointed at the -00001-of- shard); GLM-5.3-Flash (320B, 288 routed experts, KDA linear attention on 34 of 45 trunk layers + NoPE MLA/DSA with a pooled indexer on the other 11)Image on 5.3-Flash only; 5.2 and 5.3 are text-only — the GLM-5.3 repo publishes no mmproj at any quant, and on glm-dsa an --mmproj is warned about and ignored rather than refused, so the run is text-only instead of a failureYesYes (XML tool calls)Yes on glm-dsa (GLM-5.2 and GLM-5.3 both embed a complete NextN block at blk.78, so there is nothing extra to download; --spec has to be on the command line before load, and it engages on single-device or explicit --layer-split N placement — under --tp N>1 the loader prints a line on stderr and serves standard decode, because that block ships no LM head of its own and the trunk head it borrows is column-parallel). Not implemented on glm5next (GLM-5.3-Flash)
Gemma 4gemma4gemma-4-E4B, 12B, 31B, 26B-A4B (MoE)Image, Video, AudioYesYesYes (separate draft)
Qwen 3.5 / 3.6qwen35, qwen35moe, qwen3nextQwen3.5-9B, Qwen3.5/3.6-35B-A3B (MoE), Qwen3.8-27B (dense)ImageYesYesYes on 3.6 and Qwen3.8-27B (embedded NextN — only in GGUFs that retain the NextN block, e.g. the -MTP- repos); Qwen3.8-27B also takes a DFlash2 block drafter, a separate GGUF via --draft-model
Qwen 3.8 Flash Nextqwen4expQwen3.8-Flash-Next (hybrid MoE: GatedDeltaNet recurrent layers interleaved with full attention, some behind Qwen Sparse Attention's indexer, a PLE n-gram embedding block, ×4 hyper-connection streams, 512 experts / 10 used)Image, video (video_url)Yes (switchable)Yes (Qwen XML / JSON tool calls)Yes (separate shared MTP head GGUF via --draft-model, GGML backends)
GPT OSSgptoss, gpt-ossgpt-oss-20b (MoE)Text onlyYes (always)Yes—
Nemotron-Hnemotron_h, nemotron_h_moe, nemotron_h_omniNemotron-H-8B, 47B, Nemotron 3 Nano Omni, Nemotron 3.5 Lightning 30B-A3BImage (Omni); audio only with a separately prepared Parakeet companion GGUF, otherwise refusedYesYesNo (every speculator is refused: the verify and decode kernels differ, so speculation would change the output)
Hunyuan Densehunyuan-denseTencent dense Hunyuan decoders such as the Hy-MT2 releases — GQA whose per-head QK-norm runs after NeoX RoPE, then SwiGLU. Single deviceText onlyNoNo—
Mistral 3mistral3; also llama-labelled Mistral Small 3.x files (Tekken tokenizer, [INST] / [SYSTEM_PROMPT] tokens)Mistral-Small-3.1-24B-InstructImageNoNo—
Muse-Glimmermuse-glimmer, muse_glimmerMuse-Glimmer-30B (interleaved SWA + NoPE, attention output gate)ImageYesYes (ATEM)Yes (DFlash block drafter, separate GGUF)
Bonsai2 27Bqwen35 plus prism.hadamard.* metadata and PQ2_0 / PTQ1_0 tensorsTernary-Bonsai-2-27B-PQ2_0.gguf / -PTQ1_0.gguf (dense Qwen 3.5 hybrid, 48 GatedDeltaNet + 16 full-attention layers). Weights are transcoded losslessly to GGML Q2_0 at load and PRISM's signed Hadamard transforms are applied; single-device GGML backends only (cpu, cuda, mlx and --tp are refused). Validated on Metal (M5 Pro) onlyImage with a companion mmprojYesYes—
DiffusionGemmadiffusion-gemma, diffusion_gemmadiffusion-gemma text-diffusion GGUFs; also serves the Jev typed-decision endpoint POST /v1/systemoneImage (chat and Jev decisions), through the Gemma 4 vision tower loaded from an mmproj GGUF or the upstream model-00011-of-00011.safetensors shard; audio and video are refusedNo (not prompted; a spontaneous thought block is returned as reasoning only with think: true)No (requests with tools are refused)—
Qwen-Image-2.1qwen_image, qwen-image (metadata-free GGUFs are recognised from their tensors)Qwen-Image-2.1 DiT (+ dedicated 2.1 VAE & Qwen3-VL-8B companions)Text → image and image editing (RGBA); LoRA plug-insNoNo—
MiniMax-H3minimax-h3, minimax_h3MiniMax-H3 FL2VA / Ref2VA denoisers (+ Qwen3-VL-32B text encoder, video VAE & audio VAE companions). The published GGUFs carry no metadata, so they are recognised by their tensorsVideo + 32 kHz stereo audio out (text → video, image → video, first/last frame, reference → video)NoNo—
Wan videowan, wan2.1, wan2.2Wan 2.1 T2V 1.3B/14B, Wan 2.2 TI2V-5B, Wan 2.2 A14B T2V/I2V (two 14B experts) (+ UMT5-XXL & video-VAE companions)Video out (text → video, image → video)NoNo—

Detailed per-model architecture cards (forward graph, components, parameters, and how TensorSharp optimizes prefill/decode) live under docs/models/ in the repository.

Which one is fastest — the lever that matters per family

Every family has one setting that dominates its wall-clock time. The two video families are the ones people miss most often, and they miss them differently: MiniMax-H3 arrives already CFG-distilled, so the lever is how many steps and how many pixels you ask for rather than which file you fetch — while on Wan, the checkpoint you download decides whether a 5-second 720p video takes three and a half hours or seventeen minutes, and it needs no flag at all.

FamilyFast laneMeasured effect
MiniMax-H3It is already CFG-distilled — there is no faster file to fetch. Run --cfg 1.0 at 4–8 steps against the 20-step default, and spend what you save on --width / --heightM5 Pro / ggml_metal, 22 frames, 8 steps, identical seed, against stable-diffusion.cpp: 256×256 49.3 s → 20.9 s (2.4× faster), 640×384 108.5 s → 63.1 s (1.7×). Steps buy quality, not just time — the same 640×384 image-to-video request is 79 s at 8 steps with coloured fringing around moving hands, and 212 s at 24 steps, clean
Wan videoLoad a step-distilled checkpoint — a DiT file name containing Turbo, distill, Lightning, lightx2v, FastWan, -dmd or …-4steps-… is auto-detected. No flag100 DiT passes → 4, guidance off. The same 1088×832×121f image-to-video request on M5 Pro / ggml_metal: ≈ 3 h 30 m → 17 m 30 s
Wan video--cfg-cache-stride 2 / 3 — base (non-distilled) checkpoints only1.30× / 1.43× at 50 steps (77 / 70 of the 100 passes). Approximate; pointless on a distilled checkpoint, which is already guidance-free
Qwen-Image-2.1Load a step-distillation LoRA plug-in (--lora config/lora/qwen-image-2.1-viggle-turbo.json and the other shipped 4–8-step recipes), and draft at --width 1024 --height 1024 — its CFG 1 default already runs one transformer prediction per stepA 1K square has a quarter of the 2K default's latent image tokens. M5 Pro (ggml_metal, Q4_K_M): 1024×1024 at the 40-step default in 327.6 s wall (330.5 s before the prefix KV cache; generation gains little from it); 1024×1024 at Viggle Turbo's 6 steps in 54.5 s wall (7.94 s/step). An earlier run took 1548.7 s for 2048×2048 at 25 steps. Editing gains most from the default-on prefix KV cache: 1.73–2.91× per step at 1024² on the M5 Pro
DeepSeek V4DSpark block drafter via --draft-model (cuda / ggml_cuda only)4×A40, 200 greedy tokens: decode 26.0 → 34.0 tok/s (cuda) and 26.4 → 37.1 tok/s (ggml_cuda) at 69% acceptance; 1.5–2.0× across a 5-turn chat. Output is unchanged
Muse-GlimmerDFlash block drafter via --draft-model (pass no sampler flags — it needs pure greedy)RTX PRO 6000, 128 greedy tokens: 35.0 → 50.9 tok/s at a 60-token prompt and 33.5 → 43.5 at 2 050. On Apple Silicon it does not pay today (20.7 → 13.9 tok/s) — run plain decode there
Qwen 3.6 / Gemma 4--spec on the CLI or the server — Qwen 3.6 uses the NextN block embedded in an -MTP- GGUF, Gemma 4 a separate --draft-model draft (naming --draft-model turns speculation on by itself)Engages on solo (non-concurrent) sequences, and only where profitable: Qwen 3.6 reports its embedded NextN block profitable on every backend, while Gemma 4's separate draft head engages only on the GGML backends (ggml_cpu included) and the direct cuda backend — on the pure-C# cpu backend and MLX Gemma 4 serves standard decode. Tune with --spec-draft (8) and --spec-pmin (0.15)
DiffusionGemma and Qwen-Image-2.1 on a CPU--backend cpu rather than ggml_cpu — the exception to the next rowi7-11800H, 8 cores: a DiffusionGemma-26B-A4B Jev read of a new 54-token prompt 0.92 s vs 1.57 s on ggml_cpu, and 0.22–0.25 s for further reads of the same prompt (cpu caches the prompt K/V, ggml_cpu does not); Qwen-Image-2.1 512×512 with the Pruna 5-step LoRA 118 s vs 295.7 s
Every text familyPick the right --backend: ggml_cuda on NVIDIA, ggml_metal on Apple Silicon, ggml_cpu (not cpu) on CPU — except DiffusionGemma, aboveRTX 3080 Laptop, gemma-4-26B-A4B QAT: ggml_cuda decode 78.7 vs the direct cuda backend's 35.3 tok/s, prefill 1832 vs 128. Apple Silicon, Muse-Glimmer-30B: ggml_metal prefill 413.6 vs mlx's 29.0
Any MoE that does not fit--n-cpu-moe N / --cpu-moe — keeps the first N layers' routed experts in system RAMTrades decode for VRAM, but wins outright when it stops a spill: gpt-oss-20b 16.2 → 2.9 GB on a 16 GB card, and --n-cpu-moe 12 turns 0.3 tok/s of WDDM paging into 25.4
GLM 5.xTS_GLM_UBATCH=2048 — a bigger prefill micro-batch, VRAM permitting3× RTX PRO 6000, GLM-5.2 UD-IQ2_XXS: pp2048 918.9 → 1145.8 t/s and pp4096 864.7 → 1048.7 t/s; decode is memory-bound and unchanged at ~43.9 tok/s. With 256 experts at top-8, a small chunk leaves most of every expert-GEMM tile as padding
Multiple GPUs--tp N (env TENSORSHARP_TP_DEGREE) on the direct cuda backend and on GGML CUDA / VulkanMuse-Glimmer-30B UD-IQ2_XXS on 2× RTX PRO 4000: prefill 1171 → 1569 (1.34×), decode 40.2 → 63.2 tok/s (1.57×). It also runs models that fit on no single card at all — but only where the interconnect can pay for it: on a PCIe host with no NVLink the per-layer all-reduces can cost more than the split saves, and GLM-5.2 on 3× RTX PRO 6000 drops from pp2048 915.9 / tg64 43.9 to 505.6 / 17.6 under --tp 3. GLM 5.x native TP is local/single-process for GLM-5.2, GLM-5.3 and GLM-5.3-Flash alike — --tp-node-id / --tp-peers are refused for the whole family before the model is built. On GLM-5.3 nothing above --tp 1 has ever been run: the mode is accepted, but the MLA and indexer caches are replicated per rank, so the cache footprint multiplies by N and the context that fits gets shorter, and the only arithmetic on record is a non-fit — --tp 8 wants 41.7 GiB per rank against 46 GB A40s

Tool support requires both a prompt renderer that declares the tools and a matching structured-output parser. Qwen 3.8 Flash Next (qwen4exp) returns structured calls through the Qwen 3.5 parser, so caller-defined tools, Agent Skills tools, code-execution tools and, on the server, sub-agent delegation are all offered to it. Mistral 3 and Hunyuan Dense render neither tool declarations nor tool results, and DiffusionGemma renders no declarations and refuses requests that carry tools, so none of those tools, and no sub-agent delegation, is offered to these three families.

MiniMax-H3, Wan, DiffusionGemma and Hunyuan Dense reject multi-GPU modes. Qwen 3.8 Flash Next supports local eligible-quantization GGML CUDA TP as well as --layer-split N; DeepSeek V4 uses whole-layer placement. GLM supports both explicit local modes; Qwen-Image supports local tensor parallelism of its diffusion transformer only. DeepSeek V4.1 supports --layer-split N and experimental routed-MoE --tp N.