Features

A complete catalog of what TensorSharp does today. Each item links to the page where you can use it.

Sentence encoders: Snowflake Arctic Embed L v2.0 / all-MiniLM-L6-v2, with pure C# CPU and native GGML CPU/Metal/CUDA execution, served through --embeddings and OpenAI/Ollama embedding APIs. Embedding models, supported backends, and validation โ†’

Highlights

๐Ÿง 

Multi-architecture

DeepSeek V4 Flash, DeepSeek V4.1 Flash, GLM 5.2, GLM-5.3 & GLM-5.3-Flash, Gemma 4, Qwen 3.5 / 3.6 / 3.8 Flash Next, Bonsai2, GPT OSS, Nemotron-H, Mistral 3, Hunyuan Dense, Muse-Glimmer, DiffusionGemma, Qwen-Image-2.1, MiniMax-H3 audio+video, Wan video.

๐Ÿ–ผ๏ธ

Multimodal

Image, video, and audio inputs (Gemma 4); image input for several others.

๐Ÿ“„

PDF documents

Upload PDFs in the Web UI or pass --pdf on the CLI โ€” text PDFs are inlined, scanned pages go to vision models.

๐ŸŽจ

Image generation & editing

Qwen-Image-2.1 generates images, edits photos with references, and preserves pixels outside a painted selection. The Web UI and TensorAgent share the selection editor; LoRA plug-ins, prefix K/V caching and supported multi-GPU execution are available.

๐ŸŽฌ

Video generation

MiniMax-H3 generates video and native 32 kHz stereo audio together in one packed latent, CFG-free at 4โ€“8 steps. Wan 2.1 / 2.2 generate video alone, where a step-distilled checkpoint cuts a 5-second 720p clip from hours to minutes.

๐Ÿ’ญ

Thinking mode

Structured chain-of-thought, separated from the visible answer.

๐Ÿ› ๏ธ

Tool calling

Multi-turn function calling across all three API styles.

๐Ÿงฉ

Agent Skills

Folders of model-facing instructions that load only when a task needs them โ€” selected with --skill or "skills": ["pdf"], fetched through built-in tools TensorSharp answers itself.

๐Ÿค–

Agentic work

A bounded model/tool loop with four code tools, isolated workspaces, sandboxed commands, dependency installation, bounded sub-agents on the server, live progress, and downloadable artifacts.

๐Ÿ“ฆ

Native quantized compute

Q4_K_M, Q8_0, MXFP4, IQ2_XXS and more run in matmul without dequantizing to FP32.

๐Ÿ”€

Continuous batching

vLLM-style paged KV cache with cross-request prefix sharing.

โฉ

Speculative decoding

MTP / NextN draft heads, DeepSeek V4's DSpark and the DFlash block drafters (Muse-Glimmer, Qwen 3.5 family), and n-gram lookup on supported models accelerate solo decode.

๐Ÿ”—

Multi-GPU & multi-node

Tensor parallelism shards every layer across GPUs โ€” CUDA and GGML alike โ€” and across machines over a TCP mesh; the whole-model executors take a layer split instead.

๐Ÿ”Œ

Ollama & OpenAI APIs

Drop-in endpoints for existing tooling, plus a browser chat UI.

TensorAgent: local text, images, video, files, code, skills and desktop browser workflows in one app for iPhone, iPad, macOS and Windows. Eight interface languages; device memory and platform policy determine which workflows are available. Capabilities and validation limits โ†’

Models & modalities

Generation & control

Agent Skills

Performance & scale

Interfaces & integration