Using TensorSharp from C#
TensorSharp is a real .NET library, not just an executable. Reference its projects from a source checkout and drive inference directly from your own code when you do not want an HTTP hop.
Sentence encoders: Snowflake Arctic Embed L v2.0 / all-MiniLM-L6-v2, with pure C# CPU and native GGML CPU/Metal/CUDA execution, served through --embeddings and OpenAI/Ollama embedding APIs. Embedding models, supported backends, and validation →
Project and package boundaries
The repository is split along publishable package boundaries. The source tree's publish set is thirteen packages — every row below — verified by eng/verify-packages.ps1, which gates the release workflow.
Package versions follow their release tag. Work merged after v2026.09.01 — including Qwen-Image-2.1, mask editing and LoRA plug-ins, the Jev endpoint, DiffusionGemma image input and sub-agents — requires a current source build. Use project references when you need those APIs.
| Project / package | Namespace | Responsibility |
|---|---|---|
TensorSharp.Core / TensorSharp.Tensors | TensorSharp | Tensor primitives, ops, allocators, storage, device abstraction. The NuGet ID differs from the project name. |
TensorSharp.Runtime.Logging | TensorSharp.Runtime.Logging | Logging abstractions and sinks shared by every host, kept out of the engine layers. |
TensorSharp.Runtime | TensorSharp.Runtime | GGUF parsing, tokenizers, prompt rendering, sampling, paged KV cache, continuous-batching scheduler. |
TensorSharp.Models | TensorSharp.Models | ModelBase, architecture implementations, multimodal encoders, batched/paged forward passes. |
TensorSharp.Backends.GGML | TensorSharp.GGML | GGML-backed execution and native interop. |
TensorSharp.Backends.Cuda | TensorSharp.Cuda | Direct CUDA allocator, storage, cuBLAS GEMM, PTX kernels, quantized CUDA ops. |
TensorSharp.Backends.MLX | TensorSharp.MLX | Apple-Silicon MLX backend (mlx-c / Metal). |
TensorSharp.Chat | TensorSharp.Chat | Host-neutral chat pipeline — ModelService, sessions, generation, the skills loop, the Web UI request/stream contract. No ASP.NET Core; shared by the server, the CLI, and the iOS app. |
TensorSharp.Server | TensorSharp.Server | ASP.NET Core library: OpenAI/Ollama adapters, HTTP transport over TensorSharp.Chat, web UI assets. A library — it produces no executable on its own. |
TensorSharp.Server.Host | TensorSharp.Server.Host | The runnable web application built on that library: Program.cs, host wiring, wwwroot/, and the command line. This is what you run. |
TensorSharp.Cli | TensorSharp.Cli | Console host and debugging / batch tooling. |
TensorSharp.AgentHost | TensorSharp.AgentHost.* | Skills, bounded agent loop, sub-agent coordination, four code tools, sandboxes, workspaces, and artifacts. |
TensorSharp.Distributed | TensorSharp.Distributed | Peer-to-peer TCP distributed tensor parallelism. |
For a typical embedding scenario, reference TensorSharp.Models.csproj; it already references Core, Runtime, and the backend projects. From a sibling application beside the TensorSharp checkout:
dotnet add reference ../TensorSharp/TensorSharp.Models/TensorSharp.Models.csproj
The NuGet packages are a snapshot of the v2026.09.01 release; use project references when you need the current implementation. Source builds compile native GGML/MLX by default; for managed-CPU development pass -p:TensorSharpSkipGgmlNative=true -p:TensorSharpSkipMlxNative=true, or build the native libraries described under Backends.
Minimal example — load, generate, decode
Every model is loaded the same way: ModelBase.Create() reads the GGUF metadata and instantiates the right architecture. From there you tokenize, run forward passes, sample, and decode.
using System;
using System.Collections.Generic;
using System.Linq;
using TensorSharp.Models;
using TensorSharp.Runtime;
// 1. Load any supported GGUF — the architecture is detected from metadata.
// Pick the BackendType your build supports: GgmlCuda, GgmlMetal, GgmlVulkan, or GgmlCpu.
using var model = ModelBase.Create("gemma-4-E4B-it-Q8_0.gguf", BackendType.GgmlCuda);
// 2. Configure sampling (defaults match Ollama: temp 0.8, top_k 40, top_p 0.9).
var sampling = new SamplingConfig { Temperature = 0.7f, TopP = 0.9f, TopK = 40 };
// 3. Tokenize the prompt.
var tokens = model.Tokenizer
.Encode("Explain mixture-of-experts in one sentence.", addSpecial: true)
.ToList();
var generated = new List<int>();
// 4. Prefill the complete prompt once.
float[] logits = model.Forward(tokens.ToArray());
// 5. Decode one new token at a time. Forward() owns the KV-cache position,
// so after prefill pass only the newly sampled token, not the full history.
for (int step = 0; step < 200; step++)
{
int next = model.Sample(logits, sampling, generated); // applies penalties + sampling
if (model.Tokenizer.IsEos(next)) break;
generated.Add(next);
logits = model.Forward(new[] { next });
}
// 6. Detokenize the result.
Console.WriteLine(model.Tokenizer.Decode(generated));
For greedy/deterministic decoding, call model.SampleGreedy(logits) instead of Sample.
Smoke-test variant
A one-shot sanity check that loads a model, runs a single forward pass, and prints the top token:
using var model = ModelBase.Create(modelPath, backend);
var tokenIds = model.Tokenizer.Encode("Hello", addSpecial: true);
float[] logits = model.Forward(tokenIds.ToArray());
int topToken = model.SampleGreedy(logits);
Console.WriteLine($"vocab={model.Config.VocabSize}, tokens={tokenIds.Count}, topToken={topToken}");
DiffusionGemma — text diffusion
DiffusionGemma is a block text-diffusion model, not an autoregressive one. ModelBase.Create() still loads it (the GGUF architecture is diffusion-gemma / diffusion_gemma) and returns a DiffusionGemmaModel, but its Forward(int[]) intentionally throws — generation runs through DiffusionGemmaSampler, which iteratively denoises fixed-length canvas blocks over a [prompt | canvas] sequence. See the DiffusionGemma model card for the architecture.
using System;
using System.Collections.Generic;
using System.Linq;
using TensorSharp.Models; // DiffusionGemmaModel, DiffusionGemmaSampler, DiffusionEbParams
using TensorSharp.Runtime;
// Load a diffusion-gemma GGUF — the architecture is auto-detected.
using var model = (DiffusionGemmaModel)ModelBase.Create("diffusion-gemma.gguf", BackendType.GgmlCuda);
// Render the prompt with the model's chat template, then tokenize.
var messages = new List<ChatMessage> { new() { Role = "user", Content = "Write a haiku about winter." } };
var renderer = new GgufPromptRenderer();
string rendered = renderer.Render(model.Config.ChatTemplate, messages,
addGenerationPrompt: true, architecture: model.Config.Architecture);
int[] promptTokens = model.Tokenizer.Encode(rendered, addSpecial: true).ToArray();
// Configure the EntropyBound denoising sampler.
var p = new DiffusionEbParams
{
MaxDenoisingSteps = 48, // refinement steps per canvas block
Seed = 0, // deterministic
MaxBlocks = 1, // block-autoregressive canvas blocks
};
var sampler = new DiffusionGemmaSampler(model);
// Generate. The optional callback fires after every denoising step with
// (blockIndex, step, totalSteps, previewTokens) — useful for a live UI.
List<int> generated = sampler.Generate(promptTokens, p,
(block, step, total, preview) => Console.Write($"\rblock {block + 1} step {step + 1}/{total} "));
Console.WriteLine();
Console.WriteLine(model.Tokenizer.Decode(generated));
Key parameters on DiffusionEbParams: MaxDenoisingSteps (48), TMin/TMax temperature schedule (0.4 / 0.8), EntropyBound (0.1), StabilityThreshold / ConfidenceThreshold early-stop, Seed, and MaxBlocks. model.CanvasLength reports the per-block canvas size.
Image input and typed decisions. model.LoadVisionEncoder(path) loads the Gemma 4 vision tower from an mmproj GGUF or straight from the Hugging Face shard model-00011-of-00011.safetensors; audio and video are refused. With a tower loaded, set ImagePaths on the user message and pass the tokenized prompt through model.MultimodalInjector.ProcessPromptTokens(messages, tokens) and QueuePromptEmbeddings(0) before Generate, as the CLI does. For Jev-style typed decisions, reference TensorSharp.Chat and call ModelService.JevAsync(JevRequest.Parse(json.RootElement)) (namespaces TensorSharp.Server and TensorSharp.Server.Jev) — the validated path behind POST /v1/systemone. The low-level read is DiffusionGemmaModel.ReadStructured(promptTokens, seedCanvas, positions, tokenIds), called while holding GpuComputeLock. See docs/models/jev.md.
On the GPU backends and the pure-C# cpu backend the prompt K/V is cached once per block and reused across denoising steps; GGML backends default to a fused whole-model decode plus a fused lm-head tail. Tunables (DIFFUSION_STEPS, DIFFUSION_NO_SC, …) are in the Advanced page.
Qwen-Image-2.1 — image generation & editing
Qwen-Image-2.1 turns a prompt into an image, or a prompt + one or more reference images into an edited image. The loaded qwen_image GGUF is the diffusion transformer; the model also resolves its companions next to the DiT file (or via the TS_QWEN_IMAGE_VAE / TS_QWEN_IMAGE_TE / TS_QWEN_IMAGE_MMPROJ environment variables): the dedicated 2.1 VAE, the Qwen3-VL-8B text encoder, and its mmproj vision encoder, which editing requires. Like DiffusionGemma it is not an autoregressive text model — the autoregressive entry points throw, and images come from GenerateImage() and EditImage().
using TensorSharp.Models;
using TensorSharp.Models.QwenImage;
using TensorSharp.Runtime;
// Load the Qwen-Image-2.1 DiT GGUF (architecture = qwen_image). The 2.1 VAE and the
// Qwen3-VL-8B text encoder + mmproj are resolved from the same directory (or the
// TS_QWEN_IMAGE_* env vars). Any GGML backend works, and so does BackendType.Cpu (pure C#).
using var model = (QwenImageModel)ModelBase.Create("qwen_image_2.1_Q4_K_M.gguf", BackendType.GgmlMetal);
var p = new QwenImageParams
{
Steps = 0, // 0 = model default: 40 FlowMatch-Euler steps
CfgScale = 0f, // 0 = model default: 1.0, one transformer prediction per step
Seed = 42,
// Width = 1024, Height = 1024, // explicit size in multiples of 32 (faster drafts)
};
// Text-to-image (default 2048x2048)
RgbImage generated = model.GenerateImage("A small orange cat beside a blue ceramic vase, soft daylight", p);
ImageIO.SavePng("generated.png", generated);
// Editing: one reference here; pass an IReadOnlyList<RgbImage> for several, in order
RgbImage input = ImageIO.Load("generated.png", preserveAlpha: true);
RgbImage edited = model.EditImage("Change the blue vase to a red vase.", input, p);
ImageIO.SavePng("edited.png", edited);
Console.WriteLine($"Saved {edited.Width}x{edited.Height} edited image.");
For a live UI, set p.OnStep = (step, total, preview) => { … } and p.PreviewCount to receive decoded RGB snapshots of the partially denoised latent on throttled steps. RgbImage exposes Width, Height, and a planar/interleaved Pixels buffer; ImageIO provides Load, Decode(byte[]), EncodePng, SavePng, and resize helpers. PNG output preserves transparency.
LoRA plug-ins. model.SetLoras(new[] { new LoraSpec("config/lora/qwen-image-2.1-viggle-turbo.json") }) replaces the set applied to every later request; LoraSpec(Path, Scale, ConfigPath) takes the same three values as --lora / --lora-scale / --lora-config, and an empty list removes the set. The adapters are validated against the transformer immediately, and a failure leaves the previous set in place. A plug-in's sampling recipe supplies the steps and CFG whenever Steps / CfgScale are 0. See LoRA plug-ins.
Image generation is compute-heavy: the default is 2048×2048 at 40 steps. Omitting Width/Height selects 2048×2048 for generation, or about that area at the first reference's aspect ratio for editing; Width = 1024, Height = 1024 quarters the latent image tokens for drafts. The native path runs a GGML graph on GgmlMetal, GgmlCuda, GgmlVulkan or GgmlCpu. BackendType.Cpu runs the whole pipeline in managed C#, defaults to a 1024×1024 area and checks estimated memory before inference. See Image Generation for the companion files.
Local edits: set p.Mask = ImageIO.Load("selection.png", preserveAlpha: true), then MaskMode, MaskInvert, MaskFeather, MaskCrop and MaskCropPadding as needed before EditImage(). QwenImageMaskMode.Grayscale (default) uses white to edit; Alpha uses transparency. The mask matches the first input image, and the result retains its canvas and protected pixels. Mask semantics →
SamplingConfig
The sampling knobs map one-to-one to the CLI flags and API options. Defaults match Ollama.
| Property | Type | Default | Meaning |
|---|---|---|---|
Temperature | float | 0.8 | Randomness; 0 = greedy/deterministic. |
TopK | int | 40 | Limit to the top-K most probable tokens; 0 = disabled. |
TopP | float | 0.9 | Nucleus sampling; 1.0 = disabled. |
MinP | float | 0 | Minimum probability threshold relative to the max. |
RepetitionPenalty | float | 1.1 | Multiplicative penalty; >1 discourages repetition. |
PenaltyLastN | int | 64 | How many recent tokens the repetition / presence / frequency penalties consider (--repeat-last-n on both hosts; request field repeat_last_n). |
PresencePenalty | float | 0 | Additive penalty for tokens already present. |
FrequencyPenalty | float | 0 | Additive penalty proportional to token frequency. |
Seed | int | -1 | Reproducible sampling; -1 = time-based. |
StopSequences | List<string> | null | Stop when any of these strings is produced. |
MaxTokens | int | 0 | Maximum tokens to generate; 0 = use the caller's default. |
Agent Skills — SkillsChatClient
An Agent Skill is a folder of model-facing instructions — a SKILL.md plus the scripts, reference documents and assets it refers to — that a model loads only when a task needs it. SkillsChatClient (namespace TensorSharp.AgentHost.Skills, project TensorSharp.AgentHost) is how a .NET application gets them. It calls an OpenAI-compatible chat endpoint over HTTP rather than loading a GGUF in process, so unlike the rest of this page it needs no backend and no TensorSharp.Models reference.
On TensorSharp's tool-capable chat formats, selected and discovered skills begin as metadata only; the model reads instructions and bundled files on demand. If the model has no usable tool parser, TensorSharp instead inlines selected skill bodies and withholds the built-in skill/code tools, so the request remains useful without pretending a tool round trip will work.
It covers the two situations that actually arise, selected with SkillsChatClientOptions.Delivery.
Against the TensorSharp server (TensorSharp.Server.Host) — SkillDelivery.Server
Name the skills and the server does everything, including the progressive-disclosure loop, next to the model. Nothing is uploaded per request and the skill files never leave the server.
using System;
using System.Threading.Tasks;
using TensorSharp.AgentHost.Skills;
using var client = new SkillsChatClient(new SkillsChatClientOptions
{
Endpoint = "http://localhost:5000", // with or without a trailing /v1
DefaultModel = "gemma-4-E4B-it-Q8_0.gguf",
Delivery = SkillDelivery.Server,
});
// Convenience for the common single-turn case: prompt + the skills to use.
SkillsChatResponse reply = await client.CompleteAsync(
SkillsChatRequest.User("Extract the tables from report.pdf", "pdf"));
Console.WriteLine(reply.Content);
Console.WriteLine($"{reply.Rounds} round(s), {reply.PromptTokens} prompt + {reply.CompletionTokens} completion tokens");
Under server delivery the progressive-disclosure loop runs inside the server, so reply.SkillInvocations comes back empty and reply.Rounds is 1 — this client made exactly one request.
Against any other OpenAI-compatible endpoint — SkillDelivery.Local
Point the client at a local SkillRegistry and it builds the prompt block, declares the skill tools and runs the loop in this process, so an endpoint that has never heard of skills behaves as though it had. The cost is one extra round trip per file the model reads.
using System;
using System.Collections.Generic;
using System.Threading.Tasks;
using TensorSharp.Runtime; // ChatMessage
using TensorSharp.AgentHost.Skills;
var registry = new SkillRegistry(new SkillRegistryOptions
{
Roots = new[] { "/srv/skills" }, // scanned up to MaxDepth (3) levels deep
});
using var client = new SkillsChatClient(new SkillsChatClientOptions
{
Endpoint = "https://api.example.com/v1",
ApiKey = Environment.GetEnvironmentVariable("EXAMPLE_API_KEY"),
DefaultModel = "some-hosted-model",
Delivery = SkillDelivery.Local,
Registry = registry,
Discovery = true, // also advertise the skills a request did not select
});
SkillsChatResponse reply = await client.CompleteAsync(new SkillsChatRequest
{
Messages = { new ChatMessage { Role = "user", Content = "Fill in this AcroForm and tell me what you set." } },
Skills = { "pdf" },
MaxTokens = 800,
});
Console.WriteLine(reply.Content);
// Every skill file the model read, in order.
foreach (SkillToolInvocation call in reply.SkillInvocations)
Console.WriteLine($"round {call.Round}: {call.Tool} {call.SkillId}/{call.ResourcePath} ok={call.Ok}");
SkillDelivery.Auto — the default — probes /v1/skills once and caches the answer for the client's lifetime: server delivery if the endpoint answers, local delivery if it does not and this client has a registry. await client.ResolveDeliveryAsync() reports which one won, and await client.ListServerSkillsAsync() asks an endpoint what it has, returning an empty list rather than throwing when it implements no skills API.
SkillsChatResponse.ToolCalls holds calls to the caller's own tools (those passed in SkillsChatRequest.Tools), which the client never executes — only the client knows what they do. A non-empty list means you must service them and send the results back in a follow-up request; the skill tools have already been answered, and reply.Messages is the full transcript to continue from. A failed request throws SkillsChatException.
SkillsChatClientOptions
| Property | Type | Default | Meaning |
|---|---|---|---|
Endpoint | string | required | API root, with or without a trailing /v1. |
ApiKey | string? | null | Bearer token, when the endpoint wants one. TensorSharp.Server.Host does not. |
DefaultModel | string? | null | Model name sent with every request unless one is set per request. |
Delivery | SkillDelivery | Auto | Who resolves skills: Auto, Server, or Local. |
Registry | SkillRegistry? | null | The skills used under local delivery. Ignored under server delivery, where the endpoint owns them. |
PromptOptions | SkillPromptOptions | .Default | Prompt budgets for local delivery. |
LoopOptions | SkillAgentLoopOptions | .Default | Loop bounds for local delivery — MaxRounds 8, MaxCallsPerRound 8, and an OnInvocation callback fired after each executed tool call. |
MultiAgent | MultiAgentOptions | new() (enabled) | Sub-agent delegation under local delivery: the model may spawn child agents, each run as its own conversation with the same endpoint. LoopOptions.MultiAgent, when set, takes precedence. When the endpoint is detected as TensorSharp or as having a skills API, local delivery sends "multi_agent": false so the server does not delegate as well. |
Discovery | bool | true | Advertise skills the request did not select, so the model can pick one up. SkillsChatRequest.Discovery overrides it per request. |
Timeout | TimeSpan | 10 minutes | Per-request timeout. Only used when the client creates its own HttpClient — pass one from IHttpClientFactory to the constructor instead and it is left alone. |
SkillsChatRequest and SkillsChatResponse
| Member | Type | Meaning |
|---|---|---|
SkillsChatRequest.Messages | List<ChatMessage> | The conversation. A leading system message is merged with the skill block rather than displaced. |
SkillsChatRequest.Skills | List<string> | Skill names to use, as they appear in the registry. |
SkillsChatRequest.Tools | List<ToolFunction>? | The caller's own tools. Never executed by the client; returned for the caller to service. |
SkillsChatRequest.Model · MaxTokens · Temperature · TopP · Think · Discovery | string? · int? · double? · double? · bool · bool? | Per-request overrides; a null leaves the setting to the endpoint (or, for Discovery, to the client's own option). |
SkillsChatRequest.MultiAgent | bool? | false turns delegation off for this request — locally, or by sending "multi_agent": false under server delivery. Null follows the client's (or server's) policy. |
SkillsChatRequest.User(prompt, params skills) | static | Convenience for the common single-turn case. |
SkillsChatResponse.Content · Thinking | string · string? | The assistant's answer, and its reasoning when the model exposed any. |
SkillsChatResponse.ToolCalls | IReadOnlyList<ToolCall> | Calls to the caller's own tools, which the client never executes. |
SkillsChatResponse.Messages | List<ChatMessage> | The full transcript, including anything the disclosure loop appended. |
SkillsChatResponse.SkillInvocations | IReadOnlyList<SkillToolInvocation> | Every host tool call, in completion order, across the parent and any sub-agents — Round, Tool, SkillId, ResourcePath, Ok, ResultBytes, and AgentId (/root for the parent). Empty under server delivery. |
SkillsChatResponse.FinishReason · PromptTokens · CompletionTokens · Rounds | string? · int · int · int | Why generation stopped, tokens summed over every round, and how many generations ran. Rounds is 1 when the model answered without reading anything. |
SkillRegistry and SkillRegistryOptions
The registry owns discovery (walking configured roots for SKILL.md), installation, and lookup; it is the only component that touches skill storage, so the containment rules live in one place. It is also what the CLI and the server build from --skills-dir.
| Member | Type / default | Meaning |
|---|---|---|
SkillRegistryOptions.Roots | IReadOnlyList<string>, empty | Directories to scan, in precedence order. A root may be a single skill (it holds SKILL.md directly) or a directory of them. |
SkillRegistryOptions.InstallDirectory | string?, null | Where skills installed at runtime are written; also scanned, and always first in precedence. Null leaves the registry read-only. |
SkillRegistryOptions.MaxDepth | int, 3 | How deep to walk under a root looking for SKILL.md. |
SkillRegistryOptions.MaxSkills · MaxManifestBytes · MaxSkillBytes · MaxSkillFiles | 512 · 4 MB · 256 MB · 4096 | Ceilings so a root pointed at the wrong directory fails visibly instead of exhausting memory. |
Skills · Errors · Roots | IReadOnlyList<Skill> · IReadOnlyList<SkillLoadError> · IReadOnlyList<string> | What loaded, what failed to load (with Path and Message), and the roots that were scanned. |
CanInstall · InstallDirectory | bool · string? | Whether runtime installation is configured, and where it lands. |
TryGet(id, out Skill) · Resolve(ids, out unknown) | bool · IReadOnlyList<Skill> | Look one skill up by name, or resolve a selection and learn which names were not found. |
Refresh() | SkillScanResult | Re-scan every root — what POST /api/skills/rescan calls. |
InstallFromDirectory(path, overwrite) · InstallFromZip(stream, overwrite, limits) · Remove(id) | Skill · Skill · bool | Install a skill from a folder or an uploaded archive, or remove an installed one. Every archive entry is resolved through the same path guard the model's reads go through; a rejected archive throws SkillInstallException. |
Full reference, including the frontmatter fields, the prompt budget and the security model: docs/agent_skills.md. Open-source skills to start from: github.com/anthropics/skills.
AgentHost runtime layer
TensorSharp.AgentHost is the optional agentic layer over TensorSharp.Runtime; Runtime has no dependency back to it. CodeExecOptions, ShellRunner, and CodeRunnerAdapter implement the four host-owned tools read_file, write_file, shell, and apply_patch. File tools are declared only when the host supplies a persistent workspace; a direct caller without one gets shell alone.
SkillAgentLoop runs a bounded loop, executes only TensorSharp-owned skill/code calls, and returns calls to the application's own tools in SkillLoopResult.PendingClientToolCalls. Setting SkillAgentLoopOptions.MultiAgent together with SubagentGeneratorFactory — a Func<string, SkillTurnGenerator> that returns an independent generator for each child — lets the loop coordinate a request-scoped tree of sub-agents through spawn_agent, wait_agent, send_input, close_agent and list_agents; an enabled MultiAgent (its Enabled defaults to true) without a factory throws ArgumentException. MultiAgentOptions (in TensorSharp.AgentHost.Agents) bounds that tree: MaxConcurrentAgents 3, MaxAgents 8, MaxDepth 2, MaxRoundsPerAgent 8, MaxTotalChildGenerations 48, AgentTimeoutSeconds 180, MaxResultCharacters 8000, MaxTaskCharacters 16000, and AllowWorkerTools off. Callbacks carry SkillToolInvocation.AgentId. Hosts default to 8 generations, raised to 24 when code tools are offered unless an operator set the bound. SessionWorkspaceManager gives each server session a private workspace; request-only callers get one workspace for that request. Working files, allowed exported environment state, and installed packages persist in that scope, while PATH is rebuilt for every shell call.
Code tools are off by default and sandbox mode defaults to required. macOS Seatbelt cannot guarantee cleanup of a deliberately detached child; Linux needs bubblewrap 0.12+; Windows job objects cannot confine files or sockets, so execution requires explicit unconfined mode. These are startup permissions, not per-command approvals. Sub-agents are read-only unless AllowWorkerTools is set (and then only worker children gain the parent's mutable tools), never receive the application's own tools, and share one host-tool gate with their parent, so tool calls in a tree run one at a time. Structured-output requests inline selected skills and suppress built-in tools. See Agentic Work for the complete boundary.
Key types & interfaces
| Type | Role |
|---|---|
ModelBase | Abstract base for every architecture. Create(path, backend), Forward(int[]), Sample(...), SampleGreedy(...), plus Config and Tokenizer. |
BackendType | Enum: Cpu, GgmlCpu, GgmlMetal, GgmlCuda, GgmlVulkan, Cuda, Mlx. |
SamplingConfig | Sampling configuration (table above). |
ITokenizer | Encode(text, addSpecial), Decode(ids), IsEos(id), EosTokenIds (BPE & SentencePiece implementations). |
ModelConfig | Architecture metadata: VocabSize, context length, and more. |
IBatchedPagedModel | Optional batched/paged forward (ForwardBatch) implemented by most architectures for continuous batching. |
DiffusionGemmaModel + DiffusionGemmaSampler | Text-diffusion model and its EntropyBound denoising sampler (Generate(promptTokens, DiffusionEbParams, …)), with image input through LoadVisionEncoder and one-step Jev reads through ReadStructured. Forward() is unsupported. |
QwenImageModel + QwenImageParams | Qwen-Image-2.1 image generator and editor. GenerateImage(prompt, QwenImageParams) and EditImage(prompt, RgbImage or IReadOnlyList<RgbImage>, QwenImageParams) return an RgbImage; ImageIO loads/saves PNG. |
SkillsChatClient + SkillsChatClientOptions | Agent Skills against an OpenAI-compatible endpoint. CompleteAsync(SkillsChatRequest) returns a SkillsChatResponse; SkillDelivery chooses whether the server or this process runs the disclosure loop (in TensorSharp.AgentHost.Skills). |
MultiAgentOptions + MultiAgentSession | Sub-agent limits and the request-scoped agent tree behind spawn_agent / wait_agent / send_input / close_agent / list_agents (in TensorSharp.AgentHost.Agents). Used by SkillAgentLoop, SkillsChatClient and the server's chat paths. |
SkillRegistry + SkillRegistryOptions | The set of skills a host knows about: scanning configured roots for SKILL.md, installing from a folder or a .zip, and lookup. The only component that touches skill storage, so containment is enforced in one place. |
InferenceEngine | Worker-thread scheduler + paged block pool that powers the server's continuous batching (in TensorSharp.Runtime.Scheduling). |
Other runtime contracts worth knowing: IModelArchitecture, IPromptRenderer (implemented by GgufPromptRenderer), IOutputProtocolParser, and IMultimodalInjector.
For most applications the easiest integration is to run TensorSharp.Server.Host and call it over the OpenAI-compatible API — you keep your app process clean and get continuous batching for free. Reach for the library API when you need in-process control or custom decoding.