API Reference

Every flag, variable, endpoint, and public type in one place. Type in the box to filter all tables below instantly — or press / for wiki-wide search.

· Matching rows are highlighted live across all sections.

CLI flags — TensorSharp.Cli

FlagDescription
--model <path>Path to a GGUF model file (required).
--input <path>Text file containing the user prompt.
--input-jsonl <path>JSONL file with batch requests (one JSON per line).
--multi-turn-jsonl <path>JSONL for multi-turn chat simulation with KV-cache reuse.
--output <path>Write generated text to this file.
--image / --video / --audio <path>Media for vision / video / audio inference.
--pdf <path>PDF document input (one-shot): born-digital PDFs are inlined as text; scanned PDFs become page images for a vision model (page cap: TS_PDF_MAX_PAGES).
--mmproj <path|none>Multimodal projector: an mmproj GGUF or, for the Gemma 4 family (DiffusionGemma included), a Hugging Face .safetensors shard holding the vision tower. Pass it explicitly; when it is omitted and an image, audio or video is given, the CLI looks beside the model only for the architecture's known projector file names (a .safetensors shard is never auto-detected). none (any case) loads no projector and skips that lookup, as on the server.
--max-tokens <N>Maximum tokens to generate (default 100).
--backend <type>cpu, cuda, mlx, ggml_cpu, ggml_metal, ggml_cuda, ggml_vulkan. Default ggml_cpu on every OS.
--gpu-device <N>Vulkan device index for ggml_vulkan on multi-GPU hosts (default 0; env TS_GGML_VULKAN_DEVICE).
--list-gpusList visible Vulkan devices (index + adapter name) and exit.
--kv-cache-dtype <type>KV cache precision: f32, f16, q8_0, q4_0 (default: auto per backend/model; q4_0 ~1/7 of f32, for very long 128K–256K contexts; the quantized tiers need the native GGML flash-attention path).
--interactive / -i / --chatStart the interactive REPL.
--system <text> / --system-file <path>Seed the system prompt.
--thinkEnable thinking / reasoning mode.
--tools <path>JSON file with tool / function definitions.
--skills-dir <path> · --skill <name> · --list-skills · --no-skills · --skills-no-discoveryDiscover/select/list/disable Agent Skills or restrict discovery. Selected skills start as metadata for tool-capable models; bodies are read on demand. Without --skills-dir, the roots are every existing .agents/skills from the working directory up to the Git root (nearest first), then skills/ beside the binary. Env TS_SKILLS_DIR / TS_NO_SKILLS.
--skills-allow-exec · --skills-sandbox <off|preferred|required> · --skills-allow-networkEnable skills_run and choose script isolation/network policy. Execution and network are off by default; sandbox mode defaults to required. macOS uses Seatbelt, Linux needs bubblewrap 0.12+, and Windows job objects cannot confine files/network.
--skills-max-rounds <1..64>Internal model/tool round bound: default 8, or 24 when code execution is offered; an explicit value is kept.
--code-execEnable the four server-owned agentic tools: read_file, write_file, shell, apply_patch. Off by default; env TS_CODE_EXEC.
--code-exec-allow-install · --code-exec-packages <list>Allow host-performed pip/npm installs and optionally restrict package names. Does not grant generated commands network access.
--code-exec-install-domains <list> · --code-exec-install-index <url>Installer host allow-list (default pypi.org,files.pythonhosted.org,registry.npmjs.org) and operator-selected index.
--code-exec-allow-networkGive generated commands unrestricted host IP networking, including LAN/loopback and listening sockets. Off by default and independent of installs/skill networking.
--code-exec-timeout <seconds> · --code-exec-max-output <bytes> · --code-exec-shell <path|name> · --code-exec-temperature <0..2>Command deadline (default 120 s), middle-truncated output cap (32768 bytes), shell override, and optional coding-turn temperature override (only when the request temperature is still default). Coding turns always disable the built-in 1.1 repetition penalty while code tools can run.
--code-exec-unconfinedRun despite inadequate OS confinement. Required on Windows; unsafe for users not trusted with the host.
--code-exec-languagesRemoved and rejected at startup; no replacement. A shell can reach every interpreter on its rebuilt PATH, so the host reports installed interpreters instead of claiming a language allow-list.
--spec / --no-specEnable / disable speculative decoding: a drafter proposes the next few tokens and the trunk verifies them in one batched forward. Every emitted token is still drawn from a trunk row with the run's own sampler, so this is a speed path only. Must be passed before the model loads — it is what tells glm-dsa to page its ~3 GiB NextN layer into VRAM, which also leaves less room for the context. Not available under --tp N>1 on a checkpoint whose draft block borrows the trunk's LM head, which includes GLM-5.2 and GLM-5.3 — on those, speculation engages on single-device or explicit --layer-split N placement (no active tensor parallelism). The explicit opt-in for drafters embedded in the trunk checkpoint (GLM-5.2, GLM-5.3, Qwen 3.6 NextN); a drafter named with --draft-model engages without it. Default off. Env TS_SPEC.
--spec-type <name>Which speculation algorithm drafts: auto (default) uses whatever drafter the checkpoint carries, draft-head pins a per-token NextN/MTP head, block pins a block drafter, and ngram needs no trained weights at all — it drafts by finding where the last few tokens occurred earlier in the context and proposing what followed, and is strongest where the answer quotes the prompt (summarizing, editing, translating, repetitive structured output, agentic loops). It only picks the algorithm: pair it with --spec, because given without --spec or --draft-model (and with TS_SPEC unset) it — like --spec-draft and --spec-pmin — leaves speculation off, with a startup warning on both hosts. It still needs a model that can act as a speculative target: GPT-OSS, Mistral 3, Hunyuan Dense and DiffusionGemma have no speculative path, and Nemotron-H refuses every speculator. Some targets are conditional and otherwise serve plain decoding: DeepSeek V4 / V4.1 and Muse-Glimmer speculate only while their own drafter (DSpark / DFlash) is loaded, Gemma 4 only on the GGML backends or cuda, and Qwen 3.8 Flash Next only on the GGML backends. Env TS_SPEC_TYPE.
--spec-draft <N>Maximum tokens drafted per speculative step (range 1–64, default 8; a block drafter defaults to — and clamps the value to — its trained block size, 5 for DSpark). Also sizes the native graph cache at load, so it belongs on the same command line as --spec. Env TS_SPEC_DRAFT.
--spec-pmin <f>Confidence gate below which drafting stops. What the number means is the algorithm's business, so each brings its own default: 0.15 for a per-token head (top-1 probability), 0.35 for a block drafter (the cumulative prefix probability, so the same number is far stricter), 0 for n-gram, where it scales the required match length instead. 0 is accepted and means never gate. Env TS_SPEC_PMIN.
--draft-model <path>Drafter GGUF for every speculator that ships as its own file (DeepSeek V4's DSpark, the DFlash / DFlash2 block drafters for Muse-Glimmer and Qwen3.8-27B, Gemma 4's gemma4-assistant draft head, Qwen 3.8 Flash Next's shared MTP head). The file's own general.architecture decides how it loads; naming the file is the request — no --spec needed (an explicit --no-spec vetoes it). A block drafter has to be resident before the model's layer split runs; DeepSeek V4's DSpark engages on cuda / ggml_cuda only. DeepSeek V4.1 accepts a deepseek41-dspark drafter on ggml_cuda / ggml_cpu only — experimental; initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified. A DFlash / DFlash2 drafter runs on one GPU: under --tp N > 1 the CLI declines it with a warning and serves standard decoding. Qwen 3.6, Qwen 3.8 27B, GLM-5.2 and GLM-5.3 embed their NextN block in the trunk GGUF and need no file (they take --spec) — GLM-5.3's is complete at blk.78, so there is nothing extra to download. Env TS_SPEC_DRAFT_MODEL.
--temperature / --top-k / --top-p / --min-pSampling controls.
--repeat-penalty / --presence-penalty / --frequency-penaltyPenalties (1.0 / 0 = off).
--repeat-last-n <N>How many recent tokens the repeat / presence / frequency penalties consider (default 64; 0 disables history penalties, -1 uses the whole history). The same spelling as the server and the repeat_last_n request field; the former CLI-only --penalty-last-n is removed.
--seed <N> / --stop <string>Random seed (-1 = random) / stop sequence (repeatable).
--dump-promptRender prompt + tokenization and exit.
--diffusion-steps / --diffusion-seed / --diffusion-blocks <N>DiffusionGemma generation controls (48 steps by default). --diffusion-steps and --diffusion-seed also set Qwen-Image-2.1's step count (auto: 40, or a LoRA recipe's) and seed.
--prompt <text>Qwen-Image-2.1 prompt: generates an image, or edits the --image reference(s); the PNG goes to --output.
--mask <path> · --mask-mode grayscale|alpha · --mask-invert · --mask-feather <px> · --mask-crop · --mask-crop-padding <px>Qwen-Image-2.1 local editing: a mask matching the first reference selects editable pixels, with inversion, inward feathering and optional region cropping. Output retains the source dimensions and protected pixels. These are CLI flags; the server takes matching per-request fields. Semantics and ranges →
--cfg <F>True-CFG guidance scale. Qwen-Image-2.1: auto 1.0 (one transformer prediction per step); <= 1 disables the negative pass. Video generation reuses the flag — MiniMax-H3 ships CFG-distilled and refuses anything above 1.0, so --cfg 1.0 is the only setting it accepts.
--qwen-image-vae / --qwen-image-vl / --qwen-image-mmproj <path>Override the resolved Qwen-Image-2.1 companions (dedicated 2.1 VAE / Qwen3-VL-8B text encoder / Qwen3-VL-8B mmproj).
--lora <path>Qwen-Image-2.1 LoRA plug-in: a LoRA .safetensors or a TensorSharp plug-in config .json (config/lora/). Repeatable to stack; applied unmerged on top of the quantized transformer; refused with any other model. → LoRA plug-ins
--lora-scale <f> / --lora-config <path>Strength (default: the plug-in's "scale", else 1.0) and companion config (a TensorSharp LoRA config, PEFT adapter_config.json or VideoX-Fun pdd_config.json) of the preceding --lora. A plug-in's sampling recipe applies unless --diffusion-steps / --cfg are given.
--qwen-image-lora · --offload-cpuRemoved and rejected at startup. Both served only the earlier Qwen-Image-Edit pipeline. --qwen-image-lora is replaced by --lora; --offload-cpu has no replacement, because Qwen-Image-2.1 keeps its DiT weights resident.
--penalty-last-nRemoved and rejected at startup, including as a --config key. It was the CLI-only name of the repeat-penalty window — use --repeat-last-n <N>.
--paged-kv* · --paged-bench* · --paged-batchingRemoved and rejected at startup, including as --config keys. The standalone paged-KV store they configured is gone: prompt reuse across requests is the engine's radix prefix cache (--no-prefix-cache turns it off), and its TS_KV_* variables are refused too. --paged-batching / --no-paged-batching were second spellings of --continuous-batching / --no-continuous-batching.
--width <px> / --height <px>Output canvas for the media families. Qwen-Image-2.1: 0 = auto (2048×2048 for generation, or about that area at the first reference's aspect ratio for editing; explicit sizes in multiples of 32). Video generation: MiniMax-H3 rounds both up to a multiple of 32 and defaults to 640×384, or that area at the conditioning image's aspect ratio. Resolution is the dominant cost lever, and faces need pixels — 640×384 is the starting point, not 256×256.
--video-frames <N> / --fps <N>Output frame count, snapped up onto the model's temporal grid — 17k+5 for MiniMax-H3 (5, 22, 39, 56, 73, 90, 107 …; default 22), 4k+1 for Wan (default 33, or 49 for Wan2.2-TI2V) — and MP4 playback rate (Wan: 16, or 24 for Wan2.2-TI2V). A model trained at a fixed rate overrides the value asked for: MiniMax-H3 is pinned to 24 fps.
--video-mode <mode>MiniMax-H3 conditioning mode, inferred from what you pass when omitted: t2v (text only), i2v (the image is the first frame and gets animated), fl2v (first and last frame), ref (the images are identity/appearance references for a new scene). i2v/fl2v need the fl2va checkpoint and ref needs ref2va — separate files, and each refuses the other's inputs by naming the file to load instead. On ref2va a plain --image is taken as a single reference.
--end-image <file>Last-frame conditioning image (MiniMax-H3 first/last-frame mode). Combined with --image the clip is steered to start and end on the two frames.
--ref-image <file>Reference still for MiniMax-H3 Ref2VA: the subject carries over while camera, background and composition come from the prompt. Repeatable, up to nine references in total; the prompt refers to them positionally as <Picture 1>, <Picture 2>, … References are only ever scaled down and keep their own aspect ratio, so the output canvas still comes from --width/--height.
--ref-video <path> / --ref-video-audio <file>Reference clip — a video file or a directory of frames, resampled onto MiniMax-H3's own 24 fps and canvas, referred to as <Video 1>, <Video 2>, … — and its soundtrack, which pairs by position: the first --ref-video-audio with the first --ref-video. Separate flags because a container's audio track is not readable through the frame decoder. A reference clip is the most expensive input H3 takes: a 22-frame 448×320 one adds 980 conditioning tokens on top of the 1680 the output needs.
--ref-audio <file>Standalone reference soundtrack (<Audio 1>, …), WAV / MP3 / Ogg, resampled to the audio VAE's 32 kHz stereo and truncated to the generated clip's duration. Repeatable; stills, clips and audio share the same nine-reference budget.
--no-audioSkip the audio decode on a model that generates a track jointly with the video (MiniMax-H3), saving the audio VAE's time and memory. Ignored by video-only models.
--flow-shift <F> / --sampler <name>FlowMatch timestep shift (0 = the model's official recipe; 12.0 for MiniMax-H3, where it shifts the video stream only) and sampler: unipc (default) or euler. --sampler is a Wan-family knob — MiniMax-H3 runs its own flow-match schedule and ignores it.
--negative-prompt <text>Classifier-free-guidance negative prompt (default: the model's official one). No unconditional pass runs at --cfg 1.0, so it is inert on MiniMax-H3, which accepts nothing else, and on a step-distilled Wan checkpoint, which runs guidance-free.
--cfg-cache-stride <N>Wan guidance cache: run the unconditional CFG pass on one step in N and reuse the cached guidance direction in between (default 0 = off). At 50 steps, 2 = 1.30x and 3 = 1.43x. Approximate, and it has no effect at --cfg 1.0 — so nothing on MiniMax-H3 or a distilled Wan checkpoint.
--video-vae / --video-text-encoder / --audio-vae <path>Override the resolved video companions: the video VAE (minimax_h3_video_vae_fp16.safetensors for MiniMax-H3, the causal 3D VAE for Wan), the text encoder (Qwen3-VL-32B for MiniMax-H3, UMT5-XXL for Wan), and the audio VAE MiniMax-H3 decodes its soundtrack with — omit it and the model still runs, just silently. Env TS_VIDEO_VAE / TS_VIDEO_TEXT_ENCODER / TS_VIDEO_AUDIO_VAE; Wan A14B's second expert is --video-dit2.
--n-cpu-moe <N|all> (-ncmoe) · --cpu-moe (-cmoe) · --cpu-moe-threads <N>Keep the routed MoE expert weights of the first N layers (or all of them) in system RAM and multiply them on the CPU. Default 0 on every architecture, DeepSeek V4 included. Host threads default to 1 with two or fewer usable CPUs, all but one with up to eight, and otherwise half, capped at 64 (DeepSeek V4 / V4.1 and GLM-5.x on their native executors use every usable CPU once experts are offloaded, GLM-5.x also on a GPU-less run). Env TS_N_CPU_MOE / TS_CPU_MOE / TS_CPU_MOE_THREADS.
--no-prefix-cacheDisable Radix prefix caching and the interactive system/tool prompt warm-up (also sets TS_SCHED_PREFIX_CACHE=0). Prefix caching is on by default and lives in memory for the CLI process.
--continuous-batching / --no-continuous-batchingEnable (default) / disable the paged continuous-batching engine.
--benchmark / --bench-prefill / --bench-decode / --bench-runsSynthetic throughput benchmark. --bench-chunked uses the server-style chunked prefill; --bench-fixed-tokens times a predetermined decode stream without host sampling; --bench-random-tokens draws the prompt and decode ids at random, as llama-bench does.
--bench-kvcache / --bench-kv-turns <N>Multi-turn KV-cache reuse benchmark.
--warmup-runs <N>Throw-away forward passes before timing (default 0).
--test / --test-templates <dir>Built-in tokenizer/template tests; validate templates against GGUF Jinja2.
--test-chunked-prefill / --correct-prefill <N> / --correct-decode <N>Prefill one prompt single-pass and chunked and compare the decoded tokens (defaults 1500 prompt tokens, 8 decode steps).
--config <path>Read options from a JSON file keyed by long option names (repeatable). A single-valued option given on the command line (or in a later file) drops the file's entry before it is resolved, so the command line wins and that entry's download is skipped, with one [config] line on stderr; the repeatable options (--stop, --skills-dir, --skill, --lora*, --image, --ref-*) add to the file's values instead. See config/README.md.
--tp <N>Tensor parallelism degree — split the model across N GPUs in one process (default 1; requires --backend cuda, ggml_cuda, or ggml_vulkan). Also accepted by the server.
--tp-node-id <N> / --tp-peers <list>Multi-node distributed TP: this node's 0-based ID and the shared comma-separated host:port peer list.
--log-level / --log-dir / --log-file / --log-consoleLogger level, directory, and file/console toggles. CLI only: the server has no --log-* flags and reads the TENSORSHARP_LOG_* variables instead.

Server flags — TensorSharp.Server.Host

FlagDescription
--model <path>GGUF file to host. Required at startup; a model-less process cannot select a GGUF through /api/models/load.
--mmproj <path|none>Explicit multimodal projector: an mmproj GGUF or, for the Gemma 4 family (DiffusionGemma included), a Hugging Face .safetensors vision-tower shard. A bare file name is resolved next to the model; none disables it. The server does not auto-detect it.
--embeddings · --embedding-threads <N> · --embedding-context-size <N>Host a GGUF embedding encoder instead of a chat model (/v1/embeddings, /api/embed). Requires --model, forbids --mmproj, and runs on cpu, ggml_cpu, ggml_metal or ggml_cuda; the chat, Jev, image and video routes answer 400 in this mode, and an explicit --backend the machine lacks exits 2. See Embeddings.
--port <N> · --host <address> · --urls <urls>Listen port (default 5000; env PORT), bind address (default 0.0.0.0; env HOST), or full listen URLs (falls back to ASPNETCORE_URLS).
--no-webuiDo not serve the bundled Web UI; the HTTP APIs and /uploads stay up. Env TS_NO_WEBUI.
--backend <type>Compute backend; defaults to ggml_metal on macOS and ggml_cpu elsewhere.
--gpu-device <N> / --list-gpusVulkan device selection for ggml_vulkan / list visible Vulkan devices and exit.
--tp <N> · --tp-node-id <N> · --tp-peers <list>Tensor parallelism, same semantics as the CLI. Env TENSORSHARP_TP_DEGREE / TENSORSHARP_TP_NODE_ID / TENSORSHARP_TP_PEERS.
--helpPrint the parameter reference and exit (also shown when started with no arguments). An unknown option is refused at startup, with a did-you-mean suggestion when a known flag is close.
--config <path>Read options from a JSON file, as on the CLI (repeatable).
--max-tokens <N>Generation limit for every endpoint: fills in when a request omits it and caps a larger request (default 20000, which only fills in).
--temperature / --top-k / --top-p / --min-pDefault sampling values.
--repeat-penalty / --presence-penalty / --frequency-penalty / --seedDefault penalties and seed.
--repeat-last-n <N>Default penalty window (64 tokens); env TENSORSHARP_REPEAT_LAST_N. The CLI uses the same spelling; the former CLI-only --penalty-last-n is removed and refused on both hosts.
--stop <string>Stop sequence (repeatable); merged with a per-request list under config precedence, replaced by it under request.
--sampling-precedence <config|request>Whether configured sampling parameters outrank the ones a request sends (default config). Parameters left unconfigured always come from the request.
--continuous-batching / --no-continuous-batchingEnable (default) / disable iteration-level paged batching.
--no-prefix-cacheDisable the Radix prefix cache (on by default), the startup warm-up of the shared prompt and the prefix checkpoints saved between launches in prefix-cache/ beside the binary (TENSORSHARP_PREFIX_CACHE_DIR moves them). Also sets TS_SCHED_PREFIX_CACHE=0.
--n-cpu-moe <N|all> · --cpu-moe · --cpu-moe-threads <N>MoE expert offload to system RAM; same semantics and defaults as the CLI.
--spec / --no-specEnable / disable speculative decoding (default off); engages for solo, non-concurrent sequences. The explicit opt-in for drafters embedded in the trunk checkpoint; a drafter named with --draft-model engages without it.
--spec-type <name>Speculation algorithm: auto (default) uses whatever drafter the checkpoint carries, draft-head and block pin one explicitly, and ngram needs no trained weights at all — it drafts by suffix match over the context — but still only runs on a speculative-target model (not GPT-OSS, Mistral 3, Hunyuan Dense or DiffusionGemma; Nemotron-H refuses every speculator; DeepSeek V4 / V4.1 and Muse-Glimmer only while their own drafter is loaded, Gemma 4 only on the GGML backends or cuda, Qwen 3.8 Flash Next only on the GGML backends). It never turns speculation on by itself: given without --spec or --draft-model, startup logs a WARNING that speculation stays off.
--spec-draft <N>Max tokens drafted per speculative step (1–64, default 8; a block drafter defaults to — and clamps the value to — its trained block size).
--spec-pmin <f>Draft-confidence gate in [0, 1] (0 = never gate); drafting stops below it. Default per algorithm — 0.15 for a per-token head, 0.35 for a block drafter (where the gate is the cumulative prefix probability), 0 for n-gram.
--draft-model <path>Drafter GGUF for every speculator that ships as its own file (DeepSeek V4's DSpark, the DFlash / DFlash2 block drafters for Muse-Glimmer and Qwen3.8-27B, Gemma 4 gemma4-assistant, Qwen 3.8 Flash Next's shared MTP head). The file's own general.architecture decides how it loads; naming the file is the request, so it needs no --spec (an explicit --no-spec vetoes it). Qwen 3.6, Qwen 3.8 27B, GLM-5.2 and GLM-5.3 embed their NextN block and need no draft file; an explicit drafter that cannot activate (a DFlash / DFlash2 drafter under --tp N > 1, any drafter on Nemotron-H) stops startup with exit code 2; DeepSeek V4's DSpark serves solo sequences on cuda / ggml_cuda. DeepSeek V4.1 accepts a deepseek41-dspark drafter on ggml_cuda / ggml_cpu only — experimental; initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified. Env TS_SPEC_DRAFT_MODEL.
--prefill-chunk-size <N>Maximum prefill tokens per scheduler step (sets TS_SCHED_PREFILL_CHUNK).
--kv-cache-dtype <type>KV cache precision: f32, f16, q8_0, q4_0 (default: auto per backend/model; env KV_CACHE_DTYPE).
--skills-dir · --skill · --list-skills · --no-skills · --skills-no-discoveryServer-side Agent Skills registry, startup selection and discovery policy; same semantics as the CLI, except that skills/ beside the binary (where POST /api/skills installs uploads) is always scanned first, even with explicit roots, so it wins a name clash.
--skills-allow-exec · --skills-sandbox · --skills-allow-network · --skills-max-roundsScript opt-in, isolation/network policy and 1–64 round cap. Defaults: execution/network off, sandbox required, rounds 8 or 24 with code execution.
--code-exec · --code-exec-allow-install · --code-exec-allow-network · --code-exec-unconfinedEnable built-in code tools and their separate install/network/unconfined permissions. All are operator startup decisions; there is no per-command approval endpoint.
--code-exec-packages · --code-exec-install-domains · --code-exec-install-index · --code-exec-timeout · --code-exec-max-output · --code-exec-shell · --code-exec-temperaturePackage/install source policy, command limits, shell, and optional coding-turn temperature. When code tools can run, coding turns disable the built-in 1.1 repetition penalty independently of that temperature flag. See Agentic Work.
--code-exec-languagesRemoved and rejected at startup; no replacement. The shell reaches every interpreter on its rebuilt PATH.
--no-multi-agentTurn off sub-agent delegation, which is on by default for tool-capable models on /v1/chat/completions, /v1/responses, /api/chat/ollama and /api/chat. Env TS_NO_MULTI_AGENT (any value but 0). A request can only opt out, with "multi_agent": false. See Sub-agents.
--agents-max-concurrent · --agents-max-count · --agents-max-depth · --agents-max-rounds · --agents-max-generations · --agents-timeout · --agents-max-result-charsPer-request delegation limits: active children across the tree (default 3, range 1–32), children per request (8, 1–128), depth (2, 1–8), tool-loop rounds per child (8, 1–64), shared child-generation budget (48, 1–1024), seconds per child run (180, 1–3600), and characters per child report (8000, 256–64000).
--agents-allow-worker-toolsLet worker children use the parent's enabled mutable tools. Without it every child is read-only; explorer and reviewer children always are.
--upload-max-mb <N> · --upload-quota-mb <N> · --upload-ttl-hours <N>Upload-directory limits: per-file cap on client-supplied files (default 500 MB; 413 when exceeded; POST /api/upload's request-body limit follows it and never drops below 500 MB, while every other route keeps 500 MB), a total quota that also counts generated outputs (507), and deletion of files older than N hours. Quota and TTL are off by default. Env TS_UPLOAD_MAX_MB / TS_UPLOAD_QUOTA_MB / TS_UPLOAD_TTL_HOURS.
--qwen-image-vae / --qwen-image-vl / --qwen-image-mmproj <path>Override the resolved Qwen-Image-2.1 companions (2.1 VAE / Qwen3-VL-8B text encoder / mmproj).
--lora · --lora-scale · --lora-configQwen-Image-2.1 LoRA plug-ins; same semantics as the CLI. Checked at startup and applied to every image request (no per-request selection); a request's steps / cfg still override a plug-in's recipe.
--qwen-image-lora · --offload-cpuRemoved and rejected at startup. Both served only the earlier Qwen-Image-Edit pipeline. --qwen-image-lora is replaced by --lora; --offload-cpu has no replacement, because Qwen-Image-2.1 keeps its DiT weights resident.
--penalty-last-nRemoved and rejected at startup, including as a --config key. It was the CLI-only name of the repeat-penalty window — use --repeat-last-n <N>.
--paged-kv* · --paged-bench* · --paged-batchingRemoved and rejected at startup, including as --config keys. The standalone paged-KV store they configured is gone: prompt reuse across requests is the engine's radix prefix cache (--no-prefix-cache turns it off), and its TS_KV_* variables are refused too. --paged-batching / --no-paged-batching were second spellings of --continuous-batching / --no-continuous-batching.
--video-width <px> / --video-height <px>Default output size when a request omits width/height — the main quality lever, and a startup flag because the browser sends no size of its own. Aliases --width / --height, which additionally set the default Qwen-Image-2.1 image size (TS_QWEN_IMAGE_WIDTH / TS_QWEN_IMAGE_HEIGHT; used only when both are given, for image requests that name neither a size nor an area; an off-grid value is rounded down to a multiple of 32 with a one-time warning, and one side alone is ignored with a warning); rounded up to the model's grid (a multiple of 32 for MiniMax-H3). 640×384 is the documented starting point for MiniMax-H3; give only one of the two and H3 takes the other from the conditioning image's aspect ratio.
--video-steps <N>Default denoising steps when a request omits steps — the quality/time trade-off after resolution. MiniMax-H3 defaults to 20, runs its fast operating point at 4–8, is visibly cleaner at 16–24, and gains little past ~30. The server has no --cfg at all.
--video-mode <mode>Default conditioning mode when a request omits videoMode: t2v, i2v, fl2v or ref. Omit it and the mode is inferred per request, which is usually what you want; pin it for a deployment that only offers one.
--video-frames <N> / --fps <N>Default frame count and MP4 playback rate when a request omits frames/fps. The count is snapped to the model's temporal grid — 17k+5 for MiniMax-H3 (default 22), 4k+1 for Wan (default 33, or 49 for Wan2.2-TI2V). MiniMax-H3 is pinned to 24 fps whatever is asked for.
--video-vae / --video-text-encoder / --audio-vae <path>Video companions: the video VAE, the text encoder (Qwen3-VL-32B for MiniMax-H3, UMT5-XXL for Wan), and the audio VAE MiniMax-H3 decodes its soundtrack with — without it the model still runs and produces silent video. Env TS_VIDEO_VAE / TS_VIDEO_TEXT_ENCODER / TS_VIDEO_AUDIO_VAE; Wan A14B's second expert is --video-dit2.
--redis-url <url>Sets TS_RESPONSES_STORE_REDIS_URL (only when unset), which moves the /v1/responses store to Redis (--help lists it under "Responses API store").

Environment variables

VariableDescription
BACKENDDefault backend (ggml_metal on macOS, ggml_cpu elsewhere).
MAX_TOKENSDefault max generation length (20000).
TS_PDF_MAX_PAGESCap on PDF pages read (upload and CLI --pdf); 0 = all pages (default).
VIDEO_SAMPLE_FPS / VIDEO_MAX_FRAMESVideo frame sampling rate / cap.
TENSORSHARP_TEMPERATURE / _TOP_K / _TOP_P / _MIN_PDefault sampling values.
TENSORSHARP_REPEAT_PENALTY / _REPEAT_LAST_N / _PRESENCE_PENALTY / _FREQUENCY_PENALTY / _SEEDDefault penalties, penalty window (64) and seed.
TENSORSHARP_SAMPLING_PRECEDENCEconfig (default) or request: whether the values above outrank the ones a client sends (= --sampling-precedence).
TENSORSHARP_LOG_LEVEL / _LOG_DIR / _LOG_FILELogging level, directory, file toggle (CLI + server).
TS_SKILLS_DIR / TS_NO_SKILLS / TS_SKILLS_ALLOW_EXEC / TS_SKILLS_MAX_ROUNDS / TS_SKILLS_SANDBOX / TS_SKILLS_ALLOW_NETWORKAgent Skills roots, disable switch, script opt-in, round bound, sandbox mode and script network opt-in.
TS_CODE_EXEC / TS_CODE_EXEC_ALLOW_INSTALL / TS_CODE_EXEC_ALLOW_NETWORK / TS_CODE_EXEC_INSTALL_DOMAINS / TS_CODE_EXEC_INSTALL_INDEXCode-tool opt-in, separate install/network grants, and host-installer source policy.
TS_NO_MULTI_AGENTServer only: any value but 0 turns sub-agent delegation off (= --no-multi-agent).
TS_UPLOAD_MAX_MB / TS_UPLOAD_QUOTA_MB / TS_UPLOAD_TTL_HOURSUpload-directory per-file cap (default 500), total quota and file lifetime (both off by default) (= --upload-max-mb / --upload-quota-mb / --upload-ttl-hours).
TENSORSHARP_UPLOAD_DIRWhere the server stores uploads and generated files (default uploads/ beside the binary).
DIFFUSION_STEPS / DIFFUSION_MAX_BATCHDiffusionGemma steps per block / max batched requests.
TS_JEV_MAX_BODY_MB / TS_JEV_MAX_CANVAS / TS_JEV_MAX_PENDING/v1/systemone limits: request-body cap in MiB (default 8, range 1–64), answer-canvas width per question chunk (default 64, range 8–4096, also bounded by the checkpoint), and admitted requests before 529 (default 32, range 1–1024).
KV_CACHE_DTYPEKV cache precision (CLI + server): f32, f16, q8_0, q4_0; default auto (= --kv-cache-dtype).
TS_SCHED_DISABLE_BATCHED1 forces per-sequence KV-swap (= --no-continuous-batching).
TS_SCHED_MAX_BATCHED_TOKENSPer-step token budget (4096).
TS_SCHED_MAX_RUNNING_SEQSMax in-flight sequences (16).
TS_SCHED_PREFILL_CHUNKPer-request mixed-step prefill cap (256); prefill-only steps fill the batch budget.
TS_SCHED_SOLO_PREFILL_CHUNKPrefill chunk for a solo / uncontended request (8192).
TS_SCHED_DECODE_QUANTUMDecode tokens before a sequence switch (256 = block size).
TS_SCHED_NUM_BLOCKS / TS_SCHED_BLOCK_SIZEEngine block-pool size (256) / tokens per block (16 for a family that reuses only whole pages, otherwise 256).
TS_SCHED_PREFIX_CACHE0 disables all prefix reuse (--no-prefix-cache sets it).
TENSORSHARP_PREFIX_CACHE_DIRServer: root for the prefix checkpoints saved between launches, one subdirectory per model (default prefix-cache/ beside the binary).
TS_SPEC / TS_SPEC_TYPE / TS_SPEC_DRAFT / TS_SPEC_PMIN / TS_SPEC_DRAFT_MODELSpeculative-decoding knobs, mirroring the --spec* flags: on/off, algorithm (auto, draft-head, block, ngram), draft length, confidence gate, and the separate draft-head GGUF.
TS_DSV4_UBATCH / TS_DSV4_PERFDeepSeek V4: prefill micro-batch, and throughput/stage logging.
TS_DSV4_THREADS / TS_DSV4_CPU_TRACE_DIRDeepSeek V4.1 Flash on --backend cpu, the pure-C# DeepSeek4CpuExecutor — no ggml, no native library and no GPU, so it runs anywhere .NET does. TS_DSV4_THREADS is the compute-thread count and defaults to Environment.ProcessorCount on this backend, against min(cores, 32) elsewhere; TS_DSV4_CPU_TRACE_DIR writes the same per-tensor files eng/dsv41-reference.py --output writes, so the two directories diff tensor by tensor. This is a correctness and portability path, not a serving path. Engram comes directly from the GGUF metadata, and MAX_CONTEXT caps to 65,536 by default here.
TS_DSV41_ENGRAM_DEVICENative DeepSeek V4.1 loader placement: which device holds the Engram tables. On --backend cpu only 0 is accepted, and that backend also refuses --tp, distributed TP groups and any draft model, all before any weight is read. On V4.1 a drafter is accepted only on ggml_cuda or ggml_cpu and only as a deepseek41-dspark file (experimental: initial text/image HTTP probes with trained weights passed using two-GPU layer split on ggml_cuda; broad quality and throughput remain unqualified); anywhere else naming one is a hard refusal rather than DeepSeek V4's warn-and-continue. TS_DSV41_ENGRAM_WARM / _THREADS / _RANDOM, TS_DSV41_SPARSE_FA and TS_DSV41_COMPACT_RAW_GATHER are native-loader only and inert on cpu.
TS_QWEN_IMAGE_VAE / TS_QWEN_IMAGE_TE / TS_QWEN_IMAGE_MMPROJQwen-Image-2.1 companion paths (dedicated 2.1 VAE / Qwen3-VL-8B text encoder / Qwen3-VL-8B mmproj).
TS_QWEN_IMAGE_WIDTH / TS_QWEN_IMAGE_HEIGHTQwen-Image-2.1 default output size for an image request that gives neither a width/height nor an explicit targetArea (a request with its own area keeps its own geometry). Used only when both are set; a value off the 32-pixel grid is rounded down to a multiple of 32 (minimum 32) with a one-time warning, and one alone — or an unparsable or negative value — is ignored with a one-time warning. The server's --width / --height set them.
TS_QWEN21_PREFIX_CACHE / _PREFIX_CACHE_TYPE / _PREFIX_CACHE_MAX_MIBQwen-Image-2.1 prefix KV cache: the prompt and reference tokens' K/V are stored at step 1 and reused by every later step. On by default (0/false/off/no disables it); storage auto (what attention reads), f32, f16, q8_0 or q8_0_v; a MiB cap on top of the built-in limit of half the device's free memory. The model card measures 1024² single-reference edits at 1.73–1.86× per step on an M5 Pro (Metal) and 1.92× on an A40 (CUDA), with byte-identical PNGs; generation gains 1–3%.
TS_QWEN21_GRAPH_REUSE / TS_QWEN21_FLASH / TS_QWEN21_PAD_MASKQwen-Image-2.1 DiT diagnostics: 0 stops retaining graphs between steps (retention is on by default and applies on CUDA and Metal only), 0 turns flash attention off, and 1 selects padded attention masks (off by default).
TS_QWEN21_VAE_FUSED / TS_QWEN21_VISION_FUSEDQwen-Image-2.1 whole-VAE decode graph (default on CUDA and Metal, never on Vulkan; 0 forces the per-convolution path, 1 opts in on another GGML backend) and the fused vision encoder on GGML CUDA (0 disables).
TS_LORASThe channel through which both hosts hand the parsed --lora / --lora-scale / --lora-config list to the Qwen-Image-2.1 model: a JSON array of {"path", "scale", "config"} objects. The retired TS_QWEN_IMAGE_LORA is refused at load, with advice to use --lora.
TS_VIDEO_VAE / TS_VIDEO_TEXT_ENCODER / TS_VIDEO_AUDIO_VAE / TS_VIDEO_DIT2Video companion paths (= --video-vae / --video-text-encoder / --audio-vae / --video-dit2): the video VAE, the text encoder (Qwen3-VL-32B for MiniMax-H3, UMT5-XXL for Wan), the audio VAE MiniMax-H3 decodes its soundtrack with, and Wan 2.2 A14B's second high/low-noise expert GGUF.
TS_VIDEO_TOKENIZERFolder holding MiniMax-H3's vocab.json and merges.txt. The text-encoder GGUF carries no tokenizer, and auto-download can only fill in options that are flags, so take the pair from MiniMaxAI/MiniMax-H3 under processor/ and either drop it beside the encoder or point this variable at it.
TS_H3_TRACE1 prints MiniMax-H3's latent and velocity magnitudes for every denoise step. The sampler already fails a request whose velocity goes non-finite — naming the step it appeared at, rather than writing a black clip — and this shows where the magnitudes left the FP16 range.
TS_H3_PREFAULTSequentially reads the MiniMax-H3 denoiser file so its first upload is a copy out of the page cache rather than a page-fault storm inside the host-to-device transfer (0.91 GB/s faulting as it copies against 5.97 GB/s from resident pages, RTX 3080 Laptop). Default 3: the read starts the moment the text trunk hands over its hidden states and is pipelined with the upload rather than joined before it. 1 is serial, 2 overlaps it with text conditioning — measured worse, because the encoder streams its own 17 GB through the same page cache and evicts the pages just placed — and 0 disables it. Output is byte-identical in every mode.
TS_H3_PREFAULT_THREADSRead streams the prefault uses (default 1). More measured worse here — 640x384 best of three on an RTX 3080 Laptop: 1 stream 63.9 s, 4 streams 64.9 s, 16 streams 66.6 s — because this read runs concurrently with the encoder teardown and the upload it is warming, not with the machine to itself.
TS_H3_PHASE1 prints a per-stage MiniMax-H3 breakdown: text-encoder open / trunk / teardown, the prefault, every denoise step, and VAE open / decode. The one-line summaries tell you a phase was slow; this tells you which half of it.
TS_H3_TE_GROUP<n> runs MiniMax-H3's 50-layer text-encoder trunk in groups of n layers, releasing each group's device copy. Off by default: on a 16 GB RTX 3080 Laptop it does remove the trunk's spill (peak 16 041 -> 12 981 MiB) and is bit-identical, but the trunk is a one-shot prefill over a short prompt, so it was 3 s slower.
TS_H3_KEEP_RESIDENT1 keeps MiniMax-H3's denoiser and both VAEs resident between clips on Metal, the old behaviour. By default they are released before each clip's text encoder, which held peak wired memory at 19.1 GB instead of 33.3 GB in the TensorAgent Mac app on an M5 Pro, for about 1.3 s more text conditioning. Discrete GPUs are unaffected.
TS_WAN_DIT_KV_F160 restores F32 attention keys/values in the Wan DiT. F16 is the default and is 2.02x on a single 27k-token self-attention at the same 0.999964 DiT cosine.
TS_WAN_VAE_MPS_CONVMetal only: 0 restores ggml's im2col+GEMM lowering for the Wan VAE convolutions. MPSGraph is the default (a 736x544x81f decode: 159 s -> 80 s, numerics unchanged).
TS_WAN_VAE_GEMM_MAX_MB / TS_WAN_VAE_TILEWan VAE im2col GEMM budget in MB (default: derived from free device memory) and 0 to disable the horizontal-band tiling used above ~0.5 MP.
TS_WAN_METAL_TENSOR_APIForce the Metal 4 tensor API on (1) or off (0) for the Wan DiT. Default: on for A14B/14B-class DiTs, off for smaller ones such as TI2V-5B.
TS_WAN_DIT_CAPTURE / TS_WAN_DIT_FLASH0 disables the persistent CUDA-graph-captured Wan DiT graph / forces materialized attention. Both are slower — A/B and debugging only.
TS_WAN_HEARTBEAT_SWan progress heartbeat interval in seconds (default 30; 0 silences it).
TS_FFMPEGPath to the ffmpeg used to write the near-lossless MP4 (default: ffmpeg on PATH).
TENSORSHARP_TP_DEGREELocal tensor-parallel degree — GPUs to split the model across (= --tp on the CLI and the server).
TENSORSHARP_TP_DEVICESGPU ordinals the TP ranks map to, e.g. 0,2 (default 0..tp-1). GGML backends.
TENSORSHARP_TP_NODE_ID / TENSORSHARP_TP_PEERSMulti-node distributed TP: 0-based node ID and the shared host:port peer list (= --tp-node-id / --tp-peers).
TENSORSHARP_TP_CONNECT_TIMEOUT_SECONDS / _RECV_TIMEOUT_SECONDSPeer connect retry window (120 s) and per-receive timeout (300 s) for the distributed TP mesh.
TENSORSHARP_TP_DISABLE_P2P / TENSORSHARP_TP_HOST_ALLREDUCE1 forces cross-GPU copies / the local AllReduce through host memory instead of CUDA P2P (diagnostics).
TS_GGML_TP_DEVICE_AR_THRESHOLD / TS_GGML_TP_PARALLELGGML TP: element count above which AllReduce uses the device collective (default 262144), and 0 to drive the ranks sequentially (diagnostic).
TS_GEMMA4_TP_FUSED_MOE0 falls back from Gemma 4's fused MoE trunk under TP to the whole-expert per-op path.
GGML_CUDA_ALLREDUCE / GGML_CUDA_AR_BF16_THRESHOLDggml collective selection (nccl / internal / none) and the payload size above which F32 collectives convert to BF16 (raised to 1 MB by TensorSharp).
TS_RESPONSES_STORE_REDIS_URLRedis-backed OpenAI Responses API store (replaces the in-memory store); the live half of --redis-url.
TS_GGML_VULKAN_DEVICEVulkan device index for ggml_vulkan (= --gpu-device).
TS_FUSED_QKNORM_ROPE0 disables the fused QK-Norm + RoPE CUDA kernel in Qwen 3.5/3.6 text prefill on the direct cuda backend (default on).
TS_MLX_* MLX backend tuning: mlock GGUF, batched MoE decode, layer-eval cadence, memory caps.
TENSORSHARP_MLX_LIBRARY / _LIBRARY_DIROverride the search path for libmlxc.
TENSORSHARP_GGML_NO_UPDATE / _GGML_GIT_REFSkip / pin the ggml source clone on native builds.
TENSORSHARP_GGML_NATIVE_ENABLE_VULKANForce the GGML Vulkan backend on/off at native build time (ON/OFF = build-script --vulkan/--no-vulkan). Auto-enabled when a Vulkan runtime is detected; an explicit choice sticks across rebuilds.

HTTP endpoints

Method & pathStylePurpose
GET / · GET /index.htmlWeb UIBrowser chat application when wwwroot is present; bare / falls back to liveness only in a headless deployment.
GET /healthUtilityStable plain-text liveness response.
POST /api/generateOllamaSingle-prompt completion (stream or not).
POST /api/chat/ollamaOllamaMulti-turn chat with optional think / tools / images.
GET /api/tagsOllamaList the hosted model.
POST /api/showOllamaModel info.
POST /v1/chat/completionsOpenAIChat Completions (stream, tools, response_format, skills, multi_agent).
POST /v1/responses · GET /v1/responses/{id}OpenAIResponses API generation and stored-response lookup. Stateless per request: previous_response_id is refused.
GET /v1/modelsOpenAIList models.
POST /v1/embeddings · POST /api/embedOpenAI / OllamaNormalized sentence embeddings from a server started with --embeddings.
POST /v1/systemoneJevTyped decisions (noul, choice, score) about a text and/or image state on a hosted DiffusionGemma model; jev-latest / jev-preview alias the loaded model. See HTTP API.
POST /api/chatWeb UISSE chat stream with session + KV-reuse fields.
POST /api/sessions · DELETE /api/sessions/{id}Web UICreate / dispose a per-tab session.
POST /api/uploadWeb UIUpload an image / audio / video / PDF / text file.
GET /v1/skills/{name?}OpenAIList registered skills or read one descriptor including its instructions.
GET /api/skills · GET /api/skills/{name} · GET /api/skills/{name}/files/{*path}Web UIList/read skills and safely fetch a bundled file.
POST /api/skills · POST /api/skills/rescan · DELETE /api/skills/{name}Web UIInstall a skill ZIP, rescan roots, or remove an installed skill.
GET /api/code/artifacts/{runId} · GET /api/code/artifacts/{runId}/{*path}Web UIList files retained from one code run or download one as a forced attachment.
POST /api/image-generateWeb UIQwen-Image-2.1 text-to-image (JSON).
POST /api/image-generate/streamWeb UISSE text-to-image with live denoising previews.
POST /api/image-editWeb UIQwen-Image-2.1 one-shot edit (multipart or JSON).
POST /api/image-edit/streamWeb UISSE image edit with live denoising previews.
POST /api/video-generateWeb UIVideo generation on the hosted model (MiniMax-H3, Wan). JSON { prompt, width, height, frames, steps, cfg, fps, imagePath, videoMode, generateAudio, endImage, referenceImages, referenceVideos, referenceAudios, referenceVideoAudios } — videoMode, generateAudio, endImage and the four reference* lists also accept a snake_case spelling (camelCase wins if both are sent); the rest are camelCase only, and an unrecognized spelling such as image_path is silently ignored. referenceVideoAudios pairs by index with referenceVideos, and each file field names something previously uploaded through /api/upload. Returns a URL to the MP4 plus audioUrl, null when the model produced no track. A model rejection (wrong checkpoint, keyframes together with references) comes back as a 400 carrying the model's own message.
POST /api/video-generate/streamWeb UISame body, streamed as SSE progress ticks per denoise step.
POST /v1/videos/generationsOpenAIOpenAI-shaped envelope over the same parser (size: "640x384"); the sidecar track comes back as audio_url.
GET /api/modelsWeb UIHosted model, supported backends, defaults. For a video model it also returns a video object — family, supportsAudio, supportsImageConditioning, supportsEndImageConditioning, supportsReferenceConditioning, maxReferenceImages — null for every non-video model, so a client tests one field instead of pattern-matching an architecture string.
POST /api/models/loadWeb UIRe-load the startup model (optionally on a different backend).
GET /api/version · GET /api/queue/statusUtilityServer version / legacy queue snapshot.
GET /uploads/*UtilityStored uploads and generated images / videos, with an allow-listed content type and nosniff; still served under --no-webui.

Sampling parameters

Ollama (options)OpenAI (top-level)DefaultMeaning
num_predictmax_tokens--max-tokens / 20000Maximum tokens to generate; the server default applies to every endpoint when omitted. OpenAI's max_completion_tokens is accepted too.
temperaturetemperature0.8Sampling temperature (0 = greedy).
top_ktop_k40Top-K filtering (0 = disabled).
top_ptop_p0.9Nucleus sampling threshold (1.0 = disabled).
min_pmin_p0Minimum probability filtering.
repeat_penaltyrepeat_penalty / repetition_penalty1.1Repetition penalty (1.0 = off).
repeat_last_nrepeat_last_n64How many recent tokens the penalties look back over.
presence_penalty / frequency_penaltypresence_penalty / frequency_penalty0Presence / frequency penalties.
seedseed-1Random seed (-1 = random).
stopstopnullStop sequences.
—response_formatnulltext, json_object, or json_schema.

C# public API

MemberSignature / valuesNotes
ModelBase.Createstatic ModelBase Create(string ggufPath, BackendType backend, int tpDegree = 1, ITensorParallelGroup tpGroup = null, string draftModelPath = null)Auto-detects architecture from GGUF metadata.
ModelBase.Forwardfloat[] Forward(int[] tokens)Returns next-token logits (length = vocab size).
ModelBase.Sampleint Sample(float[] logits, SamplingConfig config, IList<int> generated = null)Applies penalties + sampling.
ModelBase.SampleGreedyint SampleGreedy(float[] logits)Deterministic argmax.
ModelBase.Config / .TokenizerModelConfig / ITokenizerConfig.VocabSize, context length, etc.
BackendTypeCpu, Cuda, Mlx, GgmlCpu, GgmlMetal, GgmlCuda, GgmlVulkanBackend selector enum.
ITokenizer.EncodeList<int> Encode(string text, bool addSpecial = true)Text → token ids.
ITokenizer.Decodestring Decode(List<int> ids)Token ids → text.
ITokenizer.IsEos / .EosTokenIdsbool IsEos(int id) / int[] EosTokenIdsEnd-of-sequence detection.
SamplingConfigTemperature, TopK, TopP, MinP, penalties, Seed, StopSequences, MaxTokensSee C# Library.
IBatchedPagedModel.ForwardBatchbatched/paged forwardImplemented by most architectures for continuous batching.
GgufPromptRenderer.Renderstring Render(string template, List<ChatMessage> messages, bool addGenerationPrompt = true, string? architecture = null, List<ToolFunction>? tools = null, bool enableThinking = false)Renders a conversation with the model's chat template (IPromptRenderer).
MultiAgentOptionsEnabled (true), MaxConcurrentAgents 3, MaxAgents 8, MaxDepth 2, MaxRoundsPerAgent 8, MaxTotalChildGenerations 48, AgentTimeoutSeconds 180, MaxResultCharacters 8000, MaxTaskCharacters 16000, AllowWorkerTools (false)Sub-agent limits for SkillAgentLoop and SkillsChatClient (TensorSharp.AgentHost.Agents). See C# Library.

REPL commands

CommandDescription
/help, /?Show all interactive commands.
/exit, /quitLeave the session.
/reset, /newClear conversation history and KV cache.
/history · /save <file>Print / append the transcript.
/system <text>Set the system prompt (resets KV cache).
/think on|off · /multiline on|offToggle reasoning mode / multi-line input.
/info, /statusShow model, backend, architecture, context/vocab, projector, depth.
/model <path> · /backend <name> · /mmproj <path> (/projector)Hot-swap model, backend (cpu, cuda, ggml_cpu, ggml_metal, ggml_cuda, ggml_vulkan), or projector; unloading a projector takes /model <path>.
/sampling, /showPrint current sampling configuration.
/max · /temp · /topk · /topp · /minpSet reply length / temperature / top-k / top-p / min-p (also /maxtokens, /temperature).
/repeat · /presence · /frequency · /seedSet penalties and seed.
/stop <text> · /clearstopAdd / clear stop sequences.
/image · /audio · /video · /text <path> · /clearattachAttach media / text for the next turn; drop pending attachments. Aliases /img, /vid, /txt and /file.
/skills · /skill <name>List installed skills and mark the active ones / turn one on or off (either toggle resets the conversation).