HTTP API (Ollama / OpenAI)

TensorSharp.Server.Host (the server) exposes focused Ollama- and OpenAI-compatible subsets plus its Web UI protocol, all on http://localhost:5000. OpenAI Chat Completions clients can use the /v1 base URL; Ollama chat clients need the non-standard /api/chat/ollama path.

StyleEndpoints
Ollama-compatible/api/embed, /api/generate, /api/chat/ollama, /api/tags, /api/show
OpenAI-compatible/v1/embeddings, /v1/chat/completions, /v1/responses, /v1/models, /v1/videos/generations, /v1/skills
Jev-compatible/v1/systemone (typed decisions on a hosted DiffusionGemma model — below)
Web UI (SSE)/api/chat, /api/sessions, /api/models, /api/models/load, /api/upload, /api/image-generate, /api/image-generate/stream, /api/image-edit, /api/image-edit/stream, /api/video-generate, /api/video-generate/stream, /api/skills, /api/code/artifacts
Utilities/health, /api/version, /api/queue/status, /uploads/* (stored uploads and generated files)
⚠️

Ollama chat path: the Ollama-compatible chat endpoint is POST /api/chat/ollama — not /api/chat, which is the Web UI's SSE endpoint. /api/generate, /api/tags, /api/show, and /api/version are at their standard Ollama paths.

📌

Start with a required --model and an explicit --mmproj when vision/audio needs one. Requests must name that startup GGUF file or its basename. /api/models/load only reloads the same startup model/projector; it cannot load an arbitrary path or add a model to a model-less process. When a request omits max_tokens / num_predict / max_output_tokens, every endpoint — Ollama, OpenAI and the Web UI alike — falls back to --max-tokens / MAX_TOKENS (default 20000); a value pinned with that flag or variable also caps larger requests. There is no legacy /v1/completions route and no /v1/images/* route. See the copy/paste server quickstart.

🔒

The server has no API-key authentication or built-in TLS and listens on 0.0.0.0:5000. Use it on a trusted network or put an authenticating TLS reverse proxy in front of it.

Embedding service

Start with --model encoder.gguf --embeddings to host Snowflake Arctic Embed L v2.0 or MiniLM. POST /v1/embeddings and POST /api/embed return normalized vectors on pure C# cpu or native ggml_cpu / ggml_metal / ggml_cuda. Downloads, APIs, truncation, and retrieval quality →

Quick call after the server starts

The server quickstart hosts gemma-4-E4B-it-Q8_0.gguf. Copy this into a second terminal:

curl -s http://localhost:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"gemma-4-E4B-it-Q8_0.gguf","messages":[{"role":"user","content":"Reply with one short hello."}],"max_tokens":32}'

The browser UI is http://localhost:5000/: when bundled wwwroot content is present, GET / serves index.html. GET /health is the stable plain-text liveness endpoint; only a headless deployment without the UI falls back to that response at /.

1 · Ollama-compatible API

List & show models

curl http://localhost:5000/api/tags

curl -X POST http://localhost:5000/api/show \
  -H "Content-Type: application/json" \
  -d '{"model": "Qwen3.5-9B-Q8_0.gguf"}'

Generate (non-streaming)

curl -X POST http://localhost:5000/api/generate \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-Q8_0.gguf",
    "prompt": "What is 1+1?",
    "stream": false,
    "options": { "num_predict": 50, "temperature": 0.7, "top_p": 0.9 }
  }'
{
  "model": "Qwen3.5-9B-Q8_0.gguf",
  "response": "1+1 equals 2.",
  "done": true,
  "done_reason": "stop",
  "prompt_eval_count": 15,
  "eval_count": 10,
  "prompt_cache_hit_tokens": 0,
  "prompt_cache_hit_ratio": 0.0
}

prompt_cache_hit_tokens reports how many prompt tokens were served straight from the prior turn's KV cache. /api/generate always resets the session, so it is always 0; it is non-zero on /api/chat/ollama when the prompt prefix matches a previous turn.

Generate (streaming)

curl -X POST http://localhost:5000/api/generate \
  -H "Content-Type: application/json" \
  -d '{ "model": "Qwen3.5-9B-Q8_0.gguf", "prompt": "Tell me a joke.", "stream": true, "options": {"num_predict": 100} }'

Each line is a JSON object (newline-delimited JSON); the final "done": true chunk carries timing and cache fields.

Chat (multi-turn)

curl -X POST http://localhost:5000/api/chat/ollama \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-Q8_0.gguf",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the capital of France?"}
    ],
    "stream": false,
    "options": {"num_predict": 100}
  }'

Generate with an image (multimodal)

Images are sent as base64-encoded bytes in the images array:

IMG_B64=$(base64 < photo.png)
curl -X POST http://localhost:5000/api/generate \
  -H "Content-Type: application/json" \
  -d "{
    \"model\": \"gemma-4-E4B-it-Q8_0.gguf\",
    \"prompt\": \"What is in this image?\",
    \"images\": [\"$IMG_B64\"],
    \"stream\": false,
    \"options\": {\"num_predict\": 200}
  }"

Chat with thinking mode

Thinking-capable models accept "think": true and split chain-of-thought into message.thinking:

curl -X POST http://localhost:5000/api/chat/ollama \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-Q8_0.gguf",
    "messages": [{"role": "user", "content": "Solve 17 * 23 step by step."}],
    "think": true, "stream": false, "options": {"num_predict": 200}
  }'
{
  "message": {
    "role": "assistant",
    "content": "17 * 23 = 391.",
    "thinking": "17 * 20 = 340. 17 * 3 = 51. 340 + 51 = 391."
  },
  "done": true, "done_reason": "stop"
}

2 · OpenAI-compatible API

Chat Completions (non-streaming)

curl -X POST http://localhost:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-Q8_0.gguf",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is 2+3?"}
    ],
    "max_tokens": 50, "temperature": 0.7
  }'
{
  "id": "chatcmpl-abc123...",
  "object": "chat.completion",
  "model": "Qwen3.5-9B-Q8_0.gguf",
  "choices": [{
    "index": 0,
    "message": {"role": "assistant", "content": "2 + 3 = 5."},
    "finish_reason": "stop"
  }],
  "usage": {
    "prompt_tokens": 20, "completion_tokens": 8, "total_tokens": 28,
    "prompt_tokens_details": { "cached_tokens": 0 }
  }
}

usage.prompt_tokens_details.cached_tokens follows OpenAI's KV-cache-hit extension — on a follow-up turn that shares a prefix it approaches prompt_tokens.

Chat Completions (streaming)

curl -X POST http://localhost:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{ "model": "Qwen3.5-9B-Q8_0.gguf", "messages": [{"role": "user", "content": "Hello!"}], "max_tokens": 50, "stream": true }'

Each chunk is an SSE data: frame of object: "chat.completion.chunk"; the stream ends with data: [DONE].

Image input (OpenAI format)

IMG_B64=$(base64 < photo.png)
curl -X POST http://localhost:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d "{
    \"model\": \"gemma-4-E4B-it-Q8_0.gguf\",
    \"messages\": [{
      \"role\": \"user\",
      \"content\": [
        {\"type\": \"text\", \"text\": \"What is in this image?\"},
        {\"type\": \"image_url\", \"image_url\": {\"url\": \"data:image/png;base64,$IMG_B64\"}}
      ]
    }],
    \"max_tokens\": 200
  }"

Responses API (/v1/responses)

curl -X POST http://localhost:5000/v1/responses \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-Q8_0.gguf",
    "instructions": "You are a helpful assistant.",
    "input": "What is 2+3?",
    "max_output_tokens": 50
  }'

input is a string or an array of message items; only message items are read, and other item types (function_call, function_call_output, reasoning) are skipped and logged. instructions becomes the system message, max_output_tokens falls back to --max-tokens, a reasoning object turns thinking on (reasoning.effort is the Responses spelling of reasoning_effort), and text.format requests structured output. On this route text.format cannot be combined with reasoning or with tools (both 400, on every family); the delayed-grammar exception in §5 applies only to Chat Completions response_format. The sampling fields, tools, skills, skills_discovery and multi_agent work as on Chat Completions. "stream": true sends named SSE events (response.created, response.output_text.delta, …, response.completed). The server is stateless per request: previous_response_id is refused with 400, so send the whole conversation in input. store (default true) only keeps the finished response so GET /v1/responses/{id} can return it again — from memory, or from Redis when TS_RESPONSES_STORE_REDIS_URL (or --redis-url) is set.

Jev typed decisions (/v1/systemone)

When the hosted model is DiffusionGemma, POST /v1/systemone answers Jev-style typed questions about a state (a string, object or array) from one seeded denoising read per sample instead of generated text: noul (probability of yes), choice (2–26 named options) and score (ordered levels, answered as an expected zero-based index). Each question is {"type", "instructions", "criteria"}: criteria maps option names to descriptions for choice, is an ordered array of levels for score, and is an optional {"true": …, "false": …} object for noul; any other question field gets 422. model may name the loaded file, jev-latest or jev-preview — aliases for the loaded model, not separate checkpoints.

curl -X POST http://localhost:5000/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-latest",
    "state": "I was charged twice for the same subscription this month.",
    "questions": { "billing": { "type": "noul", "instructions": "Is this a billing issue?" } },
    "samples": 1, "seed": 42
  }'

The reply carries answers.<id> (noul; or choice / score with probabilities and confidence), usage and diagnostics. The probabilities are softmaxes over the listed labels only, not calibrated likelihoods of being correct. A request takes up to 64 questions and up to 8 inline images (base64 or data: URLs, stored under the upload limits; they need the vision tower passed with --mmproj, otherwise the request gets 503); the other fields are instructions, samples (1–32, or "auto", the default), auto_max, auto_threshold, seed, chunk_rows and chunk_prompt, while steps and think accept only 1 and 0. Errors: 415 when the Content-Type is not application/json (multipart uploads included), 400 for malformed JSON, 413 over the 8 MiB body cap (TS_JEV_MAX_BODY_MB, 1–64), 413 / 507 when an inline image exceeds the upload per-file cap or quota, 422 for an invalid request (audio and video included), 404 for an unknown model name, 503 when no DiffusionGemma model is loaded, and 529 with Retry-After: 1 once TS_JEV_MAX_PENDING (default 32) requests are already admitted, running or queued. → Jev guide

3 · Tool calling over HTTP

Send a tools array; the server detects the architecture's wire format and returns structured tool_calls. A hosted DiffusionGemma model has no tool-call channel and answers a request carrying tools with 400; Mistral 3 and Hunyuan Dense never render tool declarations.

curl -X POST http://localhost:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-Q8_0.gguf",
    "messages": [{"role": "user", "content": "What is the weather in Paris?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Get current weather for a city.",
        "parameters": {
          "type": "object",
          "properties": {
            "city":  {"type": "string"},
            "units": {"type": "string", "enum": ["c", "f"]}
          },
          "required": ["city"]
        }
      }
    }],
    "max_tokens": 200
  }'
{
  "choices": [{
    "message": {
      "role": "assistant", "content": null,
      "tool_calls": [{
        "id": "call_abc123", "type": "function",
        "function": { "name": "get_weather", "arguments": "{\"city\":\"Paris\",\"units\":\"c\"}" }
      }]
    },
    "finish_reason": "tool_calls"
  }]
}

Continue the loop by appending the assistant tool_calls plus a {"role": "tool", "tool_call_id": "...", "content": "..."} message with the function result, then calling the endpoint again. (The Ollama endpoint uses the same flow with a role: "tool" message.)

4 · Agent Skills over HTTP

An Agent Skill is a folder of model-facing instructions — a SKILL.md plus the scripts, reference documents and assets it refers to — that the server loads only when a task needs it. Every chat surface — /v1/chat/completions, /v1/responses, the Ollama-compatible /api/chat/ollama, and the Web UI's /api/chat — takes the same two optional fields:

"skills": ["pdf", "xlsx"],
"skills_discovery": true

skills_discovery defaults to true: the model is also shown the names and descriptions of the skills the request did not select, so it can pick up one you did not think to name. Set it to false to restrict the request to exactly the skills it listed. Naming a skill the server does not have is a 400 — GET /api/skills lists what is registered.

For a tool-capable family, the initial prompt carries metadata only — including for explicitly selected skills. The model reads a matching SKILL.md with skills_read when it activates that skill. If a family cannot render and parse a tool round trip, selected bodies are inlined and the discovery catalog is dropped instead; this is the current path for Mistral 3, Hunyuan Dense, DiffusionGemma and architectures without a registered chat protocol. Qwen 3.8 Flash Next (qwen4exp) now parses tool calls and takes the metadata path.

# OpenAI-compatible
curl -X POST http://localhost:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-E4B-it-Q8_0.gguf",
    "messages": [{"role": "user", "content": "Pull the totals table out of this statement."}],
    "skills": ["pdf"],
    "max_tokens": 600
  }'

# Ollama-compatible
curl -X POST http://localhost:5000/api/chat/ollama \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-E4B-it-Q8_0.gguf",
    "messages": [{"role": "user", "content": "Build me a budget spreadsheet."}],
    "skills": ["xlsx"],
    "skills_discovery": false,
    "stream": false
  }'

# Web UI SSE
curl -N -X POST http://localhost:5000/api/chat \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Summarize this PDF."}], "skills": ["pdf"], "maxTokens": 400}'

The reply is an ordinary completion. The model reaches the rest of a skill through built-in skills_list / skills_read tools that the server answers itself, in process and next to the weights, so the client never receives a tool call it has no implementation for and an existing OpenAI SDK needs no changes (pass the fields as extra_body). Skills also combine with your own tools: the skill tools are merged into the list you sent, the server answers only its own, and a call to one of yours comes back as usual — with whatever the model read from a skill already folded into the conversation. When scripts are enabled, reading a SKILL.md that bundles a script also returns a short host note with a ready skills_run example, and skills_run takes args as a JSON array of strings (also when the array arrives encoded in a string) or as shell-style text.

📌

Each built-in lookup or code action may cause another generation, and every round streams. The server consumes its own tool markup and forwards content/reasoning as they decode, so the client receives a finished answer rather than an internal call. --skills-max-rounds bounds the loop: default 8, automatically 24 when code execution is offered, with any explicitly configured value preserved.

The Web UI stream carries one frame per executed skill tool call, as it completes:

data: {"skill_step":"skills_read","agent_id":"/root","skill":"pdf","detail":"references/forms.md","ok":true,"round":1,"files":null}

skill_step is the built-in that ran, agent_id names the agent that made the call (/root unless a sub-agent did), skill/detail identify its target, ok reports success, round is one-based, and files carries any generated artifacts as {name, bytes, url}. → Web UI SSE protocol

Agentic code execution

Code execution is a server startup capability, not a request switch. With --code-exec, every tool-capable chat surface may use four server-owned tools: read_file, write_file, shell, and apply_patch. They run inside the same bounded loop as skills and are never forwarded to the client; caller-defined tools remain the caller's responsibility. There is no per-command approval API. → Flags, sandbox and trust model

Each OpenAI or Ollama request gets a unique private workspace shared across its internal rounds and deleted on completion, cancellation or failure. A named Web UI session keeps its workspace across chat turns; skill scripts and code tools share it. Files retained as artifacts can be listed at GET /api/code/artifacts/{runId} and downloaded from GET /api/code/artifacts/{runId}/{*path}. Downloads are forced attachments with nosniff.

🔒

Both script and command execution are off by default, and OS sandboxing defaults to required (sandbox or refuse). macOS Seatbelt and Linux bubblewrap 0.12+ confine writes, home reads and default-off networking, although macOS cannot guarantee cleanup of a deliberately detached child. Windows job objects confine the process tree but not files or network, so code needs explicit --code-exec-unconfined and skill scripts need --skills-sandbox preferred. --code-exec-allow-network grants generated commands unrestricted host IP access, including LAN/loopback; it is separate from host-performed installs and from --skills-allow-network.

Sub-agents (multi_agent)

On tool-capable families — every family that renders tool declarations and parses tool calls, so not Mistral 3, Hunyuan Dense or DiffusionGemma — /v1/chat/completions, /v1/responses, /api/chat/ollama and the Web UI's /api/chat also offer five server-owned coordination tools: spawn_agent, wait_agent, send_input, close_agent and list_agents. The model may use them to hand bounded sub-tasks to child agents that run on the same loaded model. This is on by default, independent of skills and code execution, and absent from structured-output requests and from /api/generate; the model decides whether to delegate. explorer and reviewer children are read-only (skills_list, skills_read, read_file and the coordination tools), and so is a worker unless the server starts with --agents-allow-worker-tools. Children never receive the client's own tools, host tool calls within one request's tree run one at a time, and each child works on its own copy of the KV state. The client sees only the parent's answer; reported token usage includes the children. No latency or quality figures for delegation are published.

"multi_agent": false

Only a JSON false changes anything: it turns delegation off for that request. A request cannot turn it on when the server started with --no-multi-agent (or TS_NO_MULTI_AGENT), and cannot raise any --agents-* limit. → Roles and limits

Managing the registry

Two shapes of the same data: the read-only, OpenAI-flavoured /v1/skills, and the Web UI's /api/skills, which adds load errors, single-file reads, upload and delete.

EndpointPurpose
GET /v1/skillsEvery registered skill, wrapped as {"object": "list", "data": [...]}.
GET /v1/skills/{name}One skill, with the SKILL.md body added as instructions.
GET /api/skillsThe same objects as {"enabled": …, "installable": …, "skills": [ … ], "errors": [{"path", "message"}]} — so a management UI can also show the directories that looked like a skill and failed to load, with the reason.
GET /api/skills/{name}One skill, including its instructions.
GET /api/skills/{name}/files/{*path}One bundled file, resolved through the same path guard the model's own reads go through and always served as text/plain with nosniff: a skill may ship an .html or .js file, and serving it with its real type would execute uploaded content in the server's own origin.
POST /api/skillsInstall a skill from a multipart .zip (file=@pdf.zip), optionally with overwrite=true to replace an installed skill of the same name. Answers 201 with the installed skill object.
POST /api/skills/rescanPick up changes made on disk without a restart; answers with the same body as GET /api/skills.
DELETE /api/skills/{name}Remove an installed skill — {"removed":true}. Only a skill uploaded here ("origin": "installed") can be deleted; one discovered under an operator-configured directory is that operator's own file tree and is refused.
curl http://localhost:5000/v1/skills
curl http://localhost:5000/v1/skills/pdf

curl http://localhost:5000/api/skills
curl http://localhost:5000/api/skills/pdf
curl http://localhost:5000/api/skills/pdf/files/references/forms.md

# Upload a .zip of the skill folder (pdf/SKILL.md) or of its contents (SKILL.md at the root)
curl -X POST http://localhost:5000/api/skills -F "file=@pdf.zip" -F "overwrite=true"

curl -X POST http://localhost:5000/api/skills/rescan
curl -X DELETE http://localhost:5000/api/skills/pdf   # {"removed":true}

A skill object:

{
  "id": "pdf",
  "object": "skill",
  "name": "pdf",
  "description": "Extract text and tables from PDF files, fill in PDF forms, and merge or split documents...",
  "license": "Apache-2.0",
  "compatibility": "Requires python3 with pypdf installed.",
  "files": [
    {"path": "scripts/extract_tables.py", "bytes": 4021, "kind": "script", "text": true},
    {"path": "references/forms.md", "bytes": 18233, "kind": "reference", "text": true}
  ],
  "bytes": 41288,
  "origin": "installed",
  "warnings": [],
  "modified": "2026-08-29T12:00:00Z"
}

kind is one of script / reference / asset / manifest / other; origin is discovered for a skill found by scanning a configured directory and installed for one uploaded here; warnings carries anything that loaded despite being out of spec (a name that disagrees with its directory, a description over the 1024-character limit). GET /api/models reports the feature itself as "skills": {"enabled": true, "installable": true, "allowScripts": false, "count": 7} (allowScripts is whether skills_run is offered), or null when the server has skills disabled — which is how the Web UI decides whether to show the control at all. Errors follow the server's per-prefix convention: /v1/* returns {"error": {"message": "…", "type": "invalid_request_error"}} and /api/* returns {"error": "…"}.

🔒

An uploaded skill is untrusted content, and it is validated before anything lands: every ZIP entry uses the same path guard as model reads, decompressed sizes are enforced, and archives are capped at 4096 files, 64 MB per file, 256 MB total and 200× compression. skills_run remains absent unless the server starts with --skills-allow-exec; under the default required policy it runs in a qualifying OS sandbox or refuses. This still authorizes model-selected arbitrary code and belongs only on a host whose users and skills are trusted. → Skill flags

5 · Structured outputs (JSON schema)

The OpenAI response_format accepts text, json_object, and validated json_schema. The server injects strict JSON instructions and validates the output before returning it. It also compiles the schema (or a generic JSON-object grammar) into a grammar that constrains decoding token by token. A schema the grammar cannot express falls back to constraining only the first sampled token to a {-opening candidate, so chatty models still cannot ramble prose before the object.

curl -X POST http://localhost:5000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.5-9B-Q8_0.gguf",
    "messages": [
      {"role": "system", "content": "You are a concise extraction assistant."},
      {"role": "user", "content": "Extract the city and country from: Paris, France."}
    ],
    "response_format": {
      "type": "json_schema",
      "json_schema": {
        "name": "location_extraction", "strict": true,
        "schema": {
          "type": "object",
          "properties": {
            "city": { "type": "string" },
            "country": { "type": "string" },
            "confidence": { "type": ["string", "null"] }
          },
          "required": ["city", "country", "confidence"],
          "additionalProperties": false
        }
      }
    },
    "max_tokens": 120
  }'
⚠️

json_schema cannot be combined with caller tools. think is refused too, except on families whose reasoning the grammar can wait out — Gemma 4, Qwen 3.8 Flash Next, GPT-OSS, Muse-Glimmer, DeepSeek V4.1, glm5next and Nemotron-H — where the grammar starts once the reasoning ends. It can be combined with skills, but the selected skill bodies are inlined and all built-in skill/code tools are suppressed for that request because schema-constrained output cannot emit their call markup. Invalid schemas return HTTP 400; outputs that fail validation return HTTP 422.

6 · Web UI SSE protocol (/api/chat)

This is the protocol the bundled chat UI uses; external UIs can plug into the same endpoint. Every event is a JSON object delivered as a single data: … SSE frame.

# Create / dispose a session (Web UI flow only)
curl -X POST http://localhost:5000/api/sessions          # {"sessionId":"a3b1c2..."}
curl -X DELETE http://localhost:5000/api/sessions/a3b1c2...

# Streaming chat
curl -N -X POST http://localhost:5000/api/chat \
  -H "Content-Type: application/json" \
  -d '{ "messages": [{"role": "user", "content": "Hi"}], "maxTokens": 50, "sessionId": null, "newChat": false, "think": false, "tools": [] }'

Reusing the same sessionId splices prior assistant tokens into the KV-cache prefix on the next turn. The terminal frame reports reuse:

data: {"done":true,"tokenCount":187,"elapsed":2.143,"tokPerSec":87.23,"promptTokens":512,"kvReusedTokens":420,"kvReusePercent":82.0}

Event fields include token (content), thinking (reasoning chunk), caller-owned tool_calls, replace/diffusionStep (DiffusionGemma previews), skill_step (completed built-in plus round/files), tool_progress (transient writing / running / finished activity with tool, text, seconds, detail, agents), and the terminal done summary. Every skill_step frame also carries agent_id: /root for the parent, which is the value on every frame when nothing is delegated, and a path such as /root/<task_name> for a child. Every tool_progress frame carries agents, which is null except while wait_agent is running; then it is an array with one entry per sub-agent in the request (the parent /root is not listed) with agent_id, parent_id, task, agent_type, status, tool, tool_status, detail, result and error — which the bundled UI renders as a live sub-agent panel. A one-shot {"artifact_verified": true, "files": [{name, bytes, url}]} frame names a generated deliverable once it has passed the host's checks. The browser clears completed activity rather than building a tool-history transcript, but keeps artifact download chips.

data: {"tool_progress":"writing","tool":"shell","text":"{\"command\":\"python","seconds":0,"detail":null,"agents":null}
data: {"tool_progress":"running","tool":"shell","text":"writing report.xlsx\n","seconds":2.1,"detail":"python make_report.py","agents":null}
data: {"tool_progress":"finished","tool":"shell","text":null,"seconds":2.4,"detail":null,"agents":null}

Uploads (/api/upload)

POST /api/upload takes a multipart form (first file used) and returns the stored path/URL plus per-type metadata; /api/chat messages then reference the returned server paths (imagePaths, audioPaths, …). Accepted types: images (.png .jpg .jpeg .gif .webp .bmp .heic .heif), video (.mp4 .mov .avi .mkv .webm — frames extracted at VIDEO_SAMPLE_FPS), audio (.mp3 .wav .ogg .flac .m4a), PDF, and text/code files (returned in full as textContent). Born-digital PDFs return their complete extracted textContent; scanned PDFs return page images like video frames when a vision model is hosted (otherwise a needsVision warning). TS_PDF_MAX_PAGES caps the pages read (default: all). Files land in uploads/ beside the server binary (TENSORSHARP_UPLOAD_DIR moves it) and are served back under /uploads/. --upload-max-mb caps each client-supplied file, uploads and decoded base64 attachments alike (default 500; over it answers 413); --upload-quota-mb caps the whole directory, generated outputs included (over it answers 507), and --upload-ttl-hours deletes older files — both off by default.

Image generation & editing (/api/image-generate · /api/image-edit · /stream)

Available when the hosted model is a Qwen-Image-2.1 qwen_image DiT. With any other model, /api/image-edit answers 400 with The loaded model is not a Qwen-Image-2.1 model. and /api/image-generate answers 400 with Text-to-image generation requires a Qwen-Image-2.1 model.; the /stream variants send the same message in their terminal {"done": true, "error": …} event instead. POST /api/image-generate takes JSON such as {"prompt": "...", "width": 2048, "height": 2048, "steps": 40, "cfg": 1, "seed": 42} and generates an image. POST /api/image-edit runs one edit and needs at least one reference — multipart (one or more image parts, in reference order, plus prompt) or JSON whose imagePaths (or the older single imagePath) reference files in the upload directory. Both accept negativePrompt, targetArea, width, height, steps, cfg and seed: omitted values, or steps / cfg of 0, select the model defaults (40 steps, CFG 1, and 2048×2048 or about that area at the first reference's aspect ratio — unless the server was started with both --width and --height, in which case that default is used only when the request names neither dimensions nor targetArea; off-grid startup defaults are rounded down with a warning), except that a --lora plug-in with a sampling recipe, loaded at server startup, supplies that recipe's steps and CFG (see LoRA plug-ins); explicit dimensions must be multiples of 32 and take precedence over targetArea. Both return { ok, url, width, height, elapsedSeconds } with the PNG served under /uploads/. POST /api/image-generate/stream and POST /api/image-edit/stream take the same JSON body but stream SSE denoising progress: per-step frames {"imageGenerate": true, …} or {"imageEdit": true, "step": i, "total": N, "image": "data:image/png;base64,..."} (up to 8 evenly spaced previews), then a final {"done": true, "url": "...", "width": ..., "height": ..., "elapsedSeconds": ...}, or {"done": true, "error": …}. Image requests are serialized process-wide; the server Web UI uses these image routes, while TensorAgent dispatches image turns through its own /api/chat host.

Masked edits. Multipart /api/image-edit accepts one mask file. For JSON and /api/image-edit/stream, upload the source and mask through /api/upload, then pass imagePaths and maskPath using the returned filenames. Optional fields are maskMode (grayscale, white edits; or alpha, transparent edits), maskInvert, maskFeather (0–1024), maskCrop and maskCropPadding (0–16384, default 64). The mask must match the first reference's dimensions; the result retains that canvas and protected pixels, including in previews. A mask needs an input image and Qwen-Image-2.1; generation routes refuse it, and mask options without a mask are rejected. Selection workflow and limitations →

Video generation (/api/video-generate · /api/video-generate/stream · /v1/videos/generations)

Available when the hosted model generates video — MiniMax-H3, or Wan for video alone. The gate is the IVideoGenerationModel seam, not an architecture string, so anything else answers 400 with The loaded model is not a video-generation model. A request the model itself refuses — the wrong checkpoint for the mode asked for (Ref2VA inputs against an FL2VA denoiser), keyframes and named references in the same body, a mode without its inputs — comes back from /api/video-generate as a 400 carrying the model's own message rather than a generic 500, and from the streaming route as that same message in its terminal {"done": true, "error": …} event. All three routes share one parser, so they take the same fields, and generations are serialized process-wide.

# MiniMax-H3: prompt -> MP4 plus a 32 kHz stereo WAV, generated together
curl -s http://localhost:5000/api/video-generate \
  -H "Content-Type: application/json" \
  -d '{
        "prompt": "a red fox trotting through falling snow, cinematic",
        "width": 640, "height": 384, "frames": 22, "fps": 24,
        "steps": 8, "cfg": 1.0, "seed": 42,
        "videoMode": "t2v", "generateAudio": true
      }'
{ "ok": true, "url": "/uploads/video-8f3c….mp4", "audioUrl": "/uploads/video-8f3c….wav",
  "width": 640, "height": 384, "frames": 22, "fps": 24,
  "seed": 42, "codec": "h264", "elapsedSeconds": 63.1 }

Request fields (camelCase): prompt (required), width, height, frames, steps, cfg, cfg2, seed, fps, flowShift, negativePrompt, sampler, cfgCacheStride, videoMode (t2v / i2v / fl2v / ref), generateAudio, imagePath, image, endImage, referenceImages, referenceVideos, referenceAudios, referenceVideoAudios. The seven added for joint audio-video and reference conditioning also accept snake_case — video_mode, generate_audio, end_image, reference_images, reference_videos, reference_audios, reference_video_audios — with camelCase winning when both are present; the rest are camelCase only.

image is the inline base64 form (a data:…;base64, prefix is accepted). imagePath, endImage and every reference* entry must name a file previously returned by /api/upload and are confined to the upload directory — anything outside it is rejected with a message such as referenceImages entries must reference previously uploaded files. referenceVideoAudios pairs by index with referenceVideos: entry i is clip i's soundtrack. audioUrl is null when the model produced no track — the audio is written as a sidecar WAV rather than muxed into the MP4.

POST /api/video-generate/stream takes the same body and streams SSE: {"videoGen": true, "step": i, "total": N, "phase": …, "detail": …, "elapsedSeconds": …, "etaSeconds": …} per tick, then a terminal {"done": true, "url": …, "audioUrl": …, "width": …, "height": …, "frames": …, "fps": …, "seed": …, "codec": …, "elapsedSeconds": …} — or {"done": true, "error": …}.

POST /v1/videos/generations is the OpenAI-images-shaped envelope over the same job. On top of the fields above it accepts "size": "832x480" (split into width/height), negative_prompt, and response_format of "url" or "b64_json", and answers { created, data: [{ url, b64_json }], audio_url, width, height, frames, fps, seed, codec, elapsed_seconds } with the MP4 served under /uploads/.

GET /api/models is how a client learns what the hosted video model will accept before it builds a request: alongside the model lists it returns a video object carrying family ("minimax-h3" or "wan"), supportsAudio, supportsImageConditioning, supportsEndImageConditioning, supportsReferenceConditioning and maxReferenceImages (9 on MiniMax-H3 Ref2VA) — and null for every non-video model. The bundled Web UI reads exactly that block to decide which attachment controls to offer — a first frame, a last frame, or reference stills capped at maxReferenceImages — instead of pattern-matching an architecture string; with no block at all (an older server, or a model that reports none) it falls back to the single-image request it has always sent.

7 · Sampling parameters

Ollama-style (inside the options object)

The defaults shown apply when a request omits the field — they are the server-wide defaults, configurable via server flags / TENSORSHARP_* env vars (see Server options).

ParameterTypeDefaultDescription
num_predictint--max-tokens (20000)Maximum tokens to generate
temperaturefloat0.8Sampling temperature (0 = greedy)
top_kint40Top-K filtering (0 = disabled)
top_pfloat0.9Nucleus sampling threshold (1.0 = disabled)
min_pfloat0Minimum probability filtering
repeat_penaltyfloat1.1Repetition penalty (1.0 = none)
repeat_last_nint64How many recent tokens the repeat / presence / frequency penalties look back over
presence_penalty / frequency_penaltyfloat0Presence / frequency penalties
seedint-1Random seed (-1 = random)
stoparraynullStop sequences

OpenAI-style (top-level)

max_tokens or max_completion_tokens (default --max-tokens, 20000, when omitted), temperature, top_p, presence_penalty, frequency_penalty, seed, stop (string or array), and response_format (text / json_object / json_schema). Beyond the OpenAI spec, the same top level also accepts top_k, min_p, repeat_penalty (or repetition_penalty) and repeat_last_n, with the defaults in the table above.

8 · Python client examples

Using requests (Ollama-style)

import requests

resp = requests.post("http://localhost:5000/api/generate", json={
    "model": "Qwen3.5-9B-Q8_0.gguf",
    "prompt": "What is machine learning?",
    "stream": False,
    "options": {"num_predict": 100, "temperature": 0.7}
})
print(resp.json()["response"])

Using the openai SDK

from openai import OpenAI

client = OpenAI(base_url="http://localhost:5000/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="Qwen3.5-9B-Q8_0.gguf",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "What is 2+3?"}
    ],
    max_tokens=50, temperature=0.7
)
print(response.choices[0].message.content)

Streaming with the openai SDK

stream = client.chat.completions.create(
    model="Qwen3.5-9B-Q8_0.gguf",
    messages=[{"role": "user", "content": "Tell me about Python."}],
    max_tokens=200, stream=True
)
for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
print()