Embedding Models & Semantic Search

Turn text and code into vectors for semantic search and RAG. TensorSharp runs GGUF BERT / XLM-R encoders locally and serves embeddings through OpenAI- and Ollama-compatible HTTP APIs. This feature is in current source; older release archives may not include it.

Models

ModelDimensionsGGUF token limitPooling
Snowflake Arctic Embed L v2.0 Q8_010248192CLS
all-MiniLM-L6-v2 Q8_0384512Mean

Pinned downloads, checksums, and the complete protocol: embedding guide · model downloads.

Start the service

Embedding execution supports 100% pure C# cpu, plus native ggml_cpu, ggml_metal, and ggml_cuda. Pure C# execution needs no native inference library; native execution requires the corresponding backend build. Each process keeps one encoder resident; a chat service uses a separate port.

dotnet build TensorSharp.Server.Host -c Release \
  -p:TensorSharpSkipGgmlNative=true -p:TensorSharpSkipMlxNative=true
dotnet TensorSharp.Server.Host/bin/TensorSharp.Server.Host.dll \
  --model models/embeddings/snowflake-arctic-embed-l-v2.0-q8_0.gguf \
  --embeddings --backend cpu --embedding-threads 8 \
  --host 127.0.0.1 --port 5000 --no-webui

--embedding-context-size N reduces the per-input token limit; 0 uses model metadata. /v1/models, /api/tags, /api/show report the hosted model and its embedding capability. An embedding server does not generate: a POST to /v1/chat/completions, /v1/responses, /v1/systemone, /v1/videos/generations, /api/generate, /api/chat, /api/chat/ollama, /api/models/load, or /api/image-generate, /api/image-edit and /api/video-generate (and their /stream forms) returns HTTP 400 pointing at /v1/embeddings or /api/embed. An explicit --backend the machine does not have exits 2 with error: model load refused: …. --embeddings also requires --model and rejects --mmproj.

HTTP APIs

curl http://127.0.0.1:5000/v1/embeddings \
  -H 'Content-Type: application/json' \
  -d '{"model":"snowflake-arctic-embed-l-v2.0-q8_0","input":["query: read a file","def read_file(path): return open(path).read()"],"dimensions":256}'

curl http://127.0.0.1:5000/api/embed \
  -H 'Content-Type: application/json' \
  -d '{"model":"snowflake-arctic-embed-l-v2.0-q8_0","input":["query: read a file","open a document"],"truncate":false}'

OpenAI input accepts a string, a string array, one token-ID array, or an array of token-ID arrays; token inputs must already include this model’s special tokens. encoding_format is float or base64 (little-endian float32). Responses include indices, vectors, and token usage. A request accepts up to 2048 inputs and 262144 total tokens. Embedding and model-show JSON bodies are limited to 16 MiB; oversized bodies return HTTP 413.

OpenAI rejects inputs beyond the context limit. Ollama /api/embed accepts a string or string array; truncate defaults to true and preserves the final special token, while false rejects overflow. Its response contains embeddings, nanosecond timings, and prompt_eval_count.

Concurrent callers are grouped from requests already waiting, in FIFO order, up to 64 sequences and 4096 tokens per encoder call. The first request starts immediately; there is no batching delay. Larger whole requests run alone. Each caller retains its own result order and options; canceling one caller does not cancel healthy callers. See the request scheduling contract.

Retrieval quality and performance

Prefix Snowflake queries with query: and leave documents unprefixed. Vectors have unit L2 norm, so a dot product gives cosine similarity. dimensions crops the vector and normalizes it again; Snowflake’s Matryoshka training supports reduction to 256 dimensions, while arbitrary reductions on other models can lower quality. MiniLM’s upstream default sentence length is shorter than this GGUF’s 512-position capacity. Chunk code at function/class boundaries and rebuild the index when changing models or chunking.

Both implementations use full bidirectional attention and model-declared pooling. Each pure C# model owns a persistent worker pool; ARM Q8 projections use four-token × four-output signed-dot SIMD tiles. Managed FP32 attention transposes keys, computes four queries with adjacent keys in SIMD lanes, and packs values into channel tiles (four channels on ARM) for fused multiply-add. Portable vector-width and four-query dot-product fallbacks cover other processors and head sizes; padded key strides cover tails; the long path uses 64-query × 128-key tiles with ARM four-query × sixteen-output SIMD reuse and online FP32 softmax and bounded per-head scratch. Short and fallback paths keep four temporary score rows. The native executor packs heterogeneous GPU batches with sequence-isolating masks, packs CPU projections with independently padded attention, and reuses graphs.

Tokenization is checked against independent HuggingFace and llama.cpp references, including multilingual Unicode; the full guide records upstream differences. The native-free host check removes custom native assets and inspects loaded libraries after inference. The managed x86 path has guarded AVX/AVX2/AVX-512 support; the four-query × sixteen-column kernel is ARM-specific. Current timing evidence covers Apple Silicon, with x86 performance unmeasured. Performance conclusions come from measurements on identical GGUFs, hardware, and requests.

Validation results and limitations: docs/validation/embeddings-2026-09/README.md (local validation evidence, not committed) · Reproduce the HTTP benchmark

C# integration

Backend = "CPU" is the default pure C# encoder. Use "GGML_CPU" for native CPU execution, "METAL" for native Metal, or "CUDA" for native CUDA.

using TensorSharp.Models.Embeddings;

using var model = EmbeddingModel.Load("models/embeddings/all-MiniLM-L6-v2-Q8_0.gguf",
    new EmbeddingModelOptions { Backend = "CPU", Threads = 8 });
EmbeddingBatchResult batch = await model.EmbedAsync(new[] { "read a file", "open a document" });
float[] vector = batch.Embeddings[0];