Embedding Models & Semantic Search
Turn text and code into vectors for semantic search and RAG. TensorSharp runs GGUF BERT / XLM-R encoders locally and serves embeddings through OpenAI- and Ollama-compatible HTTP APIs. This feature is in current source; older release archives may not include it.
Models
| Model | Dimensions | GGUF token limit | Pooling |
|---|---|---|---|
| Snowflake Arctic Embed L v2.0 Q8_0 | 1024 | 8192 | CLS |
| all-MiniLM-L6-v2 Q8_0 | 384 | 512 | Mean |
Pinned downloads, checksums, and the complete protocol: embedding guide · model downloads.
Start the service
Embedding execution supports 100% pure C# cpu, plus native ggml_cpu, ggml_metal, and ggml_cuda. Pure C# execution needs no native inference library; native execution requires the corresponding backend build. Each process keeps one encoder resident; a chat service uses a separate port.
dotnet build TensorSharp.Server.Host -c Release \
-p:TensorSharpSkipGgmlNative=true -p:TensorSharpSkipMlxNative=true
dotnet TensorSharp.Server.Host/bin/TensorSharp.Server.Host.dll \
--model models/embeddings/snowflake-arctic-embed-l-v2.0-q8_0.gguf \
--embeddings --backend cpu --embedding-threads 8 \
--host 127.0.0.1 --port 5000 --no-webui
--embedding-context-size N reduces the per-input token limit; 0 uses model metadata. /v1/models, /api/tags, /api/show report the hosted model and its embedding capability. An embedding server does not generate: a POST to /v1/chat/completions, /v1/responses, /v1/systemone, /v1/videos/generations, /api/generate, /api/chat, /api/chat/ollama, /api/models/load, or /api/image-generate, /api/image-edit and /api/video-generate (and their /stream forms) returns HTTP 400 pointing at /v1/embeddings or /api/embed. An explicit --backend the machine does not have exits 2 with error: model load refused: …. --embeddings also requires --model and rejects --mmproj.
HTTP APIs
curl http://127.0.0.1:5000/v1/embeddings \
-H 'Content-Type: application/json' \
-d '{"model":"snowflake-arctic-embed-l-v2.0-q8_0","input":["query: read a file","def read_file(path): return open(path).read()"],"dimensions":256}'
curl http://127.0.0.1:5000/api/embed \
-H 'Content-Type: application/json' \
-d '{"model":"snowflake-arctic-embed-l-v2.0-q8_0","input":["query: read a file","open a document"],"truncate":false}'
OpenAI input accepts a string, a string array, one token-ID array, or an array of token-ID arrays; token inputs must already include this model’s special tokens. encoding_format is float or base64 (little-endian float32). Responses include indices, vectors, and token usage. A request accepts up to 2048 inputs and 262144 total tokens. Embedding and model-show JSON bodies are limited to 16 MiB; oversized bodies return HTTP 413.
OpenAI rejects inputs beyond the context limit. Ollama /api/embed accepts a string or string array; truncate defaults to true and preserves the final special token, while false rejects overflow. Its response contains embeddings, nanosecond timings, and prompt_eval_count.
Concurrent callers are grouped from requests already waiting, in FIFO order, up to 64 sequences and 4096 tokens per encoder call. The first request starts immediately; there is no batching delay. Larger whole requests run alone. Each caller retains its own result order and options; canceling one caller does not cancel healthy callers. See the request scheduling contract.
Retrieval quality and performance
Prefix Snowflake queries with query: and leave documents unprefixed. Vectors have unit L2 norm, so a dot product gives cosine similarity. dimensions crops the vector and normalizes it again; Snowflake’s Matryoshka training supports reduction to 256 dimensions, while arbitrary reductions on other models can lower quality. MiniLM’s upstream default sentence length is shorter than this GGUF’s 512-position capacity. Chunk code at function/class boundaries and rebuild the index when changing models or chunking.
Both implementations use full bidirectional attention and model-declared pooling. Each pure C# model owns a persistent worker pool; ARM Q8 projections use four-token × four-output signed-dot SIMD tiles. Managed FP32 attention transposes keys, computes four queries with adjacent keys in SIMD lanes, and packs values into channel tiles (four channels on ARM) for fused multiply-add. Portable vector-width and four-query dot-product fallbacks cover other processors and head sizes; padded key strides cover tails; the long path uses 64-query × 128-key tiles with ARM four-query × sixteen-output SIMD reuse and online FP32 softmax and bounded per-head scratch. Short and fallback paths keep four temporary score rows. The native executor packs heterogeneous GPU batches with sequence-isolating masks, packs CPU projections with independently padded attention, and reuses graphs.
Tokenization is checked against independent HuggingFace and llama.cpp references, including multilingual Unicode; the full guide records upstream differences. The native-free host check removes custom native assets and inspects loaded libraries after inference. The managed x86 path has guarded AVX/AVX2/AVX-512 support; the four-query × sixteen-column kernel is ARM-specific. Current timing evidence covers Apple Silicon, with x86 performance unmeasured. Performance conclusions come from measurements on identical GGUFs, hardware, and requests.
Validation results and limitations: docs/validation/embeddings-2026-09/README.md (local validation evidence, not committed) · Reproduce the HTTP benchmark
C# integration
Backend = "CPU" is the default pure C# encoder. Use "GGML_CPU" for native CPU execution, "METAL" for native Metal, or "CUDA" for native CUDA.
using TensorSharp.Models.Embeddings;
using var model = EmbeddingModel.Load("models/embeddings/all-MiniLM-L6-v2-Q8_0.gguf",
new EmbeddingModelOptions { Backend = "CPU", Threads = 8 });
EmbeddingBatchResult batch = await model.EmbedAsync(new[] { "read a file", "open a document" });
float[] vector = batch.Embeddings[0];