Quick Run jina-embeddings-v5-text-nano Windows 11 Quantized GGUF Windows

Using a native PowerShell script is the absolute quickest way to install this model.

Proceed by following the technical instructions below.

An automated background process downloads all required large-scale files.

The configuration wizard runs silently to set up the model for peak performance.

🛡️ Checksum: 14634fb8b7f9677717e1c4d7d33ddb0a — ⏰ Updated on: 2026-06-23



  • CPU: AVX2/AVX-512 instruction set required for llama.cpp
  • RAM: 32 GB highly recommended for 26B+ GGUF models
  • Storage: extra room for future model updates and datasets
  • Graphics: 12 GB VRAM minimum required for basic quantization

The jina-embeddings-v5-text-nano model delivers compact yet high‑quality text embeddings optimized for edge devices. With only 2 million parameters, it achieves competitive performance on semantic similarity tasks while maintaining a small memory footprint. Its inference latency is under 5 ms on typical CPUs, making it ideal for real‑time applications that require fast processing. The model supports multiple languages and preserves contextual nuances better than earlier nano‑sized alternatives. Key metrics are summarized in the following table:

Parameters 2 million
Size (MB) 7.8
Latency (ms) <5
Throughput (tokens/s) 2000
Supported Languages 30