Development and Fine-Tuning of an Italian Language Text Embedder

DCube developed and fine-tuned a lightweight, high-performance Italian language text embedding model optimized for retrieval, search, and document similarity tasks. The solution delivers near state-of-the-art performance on established Italian MTEB benchmarks while remaining fully on-premise, efficient, and production-ready.

The Client

Client Name: GalaxAI
Industry: Technology
Region: Global
Company Size: Enterprise

The Challenge

Most high-performing embedding models are large, resource-intensive, and dependent on third-party APIs, making them unsuitable for on-premise or privacy-sensitive deployments. Additionally, Italian language tasks are often underserved compared to English, requiring specialized fine-tuning to achieve competitive performance across retrieval and semantic similarity benchmarks.

The Solution

DCube fine-tuned a compact MiniLM-based student model using knowledge distillation from stronger teacher models, optimizing it specifically for Italian-language tasks. The resulting embedder balances performance and efficiency, integrates seamlessly with standard inference frameworks, and runs reliably on modern GPU hardware without external dependencies.

Key Features

  • Italian-specific text embedding optimized for retrieval and similarity
  • Strong performance across 7 MTEB benchmarks
  • Knowledge distillation from larger embedding models
  • Fully compatible with SentenceTransformers and Infinity Server
  • On-premise, API-free deployment for enterprise environments
  • Optimized for NVIDIA GPUs and Apple Silicon

Technologies Used

  • Transformer-based Language Models
  • SentenceTransformers
  • Knowledge Distillation
  • Python
  • GPU Acceleration (NVIDIA, Apple Silicon)

Results & Impact

  • Student MiniLM achieved near E5-Small performance across 7 MTEB tasks
  • Retrieval performance improved dramatically (nDCG@10: 0.54 → 0.92) compared to untrained baseline
  • Outperformed E5-Small on clustering tasks (V-measure: 0.395 vs 0.379)
  • Competitive performance on semantic similarity despite compact model size
  • Enabled high-quality Italian RAG, search, and similarity pipelines
  • Reduced infrastructure and inference costs compared to large models
  • Supported fully private, on-premise enterprise deployments
  • Demonstrated effective model compression without major performance loss