Development and Fine-Tuning of an Italian Language Text Embedder
DCube developed and fine-tuned a lightweight, high-performance Italian language text embedding model optimized for retrieval, search, and document similarity tasks. The solution delivers near state-of-the-art performance on established Italian MTEB benchmarks while remaining fully on-premise, efficient, and production-ready.
The Client
Client Name: GalaxAI
Industry: Technology
Region: Global
Company Size: Enterprise
The Challenge
Most high-performing embedding models are large, resource-intensive, and dependent on third-party APIs, making them unsuitable for on-premise or privacy-sensitive deployments. Additionally, Italian language tasks are often underserved compared to English, requiring specialized fine-tuning to achieve competitive performance across retrieval and semantic similarity benchmarks.
The Solution
DCube fine-tuned a compact MiniLM-based student model using knowledge distillation from stronger teacher models, optimizing it specifically for Italian-language tasks. The resulting embedder balances performance and efficiency, integrates seamlessly with standard inference frameworks, and runs reliably on modern GPU hardware without external dependencies.
Key Features
- Italian-specific text embedding optimized for retrieval and similarity
- Strong performance across 7 MTEB benchmarks
- Knowledge distillation from larger embedding models
- Fully compatible with SentenceTransformers and Infinity Server
- On-premise, API-free deployment for enterprise environments
- Optimized for NVIDIA GPUs and Apple Silicon
Technologies Used
- Transformer-based Language Models
- SentenceTransformers
- Knowledge Distillation
- Python
- GPU Acceleration (NVIDIA, Apple Silicon)
Results & Impact
- Student MiniLM achieved near E5-Small performance across 7 MTEB tasks
- Retrieval performance improved dramatically (nDCG@10: 0.54 → 0.92) compared to untrained baseline
- Outperformed E5-Small on clustering tasks (V-measure: 0.395 vs 0.379)
- Competitive performance on semantic similarity despite compact model size
- Enabled high-quality Italian RAG, search, and similarity pipelines
- Reduced infrastructure and inference costs compared to large models
- Supported fully private, on-premise enterprise deployments
- Demonstrated effective model compression without major performance loss