Apple Machine Learning researchers have shown that enhancing a model's capacity for language discrimination during pretraining improves linguistic performance in multilingual self-supervised speech models.
Model ReleasesAppleApple Machine Learning
Language Discrimination Improves Linguistic Learning in Multilingual Speech Models
While multilingual models often struggle to match the performance of monolingual models within the same pretraining budget, the study demonstrates that language discrimination can reduce or even close this gap. Using a controlled English/French HuBERT setting, researchers tested two interventions: an auxiliary language classifier and per-language k-means targets.
The results showed improvements across various metrics. Phonetic phone discrimination error decreased from 11.6% in the bilingual baseline to 10.4%, approaching the monolingual level of 10.8%. Lexical performance (sWUGGY) increased from 52.1% to 56.7%, while prosodic performance (ProsAudit, lexical subtask) rose from 68.9% to 72.9%, exceeding the monolingual baseline of 72.6%.
The researchers found that the most significant gains occurred when language discrimination was introduced during the first iteration of training. Later or repeated interventions resulted in smaller improvements and increased language-wise segregation. These findings suggest that language discrimination plays a causal role in reducing the additional computational cost typically required for multilingual learning.
Sources
- Language Discrimination Improves Linguistic Learning in Multilingual Speech Models (Apple Machine Learning, 2026-10-02)
- arXiv