SepRQ
Self-supervised speech mixture representation learning via mask-free, multi-scale source separation. One frozen encoder for diarization, separation, enhancement and target-speaker tasks.
speechbrain ≥ 1.0.3Two lines to load. One call to run.
Extract the 12 Conformer layer representations of any mixture, or pick a downstream task and let SepRQ do the speech processing end to end.
Hear what each upstream can do
Pick a downstream task. Then choose an input and a frozen upstream SSL model.
Available models
Pass any of these names to SepRQEncoder. Only the requested model is downloaded, and every model returns its per-layer features.
| Name | Layers | Dim | Notes |
|---|---|---|---|
SepRQ ⭐ | 12 | 576 | Ours. streams=2 (default) or streams=3 |
BestRQ_50Hz | 12 | 576 | BEST-RQ trained at 50 Hz, released with SepRQ |
HuBERT_BASE | 12 | 768 | Baseline |
WavLM_BASE | 12 | 768 | Baseline |
WavLM_BASE_PLUS | 12 | 768 | Baseline |
WavLM_LARGE | 24 | 1024 | Baseline |
huggingface-cli login (or set HF_TOKEN). The HuBERT / WavLM baselines come from the torchaudio pipelines, with nothing extra to set up.Separation as the pretext task
SepRQ swaps masked prediction for pseudo source separation. It predicts one stream of discrete units per speaker in the mixture, at several temporal resolutions.
Pseudo source separation
Each separation head predicts the frozen BEST-RQ codewords of one clean speaker from the mixture. Utterance-level PIT resolves speaker order, and no offline k-means is needed.
Mask-free
Masking can hide exactly the content needed to disentangle a mixture. SepRQ predicts units over the full utterance, which uses every frame for training.
Multi-resolution
A separation objective follows every two Conformer layers, with frame folding going from 20 ms to 320 ms and one codebook per scale. Inference runs at 50 Hz with the encoder only.
State of the art on multi-speaker benchmarks
Frozen upstreams with SUPERB / TS-SUPERB downstream heads. SepRQ is pre-trained on LibriSpeech 960 h mixtures only.
Multi-speaker SUPERB
| SD | SS | SE | ||||
|---|---|---|---|---|---|---|
| Model | #Param (M) | Data (h) | DER ↓ | SI-SDRi ↑ | PESQ ↑ | STOI ↑ |
| Base class (≈95M) | ||||||
| HuBERT Base | 94.68 | LS-960 | 5.88 | 9.36 | 2.58 | 93.9 |
| WavLM Base | 94.70 | LS-960 | 4.55 | 10.37 | 2.58 | 94.0 |
| WavLM Base+ | 94.70 | Mix-94k | 3.50 | 10.85 | 2.63 | 94.3 |
| C-HuBERT Base | 96.00 | LS-960 | 2.77 | 11.08 | 2.63 | 94.0 |
| SA-WavLM† | 94.97 | LS-960 | 1.88 | 11.13 | 2.62 | 94.2 |
| SepRQ (ours) | 85.68 | LS-960 | 2.08 | 12.10 | 2.67 | 94.4 |
| Large class (>316M) | ||||||
| HuBERT Large | 316.62 | LL-60k | 5.75 | 10.45 | 2.64 | 94.2 |
| WavLM Large | 316.62 | Mix-94k | 3.24 | 11.19 | 2.70 | 94.5 |
| C-HuBERT Large | 318.00 | LL-60k | 2.65 | 11.24 | 2.65 | 94.3 |
TS-SUPERB
| TSE | PSE | PVAD | TS-ASR WER ↓ | ||||
|---|---|---|---|---|---|---|---|
| Upstream | SI-SDRi ↑ | STOI ↑ | SI-SDRi ↑ | STOI ↑ | mAP ↑ | w/o LM | w/ LM |
| HuBERT Base | 9.64 | 87.30 | 8.61 | 79.92 | 94.60 | 36.86 | 30.52 |
| WavLM Base | 10.26 | 88.40 | 9.65 | 81.57 | 94.40 | 27.82 | 22.68 |
| WavLM Base+ | 10.69 | 89.00 | 10.01 | 82.67 | 95.00 | 24.75 | 20.06 |
| SepRQ (ours) | 12.53 | 91.45 | 11.18 | 85.60 | 96.61 | 22.26 | 16.78 |
Out-of-domain
| DIHARD 3 | 2mix | 3mix | ||
|---|---|---|---|---|
| Upstream | FA / MD / SC | DER ↓ | SDRi ↑ | SDRi ↑ |
| w/o SSL | 6.1 / 8.1 / 4.9 | 19.3 | 16.4 | 13.1 |
| HuBERT Base | 4.3 / 8.6 / 4.7 | 17.7 | 17.0 | 12.7 |
| WavLM Base | 4.4 / 8.3 / 4.5 | 17.3 | 17.5 | 13.2 |
| WavLM Base+ | 4.8 / 7.6 / 4.2 | 16.6 | 18.2 | 13.4 |
| SepRQ (2 src) | 4.9 / 7.4 / 3.8 | 16.1 | 19.8 | 14.5 |
| SepRQ (3 src) | 5.0 / 7.3 / 4.1 | 16.4 | 19.4 | 17.2 |
Cite SepRQ
If SepRQ helps your research, please cite the paper.
@misc{baroudi2026seprqselfsupervisedspeech,
title = {SepRQ : Self-Supervised Speech Mixture Representation Learning
via Mask-Free, Multi-Scale Source Separation},
author = {Séverin Baroudi and Hervé Bredin and Ricard Marxer},
year = {2026},
eprint = {2610.04690},
archivePrefix = {arXiv},
primaryClass = {eess.AS},
url = {https://arxiv.org/abs/2610.04690}
}