Self-supervised learning for the cocktail party

SepRQ

Self-supervised speech mixture representation learning via mask-free, multi-scale source separation. One frozen encoder for diarization, separation, enhancement and target-speaker tasks.

$pip install seprq
Requires PyTorch and speechbrain ≥ 1.0.3
Séverin Baroudi1,3 · Hervé Bredin2 · Ricard Marxer1,3
1Univ Toulon, Aix Marseille Univ, CNRS, LIS, France · 2pyannoteAI, Toulouse, France · 3CNRS, ILLS, Montréal, Canada
85.68M
inference parameters, 50 Hz frame rate
12.10dB
SI-SDRi, SUPERB speech separation (Libri2Mix)
2.08%
DER, SUPERB speaker diarization
16.78%
WER, target-speaker ASR (TS-SUPERB, with LM)
Quickstart

Two lines to load. One call to run.

Extract the 12 Conformer layer representations of any mixture, or pick a downstream task and let SepRQ do the speech processing end to end.


      
Interactive demo

Hear what each upstream can do

Pick a downstream task. Then choose an input and a frozen upstream SSL model.

Checkpoints

Available models

Pass any of these names to SepRQEncoder. Only the requested model is downloaded, and every model returns its per-layer features.

NameLayersDimNotes
SepRQ ⭐12576Ours. streams=2 (default) or streams=3
BestRQ_50Hz12576BEST-RQ trained at 50 Hz, released with SepRQ
HuBERT_BASE12768Baseline
WavLM_BASE12768Baseline
WavLM_BASE_PLUS12768Baseline
WavLM_LARGE241024Baseline
The SepRQ / BEST-RQ weights repository is private for now: authenticate once with huggingface-cli login (or set HF_TOKEN). The HuBERT / WavLM baselines come from the torchaudio pipelines, with nothing extra to set up.
Method

Separation as the pretext task

SepRQ swaps masked prediction for pseudo source separation. It predicts one stream of discrete units per speaker in the mixture, at several temporal resolutions.

SepRQ architecture: clean sources are mixed; the mixture mel-spectrogram goes through a CNN front-end and 12 Conformer layers with progressive downsampling (50, 25, 12 Hz). Per-speaker prediction heads are trained with a permutation-invariant cross-entropy against targets from frozen multi-resolution random vector quantizers applied to the clean sources.
Fig. 1. SepRQ performs pseudo source separation in the discrete space. Random vector quantizers (RVQs) give speaker-specific discrete labels at several resolutions.

Pseudo source separation

Each separation head predicts the frozen BEST-RQ codewords of one clean speaker from the mixture. Utterance-level PIT resolves speaker order, and no offline k-means is needed.

Mask-free

Masking can hide exactly the content needed to disentangle a mixture. SepRQ predicts units over the full utterance, which uses every frame for training.

Multi-resolution

A separation objective follows every two Conformer layers, with frame folding going from 20 ms to 320 ms and one codebook per scale. Inference runs at 50 Hz with the encoder only.

Results

State of the art on multi-speaker benchmarks

Frozen upstreams with SUPERB / TS-SUPERB downstream heads. SepRQ is pre-trained on LibriSpeech 960 h mixtures only.

Multi-speaker SUPERB

Speaker diarization (SD), speech separation (SS) and speech enhancement (SE). LS = LibriSpeech, LL = Libri-Light, Mix = LL-60k + GigaSpeech-10k + VoxPopuli-24k.
SDSSSE
Model#Param (M)Data (h)DER ↓SI-SDRi ↑PESQ ↑STOI ↑
Base class (≈95M)
HuBERT Base94.68LS-9605.889.362.5893.9
WavLM Base94.70LS-9604.5510.372.5894.0
WavLM Base+94.70Mix-94k3.5010.852.6394.3
C-HuBERT Base96.00LS-9602.7711.082.6394.0
SA-WavLM†94.97LS-9601.8811.132.6294.2
SepRQ (ours)85.68LS-9602.0812.102.6794.4
Large class (>316M)
HuBERT Large316.62LL-60k5.7510.452.6494.2
WavLM Large316.62Mix-94k3.2411.192.7094.5
C-HuBERT Large318.00LL-60k2.6511.242.6594.3
† Uses oracle speaker embeddings of the target speakers at inference time (shown in grey). Best base-class value among enrollment-free models highlighted.

TS-SUPERB

Target speaker extraction, personalized speech extraction, personalized VAD, target-speaker ASR.
TSEPSEPVADTS-ASR WER ↓
UpstreamSI-SDRi ↑STOI ↑SI-SDRi ↑STOI ↑mAP ↑w/o LMw/ LM
HuBERT Base9.6487.308.6179.9294.6036.8630.52
WavLM Base10.2688.409.6581.5794.4027.8222.68
WavLM Base+10.6989.0010.0182.6795.0024.7520.06
SepRQ (ours)12.5391.4511.1885.6096.6122.2616.78

Out-of-domain

EEND diarization on DIHARD 3 and ConvTasNet separation on WSJ0-2/3mix.
DIHARD 32mix3mix
UpstreamFA / MD / SCDER ↓SDRi ↑SDRi ↑
w/o SSL6.1 / 8.1 / 4.919.316.413.1
HuBERT Base4.3 / 8.6 / 4.717.717.012.7
WavLM Base4.4 / 8.3 / 4.517.317.513.2
WavLM Base+4.8 / 7.6 / 4.216.618.213.4
SepRQ (2 src)4.9 / 7.4 / 3.816.119.814.5
SepRQ (3 src)5.0 / 7.3 / 4.116.419.417.2
Citation

Cite SepRQ

If SepRQ helps your research, please cite the paper.

@misc{baroudi2026seprqselfsupervisedspeech,
  title         = {SepRQ : Self-Supervised Speech Mixture Representation Learning
                   via Mask-Free, Multi-Scale Source Separation},
  author        = {Séverin Baroudi and Hervé Bredin and Ricard Marxer},
  year          = {2026},
  eprint        = {2610.04690},
  archivePrefix = {arXiv},
  primaryClass  = {eess.AS},
  url           = {https://arxiv.org/abs/2610.04690}
}