espnet2.tasks.rst_vocoder.RestorationVocoderCollateFn
espnet2.tasks.rst_vocoder.RestorationVocoderCollateFn
class espnet2.tasks.rst_vocoder.RestorationVocoderCollateFn(context_samples: int, segment_frames: int, hop: int, input_sr: int, output_sr: int, use_predicted_feat: bool, noise_dir: str, rir_dir: str, degrade_prob: float, online_degradation: bool, train: bool = True, stats_only: bool = False)
Bases: object
Cut a context window from the 48 kHz reference and pick the excerpt.
Emits speech_ref1 (48 kHz context), vocoder_crop_start (the SSL frame at which the model cuts its fixed-length training excerpt) and, for finetuning only, noisy_speech (16 kHz degraded context, the input to the frozen feature predictor). Pretraining needs no 16 kHz copy here: the model resamples on the GPU.
With stats_only the reference passes through untouched so the shape file collected in stage 6 records true utterance lengths.
