espnet2.rst.decoder.dac_vocoder.HiFiGANVocoder
espnet2.rst.decoder.dac_vocoder.HiFiGANVocoder
class espnet2.rst.decoder.dac_vocoder.HiFiGANVocoder(input_dim: int = 1024, channels: int = 512, kernel_size: int = 7, upsample_scales: List[int] = [8, 5, 4, 3, 2], upsample_kernel_sizes: List[int] = [16, 10, 8, 6, 4], resblock_kernel_sizes: List[int] = [3, 7, 11], resblock_dilations: List[List[int]] = [[1, 3, 5], [1, 3, 5], [1, 3, 5]], nonlinear_activation: str = 'LeakyReLU', nonlinear_activation_params: Dict | None = None, use_weight_norm: bool = True)
Bases: Module
HiFi-GAN generator (ESPnet’s own) driven by SSL features.
The alternative vocoder of the recipe: ESPnet’s HiFiGANGenerator with the same 8-5-4-3-2 upsampling geometry as the DAC decoder, so it also turns one 50 Hz frame into 960 samples at 48 kHz, but with HiFi-GAN v1 residual blocks and LeakyReLU instead of Snake. About 14M parameters at 512 channels. Same calling convention as DACVocoder.
Initialize internal Module state, shared by both nn.Module and ScriptModule.
forward(x: Tensor) → Tensor
(B, D, T_frames) -> (B, 1, T_frames * upsample_factor).
generate(ssl_feat: Tensor) → Tensor
remove_weight_norm() → None
