espnet2.rst.decoder.dac_vocoder.DACVocoder
espnet2.rst.decoder.dac_vocoder.DACVocoder
class espnet2.rst.decoder.dac_vocoder.DACVocoder(input_dim: int = 1024, channels: int = 1536, rates: List[int] = [8, 5, 4, 3, 2], d_out: int = 1)
Bases: Module
DAC decoder generator.
- Parameters:
- input_dim – SSL feature dimension (1024 for w2v-BERT 2.0).
- channels – width after the input convolution; halves at every block. The published Sidon vocoder uses 1536.
- rates – transposed-convolution strides; their product is the number of output samples per input frame (960 = 48 kHz / 50 Hz).
- d_out – output channels (1, mono).
Initialize internal Module state, shared by both nn.Module and ScriptModule.
forward(x: Tensor) → Tensor
(B, D, T_frames) -> (B, 1, T_frames * upsample_factor).
Same calling convention as the official TorchScript decoder, so the two are interchangeable wherever one is used.
generate(ssl_feat: Tensor) → Tensor
(B, T_frames, D) features -> (B, T_wav) waveform.
load_official_torchscript(path: str, verify: bool = True) → None
Load the published decoder_{cpu,cuda}.pt into this module.
The release is a frozen TorchScript graph with no state_dict; see _official_weights for how its parameters are recovered. Weight norm is removed here first, since the graph holds folded weights. By default the result is verified end to end by running both decoders on the same input.
remove_weight_norm() → None
Fold weight norm into plain weights for inference.
