espnet2.tok.output.SpeechTokenizerOutput
Less than 1 minute
espnet2.tok.output.SpeechTokenizerOutput
class espnet2.tok.output.SpeechTokenizerOutput(continuous: Tensor, assignment: Tensor, token_ids: Tensor, lengths: Tensor, soft_assignment: Tensor | None = None, quantized_features: Tensor | None = None)
Bases: object
Outputs of a trainable speech tokenizer.
- Parameters:
- continuous – Continuous frontend features of shape
(B, T, D). - assignment – Hard straight-through cluster assignments of shape
(B, T, K). The forward values are one-hot during training, while gradients propagate through their soft counterparts. - token_ids – Integer cluster indices of shape
(B, T). - lengths – Valid token lengths of shape
(B,). - soft_assignment – Optional soft cluster assignments of shape
(B, T, K)for analysis and auxiliary regularization. - quantized_features – Optional differentiable quantized features of shape
(B, T, D). A future VQ implementation can expose its straight-through vectors here so downstream losses can reach the frontend.
- continuous – Continuous frontend features of shape
assignment : Tensor
continuous : Tensor
lengths : Tensor
quantized_features : Tensor | None = None
soft_assignment : Tensor | None = None
token_ids : Tensor
