espnet2.tok.tokenizer.SpeechTokenizer
espnet2.tok.tokenizer.SpeechTokenizer
class espnet2.tok.tokenizer.SpeechTokenizer(frontend: AbsFrontend, quantizer: AbsSpeechTokenizerQuantizer, freeze_epochs: int = 0)
Bases: Module
Convert speech into differentiable assignments and discrete token IDs.
This class combines a speech frontend and quantizer for joint training with downstream objectives. Its training output preserves gradients through hard token assignments, while inference exposes integer token IDs.
AbsGANCodec encodes waveforms into codes and decodes codes back into waveforms. This tokenizer only maps speech to units for downstream objectives, so it does not implement the codec’s waveform decoder.
The similarly named classes serve different purposes: speechlm.tokenizer.AbsTokenizer is intended for no-grad SpeechLM token postprocessing, AudioTokenizer extracts discrete BEATs codes for offline use, and BeatsTokenizer provides a BEATs-specific encoder and VQ for pretraining. Use this class when downstream losses must train the frontend and quantizer together.
Initialize the tokenizer from a continuous frontend and quantizer.
- Parameters:
- frontend – Continuous speech frontend such as S3PRL.
- quantizer – Quantizer applied to frontend representations.
- freeze_epochs – Number of initial epochs that use a frozen frontend and frozen deterministic quantization.
encode(speech: Tensor, speech_lengths: Tensor) → SpeechTokenizerOutput
Extract deterministic nearest-centroid speech tokens.
forward(speech: Tensor, speech_lengths: Tensor) → SpeechTokenizerOutput
Extract frontend features and differentiably quantize them.
property is_frozen : bool
Return whether tokenizer parameters are frozen in the current epoch.
property num_clusters : int
Return the tokenizer vocabulary size.
set_epoch(epoch: int) → bool
Set the current epoch and report a frozen-to-trainable transition.
- Parameters:epoch – One-based training epoch.
- Returns:
Trueonly when this call crosses the unfreeze boundary.
