CMU Speech Technology 26S
9/24/26Less than 1 minute
CMU Speech Technology 26S
CMU 11492/11692/18495, Speech Technology for Conversational AI, Spring 2026.
The demonstration notebooks from the course. Open one in Colab, choose a GPU runtime, and run it top to bottom. They work on CPU — that is what the weekly run checks — but fine-tuning and the larger models are much slower there.
| Notebook | What it does | |
|---|---|---|
speaker_verification.ipynb | Speaker embeddings with ESPnet-SPK, verification, and a simple diarization | |
speech_enhancement.ipynb | Enhancement and separation, scored with VERSA and a pretrained ASR model | |
text_to_speech.ipynb | Single-speaker and multi-speaker synthesis, and VERSA scores | |
neural_codec.ipynb | Three pretrained neural codecs and the bitrate trade between them | |
speech_translation.ipynb | Offline and simultaneous speech translation with ESPnet-ST-v2 | |
speech_recognition.ipynb | Fine-tune OWSM on one language of FLEURS with the ESPnet3 trainer |
Credit
The notebooks were written by the course's instructors and teaching assistants; each keeps its author line. They are kept here with the demonstrations intact and the grading removed.
For maintainers
MAINTAINING.md has what the badges check and what they do not, how to run these outside Colab, what was changed from the course originals, and which recordings had to be replaced and why.
