Speech
Research at Lab4SLC on speech recognition, speech synthesis, and related speech technologies.
Speech Recognition
Animal Vocalization Recognition
Research Highlights
- Analyze animal vocalization recognition results represented using the Inter-Species Phonetic Alphabet (ISPA)
- Evaluate animal-species classification performance using ISPA-based recognition results
- Cluster ISPA symbols using topic models
- Next step: conduct a more detailed evaluation using more data
Converting Speech Recognition Output into Written Style
Research Highlights
- Automatically remove fillers such as “um” and “uh” during speech recognition
- Fine-tune the large-scale Open Whisper-style Speech Model (OWSM)
- Teach the model not to produce fillers in its recognition output
- Next step: extend the approach beyond filler removal to broader transcript editing
Speech Synthesis
Controlling Speech Synthesis with Musical Constraints
Research Highlights
- Partially control pitch during text-to-speech synthesis
- Add symbols to the input text to specify the desired pitch
- Learn the relationship between the symbols and pitch using the PJS Japanese singing-voice corpus
- Next steps: expand the controllable pitch range and incorporate musical constraints other than pitch
For related publications, see the publication list.