Speech

Research at Lab4SLC on speech recognition, speech synthesis, and related speech technologies.

Speech Recognition

Animal Vocalization Recognition

Research Highlights

  • Analyze animal vocalization recognition results represented using the Inter-Species Phonetic Alphabet (ISPA)
  • Evaluate animal-species classification performance using ISPA-based recognition results
  • Cluster ISPA symbols using topic models
  • Next step: conduct a more detailed evaluation using more data

Converting Speech Recognition Output into Written Style

Research Highlights

  • Automatically remove fillers such as “um” and “uh” during speech recognition
  • Fine-tune the large-scale Open Whisper-style Speech Model (OWSM)
  • Teach the model not to produce fillers in its recognition output
  • Next step: extend the approach beyond filler removal to broader transcript editing

Speech Synthesis

Controlling Speech Synthesis with Musical Constraints

Research Highlights

  • Partially control pitch during text-to-speech synthesis
  • Add symbols to the input text to specify the desired pitch
  • Learn the relationship between the symbols and pitch using the PJS Japanese singing-voice corpus
  • Next steps: expand the controllable pitch range and incorporate musical constraints other than pitch

For related publications, see the publication list.