Language

Research at Lab4SLC on machine translation, language analysis, and applications of natural language processing.

Machine Translation

The Relationship between Latency and Accuracy in Simultaneous Machine Translation

Research Highlights

  • Study low-latency simultaneous translation between Japanese and Korean, which have similar word order
  • Use BS-SiMT, which learns from data when to wait before generating a translation
  • Show that some delay is necessary to maintain accuracy even when the source and target languages have similar word order
  • Next step: develop more advanced translation strategies, including paraphrasing, to reduce latency

Japanese-to-English Machine Translation with Reading Information

Research Highlights

  • Use large language models to translate descriptions of monuments to poems from the Man'yōshū
  • Improve translation by providing readings for difficult personal and place names that frequently occur in the descriptions
  • Use a large language model to identify difficult words and manually annotate their readings
  • Next step: automatically create and expand reading-information resources for difficult words

Language Analysis

Data Augmentation for Japanese Named Entity Recognition

Research Highlights

  • Identify key named entities in text, including proper names and numerical expressions
  • Automatically create additional training data by replacing named entities with others of the same type
  • Use large language models to generate effective replacement data
  • Next step: improve data augmentation by accounting for writing style and context

Applications of Natural Language Processing

Japanese Grammatical Error Correction

Research Highlights

  • Automatically correct grammatical errors made by beginning learners of Japanese
  • Analyze accuracy in correcting inflectional-ending errors for na-adjectives (adjectival nouns)
  • Find that correction accuracy tends to decrease for words written in hiragana
  • Next step: evaluate Japanese grammatical error correction using large language models

Research Highlights

  • Identify language in social media posts that is associated with mental health difficulties
  • Use large language models to extract keywords that may be characteristic of people experiencing mental distress
  • Analyze the extracted expressions in a large collection of social media posts
  • Next steps: extract longer expressions and infer the meaning of entire posts
In collaboration with Professor Yoshinobu Kano, Shizuoka University

Suggestions for Revising Written Text

Research Highlights

  • Use large language models to help people revise their writing
  • Ask a large language model to identify points for improvement in Japanese research paper abstracts
  • Evaluate three types of prompts using nearly 100 abstracts
  • Next step: extend the approach to types of writing other than research papers

For related publications, see the publication list.