Corpus Linguistics and Generative AI Tools for Bilingual Terminology Extraction: Two Case Studies in Sports Interpreting
Files
MA_09542200_2025.pdf
Closed access - Adobe PDF
- 4.65 MB
MA_09542200_2025_Declaration of Honor
Closed access - Adobe PDF
- 113.88 KB
Details
- Supervisors
- Faculty
- Degree label
- Abstract
- This dissertation investigates the effectiveness of corpus linguistics and generative AI tools in bilingual terminology extraction, with a focus on interpreter preparation in domain-specific contexts. Bridging theoretical and applied translation studies, the research compares two methodological paradigms: a corpus-based approach using Sketch Engine, and a generative AI-based approach employing ChatGPT and DeepSeek. The study focuses on table tennis and billiards as case studies. Both sports involve an extensive set of specialized terms, and my preliminary search found that bilingual glossaries for them are rare in terms of their bilingual resources. The study first constructs domain-specific corpora and extracts English-Chinese term pairs via Sketch Engine, which was followed by manual categorization and validation using generative AI. In the second case, both ChatGPT and DeepSeek extracted and translated terms directly from expert texts with certain prompts. The outputs are then evaluated in terms of relevance, accuracy, domain specificity, and consistency with the use of both gold-standard references and external validation tools (e.g., Baidu, Google Images, WeChat). The results indicate that while corpus-based methods offer greater control and replicability, generative AI tools demonstrate notable advantages in speed and contextual flexibility. However, the generative approach illustrates risks such as domain drift, hallucinations, and translation inconsistency. This dissertation contributes to the emerging field of AI-assisted interpreter preparation. It examines the strengths and limitations of current tools and proposes hybrid workflows that combine corpus reliability with generative AI efficiency. It also offers practical implications for interpreters operating under time constraints in low-resource domains, and lays the groundwork for further research into the integration of LLMs into professional terminology management practices.