نامه فرهنگستان

نامه فرهنگستان

طراحی یک پیکرۀ گفتاری در مطالعات آواشناسی تجربی: تصمیم‌ها و چالش‌های روش‌شناختی

نوع مقاله : مقاله پژوهشی

نویسنده
دپارتمان علم گفتار و آواشناسی،‌ دانشگاه مارتین لوتر هاله-ویتنبرگ و مؤسّسۀ زیبایی‌شناسی تجربی ماکس پلانک، فرانکفورت
10.22034/nf.2026.577011.1518
چکیده
هدف این مقاله ارائۀ مروری روش‌شناختی بر مراحل طراحی و ساخت یک پیکرۀ گفتاری با تمرکز بر تصمیم‌ها و چالش‌های فرایند گردآوری داده‌های گفتاری استاندارد در مطالعات آواشناسی تجربی است. این مرور بر پایۀ تجربه‌های حاصل از طراحی و پیاده‌سازی یک پیکرۀ گفتاری در چارچوب رسالۀ دکتری نویسنده انجام شده است. پیکرۀ اصلی ماهیتی دوزبانه دارد، امّا تمرکز این مقاله صرفاً بر بخش فارسی آن است. در این نوشتار، مراحل مختلف ساخت پیکره، ازجمله انتخاب و طراحی وظایف برانگیزش گفتار، نمونه‌گیری از گویندگان، طراحی محیط و شرایط ضبط، سازمان‌دهی و نام‌گذاری داده‌ها و، درنهایت، پیش‌پردازش و تقطیع گفتار،‌ به‌صورت نظام‌مند مورد بحث قرار می‌گیرند. همچنین، نشان داده می‌شود که طراحی پیکره فرایندی صرفاً اجرایی نیست، بلکه مجموعه‌ای از تصمیم‌های روش‌شناختی آگاهانه را در بر می‌گیرد که مستقیماً بر کیفیت داده‌ها، تفسیرپذیری نتایج و امکان بازتولید تحلیل‌ها اثر می‌گذارند. همچنین، مقاله به چالش‌های مرتبط با استفاده از روش‌های خودکار در تقطیع گفتار پیکره‌های فارسی می‌پردازد و محدودیت‌های ابزارهای بازشناسی خودکار گفتار را در بافت زبان‌های کم‌منبع برجسته می‌کند. در پایان، استدلال می‌شود که مستندسازیِ دقیق این تصمیم‌ها و مراحل خود بخشی اساسی از فرایند ساخت پیکره است و نقشی تعیین‌کننده در شفافیت روش، ارزیابی نتایج و امکان استفادۀ علمی از داده‌ها ایفا می‌کند.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Designing a Speech Corpus for Studies in Experimental Phonetics: Methodological Decisions and Challenges

نویسنده English

Neda Mousavi
دپارتمان علم گفتار و آواشناسی،‌ دانشگاه مارتین لوتر هاله-ویتنبرگ و مؤسّسۀ زیبایی‌شناسی تجربی ماکس پلانک، فرانکفورت
چکیده English

This paper provides a methodological overview of the stages involved in designing and constructing a speech corpus, with particular emphasis on the decisions and challenges that arise when collecting standardized speech data for experimental phonetics. The discussion draws on the author’s experience developing a speech corpus as part of a doctoral dissertation. Although the original corpus is bilingual, containing Persian and German data, the present article focuses exclusively on the Persian component. The paper describes the key stages of corpus construction, including the design and selection of speech elicitation tasks, speaker sampling, recording environment and conditions, data organization and naming conventions, and procedures for speech preprocessing and segmentation. It argues that corpus construction is not merely a technical or operational task, but a sequence of deliberate methodological decisions that directly affect data quality, the interpretability of results, and the reproducibility of analyses. In addition, the paper discusses the challenges of applying automatic segmentation methods to Persian speech, highlighting the limitations of automatic speech recognition and romanization tools in low-resource language contexts. It further argues that systematically documenting methodological decisions is itself an essential component of corpus construction, as such documentation promotes transparency, enables critical evaluation of results, and enhances the long-term scientific usability of the data.

کلیدواژه‌ها English

speech corpus
speech elicitation
speech segmentation
Automatic Speech Recognition (ASR)
low-resource languages
بی‌جن‌خان، محمود (۱۳۹۵)، «پیکرۀ گفتار محاوره‌ای زبان فارسی امروز»، مجموعه‌مقالات دومین همایش ملّی زبان‌شناسی پیکره‌ای، تهران، نشر نویسه پارسی.
Anderson, A., A. Clark, & J. Mullin (1991), “Introducing Information in Dialogues: How young speakers refer and how young listeners respond”, Journal of Child Language, 18, pp. 663-687.
Baker, R. and V. Hazan (2011), “DiapixUK: Task Materials for the Elicitation of Multiple Spontaneous Speech Dialogs”, Behavior Research Methods, 43, pp. 761-770.
Batliner, A., Neumann, M., Burkhardt, F., Baird, A., Meyer, S., Vu, N. T., and Schuller, B. W. (2023), “Ethical Awareness in Paralinguistics: A Taxonomy of Applications”, International Journal of Human-Computer Interaction, 39(9), pp. 1904-1921.
Bijankhan, M., Sheikhzadegan, J., Roohani, M. R., Samareh, Y., Lucas, C., and Tebyani, M. (1994), “FARSDAT-The Speech Database of Farsi Spoken Language”,The Proceedings of the Australian Conference on Speech Science and Technology, 2, pp. 828-831.
Boersma, P. and Weenink, D. (2024), Praat: Doing Phonetics by Computer (Version 6.4.23) [Computer program], Retrieved October 27, 2024.              
Brown, R. (1973), “Development of the First Language in the Human Species”, American Psychologist, 28(2), 97, American Psychological Association.
De Jong, N. H. and T. Wempe (2009), “Praat Script to Detect Syllable Nuclei and Measure Speech Rate Automatically”, Behavior Research Methods, 41(2), pp. 385-390.
Faitaki, F. and V. A. Murphy (2019), “Oral Language Elicitation Tasks in Applied Linguistics Research”, The Routledge Handbook of Research Methods in Applied Linguistics, pp. 360-369, Routledge.
Godfrey, J. J., E. C. Holliman & J. McDaniel (1992), “SWITCHBOARD: Telephone Speech Corpus for Research and Development”, IEEE International Conference on Acoustics, Speech, and Signal Processing, Vol. 1, pp. 517-520, IEEE Computer Society.
Grosman, J. (2021), “Fine-tuned XLSR-53 Large Model for Speech Recognition in Persian”, Hugging Face, pp. 410-411.
Kisler, T., U. Reichel & F. Schiel (2017), “Multilingual Processing of Speech via web Services”, Computer Speech & Language, 45, pp. 326-347.
Laver, J. (1994), Principles of Phonetics, Cambridge, Cambridge University Press.
Mousavi N., S. Grawunder (2024), “The Influence of Signal Segmentation Methods on Rhythm-based Speaker Recognition”, Baumann, T. (ed.), Studientexte zur Sprachkommunikation: Elektronische Sprachsignalverarbeitung 2024, Dresden, TUDpress, pp. 225-232.
Mousavi, N. (2025, PhD Thesis), “Individual Speech Rhythm in Persian and German: An Acoustic–Cognitive Approach of Rhythm Production in Different Speaking Tasks”, Martin-Luther Universität Halle-Wittenberg, Halle (Saale), Germany.
Nezami, M.O., P. Jamshid Lou & M. Karami (2019),“ShEMO: A Large-scale Validated Database for Persian Speech Emotion Detection”, Language Resources & Evaluation, 53(1), pp. 1-16.
Niebuhr, O. and A. Michaud (2015), “Speech Data Acquisition: The Underestimated Challenge”, KALIPHO - Kieler Arbeiten zur Linguistik und Phonetik, 3, pp. 1-42.
Pilehvar, A. & M. Sedaaghi (2008),Documentation of the Sahand Accented Speech Database (SES), Department of Electrical Eng., Sahand University of Tech, Iran.
Radford, A., J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever (2023), “Robust Speech           Recognition via Large-scale Weak Supervision”, International Conference on Machine Learning, pp. 28492-28518, PMLR.
Senft, G. (1995), “Elicitation”, Handbook of Pragmatics: Manual, pp. 577-581, John Benjamins.
Wagener, P. (1986), Sind Spracherhebungen Paradox? Über die Möglichkeit, Natürliches Sprachverhalten Wissenschaftlich zu Erfassen, A. Schöne (ed.), Tübingen, Niemeyer, Akten des VII. IVG-Kongresses, Vol. 4, pp. 319-327.
Wagner, P., J. Trouvain & F. Zimmerer (2015), “In Defense of Stylistic Diversity in Speech Research”, Journal of Phonetics, 48, pp. 1-12.
Xu, Y. (2010), “In Defense of Lab Speech”, Journal of Phonetics, 38(3), pp. 329-336.
دوره 25، شماره 3 - شماره پیاپی 103
نامه فرهنگستان (آواشناسی و واج‌شناسی زبان‌های ایرانی 2)
مرداد و شهریور 1405
صفحه 105-122

  • تاریخ دریافت 03 اسفند 1404
  • تاریخ بازنگری 13 فروردین 1405
  • تاریخ پذیرش 30 فروردین 1405