KOMPYUTER LINGVISTIKASI VA SAVOL-JAVOB TIZIMLARINING RIVOJLANISH TARIXI: JAHON TAJRIBASI VA O‘ZBEK TILI UCHUN XULOSALAR
DOI:
https://doi.org/10.67227/dywqyx81Kalit so‘zlar:
kompyuter lingvistikasi, mashina tarjimasi, korpus lingvistikasi, savol-javob tizimlari, RAG, agglyutinativ tillar, til modellari, o‘zbek tili, tokenizatsiya.Abstrak
Ushbu maqolada kompyuter lingvistikasining shakllanish jarayoni doirasida dastlab mashina tarjimasi va korpus lingvistikasi haqida qisqacha tarixiy ma’lumot beriladi. So‘ngra savol-javob tizimlarining rivojlanish bosqichlari ko‘rib chiqiladi. Bunda qat’iy qoidalarga asoslangan dastlabki dasturlardan boshlab, neyron tarmoqlarga asoslangan “o‘quvchi” modellar va Retrieval-Augmented Generation (RAG) deb nomlanuvchi zamonaviy arxitekturagacha bo‘lgan taraqqiyot yo‘li tahlil qilinadi. Jahon tajribasi bilan bir qatorda turkiy tillar guruhi hamda boshqa agglyutinativ tuzilishga ega tillar misolida morfologik jihatdan murakkab tillarga xos yechimlarning qanday sinovdan o‘tkazilgani ko‘rib chiqiladi. Shuningdek, o‘zbek tilida amalga oshirilgan kompyuter lingvistikasi sohasidagi ishlar tizimlashtirilib, ularning savol-javob tizimlarini yaratish uchun qay darajada asos bo‘la olishi baholanadi.
Yuklashlar
Havolalar
1. Weaver W. Translation // Machine Translation of Languages: Fourteen Essays / ed. by W. N. Locke, A. D. Booth. – Cambridge, MA: The Technology Press of MIT; New York: John Wiley & Sons, 1955. – P. 15–23. – URL: https://aclanthology.org/www.mt-archive.info/srch/genmisc-50.htm ACL Anthology
2. Kučera H., Francis W. N. Computational Analysis of Present-Day American English. – Providence, RI: Brown University Press, 1967. – 424 p. – URL: https://books.google.com/books?id=Gb55AAAAIAAJ Google Books
3. Green B. F., Wolf A. K., Chomsky C., Laughery K. Baseball: An Automatic Question-Answerer // Proceedings of the Western Joint IRE-AIEE-ACM Computer Conference. – 1961. – P. 219–224. – DOI: 10.1145/1460690.1460714. – URL: https://doi.org/10.1145/1460690.1460714 DOI
4. Chen D., Fisch A., Weston J., Bordes A. Reading Wikipedia to Answer Open-Domain Questions // Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Vol. 1: Long Papers. – Vancouver, Canada, 2017. – P. 1870–1879. – DOI: 10.18653/v1/P17-1171. – URL: https://aclanthology.org/P17-1171/ ACL Anthology
5. Lewis P., Perez E., Piktus A., Petroni F., Karpukhin V., Goyal N., Küttler H., Lewis M., Yih W., Rocktäschel T., Riedel S., Kiela D. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks // Advances in Neural Information Processing Systems. – 2020. – Vol. 33. – P. 9459–9474. – URL: https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
6. Clark J. H., Choi E., Collins M., Garrette D., Kwiatkowski T., Nikolaev V., Palomaki J. TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages // Transactions of the Association for Computational Linguistics. – 2020. – Vol. 8. – P. 454–470. – DOI: 10.1162/tacl_a_00317. – URL: https://aclanthology.org/2020.tacl-1.30/ ACL Anthology
7. Yeshpanov R., Efimov P., Boytsov L., Shalkarbayuli A., Braslavski P. KazQAD: Kazakh Open-Domain Question Answering Dataset // Proceedings of LREC-COLING 2024. – Torino, Italy, 2024. – P. 9645–9656. – URL: https://aclanthology.org/2024.lrec-main.843/ ACL Anthology
8. Toraman C., Yilmaz E. H., Şahinuç F., Ozcelik O. Impact of Tokenization on Language Models: An Analysis for Turkish // ACM Transactions on Asian and Low-Resource Language Information Processing. – 2023. – Vol. 22, No. 4. – Article 116. – P. 1–21. – DOI: 10.1145/3578707. – URL: https://doi.org/10.1145/3578707 OpenMETU
9. Salaev U. UzMorphAnalyser: A Morphological Analysis Model for the Uzbek Language Using Inflectional Endings // AIP Conference Proceedings. – 2024. – Vol. 3244, No. 1. – Article 030058. – DOI: 10.1063/5.0241461. – URL: https://doi.org/10.1063/5.0241461 PyPI
10. Bobojonova L., Akhundjanova A., Ostheimer P. S., Fellenz S. BBPOS: BERT-Based Part-of-Speech Tagging for Uzbek // Proceedings of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025). – Abu Dhabi, UAE, 2025. – P. 287–293. – URL: https://aclanthology.org/2025.loreslm-1.23/ ACL Anthology
11. Voorhees E. M., Tice D. M. The TREC-8 Question Answering Track // Proceedings of the Second International Conference on Language Resources and Evaluation (LREC 2000). – Athens, Greece: ELRA, 2000. – URL: https://aclanthology.org/L00-1018/ ACL Anthology
12. Ferrucci D., Brown E., Chu-Carroll J., Fan J., Gondek D., Kalyanpur A. A., Lally A., Murdock J. W., Nyberg E., Prager J., Schlaefer N., Welty C. Building Watson: An Overview of the DeepQA Project // AI Magazine. – 2010. – Vol. 31, No. 3. – P. 59–79. – DOI: 10.1609/aimag.v31i3.2303. – URL: https://doi.org/10.1609/aimag.v31i3.2303 ResearchGate
13. Karpukhin V., Oğuz B., Min S., Lewis P., Wu L., Edunov S., Chen D., Yih W. Dense Passage Retrieval for Open-Domain Question Answering // Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). – 2020. – P. 6769–6781. – DOI: 10.18653/v1/2020.emnlp-main.550. – URL: https://aclanthology.org/2020.emnlp-main.550/ ACL Anthology
14. Rajpurkar P., Zhang J., Lopyrev K., Liang P. SQuAD: 100,000+ Questions for Machine Comprehension of Text // Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. – Austin, Texas, 2016. – P. 2383–2392. – DOI: 10.18653/v1/D16-1264. – URL: https://aclanthology.org/D16-1264/ ACL Anthology
15. Murat A., Ali S. Low-Resource POS Tagging With Deep Affix Representation and Multi-Head Attention // IEEE Access. – 2024. – Vol. 12. – P. 66495–66504. – DOI: 10.1109/ACCESS.2024.3395454. – URL: https://doi.org/10.1109/ACCESS.2024.3395454 ResearchGate
16. Novák A., Novák B., Zombori T., Szabó G., Szántó Z., Farkas R. A Question Answering Benchmark Database for Hungarian // Proceedings of the 17th Linguistic Annotation Workshop (LAW-XVII). – Toronto, Canada, 2023. – P. 188–198. – DOI: 10.18653/v1/2023.law-1.19. – URL: https://aclanthology.org/2023.law-1.19/ ACL Anthology
17. Kuriyozov E., Vilares D., Gómez-Rodríguez C. BERTbek: A Pretrained Language Model for Uzbek // Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages @ LREC-COLING 2024. – Torino, Italy, 2024. – P. 33–44. – URL: https://aclanthology.org/2024.sigul-1.5/ ACL Anthology
18. Yusupkhujaev K. Benchmarking Pre-Trained Open-Source Large Language Models for Uzbek: Evaluating Performance in a Low-Resource Setting Across Translation, Comprehension, and Generation // Annali d’Italia. – 2025. – No. 71. – P. 67–73. – DOI: 10.5281/zenodo.17223973. – URL: https://zenodo.org/records/17223973 zenodo.org
19. Solidjonov D., Najmiddinov M. Investigating Linguistic Errors in Large Language Model Generation of Uzbek Text // Cogent Arts & Humanities. – published online 2025; Vol. 13, Issue 1, 2026. – Article 2600519. – DOI: 10.1080/23311983.2025.2600519. – URL: https://doi.org/10.1080/23311983.2025.2600519 Tandfonline