INTEGRATION OF SEMANTIC WIKI TECHNOLOGIES WITH LARGE LANGUAGE MODELS AS A TECHNOLOGICAL FOUNDATION FOR EXPERIENCE ACQUISITION FROM NATURAL LANGUAGE DOCUMENTS
DOI:
https://doi.org/10.17721/3041-2323.2025.266-305Keywords:
Large Language Models, semantic wiki technologies, document knowledge acquisition, natural language documentsAbstract
The results of the research presented in the article include the formalization of a class of tasks involving the extraction of domain-specific knowledge from natural language documents, as well as the formulation of a set of requirements for a technological platform capable of solving tasks of this class. This analysis supports identification of the basic functional modules of technological platform and the sequence of information processing in it. The use of LLMs for the analysis of natural language documents has been examined, criteria for evaluating their effectiveness and directions for improving their performance have been considered. To justify the proposed approach, we consider the advantages of integrating semantic technologies (on the example of Semantic MediaWiki) with LLMs used act as tools for knowledge acquisition at different stages of document processing. The considered practical examples demonstrate significant differences between tasks of the analyzed class and the necessity of adapting the proposed platform to the specificity of the tasks.
References
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Li, K., ... Xie, X. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3), 1–45. https://doi.org/10.1145/3641289.
Geary, W. L., Styan, M. C. J., Loo, H. H. T., Thompson, D. G., Davies, P. D. A., & Whitehouse, C. D. (2020). A guide to ecosystem models and their environmental applications. Nature Ecology & Evolution, 4(11), 1459–1471.
Geifman, Y., & El-Yaniv, R. (2017). Selective classification for deep neural networks. In Advances in Neural Information Processing Systems, 30.
Gruber, T. R. (n.d.). What is an ontology? Retrieved from http://www-ksl.stanford.edu/kst/what-is-an-ontology.html.
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., & Smith, N. A. (2020). Don't stop pretraining: Adapt language models to domains and tasks. arXiv. https://arxiv.org/pdf/2004.10964.
Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. In Advances in Neural Information Processing Systems, 29.
Haryanto, C. Y. (2024). A framework for legal information retrieval using semantic search and fine-tuned large language models. International Journal of Computer Science and Network Security (IJCSNS), 24(4), 161–168.
Hatgis-Kessell, S., Knox, W. B., Booth, S., Niekum, S., & Stone, P. (2025). Influencing humans to conform to preference models for RLHF. arXiv. https://doi.org/10.48550/arXiv.2501.06416.
Howard, J., & Ruder, S. (2018). Universal language model fine-tuning for text classification. arXiv. https://doi.org/10.48550/arXiv.1801.06146.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., He, H., Chen, Y., Li, A., & Riedel, S. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems (NeurIPS 2020). https://proceedings.neurips.cc/paper_files/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html.
Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text summarization branches out (pp. 74–81). Association for Computational Linguistics. https://aclanthology.org/W04-1013.
Manikas, K., & Hansen, K. M. (2013). Software ecosystems. Journal of Systems and Software, 86(5), 1294–1306.
Misback, E., Tatlock, Z., & Tanimoto, S. L. (2024). Magic markup: Maintaining document-external markup with an LLM. In Companion proceedings of the 8th international conference on the art, science, and engineering of programming (pp. 22–35).
Musumeci, P., D'Agata, P., D'Angelo, S., & Scardapane, S. (2024). LLM based multi-agent generation of semi-structured documents from semantic templates in the public administration domain. arXiv.
Sinitsyn, I. P., Rohushyna, Yu. V., & Yurchenko, K. Yu. (2025). Integration of large language models with semantic processing tools as a knowledge digitalization instrument. Problemy Programuvannya, (2), 63–76. https://pp.isofts.kiev.ua/index.php/ojs1/article/download/838/889 [in Ukrainian].
Slyusar, V. (2024). Local large language models for confidential information processing. Ozbroiennia ta Viiskova Tekhnika, 4(44), 79–91. https://doi.org/10.34169/2414-0651.2024.4(44).79-91 [in Ukrainian].
Wei, J., Bosma, M., Zhao, V. Y., Xu, D., Schuurmans, D., Gelbart, M., Guu, K., Davies, A., Salakhutdinov, R., Le, Q. V., Chi, E. H., Dean, J., & Raffel, C. (2021). Finetuned language models are zero-shot learners. arXiv. https://doi.org/10.48550/arXiv.2110.08207.
Wu, S., Ma, X., Luo, D., Li, L., Shi, X., Chang, X., Cui, H., Tang, B., Wu, Y., Liu, Y., & Gong, J. (2023). A survey on large language model for recommendation. arXiv. https://doi.org/10.48550/arXiv.2311.13969.
Wu, X., Wu, S.-H., Wu, J., Feng, L., & Tan, K. C. (2024). Evolutionary computation in the era of large language models: Survey and roadmap. arXiv. https://doi.org/10.48550/arXiv.2401.10034
Zhu, K., Wang, J., Zhou, J., Wang, Z., Chen, H., Wang, Y., Yang, L., Ye, W., Gong, N. Z., & Zhang, Y. (2023). PromptBench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv. https://doi.org/10.48550/arXiv.2306.04528.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Applied Information Systems and Technologies in the Digital Society

This work is licensed under a Creative Commons Attribution 4.0 International License.