Please use this identifier to cite or link to this item:
http://lib.kart.edu.ua/handle/123456789/33240| Title: | Методологія автоматизованої оцінки тематичної спорідненості наукових публікацій на основі семантичного аналізу |
| Other Titles: | Methodology for automated assessment of thematic relatedness of scientific publications using semantic analysis |
| Authors: | Іванюк, Олександр Ігорович Ivaniuk, Oleksandr |
| Keywords: | тематична спорідненість ембединги косинусна подібність призначення рецензентів дисертації профілі науковців OpenAI text-embedding-3-large UMAP DBSCAN NDCG Hit@k рекомендаційні системи текстова аналітика машинне навчання thematic relatedness embeddings cosine similarity reviewer assignment dissertations researcher profiles OpenAI text-embedding-3-large UMAP DBSCAN NDCG Hit@k recommender systems text analytics machine learning |
| Issue Date: | 2025 |
| Publisher: | Національний університет "Полтавська політехніка імені Юрія Кондратюка" |
| Citation: | Іванюк О. І. Методологія автоматизованої оцінки тематичної спорідненості наукових публікацій на основі семантичного аналізу / О. І. Іванюк. Системи управління, навігації та зв'язку. 2025. Вип. 3. С. 96-100. |
| Abstract: | UA: Актуальність. Автоматизована оцінка тематичної спорідненості між дисертаційними дослідженнями та профілями потенційних експертів потрібна для прозорого й відтворюваного добору офіційних опонентів і складу разових рад; відкриті дані NAQA.Svr роблять таку оцінку технічно можливою. Об’єкт дослідження:
методи та засоби побудови семантичних профілів дисертацій і науковців та їх ранжування у спільному векторному просторі. Мета статті: розробити та емпірично перевірити методологію автоматизованої оцінки тематичної спорідненості наукових публікацій на основі семантичного аналізу. Результати дослідження. Сформовано
корпус із 259 дисертацій, 662 профілів науковців і 3345 публікацій; тексти нормалізовано, назви публікацій стандартизовано засобами великої мовної моделі. Семантичні подання отримано за допомогою моделі OpenAI. Порівнювалися два варіанти подання дисертацій (повний опис; лише ключові слова) та науковців (публікації; ключові слова). Якість оцінювалася за фактичними призначеннями опонентів, використовуючи метрики NDCG,
Hit@k і NDCG-lift. Найкращі результати стабільно показала комбінація «дисертація за ключовими словами» та
«науковець за публікаціями», яка демонструє найвищу точність у верхніх позиціях і найбільший виграш відносно випадкового відбору. Висновки. Поєднання профілю дисертації за ключовими словами з профілем науковця,
агрегованим за публікаціями, забезпечує найкращий баланс точності та стійкості і є практично доцільним для
попереднього добору опонентів. Запропонована методологія відтворювана, масштабована й узгоджена з відкритими даними; перспективні напрями – урахування часової ваги публікацій, двомовності термінів та процедурних обмежень під час остаточного призначення. EN: Relevance. Automated assessment of thematic relatedness between doctoral theses and profiles of potential experts is needed for transparent and reproducible selection of official reviewers and composition of one -time specialized councils; the open NAQA.Svr data make such assessment technically feasible. Object of research: methods and tools for constructing semantic profiles of theses and researchers and for ranking them in a shared vector space. Purpose of the article. To develop and empirically validate a methodology for automated assessment of thematic relatedness of scientific publications based on semantic analysis. Research results. We assembled a corpus of 259 theses, 662 researcher profiles, and 3345 publications; texts were normalized, and publication titles were standardized using a large language model. Semantic representations were obtained with an OpenAI model. We compared two variants of thesis profiling (full description; keywords only) and two variants of researcher profiling (by publications; by keywords). Quality was evaluated against actual reviewer assignments using the metrics NDCG, Hit@k, and NDCG-lift. The best results were consistently achieved by the combination “thesis by keywords” and “researcher by publications,” which yields the highest top-rank accuracy and the largest gain over random selection. Conclusions. Combining a keyword-based thesis profile with a publication-aggregated researcher profile provides the best balance of accuracy and robustness and is practically suitable for preliminary reviewer selection. The proposed methodology is reproducible, scalable, and aligned with open data; promising directions include incorporating temporal weights of publications, handling bilingual terminol ogy, and accounting for procedural constraints in final assignments. |
| URI: | http://lib.kart.edu.ua/handle/123456789/33240 |
| ISSN: | 2073-7394 (print) |
| Appears in Collections: | 2025 |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| Ivaniuk.pdf | 444.34 kB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.