Please use this identifier to cite or link to this item: http://lib.kart.edu.ua/handle/123456789/33240
Title: Методологія автоматизованої оцінки тематичної спорідненості наукових публікацій на основі семантичного аналізу
Other Titles: Methodology for automated assessment of thematic relatedness of scientific publications using semantic analysis
Authors: Іванюк, Олександр Ігорович
Ivaniuk, Oleksandr
Keywords: тематична спорідненість
ембединги
косинусна подібність
призначення рецензентів
дисертації
профілі науковців
OpenAI text-embedding-3-large
UMAP
DBSCAN
NDCG
Hit@k
рекомендаційні системи
текстова аналітика
машинне навчання
thematic relatedness
embeddings
cosine similarity
reviewer assignment
dissertations
researcher profiles
OpenAI text-embedding-3-large
UMAP
DBSCAN
NDCG
Hit@k
recommender systems
text analytics
machine learning
Issue Date: 2025
Publisher: Національний університет "Полтавська політехніка імені Юрія Кондратюка"
Citation: Іванюк О. І. Методологія автоматизованої оцінки тематичної спорідненості наукових публікацій на основі семантичного аналізу / О. І. Іванюк. Системи управління, навігації та зв'язку. 2025. Вип. 3. С. 96-100.
Abstract: UA: Актуальність. Автоматизована оцінка тематичної спорідненості між дисертаційними дослідженнями та профілями потенційних експертів потрібна для прозорого й відтворюваного добору офіційних опонентів і складу разових рад; відкриті дані NAQA.Svr роблять таку оцінку технічно можливою. Об’єкт дослідження: методи та засоби побудови семантичних профілів дисертацій і науковців та їх ранжування у спільному векторному просторі. Мета статті: розробити та емпірично перевірити методологію автоматизованої оцінки тематичної спорідненості наукових публікацій на основі семантичного аналізу. Результати дослідження. Сформовано корпус із 259 дисертацій, 662 профілів науковців і 3345 публікацій; тексти нормалізовано, назви публікацій стандартизовано засобами великої мовної моделі. Семантичні подання отримано за допомогою моделі OpenAI. Порівнювалися два варіанти подання дисертацій (повний опис; лише ключові слова) та науковців (публікації; ключові слова). Якість оцінювалася за фактичними призначеннями опонентів, використовуючи метрики NDCG, Hit@k і NDCG-lift. Найкращі результати стабільно показала комбінація «дисертація за ключовими словами» та «науковець за публікаціями», яка демонструє найвищу точність у верхніх позиціях і найбільший виграш відносно випадкового відбору. Висновки. Поєднання профілю дисертації за ключовими словами з профілем науковця, агрегованим за публікаціями, забезпечує найкращий баланс точності та стійкості і є практично доцільним для попереднього добору опонентів. Запропонована методологія відтворювана, масштабована й узгоджена з відкритими даними; перспективні напрями – урахування часової ваги публікацій, двомовності термінів та процедурних обмежень під час остаточного призначення.
EN: Relevance. Automated assessment of thematic relatedness between doctoral theses and profiles of potential experts is needed for transparent and reproducible selection of official reviewers and composition of one -time specialized councils; the open NAQA.Svr data make such assessment technically feasible. Object of research: methods and tools for constructing semantic profiles of theses and researchers and for ranking them in a shared vector space. Purpose of the article. To develop and empirically validate a methodology for automated assessment of thematic relatedness of scientific publications based on semantic analysis. Research results. We assembled a corpus of 259 theses, 662 researcher profiles, and 3345 publications; texts were normalized, and publication titles were standardized using a large language model. Semantic representations were obtained with an OpenAI model. We compared two variants of thesis profiling (full description; keywords only) and two variants of researcher profiling (by publications; by keywords). Quality was evaluated against actual reviewer assignments using the metrics NDCG, Hit@k, and NDCG-lift. The best results were consistently achieved by the combination “thesis by keywords” and “researcher by publications,” which yields the highest top-rank accuracy and the largest gain over random selection. Conclusions. Combining a keyword-based thesis profile with a publication-aggregated researcher profile provides the best balance of accuracy and robustness and is practically suitable for preliminary reviewer selection. The proposed methodology is reproducible, scalable, and aligned with open data; promising directions include incorporating temporal weights of publications, handling bilingual terminol ogy, and accounting for procedural constraints in final assignments.
URI: http://lib.kart.edu.ua/handle/123456789/33240
ISSN: 2073-7394 (print)
Appears in Collections:2025

Files in This Item:
File Description SizeFormat 
Ivaniuk.pdf444.34 kBAdobe PDFView/Open


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.