A comparative study of language models for named entity recognition from drug labels for polypharmacy risk assessment using bayesian networks
https://doi.org/10.25881/18110193_2026_2_18
Abstract
Background. Currently, there are no algorithms capable of accounting for interactions of three or more concomitantly administered medications, which is necessary to reduce the risks of side effects associated with polypharmacy. To develop such an algorithm, we propose using a Bayesian network – a graphical probabilistic model based on a semantic graph constructed using information extracted from medical texts. However, constructing such a network requires analyzing a vast amount of available medical texts, which is extremely time-consuming and intellectually demanding for medical professionals, necessitating the automation of this process.
Objective. To compare various natural language processing architectures and deep learning technologies for the creation a tool for automatic analysis of medical text information on pharmacokinetics and pharmacodynamics, side effects, and interactions to create semantic graphs, subsequently used to build a Bayesian network for drug interaction analysis.
Methods. Natural language processing can be used to automatically tag medical instruction texts. In this study, deep learning models based on the BERT and T5 architectures were tested using the BIO format to extract named entities.
Results. BERT-based models showed the highest accuracy: RuBioBERT achieved an accuracy of 0.9569±0.0052, slightly outperforming ruBERT, which showed comparable results. It was found that pre-training on general texts provides an advantage in the initial stages, while specialized models such as RuBioBERT become more effective with an increasing number of training epochs. This is due to the fact that general texts provide a more universal language representation, while scientific medical articles have their own specific features which do not fully correspond to the structure of drug labels.
Conclusion. BERT models achieve near-expert-quality annotation, making them an effective tool for constructing semantic graphs of drug interactions. Automated NER annotation reduces expert workload and improves text processing accuracy. Therefore, the use of BERT models, such as ruBERT and RuBioBERT, is a promising method for automating medical text analysis and building platforms for assessing complex drug interactions.
About the Authors
T. A. KuropatkinaRussian Federation
KUROPATKINA T.A., PhD.
Moscow
N. V. Kilmishkin
Russian Federation
KILMISHKIN N.V.
Moscow
V. I. Panteleev
Russian Federation
PANTELEEV V.I., PhD.
Moscow
D. D. Kubrakov
Russian Federation
KUBRAKOV D.D.
Moscow
P. M. Ivanova
Russian Federation
IVANOVA P.M.
Moscow
Yu. P. Titov
Russian Federation
TITOV YU.P., PhD.
Moscow
S. D. Valentey
Russian Federation
VALENTEY S.D., DSc, Professor
Moscow
N. L. Shimanovsky
Russian Federation
SHIMANOVSKY N.L., DSc, Professor, Corresponding Member of the Russian Academy of Sciences
Moscow
References
1. Koroleva M.V., Ilinitskiy A.N., Kudashkina E.V., Korshun E.I., Sharova A.A., Polev A.V. Sovremennye napravleniya farmakoterapii geriatricheskikh pacientov: polimorbidnost' — polipragmaziya — depreskraibing. Modern directions in pharmacotherapy for geriatric patients: multimorbidity — polypharmacy — deprescribing. Sovremennye problemy zdorov'ya i meditsinskoy statistiki. 2019; 3: 150-171. (In Russ.).
2. Cucinotta D. Multimorbidity and polypharmacy: a risk factor for older patients. Acta Biomed. 2022; 93(2): e2022137.
3. Sychev D.A., Otelenov V.A., Krasnova N.M., Ilina E.S. Polipragmaziya: vzglyad klinicheskogo farmakologa. Polypharmacy: a clinical pharmacologist's perspective. Terapevticheskii arkhiv. 2016; 12: 94-102. (In Russ.)
4. Izmozherova N.V., Popov A.A., Kuryndina A.A., Gavrilova E.I., Shambatov M.A., Bakhtin V.M. Polimorbidnost' i polipragmaziya u pacientov vysokogo i ochen' vysokogo serdechno-sosudistogo riska. Multimorbidity and polypharmacy in patients with high and very high cardiovascular risk. Ratsional'naya farmakoterapiya v kardiologii. 2022; 1: 20-26. (In Russ.)
5. drugs.com. Available at: https://www.drugs.com. Accessed 07.06.2025.
6. medscape.com. Available at: https://www.medscape.com. Accessed 05.07.2025.
7. DrugBank: Drug Interaction Checker Available at: https://www.drugbank.com/clinical/drug_drug_interaction_checker. Accessed 05.07.2025.
8. WebMD. Available at: https://www.webmd.сom. Accessed 06.07.2025.
9. Murali K., Kaur S., Prakash A, Medhi B. Artificial intelligence in pharmacovigilance: Practical utility. Indian J Pharmacol. 2019; 51(6): 373-376.
10. Li Z., Wu W, Kang H. Machine Learning-Driven Metabolic Syndrome Prediction: An International Cohort Validation Study. Healthcare. 2024; 12(24): 2527.
11. Rodrigues P.P., Ferreira-Santos D., Silva A., Polónia J., Ribeiro-Vaz I. Causality assessment of adverse drug reaction reports using an expert-defined Bayesian network. Artif. Intell. 2018; 91: 12-22.
12. Wang L., Hao H., Yan X., et al. From biomedical knowledge graph construction to semantic querying: a comprehensive approach. Sci Rep. 2025; 15: 8523.
13. Polotskaya K., Muñoz-Valencia C.S., Rabasa A., Quesada-Rico J.A., Orozco-Beltrán D., Barber X. Bayesian Networks for the Diagnosis and Prognosis of Diseases: A Scoping Review. Machine Learning and Knowledge Extraction. 2024; 6(2): 1243-1262.
14. Arora P., Boyne D., Slater J.J., Gupta A., Brenner D.R., Druzdzel M.J. Bayesian networks for risk prediction using real-world data: a tool for precision medicine. Value Health. 2019; 22(4): 439-445.
15. Sarem M., Saeed F., Alsaeedi A., et al. A survey on text preprocessing techniques based on NLP tasks. Multimedia Tools and Applications. 2022; 81(20): 28443-28470.
16. Единая государственная система регистрации лекарственных средств и медицинских изделий. [доступ от 07.07.2025]. Доступ по ссылке: https://grls.rosminzdrav.ru/Default.aspx.
17. Boytsov S.A., Bubnova M.G., Vasyuk Y.A., et al. Khronicheskaya serdechnaya nedostatochnost’. Klinicheskie rekomendatsii 2024. Rossiiskii kardiologicheskii zhurnal. 2024; 29(11): 6162. (In Russ.)
18. Международная классификация болезней 10-го пересмотра (МКБ-10) [интернет]. [доступ от 07.07.2025]. Доступ по ссылке: https://mkb-10.com.
19. export_from_pdf.py. GitHub. Available at: https://github.com/kalengul/graph_model_midicine/blob/main/doccano/data_jsonl_export/export_from_pdf.py — Emhyr713 / graph_model_midicine. Accessed 30.06.2025.
20. Finding. GitHub Available at: https://github.com/kalengul/graph_model_midicine/tree/main/semantic_graph/finding. — Emhyr713 / graph_model_midicine. Accessed 30.06.2025.
21. Tabassum A., Patil D.R. A Survey on Text Pre-Processing & Feature Extraction Techniques in Natural Language Processing. IRJET. 2020; 7(6).
22. Python 3.11 [интернет]. Версия 3.11 [доступ от 08.07.2025]. Доступ по ссылке: https://docs.python.org/3.11/
23. Кукушкин А.Е. Razdel — сегментация русскоязычного текста на токены и предложения [интернет]. [доступ от 11.07.2025]. Доступ по ссылке: https://github.com/natasha/razdel
24. Alshammari N., Alanazi S. The impact of using different annotation schemes on named entity recognition. Egypt Inform J. 2021; 22: 295-302.
25. Creater_graph. GitHub. Available at: https://github.com/kalengul/graph_model_midicine/blob/main/Creater_graph/mark_BIO_spacy.py. — Emhyr713 / graph_model_midicine. Accessed 30.06.2025.
26. Библиотека spaCy [интернет]. [доступ от 10.07.2025]. Доступ по ссылке: https://github.com/explosion/spaCy.
27. Кукушкин А.Е. Razdel — сегментация русскоязычного текста на токены и предложения [интернет]. [доступ от 11.07.2025]. Доступ по ссылке: https://github.com/natasha/razdel.
28. mark_sent_1.py. GitHub. Available at: https://github.com/kalengul/graph_model_midicine/blob/ main/semantic_graph/full_pipeline/mark_sent_1.py . — Emhyr713 / graph_model_midicine. Accessed 30.06.2025.
29. Devlin J., Chang M.-W., Lee K., Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. In: Burstein J., Doran C., Solorio T., eds. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol. 1 (Long and Short Papers). Minneapolis, MN: Association for Computational Linguistics; 2019: 4171-4186.
30. Kuratov Y., Arkhipov M. Adaptation of deep bidirectional multilingual transformers for Russian language. arXiv preprint arXiv: 1905.07213. 2019.
31. Bobrova E.V., Makanov A.Z., Osnovin S.S., et al. Generating medical opinions and classification according to Bethesda using deep learning. Int J Open Inf Technol. 2023; 11(10): 119-129.
32. Ni J., Hernandez Abrego G., Constant N., Ma J., Hall K., Cer D., Yang Y. Sentence-T5: scalable sentence encoders from pre-trained text-to-text models. In: Findings of the Association for Computational Linguistics: ACL 2022. Dublin, Ireland: Association for Computational Linguistics; 2022: 1864-1874.
33. ruT5-base: модель для генерации текста на русском языке [интернет]. [доступ от 12.07.2025]. Доступ по ссылке: https://huggingface.co/ai-forever/ruT5-base.
34. ruT5-base-multitask: модель для выполнения нескольких текстовых задач на русском языке [интернет]. [доступ от 14.07.2025]. Доступ по ссылке: https://huggingface.co/cointegrated/rut5-base-multitask.
35. Doccano: инструмент с открытым исходным кодом для аннотации текстовых данных в задачах NLP [интернет]. [доступ от 14.07.2025]. Доступ по ссылке: https://github.com/doccano/doccano.
36. Nadeau D., Sekine S. A survey of named entity recognition and classification. Lingvisticæ Investigationes. 2007; 30(1): 3-26.
37. Powers D. Evaluation: From precision, recall and F-factor to ROC, informedness, markedness & correlation. Machine Learning Technology. 2008; 2: 2229-3981.
Review
For citations:
Kuropatkina T.A., Kilmishkin N.V., Panteleev V.I., Kubrakov D.D., Ivanova P.M., Titov Yu.P., Valentey S.D., Shimanovsky N.L. A comparative study of language models for named entity recognition from drug labels for polypharmacy risk assessment using bayesian networks. Medical Doctor and Information Technologies. 2026;(2):18-35. (In Russ.) https://doi.org/10.25881/18110193_2026_2_18
JATS XML