TY - JOUR
T1 - Feasibility of retrieval-augmented generation for large language models with Japanese input in radiotherapy
AU - Takahashi, Yoshiyuki
AU - Kadoya, Noriyuki
AU - Arai, Kazuhiro
AU - Tanno, Hikaru
AU - Tanaka, Shohei
AU - Katsuta, Yoshiyuki
AU - Hoshino, Taichi
AU - Harada, Hinako
AU - Omata, So
AU - Yamamoto, Takaya
AU - Umezawa, Rei
AU - Yasui, Keisuke
AU - Hayashi, Naoki
AU - Jingu, Keiichi
N1 - Publisher Copyright:
© The Author(s) 2026. Published by Oxford University Press on behalf of The Japanese Radiation Research Society and Japanese Society for Radiation Oncology. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
PY - 2026/5
Y1 - 2026/5
N2 - Large language models (LLMs) have recently gained attention for their potential. However, concerns remain regarding their reliability due to limitations such as hallucinations and insufficient domain-specific knowledge. Retrieval-augmented generation (RAG) has emerged as a promising approach, enabling LLMs to reference external knowledge sources and generate accurate outputs. We aimed to clarify the potential of RAG-enhanced LLMs with Japanese input in the field of radiotherapy. This was assessed by evaluating performance on three certification examinations in Japan: the Japanese Medical Physicist Examination, the Japanese Board Examination for Radiologists, and the Japanese Board Examination for Radiation Oncologists. In this study, we constructed a RAG system named Rad-Hub, consisting of a Japanese Radiotherapy Knowledge Database (JRKD) and a retrieval framework built on Microsoft Azure. The JRKD was populated with 32 Japanese radiotherapy textbooks and clinical guidelines. We assessed its utility by inputting all multiple-choice questions from the three examinations into ChatGPT-4o, both with and without Rad-Hub, and recording the answers. They were then compared with reference answers determined by experienced medical physicists and radiation oncologists. Rad-Hub improved accuracy across all examinations. Accuracy increased from 77.0% ± 2.6% to 84.6% ± 1.5% in the Medical Physicist examination, from 74.9% ± 2.0% to 82.1% ± 1.1% in the Radiologist examination, and from 55.6% ± 4.4% to 71.7% ± 4.5% in the Radiation Oncologist examination. Performance gains ranged from 7.2% to 16.1%. These findings highlight the potential of RAG-enhanced LLMs, particularly ChatGPT-4o with Rad-Hub, for integration into radiotherapy applications, such as educational and clinical decision assistance.
AB - Large language models (LLMs) have recently gained attention for their potential. However, concerns remain regarding their reliability due to limitations such as hallucinations and insufficient domain-specific knowledge. Retrieval-augmented generation (RAG) has emerged as a promising approach, enabling LLMs to reference external knowledge sources and generate accurate outputs. We aimed to clarify the potential of RAG-enhanced LLMs with Japanese input in the field of radiotherapy. This was assessed by evaluating performance on three certification examinations in Japan: the Japanese Medical Physicist Examination, the Japanese Board Examination for Radiologists, and the Japanese Board Examination for Radiation Oncologists. In this study, we constructed a RAG system named Rad-Hub, consisting of a Japanese Radiotherapy Knowledge Database (JRKD) and a retrieval framework built on Microsoft Azure. The JRKD was populated with 32 Japanese radiotherapy textbooks and clinical guidelines. We assessed its utility by inputting all multiple-choice questions from the three examinations into ChatGPT-4o, both with and without Rad-Hub, and recording the answers. They were then compared with reference answers determined by experienced medical physicists and radiation oncologists. Rad-Hub improved accuracy across all examinations. Accuracy increased from 77.0% ± 2.6% to 84.6% ± 1.5% in the Medical Physicist examination, from 74.9% ± 2.0% to 82.1% ± 1.1% in the Radiologist examination, and from 55.6% ± 4.4% to 71.7% ± 4.5% in the Radiation Oncologist examination. Performance gains ranged from 7.2% to 16.1%. These findings highlight the potential of RAG-enhanced LLMs, particularly ChatGPT-4o with Rad-Hub, for integration into radiotherapy applications, such as educational and clinical decision assistance.
KW - artificial intelligence
KW - large language model
KW - medical physicist
KW - radiation oncologist
KW - radiotherapy
KW - retrieval-augmented generation
UR - https://www.scopus.com/pages/publications/105040358661
UR - https://www.scopus.com/pages/publications/105040358661#tab=citedBy
U2 - 10.1093/jrr/rrag019
DO - 10.1093/jrr/rrag019
M3 - Article
C2 - 42105267
AN - SCOPUS:105040358661
SN - 0449-3060
VL - 67
SP - 420
EP - 429
JO - Journal of Radiation Research
JF - Journal of Radiation Research
IS - 3
ER -