Scholarly article · 2025
Public Perceptions and Barriers to Tuberculosis Treatment in Korea: A Large Language Model-Based Analysis of Naver Knowledge-iN Data from 2002 to 2024.
Park H, Kim S, Kim G, Chang S, Shin JG, Ahn S
원문 정보 보기의료 AI 및 데이터과학 · Medical AI & Data Science
한국어 요약
이 연구는 2002~2024년 한국 네이버 지식iN의 결핵 관련 질문 44,174건을 대규모언어모델(LLM)로 분석하여 결핵 치료에 대한 대중의 인식과 우려를 파악하고, 비정형 보건의료 데이터 분석 도구로서 LLM의 효과를 평가했다.
핵심 내용
- 특정 항결핵제가 언급된 질문 919건에서는 리팜핀이 31.8%, 이소니아지드가 31.6%로 가장 자주 언급되었다.
- 항결핵제 관련 질문 10,044건 중에서는 치료·관리상의 어려움이 44.8%로 가장 큰 범주였다.
- 감염 가능성과 사회적 영향에 관한 질문 583건의 분석에서 헌혈과 이민 자격에 관한 기존에 확인되지 않았던 우려가 나타났으며, 고용 관련 우려가 가장 큰 독립 하위집단이었다(20.6%).
- 권장 독자
- 결핵 치료에 관한 대중의 우려, 온라인 보건의료 데이터 분석, LLM 기반 텍스트 분석에 관심이 있는 임상의, 보건의료 연구자, 데이터과학자 및 고급 학습자.
- 범위와 한계
- 분석 자료는 네이버 지식iN에 게시된 온라인 질문으로, 전체 한국 인구의 인식이나 경험을 대표한다고 단정할 수 없다. 또한 제공된 초록만으로는 질문의 표본 구성, 세부적인 인간 연구자·전통적 방법과의 비교 절차, LLM 평가의 구체적 기준을 확인할 수 없다.
- 연구 맥락
- 비정형 온라인 보건의료 데이터에 LLM, 텍스트 임베딩, 차원 축소 및 군집화를 적용한 연구로서 의료 AI 및 데이터과학 분야의 공중보건 데이터 분석 주제와 연결된다.
English summary
This study used large language models (LLMs) to analyze 44,174 tuberculosis-related questions posted on Naver Knowledge-iN in Korea from 2002 to 2024. It examined public perceptions and concerns about tuberculosis treatment and evaluated LLMs as tools for analyzing unstructured healthcare data.
Key points
- Among 919 questions that mentioned specific antitubercular medications, rifampin and isoniazid were the most frequently referenced, accounting for 31.8% and 31.6%, respectively.
- Among 10,044 questions about antitubercular medications, treatment and management challenges formed the largest category at 44.8%.
- Analysis of 583 questions concerning infectivity and social implications identified previously unrecognized concerns about blood donation and immigration eligibility. Employment-related concerns were the largest distinct subgroup at 20.6%.
- Audience
- Clinicians, public-health researchers, health data scientists, and advanced students interested in public concerns about tuberculosis treatment, online healthcare data, and LLM-based text analysis.
- Scope and limitations
- The dataset consisted of online questions posted on Naver Knowledge-iN and therefore cannot be assumed to represent the perceptions or experiences of the Korean population as a whole. Based on the supplied abstract, the sampling composition, detailed procedures for comparison with human researchers and traditional methods, and specific criteria used to evaluate the LLMs cannot be determined.
- Research context
- The study applies LLMs, text embeddings, dimensionality reduction, and clustering to unstructured online healthcare data, connecting it thematically with medical AI and data-science approaches to public-health research.
AI-assisted summary based on the cited primary source. Model gpt-5.6-luna, prompt academic-hub-v1, generated 2026-08-29T02:32:39+00:00.