Scholarly article · 2025

Public Perceptions and Barriers to Tuberculosis Treatment in Korea: A Large Language Model-Based Analysis of Naver Knowledge-iN Data from 2002 to 2024.

Park H, Kim S, Kim G, Chang S, Shin JG, Ahn S

Healthc Inform Res

원문 정보 보기
의료 AI 및 데이터과학 · Medical AI & Data Science

한국어 요약

이 연구는 2002~2024년 한국 네이버 지식iN의 결핵 관련 질문 44,174건을 대규모언어모델(LLM)로 분석하여 결핵 치료에 대한 대중의 인식과 우려를 파악하고, 비정형 보건의료 데이터 분석 도구로서 LLM의 효과를 평가했다.

핵심 내용

  • 특정 항결핵제가 언급된 질문 919건에서는 리팜핀이 31.8%, 이소니아지드가 31.6%로 가장 자주 언급되었다.
  • 항결핵제 관련 질문 10,044건 중에서는 치료·관리상의 어려움이 44.8%로 가장 큰 범주였다.
  • 감염 가능성과 사회적 영향에 관한 질문 583건의 분석에서 헌혈과 이민 자격에 관한 기존에 확인되지 않았던 우려가 나타났으며, 고용 관련 우려가 가장 큰 독립 하위집단이었다(20.6%).
권장 독자
결핵 치료에 관한 대중의 우려, 온라인 보건의료 데이터 분석, LLM 기반 텍스트 분석에 관심이 있는 임상의, 보건의료 연구자, 데이터과학자 및 고급 학습자.
범위와 한계
분석 자료는 네이버 지식iN에 게시된 온라인 질문으로, 전체 한국 인구의 인식이나 경험을 대표한다고 단정할 수 없다. 또한 제공된 초록만으로는 질문의 표본 구성, 세부적인 인간 연구자·전통적 방법과의 비교 절차, LLM 평가의 구체적 기준을 확인할 수 없다.
연구 맥락
비정형 온라인 보건의료 데이터에 LLM, 텍스트 임베딩, 차원 축소 및 군집화를 적용한 연구로서 의료 AI 및 데이터과학 분야의 공중보건 데이터 분석 주제와 연결된다.

English summary

This study used large language models (LLMs) to analyze 44,174 tuberculosis-related questions posted on Naver Knowledge-iN in Korea from 2002 to 2024. It examined public perceptions and concerns about tuberculosis treatment and evaluated LLMs as tools for analyzing unstructured healthcare data.

Key points

  • Among 919 questions that mentioned specific antitubercular medications, rifampin and isoniazid were the most frequently referenced, accounting for 31.8% and 31.6%, respectively.
  • Among 10,044 questions about antitubercular medications, treatment and management challenges formed the largest category at 44.8%.
  • Analysis of 583 questions concerning infectivity and social implications identified previously unrecognized concerns about blood donation and immigration eligibility. Employment-related concerns were the largest distinct subgroup at 20.6%.
Audience
Clinicians, public-health researchers, health data scientists, and advanced students interested in public concerns about tuberculosis treatment, online healthcare data, and LLM-based text analysis.
Scope and limitations
The dataset consisted of online questions posted on Naver Knowledge-iN and therefore cannot be assumed to represent the perceptions or experiences of the Korean population as a whole. Based on the supplied abstract, the sampling composition, detailed procedures for comparison with human researchers and traditional methods, and specific criteria used to evaluate the LLMs cannot be determined.
Research context
The study applies LLMs, text embeddings, dimensionality reduction, and clustering to unstructured online healthcare data, connecting it thematically with medical AI and data-science approaches to public-health research.

AI-assisted summary based on the cited primary source. Model gpt-5.6-luna, prompt academic-hub-v1, generated 2026-08-29T02:32:39+00:00.