Video

Can an LLM Be Used for Peer Review? Educational Material for Journal Publication Committee Member...

YouTube에서 보기
Can an LLM Be Used for Peer Review? Educational Material for Journal Publication Committee Member... 영상 썸네일
의료 AI 및 데이터과학 · Medical AI & Data Science

한국어 요약

이 영상은 대규모 언어 모델의 학술 동료심사 활용 가능성, 실제 심사에서의 AI 보조 사용 양상, 심사 결과에 미치는 영향, 그리고 학술지의 생성형 AI 정책을 검토한다. 제시된 연구에서 GPT-4의 인간 심사와의 의견 중복은 Nature 계열 학술지에서 약 30퍼센트, ICLR에서 약 40퍼센트로 보고되었으며, 서로 다른 논문에 대한 잘못 짝지은 비교에서는 약 1퍼센트로 감소했다. 이는 GPT-4의 평가가 일정 부분 논문별 내용을 반영했음을 시사한다. 한편 AI 보조 심사는 특정 학술대회에서 심사 점수와 경계선 논문의 게재 가능성 증가와 연관된 것으로 보고되었다. 영상은 AI 사용을 일률적으로 단속하기보다 최종 내용의 정확성, 인간의 검토, 기밀성 보호를 중심으로 정책을 설계해야 한다고 정리한다.

핵심 내용

  • Liang 등의 분석에서 GPT-4와 인간 심사의 의견 중복은 인간 심사자 간 중복과 비슷한 수준으로 보고되었고, 합의된 거절 논문에서 중복이 가장 높았다. 반면 높은 점수나 게재 논문에서는 인간 심사자 간 및 GPT-4와 인간 간의 일치도가 더 낮았다.
  • GPT-4는 여러 인간 심사자가 공통으로 제기한 비판과 인간 심사의 우선순위가 높은 의견을 비교적 잘 포착했지만, 문헌 인용의 누락, 연구 맥락, 기계학습의 절제 실험, 실제 과학적 신규성 평가는 상대적으로 덜 다루었다.
  • 308명의 저자 중 57.4퍼센트는 GPT-4 피드백을 유용하거나 매우 유용하다고 평가했고, 58.5퍼센트는 해당 시스템을 다시 사용하겠다고 답했다. 이 결과는 저자 대상의 전향적 설문에 근거한다.
권장 독자
학술지 편집자, 심사자, 연구자와 저자, 편집위원회 및 생성형 AI 정책을 수립하는 기관
범위와 한계
이 요약은 제공된 영상 발표의 증거 노트만을 바탕으로 한다. 주요 근거는 arXiv에 보고된 연구와 특정 학술지의 정책 사례이며, 영상 자체가 독립적인 체계적 문헌고찰은 아니다. AI 생성 여부 탐지에는 위양성과 위음성이 모두 존재하고, 비원어민 영어 작성자에서 위양성이 더 높을 수 있으므로 탐지 점수만으로 제재해서는 안 된다고 발표는 명시한다. 또한 AI 보조 심사와 점수 또는 게재 결과의 연관성이 인과관계를 확정하는 것은 아니다.
연구 맥락
이 내용은 의료 AI와 데이터과학에서 생성형 AI의 평가, 인간 검토, 재현성, 데이터 기밀성 및 연구 거버넌스를 다루는 Sangzin Ahn의 포트폴리오와 주제적으로 연결된다.

English summary

The video examines the feasibility of using large language models in academic peer review, patterns of AI-assisted reviewing, reported effects on review outcomes, and journal policies for generative AI. In the studies presented, overlap between GPT-4 and human reviews was approximately 30 percent for Nature family journals and 40 percent for ICLR, falling to approximately 1 percent when reviews were mismatched across papers. This suggests that GPT-4 critiques reflected paper-specific content to some extent. AI-assisted reviews were also reported to be associated with higher review scores and a higher acceptance probability for borderline submissions at one conference. The presentation concludes that policy should prioritize the accuracy of the final content, human oversight, and confidentiality rather than attempting to prohibit or police AI use in general.

Key points

  • In the analysis by Liang and colleagues, overlap between GPT-4 and human reviews was reported to be similar to human-to-human overlap and was highest for papers with consensual rejection decisions. Human-to-human and GPT-4-to-human agreement was lower for highly scored or accepted papers.
  • GPT-4 more often identified criticisms raised by multiple human reviewers and comments ranked as highly important in human reviews. It less consistently addressed missing citations and literature context, machine learning ablation experiments, and the assessment of true scientific novelty.
  • Among 308 authors in the prospective survey, 57.4 percent rated GPT-4 feedback as helpful or extremely helpful, and 58.5 percent said they would use the system again. These findings were based on the author survey described in the presentation.
Audience
Journal editors, peer reviewers, researchers and authors, editorial boards, and organizations developing generative AI policies
Scope and limitations
This summary is based only on the supplied evidence notes from the video presentation. Its main evidence comes from studies reported on arXiv and from a policy example at one journal; the video is not itself an independent systematic review. The presentation notes that AI detection tools have both false positive and false negative errors, with potentially higher false positive rates for non-native English writers, so sanctions should not rely solely on detector scores. In addition, associations between AI-assisted reviews and review scores or acceptance outcomes do not establish causation.
Research context
The topic has a neutral thematic connection to Sangzin Ahn's portfolio through medical AI and data science, particularly generative AI evaluation, human oversight, reproducibility, data confidentiality, and research governance.

AI-assisted summary based on the cited primary source. Model gpt-5.6-luna, prompt academic-hub-v2, generated 2026-08-29T03:48:14+00:00.