Video
연구용 AI도구 최신동향 업데이트 (2025.4)
YouTube에서 보기
의료 AI 및 데이터과학 · Medical AI & Data Science
한국어 요약
이 발표는 2025년 4월 기준 학술 및 과학 연구용 AI 도구의 최근 동향을 정리한다. 주요 내용은 추론형 대규모 언어모델과 자율형 AI 에이전트, 문헌 검색과 심층 연구, 자연어 기반 데이터 분석과 시각화, 과학적 글쓰기 지원, 그리고 연구에서의 정책과 윤리적 한계이다.
핵심 내용
- 추론형 모델은 응답 생성 단계에서 추가 계산을 사용해 복잡한 추론, 코딩, 수학 문제의 정확도를 높일 수 있지만, 더 많은 토큰과 긴 응답 시간이 필요하며 단순한 문장 다듬기에는 이점이 제한적이라고 설명된다.
- 자율형 AI 에이전트는 계획, 기억, 도구 사용, 반복적 실행을 결합한다. 발표에서는 한 에이전트가 시계 바늘 각도 문제를 11분 40초 동안 코드 작성과 웹 검색을 포함해 처리했지만 최종적으로 오답을 낸 사례를 제시한다.
- Sakana AI의 The AI Scientist는 아이디어 생성, 문헌 기반 신규성 확인, 시뮬레이션 실험, 도표 작성, 원고 작성 및 자동 심사를 포함하는 연구 자동화 파이프라인으로 소개된다. 발표에 따르면 2025년 3월 ICLR 2025 워크숍에 제출된 생성 논문 3편 중 1편이 심사 점수 6.3점으로 채택되었다고 보고되었다. 이 결과는 특정 사례에 해당한다는 점에 유의해야 한다. AI 기반 문헌 도구는 검색 증강 생성 방식으로 색인된 문헌을 검색해 인용 환각을 줄이는 것을 목표로 한다. Elicit, Consensus.app 및 여러 Deep Research 기능은 문헌 선별, 요약 표 작성, 근거 연결형 합성 보고서 생성을 지원하는 것으로 제시된다. NotebookLM은 업로드한 문서에 근거한 질의, 인용 지점, 오디오 개요, 대화형 모드 및 개념 마인드맵을 제공하는 도구로 설명된다. ChatGPT Data Analyst와 Google Colab with Gemini는 원자료에서 탐색적 분석 계획과 Python 코드를 생성하고, 시각화와 회귀분석을 수행하며, 후속 기계학습 절차를 제안할 수 있는 사례로 제시된다. 발표의 와파린 용량 자료 예시에서는 VKORC1 유전형이 가장 강한 예측변수로 확인되었고 R2 값은 0.53이었다. Claude의 SVG 생성, Napkin AI, ChatGPT 이미지 생성, Gamma.app은 각각 편집 가능한 벡터 도식, 템플릿 기반 인포그래픽, 시각적 모형과 도식, 발표자료 생성을 지원하는 선택지로 비교된다. NEJM AI의 2023년 12월 편집 입장은 언어 장벽을 낮추고 원고 가독성을 높이는 LLM 사용을 장려하지만, AI 탐지기의 오탐과 미탐 및 비원어민 영어 연구자에 대한 불균형한 영향을 이유로 전면 금지는 비현실적이라고 설명한다. 미출판 원고와 연구자료의 기밀성을 위해 모델 학습 사용 설정을 확인하고 임시 대화나 기업용 보안 설정을 활용해야 한다고 발표는 강조한다. 문법 교정과 대필의 경계도 불분명한 정책 문제로 제시된다. 발표는 AI의 유창한 설명을 개인의 깊은 이해로 오인하는 설명 깊이의 환상, AI가 가능한 모든 가설을 탐색한다고 믿는 탐색 폭의 환상, AI 결과를 객관적이라고 간주하는 객관성의 환상을 경고한다. 자동화에 대한 의존은 기초 능력의 저하, 감시 과정에서의 경계심 감소, 자동화 실패 시 필요한 수동 개입 능력의 약화 및 상황 인식 상실을 초래할 수 있으므로 AI 활용 능력과 결과 비판 교육이 필요하다고 논의한다. METR 자료를 바탕으로 발표는 복잡한 소프트웨어 및 연구 과제에서 50퍼센트 성공률을 보이는 작업 지속 시간이 약 7개월마다 두 배로 증가했다고 설명하며, 2030년경 장기간 연구 과제를 자율적으로 수행할 가능성을 외삽한다. 이는 발표의 전망이며 확정된 결과가 아니다.
- 권장 독자
- 학술 및 의료 연구자, 의학교육자, 대학원생, 교수진, 과학적 글쓰기와 데이터 분석에 관심 있는 독자를 위한 내용이다.
- 범위와 한계
- 이 요약은 제공된 발표 증거 노트만을 바탕으로 한다. 발표에서 제시된 도구 사례, 성능 수치, 문헌 인용 및 2030년 전망은 독립적인 체계적 검토나 재현성 평가로 검증된 결과로 간주할 수 없다. 도구의 성능은 모델 버전, 자료 접근성, 프롬프트, 검증 절차 및 과제에 따라 달라질 수 있으며, 발표는 모든 도구의 비교 성능이나 오류율을 제시하지 않는다.
- 연구 맥락
- 이 내용은 의료 AI와 데이터과학에서 연구 문헌 탐색, 데이터 분석, 시각화, 과학적 의사소통 및 연구 자동화가 어떻게 결합되는지를 이해하는 주제와 연결된다.
English summary
This presentation reviews developments in AI tools for academic and scientific research as of April 2025. It covers reasoning-oriented large language models and autonomous AI agents, literature search and deep research, natural-language data analysis and visualization, scientific writing support, and policy and ethical limitations in research use.
Key points
- Reasoning models use additional computation during response generation and may improve accuracy on complex reasoning, coding, and mathematical tasks. They require more tokens and longer response times, however, and offer limited benefit for simple text polishing.
- Autonomous AI agents combine planning, memory, tool use, and iterative execution. The presentation describes a test in which an agent spent 11 minutes and 40 seconds using code generation and web searches to solve a clock-hand angle problem but ultimately produced an incorrect answer.
- Sakana AI's The AI Scientist is presented as an automated research pipeline covering idea generation, literature-based novelty checking, in-silico experiments, plotting, manuscript drafting, and automated review. According to the presentation, one of three generated papers submitted to an ICLR 2025 workshop in March 2025 was accepted with a reviewer score of 6.3 out of 10. This was a specific reported case rather than a general performance estimate. Scholarly AI tools use retrieval-augmented generation to retrieve indexed literature with the aim of reducing citation hallucinations. Elicit, Consensus.app, and several Deep Research functions are presented as supporting paper screening, summary tables, and linked evidence synthesis. NotebookLM is described as supporting document-grounded questions, citation anchors, audio overviews, interactive questioning, and concept-based mind maps. ChatGPT Data Analyst and Google Colab with Gemini are presented as examples of systems that can generate exploratory analysis plans and Python code from raw data, produce visualizations and regression analyses, and suggest subsequent machine-learning workflows. In the presentation's Warfarin dosing example, VKORC1 genotype was identified as the strongest predictor, with an R2 value of 0.53. Claude SVG generation, Napkin AI, ChatGPT image generation, and Gamma.app are compared as options for editable vector diagrams, template-based infographics, visual mockups and diagrams, and presentation generation, respectively. The December 2023 NEJM AI editorial position is described as encouraging LLM use to reduce language barriers and improve manuscript readability, while viewing outright bans as impractical because AI detectors produce false positives and false negatives and can disproportionately affect non-native English writers. The presentation emphasizes protecting unpublished manuscripts and research data by checking whether data may be used for model training and by using temporary chats or enterprise security settings when appropriate. It also identifies the boundary between grammatical editing and ghostwriting as an unresolved policy issue. The presentation warns against an illusion of explanatory depth, in which fluent AI explanations are mistaken for deep personal understanding; an illusion of exploratory breadth, in which AI is assumed to search all possible hypotheses; and an illusion of objectivity, in which AI outputs are treated as unbiased. It discusses possible effects of automation dependence, including loss of foundational skills, reduced vigilance during monitoring, weakened manual intervention when automation fails, and loss of situational awareness. AI literacy and critical evaluation of outputs are therefore presented as necessary educational components. Drawing on METR data, the presentation states that the task horizon associated with a 50 percent success rate on complex software and research tasks has been doubling approximately every seven months, and extrapolates that agents may conduct month-long research tasks around 2030. This is a projection in the presentation, not an established result.
- Audience
- The material is intended for academic and medical researchers, medical educators, graduate students, faculty, and others interested in scientific writing and data analysis.
- Scope and limitations
- This summary is based only on the supplied presentation evidence notes. The tool examples, performance figures, cited literature, and 2030 projection should not be treated as independently verified findings from a systematic review or reproducibility assessment. Tool performance can vary with model version, data access, prompting, verification procedures, and task characteristics, and the presentation does not provide comparative performance or error rates for all tools.
- Research context
- The presentation is thematically connected to medical AI and data science through its focus on integrating literature retrieval, data analysis, visualization, scientific communication, and research automation.
AI-assisted summary based on the cited primary source. Model gpt-5.6-luna, prompt academic-hub-v2, generated 2026-08-29T03:38:32+00:00.