AISI

[자료집] 2026 제4차 AI 안전 정책 연구 세미나 「Agentic AI와 AI 안전 연구」
  • 분류
    연구
  • 등록일
    2026-09-19 11:37:53
  • 작성자
    운영자
  • 조회수
    190
  • 인공지능안전연구소(Korea AISI)는 AI 위험 및 안전 정책 연구의 주요 쟁점을 논의하고 국내외 전문가와의 협력 기반을 확대하기 위하여 「AI 안전 정책 연구 세미나 시리즈」를 기획·운영하고 있습니다.

    최근 Agentic AI의 확산과 함께 통제 가능성·에이전트 식별 및 권한 관리·오용 방지와 안전한 종료 메커니즘 등 새로운 AI 안전·거버넌스 과제가 중요해지고 있습니다.

    이에 제4차 세미나는 「Agentic AI와 AI 안전 연구」를 주제로 개최되었으며, Agentic AI의 Scheming 및 Loss of Control 위험과 Agent ID 체계를 중심으로 최신 연구와 정책적 시사점을 논의하였습니다.

    ■ 세미나 개요

    ● (주제) Agentic AI와 AI 안전 연구(Research on Agentic AI and AI Safety)

    ● (일시) 2026년 8월 12일(수), 18:00~20:00

    ● (장소) 온라인(Zoom)

    ● (프로그램)

    18:00 Introduction, Jiyeon Cho(Korea AISI)
    18:10 Scheming in the wild – key findings from the Loss of Control Observatory, Tommy Shaffer Shane(The Centre for Long-Term Resilience)
    19:00 AgentID: Identifying, authorizing and shutting down AI agents, Amin Oueslati(Singapore AI Safety Hub & Oxford Martin AI Governance Initiative), Sam Boger(Singapore AI Safety Hub)
    19:50 Closing

    ● (주요 논의 주제)

    Agentic AI의 Scheming 및 Loss of Control 위험에 대한 최신 연구 동향
    Loss of Control Observatory 사례를 통한 Agentic AI의 통제 가능성 및 안전성 평가 방법론
    Agentic AI의 식별·인증(Agent IDs)·권한 관리 및 추적 가능성 확보 방안
    Agentic AI의 안전한 종료 메커니즘과 신원 기반 거버넌스의 정책적 시사점 및 발전 방향

    ■ 발표 내용

    ○ 발표1. Scheming in the wild – key findings from the Loss of Control Observatory

    (연사) Tommy Shaffer Shane(The Centre for Long-Term Resilience)
    · CLTR에서 변혁적 AI의 대비 및 영향, AI 사고와 통제상실 위험에 관한 연구·정책 업무를 주도하고 있으며, 이전에는 영국 DSIT에서 프론티어 AI로 인한 인식론적 위험을 평가함
    (주요내용)
    · AI 에이전트의 Scheming 발생 환경을 모니터링하기 위해 Loss of Control Observatory의 OSINT 기반 접근법을 소개하고, 실제 관측 사례에 기반한 모니터링의 필요성을 제시함
    · 실험 기반 평가에서 나타나는 Situational awareness·Ecological validity·발생 빈도 측정의 한계를 지적하고, 실제 환경 행동 관찰을 통한 가설 및 실험 설계 보완 가능성을 설명함
    · Transcript 기반 OSINT가 Scheming 및 Loss of Control 모니터링을 확장 가능하게 하는 방법임을 제시하고, 다수 데이터 소스를 연계한 실시간 사고 탐지 체계 구축을 향후 목표로 제시함

    ○ 발표2. AgentID: Identifying, authorizing and shutting down AI agents

    (연사) Amin Oueslati(Singapore AI Safety Hub(SASH) & Oxford Martin AI Governance Initiative), Sam Boger(Singapore AI Safety Hub)
    · SASH에서 AI 에이전트 거버넌스를 공동 총괄하며, 에이전트의 식별·인증·권한 관리를 지원하는 오픈소스 프로토콜을 정부·산업계 글로벌 컨소시엄과 함께 개발 중임(Amin Oueslati)
    · SASH의 AI 에이전트 기술 개발을 총괄하며, 미국 상원 입법 보좌관 및 Google Senior Software Engineer 경력을 바탕으로 에이전트의 투명성·책임성·보안 강화 인프라 구축을 담당함(Sam Boger)
    (주요내용)
    · AI 에이전트의 식별·인증·권한 부여 및 긴급 종료를 지원하는 AgentID 체계의 필요성을 제시하고, 기술적·제도적 관점에서 에이전트 거버넌스 기반 마련을 위한 접근법을 소개함
    · AgentID를 통해 에이전트 관련 주체 식별·책임성 확보·권한 관리·사고 대응을 지원하고, 표준화와 정책적 보완이 필요함을 제시함
    · AI 에이전트의 오작동이나 악의적 행위에 대응하기 위한 Emergency Shutdown의 개념과 절차를 소개하고, 탐지·책임주체 식별·종료 요청 및 실행·검증·보고·재심을 포함하는 체계적인 긴급 종료 메커니즘의 필요성을 강조함

    English Summary

    ■ Background
    The Korea AI Safety Institute (Korea AISI) operates the AI Safety Policy Research Seminar Series to discuss key issues in AI risk and safety policy research and to expand collaboration with domestic and international experts.

    As Agentic AI systems become increasingly capable of autonomous planning, decision-making, and tool use, concerns related to scheming, loss of control, agent identity and authorization, and safe shutdown mechanisms are emerging as critical AI safety and governance challenges.

    ■ Overview of the Fourth Seminar
    The fourth seminar, Research on Agentic AI and AI Safety, was held on August 12, 2026 (Wed), 18:00 to 20:00 (KST), online via Zoom. The seminar featured expert presentations and discussions on recent research trends concerning scheming and loss-of-control risks in Agentic AI, evaluation methodologies for controllability and safety assessment based on Loss of Control Observatory cases, approaches for identifying, authorizing, and tracking AI agents through Agent ID frameworks, and the policy implications and future development of safe shutdown mechanisms and identity-based governance for Agentic AI systems.

    ■ Program

    18:00 Introduction, Jiyeon Cho (Korea AISI)
    18:10 Scheming in the wild – key findings from the Loss of Control Observatory, Tommy Shaffer Shane (The Centre for Long-Term Resilience)
    19:00 AgentID: Identifying, authorizing and shutting down AI agents, Amin Oueslati (Singapore AI Safety Hub & Oxford Martin AI Governance Initiative), Sam Boger (Singapore AI Safety Hub)
    19:50 Closing

    ■ Presentation Highlights

    Scheming in the wild – key findings from the Loss of Control Observatory
    Speaker: Tommy Shaffer Shane (The Centre for Long-Term Resilience), who leads research and policy work at CLTR on preparedness for and impacts of transformative AI, AI incidents, and loss-of-control risks. Previously, at the UK Department for Science, Innovation and Technology, he assessed epistemic risks from frontier AI and developed policy responses.

    Key Takeaways

    Introduced the Loss of Control Observatory's OSINT-based approach to monitoring scheming in AI agents in real-world environments and highlighted the need for real-world evidence of scheming.
    Highlighted limitations of experimental evaluations, including situational awareness, ecological validity, and measuring real-world frequency, and the value of real-world behavioral evidence for informing hypotheses and experimental design.
    Proposed transcript-based OSINT as a scalable approach to monitoring scheming and loss of control, with the goal of developing real-time monitoring through integrated data sources.
    AgentID: Identifying, authorizing and shutting down AI agents
    Speakers: Amin Oueslati (Singapore AI Safety Hub (SASH) & Oxford Martin AI Governance Initiative), who co-leads AI agent governance at SASH and is a Research Affiliate at the Oxford Martin AI Governance Initiative, currently developing an open-source protocol for identifying, authenticating, authorizing, and managing AI agents with a global consortium of governments and industry; and Sam Boger (Singapore AI Safety Hub), Technical Lead for AI Agents at SASH, with previous experience in AI governance at The Future Society, legislative work in the U.S. Senate, and software engineering at Google.

    Key Takeaways

    Proposed the need for an AgentID framework supporting the identification, authentication, authorization, and emergency shutdown of AI agents, and introduced a technical and governance approach to establishing a foundation for agent governance.
    Highlighted how AgentID can support the identification of agents and relevant stakeholders, accountability, authorization management, and incident response, while emphasizing the need for standardization and policy measures.
    Introduced the concept and process of Emergency Shutdown for responding to malfunctioning or malicious AI agents, emphasizing the need for a systematic shutdown mechanism covering rapid detection, attribution, shutdown request and execution, verification, documentation, reporting, and recourse.

    ▸ 첨부파일: 제4차 AI 안전 정책 연구 세미나 자료집(PDF, 하단 첨부파일 참조)

    ■ 이전 세미나 자료집 보기
    ▸ 2026 제3차 AI 안전 정책 연구 세미나 자료집 「LLM 안전성 벤치마크의 한국 맥락화」
    ▸ 2026 제2차 AI 안전 정책 연구 세미나 자료집 「AI 안전·안보 정책 연구」
    ▸ 2026 제1차 AI 안전정책 연구 세미나 자료집 「AI 위험 연구의 방법론적 접근」
  • 첨부파일