최신뉴3

Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent

7월 19, 2024

133

비공개:-interpretability-in-action:-exploratory-analysis-of-vpt,-a-minecraft-agent — 비공개: Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent

arXiv:2407.12161v1 Announce Type: new
Abstract: Understanding the mechanisms behind decisions taken by large foundation models in sequential decision making tasks is critical to ensuring that such systems operate transparently and safely. In this work, we perform exploratory analysis on the Video PreTraining (VPT) Minecraft playing agent, one of the largest open-source vision-based agents. We aim to illuminate its reasoning mechanisms by applying various interpretability techniques. First, we analyze the attention mechanism while the agent solves its training task – crafting a diamond pickaxe. The agent pays attention to the last four frames and several key-frames further back in its six-second memory. This is a possible mechanism for maintaining coherence in a task that takes 3-10 minutes, despite the short memory span. Secondly, we perform various interventions, which help us uncover a worrying case of goal misgeneralization: VPT mistakenly identifies a villager wearing brown clothes as a tree trunk when the villager is positioned stationary under green tree leaves, and punches it to death.

News Week
Magazine PRO

Company

Interpretability in Action: Exploratory Analysis of VPT, a Minecraft Agent

LEAVE A REPLY Cancel reply

About us

Company

The latest

ABB Robotics는 비전 AI를 가속화하기 위해 Landingai에 투자합니다

RETHINK RETHINK ROBOTICS가 다시 종료됩니다

원격 수술에서 자율성으로 : Boston Dynamics의 아틀라스 훈련 내부

News WeekMagazine PRO

Company

관련된 글:

관련된 글:

LEAVE A REPLY Cancel reply

About us

Company

The latest

ABB Robotics는 비전 AI를 가속화하기 위해 Landingai에 투자합니다

RETHINK RETHINK ROBOTICS가 다시 종료됩니다

원격 수술에서 자율성으로 : Boston Dynamics의 아틀라스 훈련 내부

News Week
Magazine PRO