본문 바로가기
  • 책상 밖 세상을 경험할 수 있는 Playground를 제공하고, 수동적 학습에서 창조의 삶으로의 전환을 위한 새로운 라이프 스타일을 제시합니다.

Multi-Modal34

[2026-1] 정인아 - Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models 논문 제목 : Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models (Bis-JEPA)논문 링크 : https://arxiv.org/pdf/2602.18639 Introductiongenerative world model에서 JEPA 계열이 주목받고 있다.(2026)이는 raw observation을 reconstrction하지 않고 latent representation을 prediction하는 방식이다.Bis-JEPA는 base로 DINO-WM을 둔다.DINO-WM는 JEPA 아이디어를 가지고, DINOv2 feature를 사용해 latent dynamics를 학습한 연구다.r.. 2026. 8. 23.
[2026-1] 김지원 Self-Supervised Learning from Images with aJoint-Embedding Predictive Architecture 논문 제목 : Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture논문 링크: https://arxiv.org/pdf/2301.08243 배경 I-JEPA(Image-based Joint-Embedding Predictive Architecture)는 '라벨 없이 이미지의 semantic representation을 어떻게 학습할 것인가'라는 자기 지도 학습 문제를 다룸 학계에서는 두 가지 방법으로 이 문제를 다뤄왔음:1. View-invariance(SimCLR, BYOL, DINO) 같은 이미지를 증강(crop, color jitter 등)으로 여러 view로 변형한 뒤 변형된 view의 임베딩이 서로 .. 2026. 8. 23.
[2026-1] 백승우 - How Mobile World Model Guides GUI Agents? How Mobile World Model Guides GUI Agents?Recent advances in vision-language models have enabled mobile GUI agents to perceive visual interfaces and execute user instructions, but reliable prediction of action consequences remains critical for long-horizon and high-risk interactions. Existing mobiarxiv.org 2026. 5. 19.
[2026-1] 백승우 - Agent+P: Guiding UI Agents via Symbolic Planning Agent+P: Guiding UI Agents via Symbolic PlanningLarge Language Model (LLM)-based UI agents show great promise for UI automation but often hallucinate in long-horizon tasks due to their lack of understanding of the global UI transition structure. To address this, we introduce AGENT+P, a novel framework tarxiv.org 2026. 5. 19.