NLP113 [2026-2] 정유림 - Neural Thickets:Diverse Task Experts Are Dense Around Pretrained Weights 논문제목: Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights링크: https://arxiv.org/pdf/2603.122282026년 3월 arXiv preprintICML 2026 Spotlight핵심 요약충분히 크고 잘 pretrain된 모델에서는 현재 pretrained weight 근처에 서로 다른 downstream task를 잘하는 task-specific weight solutions이 많이 존재한다. 저자들은 이런 weight-space 구조를 Neural Thicket이라고 부른다.그래서 굳이 gradient descent나 RL로 수백 step에 걸쳐 좋은 weight를 찾아갈 필요 없이,pretraine.. 2026. 9. 13. [2026-1] 정유림 - Simple and EffectiveMasked Diffusion Language Models paper : https://arxiv.org/abs/2406.07524 Simple and Effective Masked Diffusion Language ModelsWhile diffusion models excel at generating high-quality images, prior work reports a significant performance gap between diffusion and autoregressive (AR) methods in language modeling. In this work, we show that simple masked discrete diffusion is more perarxiv.org2024년 260613 기준 : 725회 인용 최근 언어 모델은 대부.. 2026. 6. 13. [2026-1] 박승원 - Hymba: A Hybrid-head Architecture for Small Language Models 논문 정보https://arxiv.org/pdf/2411.13676저자: Xin Dong*, Yonggan Fu* et al.소속: NVIDIA 핵심 아이디어Hymba는 Transformer의 attention mechanism과 Mamba 계열의 State Space Model을 결합한 hybrid architecture이다. 기존 hybrid 모델들이 attention layer와 SSM layer를 번갈아 쌓는 방식이었다면, Hymba는 한 layer 안에서 attention head와 SSM head를 병렬로 배치한다. 이를 통해 attention의 정확한 recall 능력과 SSM의 효율적인 context summarization 능력을 함께 활용한다. 또한 meta tokens, sliding.. 2026. 5. 15. [2026-1] 장인영 - Attention is All You Need https://arxiv.org/pdf/1706.037621. Introduction 1. 순환 모델의 한계 기존의 시퀀스 모델링에서는 RNN, LSTM, GRU와 같은 순환 신경망이 널리 사용되어 왔다.이러한 모델은 입력 시퀀스의 각 위치에 따라 계산을 나누어 수행하며,각 위치를 계산 시간의 단계와 정렬하여 이전 은닉 상태와 현재 입력을 기반으로 새로운 은닉 상태를 생성한다.이러한 구조는 본질적으로 순차적이기 때문에, 하나의 학습 예제 내에서 병렬 처리가 불가능하다.이 문제는 시퀀스 길이가 길어질수록 더욱 중요해지며, 메모리 제약으로 인해 여러 예제를 동시에 처리하는 데에도 한계를 발생시킨다.2. Attention의 등장과 한계Attention 메커니즘은 입력 또는 출력 시퀀스 내의 거리와 관계없이 의.. 2026. 3. 21. 이전 1 2 3 4 ··· 29 다음