본문 바로가기
  • 책상 밖 세상을 경험할 수 있는 Playground를 제공하고, 수동적 학습에서 창조의 삶으로의 전환을 위한 새로운 라이프 스타일을 제시합니다.

NLP113

[2026-1] 백승우 - Self-Improving Pretraining:using post-trained models to pretrain better models Self-Improving Pretraining: using post-trained models to pretrain better modelsEnsuring safety, factuality and overall quality in the generations of large language models is a critical challenge, especially as these models are increasingly deployed in real-world applications. The prevailing approach to addressing these issues involvearxiv.org 2026. 2. 4.
[2026-1] 백승우 - UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action Generation UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action GenerationYuanzhang Lin, Zhe Zhang, He Rui, Qingao Dong, Mingyi Zhou, Jing Zhang, Xiang Gao, Hailong Sun. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.aclanthology.org 2026. 1. 28.
[2025-2] 최민서 - SimPO: Simple Preference Optimization with a Reference-Free Reward [논문링크] https://arxiv.org/abs/2405.14734 SimPO: Simple Preference Optimization with a Reference-Free RewardDirect Preference Optimization (DPO) is a widely used offline preference optimization algorithm that reparameterizes reward functions in reinforcement learning from human feedback (RLHF) to enhance simplicity and training stability. In this work, we proposarxiv.org DPO에 대해 잘 모른다면 논문을 이해하는데 힘.. 2025. 12. 31.
[2025-2] 정유림 - Descending through a Crowded Valley —Benchmarking Deep Learning Optimizers paper link :https://arxiv.org/pdf/2007.01547 Descending through a Crowded Valley— Benchmarking Deep Learning Optimizers (ICML 2021)딥러닝에서 optimizer 선택은 중요한 결정 중 하나.Adam, SGD 부터 수많은 Adam 변형까지, 최근 수년간 제안된 optimizer는 수백개에 이른다.이 중 실제로 얼마나 의미 있는 차이가 있는지에 대한 대규모의 체계적인 optimizer 벤치마킹 연구 논문.논문 결과 요약optimizer 성능은 task-dependent어떤 optimizer도 모든 task에서 좋진않았음.여러 optimizer를 default로 설정해서 돌려보는것이 성능면에서 효율적인 선택... 2025. 12. 19.