Large Language Model Alignment
Language models do not always give answers that match what people expect.

This research teaches models to respond in safer and
more helpful ways using feedback from people.
We also design simple ways to test how well the models are
working and make the training process easier.
Yejin Kim, Dongsuk Oh, and H. Howie Huang
Expert Systems with Applications 287 (2025): 128047
Pre-trained language models (PrLMs) trained via contrastive learning methods achieved state-of-the-art performance on various natural language processing (NLP) tasks. Most PrLMs for sentence embedding focuses on context similarity as an objective function …
Keywords:
Dependency parser, Sentence embeddings, Pre-trained language models, Graph encoder, Contrastive learning
Dongryul Oh, Sujin Kang, Heejin Kim, and Dongsuk Oh
Applied Sciences 15.5 (2025): 919–933
Small language models (SLMs) are increasingly utilized for on-device applications due to their ability to ensure user privacy, reduce inference latency, and operate independently of cloud infrastructure. However, their performance is often limited when processing complex data structures …
Keywords:
small language model (SLM), on-device AI, graph neural network (GNN), graph transformer, graph convolutional network (GCN) …
Dongsuk Oh, Yejin Kim, Hodong Lee, H. Howie Huang, and Heuiseok Lim
Proceedings of COLING (2022): 4585–4592
Recent pre-trained language models (PLMs) achieved great success on many natural language processing tasks through learning linguistic features and contextualized sentence representation. Since attributes captured in stacked layers of PLMs are not clearly identified, straightforward approaches …
Keywords:
Jeong, Seungwon, Dongsuk Oh, Kinam Park, and Heuiseok Lim
Applied Sciences 12.9 (2022): 4099
Unlike previous dialogue-based question-answering (QA) datasets, DREAM, multiple-choice Dialogue-based REAding comprehension exaMination dataset, requires a deep understanding of dialogue. Many problems require multi-sentence reasoning, whereas some require commonsense …
Keywords:
dialogue-based multiple-choice QA, commonsense reasoning, semantic search, pre-trained language models, deep learning
PU-GEN: Enhancing generative commonsense reasoning for
language models with human-centered knowledge
Jaehyung Seo, Dongsuk Oh, Sugyeong Eo, Chanjun Park, Kisu Yang, Hyeonseok Moon, Kinam Park, and Heuiseok Lim
Knowledge-Based Systems 256 (2022):
Generative commonsense reasoning refers to the ability of a language model to generate a sentence with a given concept-set based on compositional generalization and commonsense reasoning. In the CommonGen challenge, which evaluates the capability of generative commonsense reasoning …
Keywords:
Text generation, Commonsense reasoning, Human-centered knowledge, Language model
Jang, Yoonna, Jungwoo Lim, Yuna Hur, Dongsuk Oh, Suhyune Son, Yeonsoo Lee, Donghoon Shin, Seungryong Kim, and Heuiseok Lim
Proceedings of the AAAI Conference on
Artificial Intelligence 36. 10 (2022): 10803-10812
Humans usually have conversations by making use of prior knowledge about a topic and background information of the people whom they are talking to. However, existing conversational agents and datasets do not consider such comprehensive information, and thus they have a limitation in generating …
Keywords:
Speech & Natural Language Processing (SNLP)
Whang, Taesun, Dongyub Lee, Dongsuk Oh, Chanhee Lee, Kijong Han, Dong-hun Lee, and Saebyeok Lee
Proceedings of the AAAI Conference on
Artificial Intelligence 35.16 (2021): 14041–14049
In this paper, we study the task of selecting the optimal response given a user and system utterance history in retrieval-based multi-turn dialog systems. Recently, pre-trained language models (e.g., BERT, RoBERTa, and ELECTRA) showed significant improvements in …
Keywords:
Conversational AI/Dialog Systems
Whang, Taesun, Dongyub Lee, Chanhee Lee, Kisu Yang, Dongsuk Oh, and Heuiseok Lim
Proceedings of Interspeec (2019): 1585-1589
We focus on multi-turn response selection in a retrieval-based dialog system. In this paper, we utilize the powerful pre-trained language model Bi-directional Encoder Representations from Transformer (BERT) for a multi-turn dialog system and propose a highly effective post-training …
Keywords:
Response selection, Human computer dialog system, Spoken language processing