Publication
Discover highlights from our published research addressing challenges in
Digital Humanities and Natural Language Processing.
Dongryul Oh, Sujin Kang, Heejin Kim, and Dongsuk Oh
Small language models (SLMs) are increasingly utilized for on-device applications due to their ability to ensure user privacy, reduce inference latency, and operate independently of cloud infrastructure. However, their performance is often limited when processing complex data structures …
Keywords:
small language model (SLM), on-device AI, graph neural network (GNN), graph transformer, graph convolutional network (GCN) …
Eunsong Lee, Hyein Do, Minsu Kim, and Dongsuk Oh
Applied Sciences 15.13 (2025): 7561
This study proposes a new benchmark to evaluate the cultural understanding and natural language processing capabilities of large language models based on Sino-Korean words and four-character idioms. Those are essential linguistic and cultural assets in Korea …
large language models evaluation, cultural contextual understanding, Sino-Korean vocabulary, four-character idioms …
Keywords:
Sungeun Kim and Dongsuk Oh
Applied Sciences 15.6 (2025): 2971
The evaluation of creative writing has long been a complex and subjective process, made even more intriguing by the rise of advanced Artificial Intelligence (AI) tools like Large Language Models (LLMs). This study evaluates the potential of LLMs as reliable …
Keywords:
large language models (LLMs) evaluation, creative writing evaluation, creativity, AI evaluation, human evaluation
Yejin Kim, Dongsuk Oh, and H. Howie Huang
Expert Systems with Applications 287 (2025): 128047
Pre-trained language models (PrLMs) trained via contrastive learning methods achieved state-of-the-art performance on various natural language processing (NLP) tasks. Most PrLMs for sentence embedding focuses on context similarity as an objective function …
Keywords:
Dependency parser, Sentence embeddings, Pre-trained language models, Graph encoder, Contrastive learning
Changwon Ok, Eunkyeong Lee, and Dongsuk Oh
Proceedings of the 31st International Conference on Computational Linguistics (2025): 5168-5180
Recently, large language models (LLMs) have made significant progress through retrieval-augmented generation (RAG) and preference learning. However, they still exhibit issues such as confirmation bias, the tendency to favor information that confirms one’s beliefs, which remains largely …
Keywords:
Heejin Kim
In/Outside 56 (2024): 18–58
This study utilizes network analysis to explore structural unity in Renaissance plays, tracing the influence of medieval touring companies on 16th-century dramatic structures. Employing digital humanities methodologies, the research applies community detection algorithms and silhouette scores …
Keywords:
Shakespeare, Character Network Analysis, Digital Humanities, Mediality, Dramatic Structure, Silhouette Score
Dongsuk Oh, Jonghyeon Moon, Kyoungtae Park, Wonjun Kim, Seungho Yoo, Hyungwoo Lee, and Jiho Yoo
Expert Systems with Applications 249 (2024): 123620
With the increase in the aging population of many countries, the prevalence of neovascular age-related macular degeneration (nAMD) is expected to increase. Morphological parameters such as intraretinal fluid (IRF), subretinal fluid (SRF), subretinal hyperreflective material (SHRM) …
Keywords:
Graph convolution network, Transformer, Multiscale skip connection, Medical image segmentation, Retinopathy
Sungeun Kim, Jakyung Kim, and Dongsuk Oh
Journal of Digital Contents Society 25.9 (2024): 2479-2490
Large language models (LLMs) have demonstrated competitive performance across various domains, particularly in tasks requiring creativity, and thus offer a wide range of applications. This study evaluates the performance of large language models (LLMs), such as GPT-4, in generating creative …
Keywords:
GPT-4, Large Language Models (LLMs), Creative Writing, Creative Contents, Creativity Evaluation
Heejin Kim
In/Outside 55 (2023): 208–230
The advent of artificial intelligence (AI), particularly Large Language Models (LLMs), is poised to bring about profound societal changes. Despite the risks associated with AI, such as the production of inaccurate information, labor market shifts, and the potential for AI to escape human control …
Keywords:
Artificial Intelligence, Large Language Model, ChatGPT, AI Regulation, Composition Artiginality, “The Death of the Author”
Heejin Kim
Papers of the Bibliographical Society of America 117.3 (2023): 271–309
In the printed texts of early modern plays, scholars have observed a number of lines bracketed by a set of duplicate lines. In 1918, J. Dover Wilson called this type of textual error a “repetition bracket” and argued that it is evidence for the insertion of additional text. In 1930, W. W. Greg adduced …
Keywords:
Dongsuk Oh, Jungwoo Lim, and Heuiseok Lim
Applied Sciences 12.19 (2022): 9424
The construction of high-quality word embeddings is essential in natural language processing. In existing approaches using a large text corpus, the word embeddings learn only sequential patterns in the context; thus, accurate learning of the syntax and semantic relationships between words …
Keywords:
neuro-symbolic, graph convolutional network, word embedding, dependency parsing, semantic role labeling, ConceptNet …
Oh, Dongsuk, Jungwoo Lim, Kinam Park, and Heuiseok Lim
Applied Sciences 12.18 (2022): 9022
Small language models (SLMs) are increasingly utilized for on-device applications due to their ability to ensure user privacy, reduce inference latency, and operate independently of cloud infrastructure. However, their performance is often limited when processing complex data structures such as …
Keywords:
abstract meaning representation, semantic representation, sub-symbolic; commonsense reasoning, ConceptNet …
Jeong, Seungwon, Dongsuk Oh, Kinam Park, and Heuiseok Lim
Applied Sciences 12.9 (2022): 4099
Unlike previous dialogue-based question-answering (QA) datasets, DREAM, multiple-choice Dialogue-based REAding comprehension exaMination dataset, requires a deep understanding of dialogue. Many problems require multi-sentence reasoning, whereas some require commonsense …
Keywords:
dialogue-based multiple-choice QA, commonsense reasoning, semantic search, pre-trained language models, deep learning
PU-GEN: Enhancing generative commonsense reasoning for
language models with human-centered knowledge
Jaehyung Seo, Dongsuk Oh, Sugyeong Eo, Chanjun Park, Kisu Yang, Hyeonseok Moon, Kinam Park, and Heuiseok Lim
Knowledge-Based Systems 256 (2022):
Generative commonsense reasoning refers to the ability of a language model to generate a sentence with a given concept-set based on compositional generalization and commonsense reasoning. In the CommonGen challenge, which evaluates the capability of generative commonsense reasoning …
Keywords:
Text generation, Commonsense reasoning, Human-centered knowledge, Language model
Dongsuk Oh, Yejin Kim, Hodong Lee, H. Howie Huang, and Heuiseok Lim
Proceedings of COLING (2022): 4585–4592
Recent pre-trained language models (PLMs) achieved great success on many natural language processing tasks through learning linguistic features and contextualized sentence representation. Since attributes captured in stacked layers of PLMs are not clearly identified, straightforward approaches …
Keywords:
Jang, Yoonna, Jungwoo Lim, Yuna Hur, Dongsuk Oh, Suhyune Son, Yeonsoo Lee, Donghoon Shin, Seungryong Kim, and Heuiseok Lim
Proceedings of the AAAI Conference on
Artificial Intelligence 36. 10 (2022): 10803-10812
Humans usually have conversations by making use of prior knowledge about a topic and background information of the people whom they are talking to. However, existing conversational agents and datasets do not consider such comprehensive information, and thus they have a limitation in generating …
Keywords:
Speech & Natural Language Processing (SNLP)
Heejin Kim
Digital Scholarship in the Humanities 36.4 (2021): 919–933
The relationship between Shakespeare’s First Folio and early printings, published in his lifetime, has been a matter of dispute for centuries. A computer program that I have developed visualizes the fluctuating quality of textual correspondences between Folio texts, Henry the Sixth …
Keywords:
Sunjae Kwon, Dongsuk Oh, and Youngjoong Ko
Information Processing & Management 58.4 (2021):
In this paper, we introduce a novel knowledge-based word-sense disambiguation (WSD) system. In particular, the main goal of our research is to find an effective way to filter out unnecessary information by using word similarity. For this, we adopt two methods in our WSD system …
Keywords:
Natural language processing, Word sense disambiguation, Knowledge-based word vector representation, Similarity-based word selection
Whang, Taesun, Dongyub Lee, Dongsuk Oh, Chanhee Lee, Kijong Han, Dong-hun Lee, and Saebyeok Lee
Proceedings of the AAAI Conference on
Artificial Intelligence 35.16 (2021): 14041–14049
In this paper, we study the task of selecting the optimal response given a user and system utterance history in retrieval-based multi-turn dialog systems. Recently, pre-trained language models (e.g., BERT, RoBERTa, and ELECTRA) showed significant improvements in …
Keywords:
Conversational AI/Dialog Systems
Jungwoo Lim, Dongsuk Oh, Yoonna Jang, Kisu Yang, and Heuiseok Lim
Proceedings of COLING (2020): 2459–2471
CommonsenseQA is a task in which a correct answer is predicted through commonsense reasoning with pre-defined knowledge. Most previous works have aimed to improve the performance with distributed representation without considering the process of predicting the answer …
Keywords:
Whang, Taesun, Dongyub Lee, Chanhee Lee, Kisu Yang, Dongsuk Oh, and Heuiseok Lim
Proceedings of Interspeec (2019): 1585-1589
We focus on multi-turn response selection in a retrieval-based dialog system. In this paper, we utilize the powerful pre-trained language model Bi-directional Encoder Representations from Transformer (BERT) for a multi-turn dialog system and propose a highly effective post-training …
Keywords:
Response selection, Human computer dialog system, Spoken language processing
Shakespeare 15.4 (2019): 356–378
Since 1928, The First Part of the Contention and Richard Duke of York (printed separately in the 1590s) have been regarded as memorial reconstructions of two texts in the Folio edition of Shakespeare’s Comedies, Histories, and Tragedies (printed in 1623) …
Keywords:
The First Part of the Contention, Richard Duke of York, Christopher Marlowe, bad quarto, collaboration, attribution