About Me
I’m a Master’s student in Computational Linguistics at the University of Washington in Seattle. I’m broadly interested in computational linguistics and NLP, especially in how linguistic structure and meaning can be modeled and used to build better language technologies.
More recently, I have become particularly interested in linguistic grounding and faithfulness in NLP: how computational models represent, preserve, or distort meaning across contexts, languages, and dialogue structure. My current work focuses on evaluating LLM outputs through linguistically informed analyses of hallucination, ambiguity, speaker roles, event structure, and discourse-level meaning.
I previously worked at Korea Electronics Technology Institute(KETI), NCSOFT, and SK Telecom. Those experiences made me care about how language models behave in the real world, not just how accurate they are, but how reliable and fair they can be.
I hold a B.A. in Linguistics & Cognitive Science and a double major in Language Science in Artificial Intelligence from Hankuk University of Foreign Studies, where I was advised by Prof. Jeesun Nam.
News
- Jun 2026 Started as Research Assistant at LT4CPR, UW with Prof. Fei Xia.
- Sep 2025 Started M.S. in Computational Linguistics at University of Washington.
Education
- Master of Science in Computational Linguistics
- Bachelor of Arts in Linguistics & Cognitive Science
- Bachelor of Language Science in Artificial Intelligence (Double Major)
- Advisor: Prof. Jeesun Nam
Research Experience
Supervisor: Prof. Fei Xia. I evaluate LLM-generated situation reports (SITREPs) for crisis response, focusing on faithfulness, information coverage, and structured content alignment.
- Developing section-level and bullet-level evaluation pipelines using ROUGE, BERTScore, and bipartite matching to compare system outputs against reference reports.
- Analyzing how well generated SITREPs preserve critical crisis information across structured report sections.
Supervisor: Prof. Jeesun Nam. I contributed to two research projects: constructing datasets for sentiment analysis of stock-market articles, and generating NLU datasets for training a big-data-driven chatbot model.
- Supported corpus collection and preprocessing pipelines, emphasizing consistent unit definitions and annotation-ready formatting.
- Assisted in annotation design and QA routines to improve label clarity and reduce ambiguity in training data.
Industry Experience
I developed trustworthiness benchmarks for Korean LLMs (hallucination, reliability, sociocultural bias) through prompt design, rubrics, and failure-mode analysis. I also validated a Korean multimodal dialogue dataset by aligning utterance-level pragmatic functions with nonverbal cue labels.
- Designed prompt suites and rubric-based evaluations for groundedness/hallucination, with failure-mode categories tied to discourse phenomena, such as attribution, consistency, and hedging.
- Developed annotation guidelines to assess socio-cultural bias and pragmatic harm, emphasizing appropriateness in Korean register and stance.
- Validated a Korean multimodal dialogue dataset by aligning utterance-level pragmatic functions, such as agreement, refusal, and hesitation, with nonverbal cue labels.
Key insight: many reliability issues are not isolated “wrong facts,” but emerge from how meaning is constructed across turns, including implicature, underspecification, and discourse coherence.
I led AI Red Team initiatives, creating robust testing protocols and diverse conversational datasets to identify and mitigate ethical and safety risks in language models. I categorized vulnerabilities into single-turn versus multi-turn threats, including discourse-level risks such as anaphoric references undetectable within a single turn.
- Created semantics-informed risk probes for single-turn inputs, including cases where meaning depends on polysemy or indirect phrasing.
- Developed multi-turn scenarios to capture escalation patterns, such as justification, concretization, and actionable requests, and measured how risk accumulates over dialogue context.
- Analyzed evasion behaviors after refusals, such as repair, paraphrase, and euphemisms, and incorporated these patterns into testing protocols.
Key insight: risk is often discourse-driven. What looks harmless in isolation can become unsafe when reference resolution, implicature, and turn-by-turn intent shifts are taken into account.
I led human-centric annotation to ensure sociolinguistic coverage (honorific levels, informal slang, code-mixing) in conversational training data. I also created a linguistically grounded ambiguity taxonomy and applied it to normalize Korean queries and improve chatbot robustness.
- Built a linguistically grounded ambiguity taxonomy, including ellipsis, anaphora, spacing/noise, and informal variants, to standardize how ambiguous inputs are handled.
- Applied text normalization and preprocessing guided by the taxonomy to reduce interpretation noise and improve NLU robustness.
- Led annotation efforts to ensure sociolinguistic coverage, including honorific levels, informal slang, and code-mixing, so the system behaves consistently across styles.
Key insight: in Korean conversational products, performance depends not only on intent classification accuracy but also on pragmatic consistency under register shifts.
Awards and Honors
- 2025 UW CLMS Scholarship ($14,500)
- 2023 HUFS Departmental Scholarship
- 2021 HUFS Departmental Scholarship
- 2022 Outstanding Undergraduate Thesis Award
- 2018 Best Composition Award
- 2018 Student Leadership Scholarship
- 2018 National Merit Scholarship
Projects
I co-developed a linguistically grounded dialogue-semantic error taxonomy to evaluate source-faithfulness in cross-lingual dialogue summarization. I diagnosed relational hallucinations in small language models through errors in speaker-role attribution, event structure, modality, and discourse outcomes.
- Evaluated English-to-Chinese dialogue summarization on XSAMSum using zero-shot pipelines with local open-weight LLMs.
- Compared Direct, Translate-then-Summarize, Summarize-then-Translate, and Semantic pipeline configurations.
- Evaluated generated Chinese summaries using ROUGE, BERTScore, and OmniScore.
- Analyzed semantic representation errors to identify information bottlenecks and model-specific failure patterns.
I modeled semantic constraints between attribute nouns (e.g., interest rates, costs) and directional predicates (rise/fall) to capture context-dependent polarity shifts in financial discourse.
- Designed pattern-based rules grounded in lexical semantics and compositionality, such as “costs rise” versus “stock prices rise.”
- Built a small, interpretable pipeline to extract attribute-predicate relations from financial text and analyzed common error patterns.
- Reflected on limitations and extensions, such as industry metadata and context windows, for more robust interpretation.
- Recognized with the Outstanding Undergraduate Thesis Award.
Relevant Coursework
- LING 473: Basics for Computational Linguistics
- LING 566: Introduction to Syntax for Computational Linguistics
- LING 570: Shallow Processing Techniques for NLP
- LING 571: Deep Processing Techniques for NLP
- LING 572: Advanced Statistical Methods in NLP
- LING 573: NLP Systems and Applications
- LING 575: Data Matters (Topics in NLP)
- Machine Learning for Language Analysis
- Big Data and Sentiment Analysis
- Language Information Processing
- Natural Language Data
- Corpus Analysis and Dictionary
- Intro to Linguistics and Language Technology
- Computer and Linguistics
- Syntactic Analysis
- Phonetics
- Pragmatics
- Language Typology
- Language and Logic
- Introduction to Linguistics
- Forensic Linguistics
- Grammar in Korean as Foreign Language
- Introduction to Cognitive Science
- Cognitive Psychology
- Neurolinguistics
- Linguistics and Psychological Experiments
- Language and Human Beings
- CSE 373: Data Structures and Algorithms
- Language and Database
- Programming Languages and Laboratory
- Introduction to Programming Languages
- Essential Programming for Linguistics
- Programming for Language Analysis
- Probability and Statistics
- Computer Mathematics
- Statistics for Language Analysis
Languages and Technical Skills
- Programming: Python, R, C++, SQL, Java
- ML/NLP: PyTorch, NLTK
- Tools: Git, Linux/CLI
- Typesetting: LaTeX
- Languages: Korean (native), English (fluent)