Research Interests
- Domain-specialized LLM agents: multi-agent orchestration, intent decomposition and routing, and grounded tool use for high-stakes domains (insurance, finance).
- Retrieval-augmented generation (RAG) systems: domain-organized index design, semantic chunking, hybrid retrieval, and principled evaluation (LLM-as-a-judge, regression testing) for production QA systems.
- LLM-driven data systems: dataset curation, augmentation, and distillation for domain adaptation and fine-tuning; tool-augmented agents for grounded reasoning and action.
- Natural-language interfaces to databases: robust dataset augmentation/cleaning and principled evaluation for NL2SQL, and multi-turn dialogue systems to databases.
- Large-scale retrieval & indexing: hierarchical and multi-vector embeddings, cross-modal (text–image) retrieval, time-aware search, and approximate nearest neighbor (ANN) methods.
- Event-centric intelligence: detecting and modeling real-world events from news, social media, and enterprise data for downstream analytics and decision support.
Education
Dissertation: “Systematic Data Augmentation and Cleaning Techniques to Improve Text-to-SQL Models.” Advisor: Prof. Wook-Shin Han.
Appointments
Research and development of LLM-based agent systems for insurance sales and claims support.
Participated in the Parallel Graph AnalytiX (PGX) project - performance analysis and tuning of graph analytics and graph queries on large-scale graphs in a distributed graph engine, and efficient graph partitioning for the distributed graph engine PGX/D.
Industry Experience — Meritz Fire & Marine Insurance
MeAI: AI Sales Assistant for Insurance Planners
MeAI is an in-production AI sales assistant that answers insurance planners’ (FPs’) questions on coverage, claims, underwriting, sales strategy, and internal operations, grounded in company policy documents and enterprise data.
- Agent architecture redesign (intent-graph multi-agent system). Redesigned the query-processing pipeline from a single unidirectional planner into an intent-graph architecture: a router decomposes a multi-intent user query into a dependency graph of intents and dispatches each intent to domain-specialized expert agents (e.g., claims/coverage, underwriting, sales support), executing independent intents in parallel while propagating results along dependencies and synthesizing a single coherent answer - enabling compound questions that the previous planner could not handle.
- Retrieval index design for domain-specialized RAG. Reorganized retrieval corpora into domain-organized indexes with semantic-unit chunking; condensed lengthy policy documents into coverage-level summary units and connected coverages to standardized disease/procedure classification codes, substantially improving retrieval precision for grounded answers.
- Evaluation and quality infrastructure. Built automated quality assurance for LLM agents: pre-deployment regression gates that evaluate the exact release artifact, and nightly automated evaluation with LLM-as-a-judge scoring and day-over-day regression detection, with results persisted for monitoring. Established label-free evaluation methodology and expert-in-the-loop gold-set construction for coverage-analysis QA.
Research Projects (Selected)
Methodology for Extracting Notable Stocks from Domestic Economic News Using Generative AI
High-level pipeline for event detection and stock candidate surfacing (research prototype). Responsibilities - project lead, problem formulation, evaluation design, and coordination with the industry partner.
SW StarLab "Development of a Conversational and Self-Tuning DBMS"
Led efforts to make DBMSs usable to non-experts by building a natural-language interface to databases and automatic physical/knob tuning that simplifies performance optimization for practitioners. Responsibilities included NL2SQL modeling, dataset curation/augmentation, and evaluation infrastructure for robust QA over RDBMS.
Smart Structuring of Large-Scale Dark Data
Developed a machine-learning-based framework to convert unstructured enterprise “Dark Data” into structured data. Delivered workflow components (preprocessing, candidate-set mapping, feature extraction, supervised learning and inference), a pilot on public job announcements, and an end-to-end automation plan including SQL/Python APIs.
Publications (Selected)
Combining Sampling and Synopses with Worst-Case Optimal Runtime and Quality Guarantees for Graph Pattern Cardinality Estimation
Natural Language to SQL: Where Are We Today?
International Patents
Method for Enhancing Learning Data Set in Natural Language Processing System
Apparatus and Method for Processing Natural Language Query about Relational Database Using Transformer Neural Network
Column and Table Prediction Method for Text-to-SQL Query Translation Based on a Neural Network
Honors & Scholarships
Google Anita Borg Scholars APAC
POSTECH–SUNY Korea IT Consilience Creative Program (ITCCP) Scholarship
Teaching & Mentoring
Instructor, Corporate Lecture Series on Retrieval-Augmented Generation and Agentic AI
Four sessions with hands-on labs - IR fundamentals; RAG foundations; Advanced RAG; Agentic RAG & Evaluation.
Teaching Assistant / Instructor Support
Data Foundation (Oct 2017), Advanced Data Expert Course (Jul–Nov 2019), Data Programming Course (Apr 2020).