Projects
Selected research and engineering projects. Full write-ups for the first three are in the project portfolio (PDF).
Projects
Audio-to-Audio Speech Unit Translation (Korean ↔ English) · Sep. 2024 – Present ECE Capstone Design, Inha University
Textless speech-to-speech translation without an intermediate text step. Prior work (Textless NLP, AV2AV) covered only SVO languages and excluded SOV languages such as Korean.
- Textless speech-to-speech pipeline: mHuBERT with a k-means quantizer (500 units) → Transformer encoder–decoder with language tags → HiFi-GAN unit vocoder conditioned on a speaker d-vector.
- Extended prior SVO-only work to Korean (SOV): reordered transcripts with Llama 3.1 8B Instruct before resynthesis, so the model aligns by time frame rather than by lexical correspondence.
- BLEU 42.8 / COMET 0.1009 (AV2AV baseline: 60.1 / 0.1587); output intelligible but unstable in pitch.
- Setup: Multilingual AIhub + LibriSpeech, Fairseq / PyTorch, NVIDIA A6000 48GB.
Scene Graph to Video Generation with Diffusion · Aug. 2024 – Dec. 2024 Machine Intelligence Lab, Inha University
Scene graphs have been applied to image generation but not to general video generation. This project targets that gap, using the structured graph as a directable condition for video.
- R-GCN scene-graph embedding aligned to a CLIP image encoder by contrastive learning (SGClip), conditioning a latent diffusion model through cross-attention with a time-extended U-Net.
- Autoregressive long-video module injecting noise only into the first frame.
- Pipeline fully implemented; results limited by scarce Scene-Graph–video paired data (Action Genome) and weak graph–video alignment.
Water Quality Prediction: Embedded Device and Sensor Analysis · Mar. 2025 – Jun. 2025 R&D Department, STS Engineering
Soft sensors that predict pollution indicators from cheap sensor data, replacing expensive physical measurement equipment and covering intermediate process stages where sensors cannot be installed.
- IoT device with water quality sensors, RS485 serial, and LTE; FreeRTOS firmware focused on fault handling and reboot recovery.
- Ensemble separating trend prediction from noise modelling for irregular sensor series.
Competitions
Korean LLM Fine-tuning for Question Answering · Jul. 2024 – Aug. 2024 2024 Inha Artificial Intelligence Challenge (Dacon)
- QA over Korean economic articles; found LoRA adaptation degraded an already well-aligned base model, so the approach shifted to minimal fine-tuning that preserves base behaviour.
Global Wildfire Detection Challenge · Mar. 2024 6th AI SPARK Challenge · github.com/hytric/Wildfire-detection
- Satellite-image segmentation with TransUNet and Attention U-Net; ~90% with a single model, improved by ensembling.
Earlier Work
- Vision-based Autonomous Human-Following Wheeled Mobile Robot — FVE Alpha Project, Inha University (Sep. – Dec. 2022). Led to the KSAE 2022 poster above.
- Model Ensemble ViT-SSD — Vision Transformer with Single Shot Detection, ECE Deep Learning course project (Nov. – Dec. 2023).
- Real-time Computer Vision on AWS + Raspberry Pi — Hanium ICT Challenge (Mar. – Aug. 2023).