Selected research and engineering projects. Full write-ups for the first three are in the project portfolio (PDF).


Projects

Audio-to-Audio Speech Unit Translation (Korean ↔ English) · Sep. 2024 – Present ECE Capstone Design, Inha University

Textless speech-to-speech translation without an intermediate text step. Prior work (Textless NLP, AV2AV) covered only SVO languages and excluded SOV languages such as Korean.

  • Textless speech-to-speech pipeline: mHuBERT with a k-means quantizer (500 units) → Transformer encoder–decoder with language tags → HiFi-GAN unit vocoder conditioned on a speaker d-vector.
  • Extended prior SVO-only work to Korean (SOV): reordered transcripts with Llama 3.1 8B Instruct before resynthesis, so the model aligns by time frame rather than by lexical correspondence.
  • BLEU 42.8 / COMET 0.1009 (AV2AV baseline: 60.1 / 0.1587); output intelligible but unstable in pitch.
  • Setup: Multilingual AIhub + LibriSpeech, Fairseq / PyTorch, NVIDIA A6000 48GB.

Scene Graph to Video Generation with Diffusion · Aug. 2024 – Dec. 2024 Machine Intelligence Lab, Inha University

Scene graphs have been applied to image generation but not to general video generation. This project targets that gap, using the structured graph as a directable condition for video.

  • R-GCN scene-graph embedding aligned to a CLIP image encoder by contrastive learning (SGClip), conditioning a latent diffusion model through cross-attention with a time-extended U-Net.
  • Autoregressive long-video module injecting noise only into the first frame.
  • Pipeline fully implemented; results limited by scarce Scene-Graph–video paired data (Action Genome) and weak graph–video alignment.

Water Quality Prediction: Embedded Device and Sensor Analysis · Mar. 2025 – Jun. 2025 R&D Department, STS Engineering

Soft sensors that predict pollution indicators from cheap sensor data, replacing expensive physical measurement equipment and covering intermediate process stages where sensors cannot be installed.

  • IoT device with water quality sensors, RS485 serial, and LTE; FreeRTOS firmware focused on fault handling and reboot recovery.
  • Ensemble separating trend prediction from noise modelling for irregular sensor series.

Competitions

Korean LLM Fine-tuning for Question Answering · Jul. 2024 – Aug. 2024 2024 Inha Artificial Intelligence Challenge (Dacon)

  • QA over Korean economic articles; found LoRA adaptation degraded an already well-aligned base model, so the approach shifted to minimal fine-tuning that preserves base behaviour.

Global Wildfire Detection Challenge · Mar. 2024 6th AI SPARK Challenge · github.com/hytric/Wildfire-detection

  • Satellite-image segmentation with TransUNet and Attention U-Net; ~90% with a single model, improved by ensembling.

Earlier Work

  • Vision-based Autonomous Human-Following Wheeled Mobile Robot — FVE Alpha Project, Inha University (Sep. – Dec. 2022). Led to the KSAE 2022 poster above.
  • Model Ensemble ViT-SSD — Vision Transformer with Single Shot Detection, ECE Deep Learning course project (Nov. – Dec. 2023).
  • Real-time Computer Vision on AWS + Raspberry Pi — Hanium ICT Challenge (Mar. – Aug. 2023).