PhD candidate / AI researcher

Multimodal generative models, multilingual AI, and LLM reasoning.

I am a PhD candidate in Computer Science and Artificial Intelligence at Oregon State University, advised by Liang Huang. My research uses sign-language translation and structured 3D human motion as demanding settings for multilingual generation, streaming inference, and evaluation. I have also worked on long-context multimodal models and text-to-SQL at Amazon and Genies.

Seeking full-time Research Scientist, Applied Scientist, Research Engineer, and ML Engineer roles starting in 2027.

01 / Selected work

Research

2026 · Preprint

Direct Translation between Sign Languages

Direct ASL, CSL, and DGS translation without a spoken-text pivot. The system uses 112K cross-lingual pairs and a shared motion representation, reducing mean motion error by 19% and running 2.3× faster than a three-stage cascade.

Multilingual generation · mBART · VQ-VAE · synthetic data · evaluation

To NAACL 2027

Simultaneous Translation between Sign Languages

Streaming sign-to-sign generation with prefix training and wait-k decoding. The approach cuts computation-aware latency by 38% while making the quality–latency trade-off explicit.

Streaming decoding · SMPL-X · wait-k · multi-path training

2025 · Preprint

Geometry-Aware Text-to-Sign Generation

Variable-length 3D sign generation with geometry-aware objectives for articulated hands and consistent skeletons. The model improves back-translation BLEU-4 by 3.08 points on PHOENIX14T development data.

3D human motion · XLM-R · OpenPose · inverse kinematics

02 / Industry

Experience

2024–2025

Research Scientist Intern · Genies

Built an LLM-based user-insight system and explored GRPO and MCTS-guided reasoning for text-to-SQL.

2023

Applied Scientist Intern · Amazon

Researched generative models for long multimodal sequences and cross-modal dependencies.

2022

Applied Scientist Intern · Amazon

Developed multilingual representation-learning methods on XLM-R for cross-lingual retrieval.

2019–2020

Machine Learning Engineer · EnjoyMusic

Built generative sequence models for music style transfer and MIDI generation.

03 / Writing

Selected publications

  1. 2026

    Direct Translation between Sign Languages
    Zetian Wu, Bowen Xie, Wuyang Meng, Milan Gautam, Stefan Lee, and Liang Huang

  2. 2025

    Geometry-Aware Losses for Structure-Preserving Text-to-Sign Language Generation
    Zetian Wu, Tianshuo Zhou, Stefan Lee, and Liang Huang

  3. 2022

    Inducing Generalizable and Interpretable Lexica
    Findings of EMNLP

  4. 2021

    MultiBench: Multiscale Benchmarks for Multimodal Representation Learning
    NeurIPS Datasets and Benchmarks

  5. 2020

    Interactive Rare-Category-of-Interest Mining from Large Datasets
    AAAI

View all publications on Google Scholar →

Education. PhD, Oregon State University · MSE, Johns Hopkins University · BS, Zhejiang University.

Tools. Python, PyTorch, Hugging Face, SQL, C/C++, multi-GPU training, LLM post-training, multimodal learning, and 3D human motion.