Researcher · NVIDIA

Wei Xiong

I am a senior research scientist at NVIDIA. I obtained my Ph.D. in Computer Science at the University of Rochester, under the supervision of Prof. Jiebo Luo.

My recent research focus is visual generative modeling, including foundational generative models like pixel-space diffusion models, generative rendering, text-to-game generation, streaming video generation, interactive world models, etc. I am also interested in image composition, relighting, shadow synthesis, and representation learning.

I am looking for motivated research interns to explore pixel-space diffusion and world modeling. Since internship positions are limited, I am also open to long-term collaborations with university students whom I can advise. Feel free to reach out with your resume if you are interested.

Portrait of Wei Xiong

News

Earlier news

Selected Research Work

Representative works are highlighted. See the full list in my Google Scholar.

MMIG-Bench teaser MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
Hang Hua*, Ziyun Zeng*, Yizhi Song*, Yunlong Tang, Liu He, Daniel Aliaga, Wei Xiong, Jiebo Luo
NeurIPS 2025
Project PagePDFCodeData

We propose Multi-Modal Image Generation Benchmark (MMIG-Bench), a comprehensive benchmark for evaluating multi-modal image generation models. MMIG-Bench unifies compositional evaluation across T2I and customized generation, introduces explainable aspect-level metrics, and provides extensive human and automatic evaluations.

Refine-by-Align teaser Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment
Yizhi Song, Liu He, Zhifei Zhang, Soo Ye Kim, He Zhang, Wei Xiong, Zhe Lin, Brian L. Price, Scott Cohen, Jianming Zhang, Daniel Aliaga
ICLR 2025
Project PagePDF

We propose an approach that improves the identity fidelity of generated objects by automatically locating and aligning visual tokens in a reference image with the target region to be refined.

GroundingBooth teaser GroundingBooth: Grounding Text-to-Image Customization
Zhexiao Xiong, Wei Xiong, Jing Shi, He Zhang, Yizhi Song, Nathan Jacobs
( Project Lead & Primary Advisor)
TMLR 2025
Project PagePDFCode

We introduce GroundingBooth, a framework that achieves zero-shot instance-level spatial grounding for both foreground subjects and background objects in text-to-image customization.

SwapAnything teaser SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing
Jing Gu, Yilin Wang, Nanxuan Zhao, Wei Xiong, Qing Liu, Zhifei Zhang, He Zhang, Jianming Zhang, HyunJoon Jung, Xin Eric Wang
ECCV 2024
Project PagePDF

We introduce SwapAnything, a framework that swaps arbitrary objects in an image with personalized concepts provided by reference images while preserving the surrounding context.

IMPRINT teaser IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation
Yizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen, Brian L. Price, Jianming Zhang, Soo Ye Kim, He Zhang, Wei Xiong, Daniel Aliaga
CVPR 2024
Project PagePDF

Our work achieves advanced image composition with strong identity preservation, automatic object viewpoint and pose adjustment, color and lighting harmonization, and shadow synthesis in a single framework.

PHOTOSWAP teaser PHOTOSWAP: Personalized Subject Swapping in Images
Jing Gu, Yilin Wang, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong, Qing Liu, Zhifei Zhang, He Zhang, Jianming Zhang, HyunJoon Jung, Xin Eric Wang
NeurIPS 2023
Project PagePDFCode

We present PHOTOSWAP, an approach for personalized subject swapping in existing images while preserving the surrounding scene.

Guidance-driven Visual Synthesis thesis teaser Guidance-driven Visual Synthesis with Generative Models
Wei Xiong
Ph.D. Thesis 2022
PDF

My Ph.D. thesis summarizes my research on guidance-driven visual content creation, including visually pleasing data synthesis and synthesis for downstream visual recognition tasks.

Image sentiment transfer teaser Image Sentiment Transfer
Tianlang Chen, Wei Xiong, Haitian Zheng, Jiebo Luo
ACM MM 2020
PDF

We introduce Image Sentiment Transfer, an important but still underexplored research task, and propose an effective and flexible framework that performs sentiment transfer at both the image and object levels.

Fine-grained image transformation teaser Fine-grained Image-to-Image Transformation towards Visual Recognition
Wei Xiong, Yutong He, Yixuan Zhang, Wenhan Luo, Lin Ma, Jiebo Luo
CVPR 2020
PDFSupplementary

We transform images from fine-grained categories to synthesize new examples that preserve the identity of the input, thereby benefiting fine-grained recognition and few-shot learning.


Mentoring

I am fortunate to have mentored and collaborated with talented researchers:

  • Yongsheng Yu2025–2026

    Pixel-space diffusion

  • Zhexiao Xiong2024–2025

    Grounded text-to-image customization
    Physics-coherent image-to-video generation

  • Portrait relighting using diffusion models

  • Customized image composition

  • Shadow detection and synthesis

  • Subject-driven image editing

  • Tumor growth prediction

  • Image transformation for visual recognition


Academic Service

Area Chair

  • ICLRInternational Conference on Learning Representations, 2026.
  • AAAIAAAI Conference on Artificial Intelligence, 2026 and 2027.

Conference Reviewer

  • CVPRIEEE/CVF Conference on Computer Vision and Pattern Recognition
  • ICCVIEEE/CVF International Conference on Computer Vision
  • ECCVEuropean Conference on Computer Vision
  • NeurIPSConference on Neural Information Processing Systems
  • ICLRInternational Conference on Learning Representations
  • AAAIAAAI Conference on Artificial Intelligence
  • ICMLInternational Conference on Machine Learning
  • MICCAIInternational Conference on Medical Image Computing & Computer Assisted Intervention
  • BMVCBritish Machine Vision Conference
  • WACVWinter Conference on Applications of Computer Vision

Journal Reviewer

  • TPAMIIEEE Transactions on Pattern Analysis and Machine Intelligence
  • TMLRTransactions on Machine Learning Research
  • TIPIEEE Transactions on Image Processing
  • TNNLSIEEE Transactions on Neural Networks and Learning Systems
  • TMMIEEE Transactions on Multimedia
  • TCSVTIEEE Transactions on Circuits and Systems for Video Technology
  • CVIUComputer Vision and Image Understanding
  • SPICSignal Processing: Image Communication

Teaching

Teaching Assistant

Spring 2019

CSC 240/440 · Data Mining

Instructor
Ted Pawlicki
Office hours
Monday & Wednesday, 2:30–4:00 PM
Fall 2018

CSC 240/440 · Data Mining

Instructor
Ted Pawlicki
Office hours
Monday & Wednesday, 2:00–3:00 PM
Spring 2018

CSC 240/440 · Data Mining

Instructor
Anand Ajay

More About Me

Outside of work, I enjoy playing tennis, basketball, and table tennis. I am also always happy to exchange ideas and chat about research.