CSC 240/440 · Data Mining
- Instructor
- Ted Pawlicki
- Office hours
- Monday & Wednesday, 2:30–4:00 PM
|
Researcher · NVIDIA Wei XiongI am a senior research scientist at NVIDIA. I obtained my Ph.D. in Computer Science at the University of Rochester, under the supervision of Prof. Jiebo Luo. My recent research focus is visual generative modeling, including foundational generative models like pixel-space diffusion models, generative rendering, text-to-game generation, streaming video generation, interactive world models, etc. I am also interested in image composition, relighting, shadow synthesis, and representation learning. I am looking for motivated research interns to explore pixel-space diffusion and world modeling. Since internship positions are limited, I am also open to long-term collaborations with university students whom I can advise. Feel free to reach out with your resume if you are interested. |
|
News
Earlier news
|
Selected Research WorkRepresentative works are highlighted. See the full list in my Google Scholar. |
|
|
PixelDiT: Pixel Diffusion Transformers for Image Generation Yongsheng Yu, Wei Xiong†, Weili Nie, Yichen Sheng, Shiqiu Liu, Jiebo Luo († Project Lead & Primary Advisor) CVPR 2026 · Best Paper Finalist Project PagePDFCode PixelDiT scales diffusion transformers directly in pixel space, pretrained at 1024×1024 resolution to deliver high-fidelity text-to-image generation. |
|
DIVE: Taming DINO for Subject-Driven Video Editing Yi Huang, Wei Xiong†, He Zhang, Chaoqi Chen, Jianzhuang Liu, Mingfu Yan, Shifeng Chen († Project Lead) ICCV 2025 Project PagePDF We propose DINO-guided Video Editing (DIVE), a framework for subject-driven editing in source videos conditioned on either target text prompts or reference images with specific identities. |
|
MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models Hang Hua*, Ziyun Zeng*, Yizhi Song*, Yunlong Tang, Liu He, Daniel Aliaga, Wei Xiong, Jiebo Luo NeurIPS 2025 Project PagePDFCodeData We propose Multi-Modal Image Generation Benchmark (MMIG-Bench), a comprehensive benchmark for evaluating multi-modal image generation models. MMIG-Bench unifies compositional evaluation across T2I and customized generation, introduces explainable aspect-level metrics, and provides extensive human and automatic evaluations. |
|
Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment Yizhi Song, Liu He, Zhifei Zhang, Soo Ye Kim, He Zhang, Wei Xiong, Zhe Lin, Brian L. Price, Scott Cohen, Jianming Zhang, Daniel Aliaga ICLR 2025 Project PagePDF We propose an approach that improves the identity fidelity of generated objects by automatically locating and aligning visual tokens in a reference image with the target region to be refined. |
|
GroundingBooth: Grounding Text-to-Image Customization Zhexiao Xiong, Wei Xiong†, Jing Shi, He Zhang, Yizhi Song, Nathan Jacobs († Project Lead & Primary Advisor) TMLR 2025 Project PagePDFCode We introduce GroundingBooth, a framework that achieves zero-shot instance-level spatial grounding for both foreground subjects and background objects in text-to-image customization. |
|
SwapAnything: Enabling Arbitrary Object Swapping in Personalized Visual Editing Jing Gu, Yilin Wang, Nanxuan Zhao, Wei Xiong, Qing Liu, Zhifei Zhang, He Zhang, Jianming Zhang, HyunJoon Jung, Xin Eric Wang ECCV 2024 Project PagePDF We introduce SwapAnything, a framework that swaps arbitrary objects in an image with personalized concepts provided by reference images while preserving the surrounding context. |
|
IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation Yizhi Song, Zhifei Zhang, Zhe Lin, Scott Cohen, Brian L. Price, Jianming Zhang, Soo Ye Kim, He Zhang, Wei Xiong, Daniel Aliaga CVPR 2024 Project PagePDF Our work achieves advanced image composition with strong identity preservation, automatic object viewpoint and pose adjustment, color and lighting harmonization, and shadow synthesis in a single framework. |
|
Relightful Harmonization: Lighting-aware Portrait Background Replacement Mengwei Ren*, Wei Xiong, Jae Shin Yoon, Zhixin Shu, Jianming Zhang, HyunJoon Jung, Guido Gerig, He Zhang (* Work done while Mengwei was an intern at Adobe) CVPR 2024 Project PagePDF We introduce Relightful Harmonization, a lighting-aware diffusion model designed to harmonize sophisticated lighting effects on a foreground portrait using any background image. |
|
InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning Jing Shi*, Wei Xiong*, Zhe Lin, HyunJoon Jung (* Equal Contribution) CVPR 2024 Project PagePDF We propose InstantBooth, an approach built on pretrained text-to-image models that enables fast personalized text-to-image generation without test-time fine-tuning. |
|
PHOTOSWAP: Personalized Subject Swapping in Images Jing Gu, Yilin Wang, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong, Qing Liu, Zhifei Zhang, He Zhang, Jianming Zhang, HyunJoon Jung, Xin Eric Wang NeurIPS 2023 Project PagePDFCode We present PHOTOSWAP, an approach for personalized subject swapping in existing images while preserving the surrounding scene. |
|
Guidance-driven Visual Synthesis with Generative Models Wei Xiong Ph.D. Thesis 2022 My Ph.D. thesis summarizes my research on guidance-driven visual content creation, including visually pleasing data synthesis and synthesis for downstream visual recognition tasks. |
|
Unsupervised Low-light Image Enhancement with Decoupled Networks Wei Xiong, Ding Liu, Xiaohui Shen, Chen Fang, Jiebo Luo ICPR 2022 Our work is among the pioneering methods for unsupervised real-world low-light image enhancement. Specifically, we tackle the problem of enhancing real-world low-light images with significant noise in an unsupervised fashion. To this end, we explicitly decouple this task into two sub-problems: illumination enhancement and noise suppression. |
|
Image Sentiment Transfer Tianlang Chen, Wei Xiong, Haitian Zheng, Jiebo Luo ACM MM 2020 We introduce Image Sentiment Transfer, an important but still underexplored research task, and propose an effective and flexible framework that performs sentiment transfer at both the image and object levels. |
|
Fine-grained Image-to-Image Transformation towards Visual Recognition Wei Xiong, Yutong He, Yixuan Zhang, Wenhan Luo, Lin Ma, Jiebo Luo CVPR 2020 PDFSupplementary We transform images from fine-grained categories to synthesize new examples that preserve the identity of the input, thereby benefiting fine-grained recognition and few-shot learning. |
|
Foreground-aware Image Inpainting Wei Xiong, Jiahui Yu, Zhe Lin, Jimei Yang, Xin Lu, Connelly Barnes, Jiebo Luo CVPR 2019 PDFData We propose a foreground-aware image inpainting system that explicitly disentangles structure inference and content completion. Our model first predicts the foreground contour and then uses it to guide the completion of the missing region. |
|
Learning to Generate Time-Lapse Videos Using Multi-Stage Dynamic Generative Adversarial Networks Wei Xiong, Wenhan Luo, Lin Ma, Wei Liu, Jiebo Luo CVPR 2018 Project PagePDFCodeData We propose a two-stage GAN model that generates vivid yet content-preserving time-lapse videos from a single starting frame by disentangling content generation from motion enhancement. |
MentoringI am fortunate to have mentored and collaborated with talented researchers:
|
Academic ServiceArea Chair
Conference Reviewer
Journal Reviewer
|
TeachingTeaching Assistant CSC 240/440 · Data Mining
CSC 240/440 · Data Mining
CSC 240/440 · Data Mining
|
More About MeOutside of work, I enjoy playing tennis, basketball, and table tennis. I am also always happy to exchange ideas and chat about research. |