Robotics & embodied intelligence

Jim Yun-Jin Li

I’m a PhD student at the Learning Systems and Robotics Lab at TUM, advised by Prof. Angela Schoellig.

I work on safe embodied interaction for humanoid robots, with a current focus on real2sim for articulated objects in our daily life.

Previously, I completed my Master’s at TUM, focusing on helping machines develop a richer understanding of the 3D world. In the Computer Vision Group, led by Prof. Dr. Daniel Cremers, I worked on three research projects: VXP (3DV 2025), TRASE (3DV 2026), and UniLoc.

Portrait of Jim Li
PhD student · TUM
National Tsing Hua University Technical University of Munich Learning Systems and Robotics Lab Munich Center for Machine Learning

Latest news

Feb 16, 2026 I joined Learning Systems and Robotics Lab (LSY) as a PhD student supervised by Prof. Dr. Angela Schoellig.
Nov 06, 2025 Our work “TRASE: Tracking-free 4D Segmentation and Editing” based on my Master’s thesis (previously named SADG) is accepted to 3DV 2026. See you in Vancouver. I can finally say that I have marked an end for my Master’s studies LoL.
Nov 28, 2024 Our work “SADG: Segment Any Dynamic Gaussian Without Object Trackers” based on my Master’s thesis is on arXiv. Check out the paper and project page.
Nov 05, 2024 Our work Voxel-Cross-Pixel Large-scale Image-LiDAR Place Recognition is accepted to 3DV 2025. See you in Singapore.
Sep 12, 2024 Successfully finished my Master’s thesis defense titled: “4DGSAM: Segment Anything in Dynamic Scene Novel View Synthesis”.
Mar 22, 2024 Our work “VXP: Voxel-Cross-Pixel Large-scale Image-LiDAR Place Recognition” is on arXiv. Check out the paper and project page.
Mar 19, 2024 Personal Github page created!!! :smile:

Selected research

  1. sadg.gif
    Yun-Jin Li, Mariia Gladkova , Yan Xia , and Daniel Cremers
    2026 International Conference on 3D Vision (3DV), 2026

    Understanding dynamic 3D scenes is crucial for extended reality (XR) and autonomous driving. Incorporating semantic information into 3D reconstruction enables holistic scene representations, unlocking immersive and interactive applications. To this end, we introduce TRASE, a novel tracking-free 4D segmentation method for dynamic scene understanding. TRASE learns a 4D segmentation feature field in a weakly-supervised manner, leveraging a soft-mined contrastive learning objective guided by SAM masks. The resulting feature space is semantically coherent and well-separated, and final object-level segmentation is obtained via unsupervised clustering. This enables fast editing, such as object removal, composition, and style transfer, by directly manipulating the scene’s Gaussians. We evaluate TRASE on five dynamic benchmarks, demonstrating state-of-the-art segmentation performance from unseen viewpoints and its effectiveness across various interactive editing tasks.

  2. vxp.gif
    Yun-Jin Li, Mariia Gladkova , Yan Xia , Rui Wang , and Daniel Cremers
    2025 International Conference on 3D Vision (3DV), 2025

    Recent works on the global place recognition treat the task as a retrieval problem, where an off-the-shelf global descriptor is commonly designed in image-based and LiDAR-based modalities. However, it is non-trivial to perform accurate image-LiDAR global place recognition since extracting consistent and robust global descriptors from different domains (2D images and 3D point clouds) is challenging. To address this issue, we propose a novel Voxel-Cross-Pixel (VXP) approach, which establishes voxel and pixel correspondences in a self-supervised manner and brings them into a shared feature space. Specifically, VXP is trained in a two-stage manner that first explicitly exploits local feature correspondences and enforces similarity of global descriptors. Extensive experiments on the three benchmarks (Oxford RobotCar, ViViD++ and KITTI) demonstrate our method surpasses the state-of-the-art cross-modal retrieval by a large margin. The code will be publicly available.

  3. uniloc.gif
    Yan Xia , Zhendong Li , Yun-Jin Li, Letian Shi , Hu Cao , João F. Henriques , and Daniel Cremers
    arXiv preprint arXiv:2412.12079, 2024

    To date, most place recognition methods focus on single-modality retrieval. While they perform well in specific environments, cross-modal methods offer greater flexibility by allowing seamless switching between map and query sources. It also promises to reduce computation requirements by having a unified model, and achieving greater sample efficiency by sharing parameters. In this work, we develop a universal solution to place recognition, UniLoc, that works with any single query modality (natural language, image, or point cloud). UniLoc leverages recent advances in large-scale contrastive learning, and learns by matching hierarchically at two levels: instance-level matching and scene-level matching. Specifically, we propose a novel Self-Attention based Pooling (SAP) module to evaluate the importance of instance descriptors when aggregated into a place-level descriptor. Experiments on the KITTI-360 dataset demonstrate the benefits of cross-modality for place recognition, achieving superior performance in cross-modal settings and competitive results also for uni-modal scenarios. The code will be available upon acceptance.

Selected projects

Jetson Nano Car

ROS, Motor Control, SolidWorks, URDF, CV, Jetson Nano, Mobile Robot, Arduino, PID.

Students & collaborators

Let's work together.

I'm open to thesis, guided research, and internship students, as well as research collaborations. If you'd like to work with me, email your interests, CV, and up-to-date transcripts.