Robotics & embodied intelligence
Jim Yun-Jin Li
I’m a PhD student at the Learning Systems and Robotics Lab at TUM, advised by Prof. Angela Schoellig.
I work on safe embodied interaction for humanoid robots, with a current focus on real2sim for articulated objects in our daily life.
Previously, I completed my Master’s at TUM, focusing on helping machines develop a richer understanding of the 3D world. In the Computer Vision Group, led by Prof. Dr. Daniel Cremers, I worked on three research projects: VXP (3DV 2025), TRASE (3DV 2026), and UniLoc.
Latest news
| Feb 16, 2026 | I joined Learning Systems and Robotics Lab (LSY) as a PhD student supervised by Prof. Dr. Angela Schoellig. |
|---|---|
| Nov 06, 2025 | Our work “TRASE: Tracking-free 4D Segmentation and Editing” based on my Master’s thesis (previously named SADG) is accepted to 3DV 2026. See you in Vancouver. I can finally say that I have marked an end for my Master’s studies LoL. |
| Nov 28, 2024 | Our work “SADG: Segment Any Dynamic Gaussian Without Object Trackers” based on my Master’s thesis is on arXiv. Check out the paper and project page. |
| Nov 05, 2024 | Our work Voxel-Cross-Pixel Large-scale Image-LiDAR Place Recognition is accepted to 3DV 2025. See you in Singapore. |
| Sep 12, 2024 | Successfully finished my Master’s thesis defense titled: “4DGSAM: Segment Anything in Dynamic Scene Novel View Synthesis”. |
| Mar 22, 2024 | Our work “VXP: Voxel-Cross-Pixel Large-scale Image-LiDAR Place Recognition” is on arXiv. Check out the paper and project page. |
| Mar 19, 2024 | Personal Github page created!!! |
Selected research
-
2026 International Conference on 3D Vision (3DV), 2026Understanding dynamic 3D scenes is crucial for extended reality (XR) and autonomous driving. Incorporating semantic information into 3D reconstruction enables holistic scene representations, unlocking immersive and interactive applications. To this end, we introduce TRASE, a novel tracking-free 4D segmentation method for dynamic scene understanding. TRASE learns a 4D segmentation feature field in a weakly-supervised manner, leveraging a soft-mined contrastive learning objective guided by SAM masks. The resulting feature space is semantically coherent and well-separated, and final object-level segmentation is obtained via unsupervised clustering. This enables fast editing, such as object removal, composition, and style transfer, by directly manipulating the scene’s Gaussians. We evaluate TRASE on five dynamic benchmarks, demonstrating state-of-the-art segmentation performance from unseen viewpoints and its effectiveness across various interactive editing tasks.
-
2025 International Conference on 3D Vision (3DV), 2025Recent works on the global place recognition treat the task as a retrieval problem, where an off-the-shelf global descriptor is commonly designed in image-based and LiDAR-based modalities. However, it is non-trivial to perform accurate image-LiDAR global place recognition since extracting consistent and robust global descriptors from different domains (2D images and 3D point clouds) is challenging. To address this issue, we propose a novel Voxel-Cross-Pixel (VXP) approach, which establishes voxel and pixel correspondences in a self-supervised manner and brings them into a shared feature space. Specifically, VXP is trained in a two-stage manner that first explicitly exploits local feature correspondences and enforces similarity of global descriptors. Extensive experiments on the three benchmarks (Oxford RobotCar, ViViD++ and KITTI) demonstrate our method surpasses the state-of-the-art cross-modal retrieval by a large margin. The code will be publicly available.
-
arXiv preprint arXiv:2412.12079, 2024To date, most place recognition methods focus on single-modality retrieval. While they perform well in specific environments, cross-modal methods offer greater flexibility by allowing seamless switching between map and query sources. It also promises to reduce computation requirements by having a unified model, and achieving greater sample efficiency by sharing parameters. In this work, we develop a universal solution to place recognition, UniLoc, that works with any single query modality (natural language, image, or point cloud). UniLoc leverages recent advances in large-scale contrastive learning, and learns by matching hierarchically at two levels: instance-level matching and scene-level matching. Specifically, we propose a novel Self-Attention based Pooling (SAP) module to evaluate the importance of instance descriptors when aggregated into a place-level descriptor. Experiments on the KITTI-360 dataset demonstrate the benefits of cross-modality for place recognition, achieving superior performance in cross-modal settings and competitive results also for uni-modal scenarios. The code will be available upon acceptance.
Selected projects
Visual-SLAM Loop Closure and Relocalization
Visual SLAM, Loop Closure, Relocalization, ORB-SLAM, Stereo Camera, Bundle Adjustment
CodeVisual-Inertial Tracking using Preintegrated Factors
Visual SLAM, IMU, CV, C++, Bundle Adjustment
CodeGraph Attention Network for Social Navigation (GAT4SN)
RL, Temporal Difference (TD) Learning, Imitation Learning (TUM ADLR Final Project)
CodeJetson Nano Car
ROS, Motor Control, SolidWorks, URDF, CV, Jetson Nano, Mobile Robot, Arduino, PID.
Text-Conditioned Face Diffusion
A weekend project for text-conditioned latent diffusion model on CelebA dataset.
Students & collaborators
Let's work together.
I'm open to thesis, guided research, and internship students, as well as research collaborations. If you'd like to work with me, email your interests, CV, and up-to-date transcripts.