Hongming Fu

Ph.D. Student in Computer Science, Shanghai Jiao Tong University

About

I am a Ph.D. student at Shanghai Jiao Tong University, advised by Prof. Bo Zhao.

My research interests center on egocentric vision, human motion, and hand-object interaction, with a broader interest in real-to-sim data acquisition methods for embodied intelligence.

Discussions and collaborations are always welcome; feel free to reach out.

Publications

*: joint first author; † Project Lead; ✉ corresponding author(s).

EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes

arXiv preprint

A unified end-to-end model for world-space egocentric hand reconstruction that predicts point-wise visibility, contact, and distance in a single pass, without separate detection, camera tracking, or contact stages.

EgoGrasp: World-Space Hand-Object Interaction Estimation from Egocentric Videos

arXiv preprint

A framework for world-space hand-object interaction reconstruction from egocentric videos, built on spatial intelligence preprocessing, upper-body HOI diffusion priors, and test-time optimization.

MergeIT: From Selection to Merging for Efficient Instruction Tuning

Findings of the Association for Computational Linguistics

A data-centric instruction tuning method that selects diverse IFT data and merges similar samples through prompting, achieving stronger fine-tuning performance with a much smaller training set.

Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond

IEEE/CVF Conference on Computer Vision and Pattern Recognition

A multimodal image fusion framework that uses Segment Anything priors, bi-level optimization, and downstream task feedback to improve visual quality and semantic perception efficiency.

Segmentation-driven Infrared and Visible Image Fusion via Transformer-enhanced Architecture Searching

IEEE International Conference on Acoustics, Speech and Signal Processing

A neural architecture search approach for infrared-visible image fusion, designed to balance upstream fusion quality and downstream semantic segmentation performance.

Hybrid-Supervised Dual-Search: Leveraging Automatic Learning for Loss-free Multi-Exposure Image Fusion

AAAI Conference on Artificial Intelligence

A dual-search framework that combines neural architecture search and loss-function search for multi-exposure image fusion, improving visual quality through automatic optimization.

Education