Computer vision & multimodal learning

Junzhe Chen.

Ph.D. student · Computer Science · Tianjin University

Understanding the visual world,
one object at a time.

Latest notes

  1. SlotNarrative is now available on arXiv.
  2. Introducing TOC-Bench for temporal object consistency in Video-LLMs.

01 / About

Seeing beyond
the frame.

I am a Ph.D. student in Computer Science at Tianjin University. My research sits at the intersection of computer vision and multimodal learning.

I am particularly interested in object-centric video understanding: preserving identity through occlusion and change, grounding temporal reasoning in visual evidence, and designing token-efficient interfaces for Video-LLMs. My earlier work spans tiny-object detection and mixed-precision model quantization.

  • 01Video understanding
  • 02Vision-language models
  • 03Object-centric learning
  • 04Efficient deep learning

02 / Research

Publications & preprints.

View on Scholar ↗
My name is underlined.

Showing all 4 publications.

Preprints are explicitly labeled. Publication details checked Aug 2026. Download all citations ↗

03 / Background

A little more
background.

Education

Current

Tianjin University

Ph.D. student in Computer Science

Johns Hopkins University

M.S. in Computer Science

2019–2023

Beijing Jiaotong University

B.Eng. in Software Engineering

Industry experience

  • Huizhi Square TechnologyAlgorithm Engineer
  • OPPODeep Learning Intern
  • Sense TechnologyComputer Vision Algorithm Intern
Honors & awards
  • 2023 Third Prize, Beijing Big Data Skills Competition
  • 2023 Outstanding Undergraduate Thesis, Beijing Jiaotong University
  • 2023 4th Place, 4th International Competition on Human Identification at a Distance

04 / Contact

Good research starts
with a conversation.

For research ideas, questions about my work,
or possible collaborations, feel free to reach out.