pipette
ESEspañol

SurgGraph: Quantitative Laparoscopic Video Understanding via Geometry-Grounded Scene Graphs

Jingying Wang, Rosiana Natalie, Marquise D Singleterry, Filippos Bellos, Brian George, Gurjit Sandhu, Jason J Corso, Anhong Guo, Vitaliy Popov, Xu Wang

PreprintReal-world use

In the authors' words

Surgical videos are a primary resource for teaching trainees anatomy, tool usage, and procedural skills. Yet learning from them at scale requires systems that understand surgical scenes. Existing approaches fall short: vision-language models lack fine-grained domain reasoning, task-specific models do not generalize, and prior scene graphs omit clinically meaningful detail. We present SurgGraph, a training-free pipeline that generates quantitative scene graphs from surgical videos. Operating on segmentation masks and depth maps, SurgGraph encodes each clinically meaningful relation (attachment, occlusion, separation, tool actions) as a <subject, verb, object, value> tuple whose numeric value quantifies the relation's extent over time. Technical evaluations show more precise scene understanding than state-of-the-art surgical VLM baselines. We then build SurgGraphQA, a proof-of-concept learning application that retrieves meaningful and boundary-case exemplars and generates visual explanations and feedback. A study with 17 medical students and 2 resident surgeons shows significant learning gains, demonstrating its educational value.

Main resultLimitation the authors admit

Appeared: Wednesday, September 23. arXiv. Preprint, not yet peer-reviewed.