TR2026-141

ORIGAMI: Object Representation Inferred Geometrically for Articulated ManIpulation


    •  Deng, Y., Nikovski, D.N., "ORIGAMI: Object Representation Inferred Geometrically for Articulated ManIpulation", IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), September 2026.
      BibTeX TR2026-141 PDF
      • @inproceedings{Deng2026sep,
      • author = {Deng, Yunfu and Nikovski, Daniel N.},
      • title = {{ORIGAMI: Object Representation Inferred Geometrically for Articulated ManIpulation}},
      • booktitle = {IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)},
      • year = 2026,
      • month = sep,
      • url = {https://www.merl.com/publications/TR2026-141}
      • }
  • MERL Contact:
  • Research Area:

    Robotics

Abstract:

We present ORIGAMI, a geometric pipeline that discovers the kinematic structure of unknown articulated mechanisms from a single RGB-D demonstration and maps each subsequent high-dimensional visual observation into a compact joint-position vector to be used for reinforcement learning. Existing approaches to visual articulated manipulation either learn policies directly from pixels, suffering from the high dimensionality of the observation space, or depend on pretrained visual representations that require large-scale data collection and offer no structural guarantee about the resulting features. ORIGAMI sidesteps both limitations: from tracked 3D keypoints, it recovers link segmentation, joint types, axis parameters, and kinematic topology through purely geometric reasoning, without any task-specific training or category-level priors. A formal consistency analysis shows that the systematic estimation errors of the geometric estimator cancel to first order in the goal-conditioned displacement observed by the policy, eliminating the need for learned correction. Experiments on revolute and prismatic mechanisms in SAPIEN simulation and on real hardware demonstrate that the discovered low-dimensional representation substantially narrows the gap between pixelbased reinforcement learning and privileged ground-truth state, while outperforming methods built on pretrained visual representations.