Tactile exploration → Object context → Adaptive action
Touch reveals hidden physical properties beyond appearance.
Properties that vision cannot infer directly, such as material, contents, or softness, can be revealed through tactile exploration. Our framework learns object representations from that exploration and uses them as object context, which determines property-dependent manipulation such as how each object is grasped and where it is placed.
a Object representation learning
Swing
Shake
Squeeze
Tapb Context-aware manipulation
Rearrange
Pour
OpenEight-seed average, decoded against text descriptors
15 → 67% unseen
without → with object context
From exploration to action
How It Works
Explore by touch
The robot probes each object with exploratory actions such as shaking, squeezing and tapping, revealing physical properties that vision alone cannot observe.
Distill into a representation
Informative tactile cues are sparse over a long exploration. The model distills the few contact moments that matter into a compact object representation that captures rigidity, material, contents, and fill level.
Act with object-centric context
The policy handles one object at a time. That object's representation and position form a fixed context for the whole manipulation.
Methods figures
Real-world manipulation
Object Properties Change Robot Behavior
Across three tasks, object context selects the manipulation strategy and stays fixed during execution as time-invariant context, while online vibrotactile feedback handles reactive adjustment. Every clip is an autonomous real-world policy rollout, not a teleoperated demonstration.
Task success, with and without object context
Successfully completing the task requires both the correct strategy and successful physical execution.
Multi-object rearrangement. Rigidity determines how the object is grasped, and contents determine where it belongs. An empty container will be placed into the recycle bin through a top grasp.
Rigidity decides the grasp
Contents decide the destination
Rearrangement failure modes
Two representative failures show what goes wrong when the strategy does not match the object. Playback is 4× speed.
Three consecutive placements from a single rearrangement episode.
Pouring & recycling. The policy selects its strategy from the inferred fill level. Filled containers are poured into the sink before recycling, while empty containers are recycled directly.
Why object context matters
Representative failure and success rollouts illustrate strategy selection; these are not synchronized, matched trials.
Box opening with the DeltaHand. For the DeltaHand, exploration also captures weight-related properties that distinguish visually identical boxes. The hand opens the box containing screws and takes no action on the empty one.
Object Representation Learning Details
We evaluate the contrastive learning and language alignment method on 9 seen and 9 unseen objects, each with 3 independently collected test trials that are never used in training. Seen means the object and its property combination appear in the training set. The same framework on the multi-finger DeltaHand reaches 96% on seen and 88% on unseen objects.
Exploratory actions. A Franka arm with fin-ray parallel-jaw fingers probes each object with six predefined actions, while vibrotactile signals from fingertip contact microphones, 6-axis end-effector forces and end-effector velocities are recorded in synchronization.
A complete exploration with sound.
Playback at 2× speed. The timer in the video shows real elapsed time. Metal, deformable, granular contents, half filled.
Learning framework. The three modalities are encoded into a multimodal token sequence, compressed by a Top-K token learner, and aligned with text embeddings of physical properties.
Learn from exploration
- Encode synchronized signalsVibrotactile spectrograms, end-effector forces, and velocities are encoded in 0.5-second windows.
- Keep informative tokensA Top-K token learner compresses the trajectory into a compact object embedding.
- Align with property descriptionsDuring training, contrastive losses align the representation with text embeddings of physical properties.
Use object context for manipulation
- Infer properties from touchAt inference, the learned embedding is compared with candidate property descriptors using cosine similarity.
- Choose the context representationThe paper evaluates both the tactile embedding itself and text embeddings of its decoded property descriptors.
- Condition the task policyThe current object's representation and position condition an Action Chunking Transformer alongside visual, proprioceptive, and online tactile observations.
Property estimation on independently collected test trials
Accuracy averaged over the four property axes and eight training seeds. Properties are decoded by cosine similarity between the tactile embedding and candidate text descriptors, so the decoding vocabulary can be changed without retraining.
Which moments does the model keep?These examples show selected tokens concentrated in particular exploration phases. Select an object to inspect its trajectory and predicted properties.

Ground truth on the left, model prediction on the right. Green marks a match.
Selected tokens occur in squeezing, swinging, cavity tapping, and shaking, with few in table tapping or sliding.
The complete walkthrough
Supplementary Video
A narrated walkthrough of exploration, representation learning, and robot manipulation.
BibTeX
@misc{yang2026objectcentric,
title = {Learning with Object-centric Representations of Tactile Interactive Perception for Robot Manipulation},
author = {Yang, Xinyi and Si, Zilin and Xu, Zhuowei and Temel, Zeynep and Kroemer, Oliver},
year = {2026},
eprint = {2609.33235},
archivePrefix = {arXiv},
primaryClass = {cs.RO},
url = {https://arxiv.org/abs/2609.33235}
}