Learning with Object-centric Representations of Tactile Interactive Perception for Robot Manipulation

CoRL 2026

Xinyi Yang*, Zilin Si*, Zhuowei Xu, Zeynep Temel, Oliver Kroemer

Robotics Institute, Carnegie Mellon University

* Equal contribution

Tactile exploration → Object context → Adaptive action

Touch reveals hidden physical properties beyond appearance.

Properties that vision cannot infer directly, such as material, contents, or softness, can be revealed through tactile exploration. Our framework learns object representations from that exploration and uses them as object context, which determines property-dependent manipulation such as how each object is grasped and where it is placed.

a Object representation learning

SwingSwing
ShakeShake
SqueezeSqueeze
Tap on cavityTap
Tactile explorationA sequence of exploratory procedures
Synchronized signalsVibrotactile, force, velocity
Top-K token selectionOnly informative contacts survive
Language-aligned spaceStructured by physical property

b Context-aware manipulation

Object contextOne representation per object, tied to its position
ACT policyPredicts an action chunk
RearrangementRearrange
Pouring and recyclingPour
Box openingOpen
Adaptive manipulationStrategy follows the inferred properties
Tactile exploration is distilled into a language-aligned object representation, which then conditions a manipulation policy. The representation captures physical object properties independently of the downstream task. Task-specific policies use this object context to adapt their manipulation strategies.
93% / 84%
Property estimation accuracy on seen / unseen objects
Eight-seed average, decoded against text descriptors
Explore the representation
41 → 83% seen
15 → 67% unseen
Rearrangement task success
without → with object context
See task success rates
47 → 77%
Pouring task success
without → with object context
See task success rates
25 → 100%
Box opening task success
without → with object context
See task success rates

From exploration to action

How It Works

1

Explore by touch

The robot probes each object with exploratory actions such as shaking, squeezing and tapping, revealing physical properties that vision alone cannot observe.

2

Distill into a representation

Informative tactile cues are sparse over a long exploration. The model distills the few contact moments that matter into a compact object representation that captures rigidity, material, contents, and fill level.

3

Act with object-centric context

The policy handles one object at a time. That object's representation and position form a fixed context for the whole manipulation.

Methods figures
Object representation learning pipeline from the paper.
Object representation learning through tactile exploration. Each object is explored with a sequence of predefined actions while vibrotactile signals, 6-axis end-effector forces, and end-effector velocities are recorded. A token learner selects the most informative K tokens, which are aligned with text embeddings of physical-property descriptions using contrastive learning.
Context-aware manipulation policy from the paper.
Object-centric and context-aware manipulation policy learning. Built on the Action Chunking Transformer, the policy conditions on the learned object context and object position to focus on target-specific visual, tactile, and proprioceptive observations.

Real-world manipulation

Object Properties Change Robot Behavior

Across three tasks, object context selects the manipulation strategy and stays fixed during execution as time-invariant context, while online vibrotactile feedback handles reactive adjustment. Every clip is an autonomous real-world policy rollout, not a teleoperated demonstration.

Task success, with and without object context

Successfully completing the task requires both the correct strategy and successful physical execution.

RearrangementSeen objects
41%
83%
RearrangementUnseen objects
15%
67%
Pouring & recyclingSeen + unseen objects
47%
77%
Box openingDeltaHand
25%
100%
Overall task success from Fig. 7 and Table 2 of the paper, rounded to whole percentages. Rearrangement uses 54 trials per split and method, pouring 30 trials per method, and box opening 48 trials per method. “Text context” uses property descriptors inferred from tactile exploration, not ground-truth labels. The baseline retains object position but has no object-property context.

Multi-object rearrangement. Rigidity determines how the object is grasped, and contents determine where it belongs. An empty container will be placed into the recycle bin through a top grasp.

Rigidity decides the grasp

Rigid top grasp Deformable side grasp

Contents decide the destination

Liquid yellow Granular purple Solid green Empty recycle bin
Rigidity
Contents
No rollout clip was recorded for this combination. The rule on the right still applies.

Rearrangement failure modes

Two representative failures show what goes wrong when the strategy does not match the object. Playback is 4× speed.

4× speed
✗ Wrong strategy. A deformable object is misread as rigid, so the policy uses a top grasp instead of a side grasp.
4× speed
✗ Failed grasp. In-hand adjustment fails to reach an accurate target and a stable grasp.

Three consecutive placements from a single rearrangement episode.

First placement. Liquid-filled rigid jar. Top grasp, yellow area.4× speed
Second placement. Liquid-filled deformable bottle. Side grasp, yellow area.4× speed
Third placement. Granular-filled deformable can. Side grasp, purple area.4× speed

Pouring & recycling. The policy selects its strategy from the inferred fill level. Filled containers are poured into the sink before recycling, while empty containers are recycled directly.

✓ Filled container, poured then recycled4× speed
✓ Filled container, poured then recycled4× speed
✓ Empty container, recycled directly4× speed

Why object context matters

Representative failure and success rollouts illustrate strategy selection; these are not synchronized, matched trials.

✗ Without context. An inappropriate action sequence spills the liquid4× speed
✓ With context (ours). An unseen container handled with the correct strategy4× speed

Box opening with the DeltaHand. For the DeltaHand, exploration also captures weight-related properties that distinguish visually identical boxes. The hand opens the box containing screws and takes no action on the empty one.

Filled box, opened via coordinated finger sliding20× speed
Empty box, no action8× speed

Object Representation Learning Details

We evaluate the contrastive learning and language alignment method on 9 seen and 9 unseen objects, each with 3 independently collected test trials that are never used in training. Seen means the object and its property combination appear in the training set. The same framework on the multi-finger DeltaHand reaches 96% on seen and 88% on unseen objects.

Exploratory actions. A Franka arm with fin-ray parallel-jaw fingers probes each object with six predefined actions, while vibrotactile signals from fingertip contact microphones, 6-axis end-effector forces and end-effector velocities are recorded in synchronization.

Swing
Shake
Squeeze
Tap on table
Tap on cavity
Slide

A complete exploration with sound.

Playback at 2× speed. The timer in the video shows real elapsed time. Metal, deformable, granular contents, half filled.

Learning framework. The three modalities are encoded into a multimodal token sequence, compressed by a Top-K token learner, and aligned with text embeddings of physical properties.

Learn from exploration

  1. Encode synchronized signalsVibrotactile spectrograms, end-effector forces, and velocities are encoded in 0.5-second windows.
  2. Keep informative tokensA Top-K token learner compresses the trajectory into a compact object embedding.
  3. Align with property descriptionsDuring training, contrastive losses align the representation with text embeddings of physical properties.

Use object context for manipulation

  1. Infer properties from touchAt inference, the learned embedding is compared with candidate property descriptors using cosine similarity.
  2. Choose the context representationThe paper evaluates both the tactile embedding itself and text embeddings of its decoded property descriptors.
  3. Condition the task policyThe current object's representation and position condition an Action Chunking Transformer alongside visual, proprioceptive, and online tactile observations.

Property estimation on independently collected test trials

Accuracy averaged over the four property axes and eight training seeds. Properties are decoded by cosine similarity between the tactile embedding and candidate text descriptors, so the decoding vocabulary can be changed without retraining.

Training promptsThe descriptors used during training
93%
84%
29 alternative descriptorsNot used during training
88%
81%
17-descriptor subsetNo shared word stem with training prompts; e.g., steel for metal
85%
79%
Property decoding generalizes to the tested alternative descriptors without retraining, with a 3 to 8 percentage-point drop in accuracy relative to training prompts.

Which moments does the model keep?These examples show selected tokens concentrated in particular exploration phases. Select an object to inspect its trajectory and predicted properties.

Selected tokens for an empty metal can across squeeze, swing, table tap, cavity tap, shake, and slide phases.

Ground truth on the left, model prediction on the right. Green marks a match.

ContentsGT empty → Pred emptyFill levelGT empty → Pred emptyMaterialGT metal → Pred metalRigidityGT deformable → Pred deformable

Selected tokens occur in squeezing, swinging, cavity tapping, and shaking, with few in table tapping or sliding.

Tactile representations meet language

Explore tactile and language alignment

Visualization of 136 training-set tactile embeddings from one frozen model run. Contrastive learning aligns tactile object representations with language descriptions of physical properties. The same 136 tactile embeddings stay in place as you color them by contents, fill level, material, or rigidity. Diamonds show the corresponding language anchors.

Frozen UMAP projection showing tactile samples and language anchors colored by contents.
136 training-set tactile embeddings from one frozen model run, colored by contents. Contents are solid, granular, liquid, or empty.

The complete walkthrough

Supplementary Video

A narrated walkthrough of exploration, representation learning, and robot manipulation.

BibTeX

@misc{yang2026objectcentric,
  title         = {Learning with Object-centric Representations of Tactile Interactive Perception for Robot Manipulation},
  author        = {Yang, Xinyi and Si, Zilin and Xu, Zhuowei and Temel, Zeynep and Kroemer, Oliver},
  year          = {2026},
  eprint        = {2609.33235},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2609.33235}
}