SemAnCorr
Under Review

SemAnCorr: Semantic Anchored Correspondence for Zero-Shot Manipulation Skill Transfer

Anonymous Author(s) Author names and affiliations withheld for double-blind review.
SemAnCorr produces dense vertex-level correspondence across geometrically diverse object instances, enabling zero-shot transfer of manipulation skills such as pouring, wiping, and grasping.

SemAnCorr computes dense, vertex-level correspondence across geometrically diverse object instances that respects semantically similar parts while smoothly spanning the object surface (left). The resulting correspondence transfers a single demonstration to previously unseen objects that share functional structure but differ in geometry (right).


Video

Overview video

A short overview of SemAnCorr: the semantic‑anchored correspondence problem, our training‑free pipeline, and real‑world manipulation skill demonstrations and transfer executions.


Abstract

Transferring manipulation skills across object instances that share functionality but differ in geometry remains a fundamental challenge in robot learning. While recent correspondence methods leverage dense visual descriptors and 3D feature fields, nearest‑neighbor feature matching often produces spatially incoherent correspondences that fail to recover the local geometric frames required for reliable skill transfer.

We introduce SemAnCorr, a training‑free framework that establishes dense correspondence by selecting semantically consistent anchor regions through joint pose‑correspondence optimization and propagating these constraints over the object surface using functional maps. The resulting correspondences preserve both semantic consistency and geometric coherence, enabling object‑centric manipulation skills to transfer across geometrically diverse instances.

We evaluate SemAnCorr on a dense correspondence benchmark built on PartNet‑Mobility, achieving 90.8% semantic accuracy while substantially improving geometric coherence over existing correspondence and affordance‑transfer baselines. Finally, we show these improvements translate directly into real‑world manipulation performance: using a single demonstration, SemAnCorr enables substantially more reliable zero‑shot manipulation skill transfer to previously unseen objects than existing correspondence methods.

90.8%
Mean semantic accuracy across 7 PartNet‑Mobility categories
0.40
Geometric Coherence Score — 2× the next best baseline
1
Demonstration needed per skill — no task‑specific training
~6s
Per object pair on a single RTX 4090

Method

Anchor semantics, then let geometry fill in the rest

Semantic correspondence tells a robot where to interact; geometric coherence tells it how. SemAnCorr treats these as complementary constraints on a single dense correspondence rather than as separate problems — semantically confident regions are anchored first, and a functional map propagates those constraints smoothly across the rest of the surface.

01

Multi‑view semantic lifting

Renders of each mesh are encoded with pretrained SigLIP2 features, lifted to 3D, and clustered into spatially coherent, semantically consistent part regions.

02

Joint pose & anchor optimization

Candidate cluster pairs are scored by relative cosine similarity and bilateral margin confidence, then refined by jointly optimizing a rigid alignment against Chamfer‑based geometric compatibility — resolving cases where semantically similar parts sit in geometrically incompatible poses.

03

Semantic‑anchored functional map

High‑confidence anchor pairs are hard‑pinned into a functional map over Laplace‑Beltrami spectral bases, then refined with a geometry‑constrained ZoomOut scheme that upsamples the basis while re‑pinning anchors to prevent drift.

04

Object‑centric skill transfer

A demonstrated skill — contact keyframes plus relative motion segments — is transferred by aligning contact regions through the dense correspondence and updating end‑effector poses via Procrustes alignment.

SemAnCorr pipeline: multi-view rendering and patch embedding, vertex embedding and clustering, pose optimization and anchor selection, sparse anchor correspondence constraining a functional map, producing dense correspondence.

Semantic features are extracted and lifted into 3D to form clusters (red); corresponding anchors are selected via joint semantic‑geometric optimization (green); anchors constrain a functional map (blue) to produce smooth, dense correspondence.


Quantitative Results

Within‑category correspondence

PartNet‑Mobility

Evaluated on a dense correspondence benchmark built from seven PartNet‑Mobility categories, against functional‑map, 2D semantic, and 3D feature‑field baselines. SemAnCorr leads on semantic accuracy in every category and is the only method that achieves both high continuity and meaningful surface coverage.

Scissors
Pliers
Eyeglasses
Knife
Suitcase
Bottle
Kettle
Source
loading…
loading…
loading…
loading…
loading…
loading…
loading…
Target
loading…
loading…
loading…
loading…
loading…
loading…
loading…
Drag to rotate · scroll to zoom — each panel is an independent live mesh viewer.

Within‑category correspondence. Top row: source coloring. Bottom row: each vertex takes the corresponding point's color from the source, rendered as interactive 3D meshes.

Source Object
Coloring
FM‑WKS
Robo‑ABC
DenseMatcher
D3Fields
SemAnCorr
(Ours)
loading…
loading…
loading…
loading…
loading…
loading…
loading…
loading…
loading…
loading…
loading…
loading…
Drag to rotate · scroll to zoom — each panel is an independent live mesh viewer.

Baseline comparison. Source coloring on the left; corresponded colorings on the right for each method, rendered as interactive 3D meshes.

MethodScissorsPliersEyeglassesKnifeSuitcaseBottleKettleAverage
FM‑WKS58.5 ± 10.763.7 ± 16.159.1 ± 15.134.3 ± 28.683.8 ± 25.162.5 ± 37.179.7 ± 14.463.1 ± 15.0
DenseMatcher71.5 ± 16.769.4 ± 17.570.7 ± 14.863.3 ± 25.554.6 ± 28.854.9 ± 27.871.5 ± 7.065.1 ± 7.1
Robo‑ABC72.9 ± 9.174.4 ± 10.962.3 ± 18.551.9 ± 26.971.5 ± 14.846.9 ± 26.462.6 ± 6.863.2 ± 9.9
D3Fields86.2 ± 4.388.9 ± 7.085.4 ± 7.885.8 ± 7.983.2 ± 13.378.9 ± 7.484.1 ± 8.184.6 ± 2.9
SemAnCorr (Ours)88.5 ± 3.094.0 ± 2.593.5 ± 2.390.8 ± 3.892.2 ± 5.589.4 ± 4.586.9 ± 9.990.8 ± 2.5

Within‑category semantic accuracy (sAcc ± std, %). SemAnCorr improves 6.2 points over the strongest baseline, D3Fields.

MethodContinuity ↑Coverage ↑GCS ↑
FM‑WKS0.96 ± 0.010.003 ± 0.0010.007 ± 0.003
DenseMatcher0.93 ± 0.070.01 ± 0.010.01 ± 0.01
Robo‑ABC0.77 ± 0.110.07 ± 0.060.11 ± 0.10
D3Fields0.61 ± 0.080.10 ± 0.070.16 ± 0.09
SemAnCorr (Ours)0.93 ± 0.030.25 ± 0.110.40 ± 0.12

Geometric Coherence Score (GCS), the harmonic mean of continuity and coverage. FM‑WKS's high continuity comes from collapsing correspondence onto a handful of vertices; SemAnCorr is the only method with both high continuity and broad coverage.

Semantic Accuracy
$$\text{sAcc} = \frac{1}{|P|}\sum_{p \in P} \mathbb{1}\!\left[\, l_p = l_{f(p)} \,\right]$$
Geometric Coherence Score
$$\text{GCS} = \frac{2 \cdot \text{Cont} \cdot \text{Cov}}{\text{Cont} + \text{Cov}}$$

sAcc measures the fraction of target vertices whose transferred label lf(p) matches ground truth — where to interact. GCS is the harmonic mean of continuity (local smoothness) and coverage (surface reached by the map), rewarding correspondences that are both smooth and non‑collapsing — how to interact.


Generalization

Beyond category boundaries

Many manipulation skills transfer across functionally related but visually distinct categories. SemAnCorr's anchor confidence also doubles as a lightweight, training‑free signal for whether a demonstrated skill is likely to transfer between two objects at all.

Source
Target
Scissors → Pliers
loading…
loading…
Kettle → Bottle
loading…
loading…
Drag to rotate · scroll to zoom — each panel is an independent live mesh viewer.
Cross‑category correspondence: Scissors → Pliers (top), Kettle → Bottle (bottom), rendered as interactive 3D meshes.
MethodScissors → PliersKettle → Bottle
FM‑WKS55.3 ± 14.772.3 ± 15.4
DenseMatcher63.5 ± 18.965.2 ± 13.8
Robo‑ABC70.3 ± 12.152.4 ± 15.6
D3Fields87.2 ± 4.683.4 ± 7.5
SemAnCorr (Ours)89.7 ± 2.287.8 ± 4.2

Cross‑category semantic accuracy (sAcc ± std, %).

Category‑level anchor confidence. Within‑category pairs (diagonal) score highest; functionally related categories — containers, hand tools, appliances — score relatively higher than unrelated ones, even without any comparable baseline metric. Rendered as a vector graphic — text and values are selectable.


Ablation

Every component earns its place

ConfigurationsAccGCS
SemAnCorr (Ours)90.8 ± 2.50.40 ± 0.12
w/o Semantic Clustering66.0 ± 12.60.05 ± 0.03
w/o Anchor Selection39.1 ± 12.10.08 ± 0.06
w/o Functional Map72.9 ± 7.40.25 ± 0.07

Removing anchor selection causes the largest drop in semantic accuracy; removing semantic clustering most damages geometric coherence.

ClusteringSuppresses noisy lifted features before correspondence estimation — without it, geometric coherence collapses.
Anchor selectionDetermines where manipulation should transfer — without it, semantic accuracy drops the most.
Functional mapProvides smooth dense propagation — without it, both semantic and geometric quality degrade.
loading…
Source Object
Coloring
loading…
w/o Semantic
Clustering
loading…
w/o Anchor
Selection
loading…
w/o Functional
Map
loading…
SemAnCorr
(Ours)
Drag to rotate · scroll to zoom — each panel is an independent live mesh viewer.

Source object (left) and target correspondence under each ablation: w/o semantic clustering, w/o anchor selection, w/o functional map, and the full pipeline, rendered as interactive 3D meshes.


Real‑World Manipulation

From one demonstration to a novel object

Five kinesthetically demonstrated skills — peeling, opening, pouring, twisting, and cutting — are transferred zero‑shot to unseen target objects using only the computed dense correspondence. A contact keyframe (end‑effector pose + contact region) and relative motion segments are transferred via Procrustes alignment to the corresponding region on the target.

Real-world manipulation transfer across five tasks: banana peeling, pen opening, kettle pouring, bottle twisting, and scissor cutting, showing observations, correspondences, and execution key frames.
For each task, a skill demonstrated on a source object (left) is transferred to a semantically related target object (right) via the computed dense correspondence (second row); remaining rows show key frames of the source demonstration and the transferred execution.
TaskD3FieldsSemAnCorr
Task 1: Banana Peeling7/108/10
Task 2: Pen Cap Opening9/109/10
Task 3: Kettle Pouring3/107/10
Task 4: Bottle Twisting1/107/10
Task 5: Scissor Cutting3/106/10

Real‑world task success rate, 10 trials per task with varied object pose and viewpoint.

SemAnCorr real-world object correspondence for all five tasks.

SemAnCorr correspondence on real‑world object pairs, Tasks 1–5.

D3Fields failure cases on Task 3 and Task 4, showing spatially incoherent correspondence leading to incorrect local frames.

D3Fields fails on Tasks 3 & 4: spatially incoherent correspondence yields incorrect local frames, causing task failure.

Both methods perform comparably on simpler tasks where the correct interaction region is relatively unambiguous. On tasks that additionally require preserving a local geometric frame — approach direction, end‑effector orientation — D3Fields' spatially incoherent correspondence produces erroneous frames despite often identifying the correct semantic region. SemAnCorr's remaining failures stem mainly from mesh‑to‑workspace alignment error rather than correspondence estimation.