-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathindex.html
More file actions
241 lines (224 loc) · 20.8 KB
/
Copy pathindex.html
File metadata and controls
241 lines (224 loc) · 20.8 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="description" content="Plan2Pose learns prescient, predicate-agnostic downward refinement of symbolic plans into continuous object poses.">
<meta name="robots" content="noindex, nofollow">
<meta name="theme-color" content="#4c6a8d">
<title>Plan2Pose — Prescient Downward Refinement of Symbolic Plans</title>
<link rel="stylesheet" href="static/css/style.css">
</head>
<body>
<div class="scroll-progress" aria-hidden="true"><span></span></div>
<header class="hero">
<div class="container hero-inner">
<h1 class="paper-title"><span>Plan2Pose</span>: Prescient Downward<br> Refinement of Symbolic Plans<br> over Long Horizons</h1>
<p class="authors">Anonymous authors</p>
<p class="venue">Under review as a conference paper at ICLR 2027</p>
<div class="link-buttons" aria-label="Project links">
<a class="btn btn-ghost" href="https://github.com/plan2pose/plan2pose-code.git" target="_blank" rel="noopener" aria-label="Plan2Pose code repository">
<svg aria-hidden="true" viewBox="0 0 24 24"><path d="M15 22v-4a4.8 4.8 0 0 0-1-3.5c3.3-.4 6.8-1.6 6.8-7A5.4 5.4 0 0 0 19.4 4 5 5 0 0 0 19.3.5S18.2.1 15 1.8a13.4 13.4 0 0 0-7 0C4.8.1 3.7.5 3.7.5A5 5 0 0 0 3.6 4a5.4 5.4 0 0 0-1.4 3.7c0 5.4 3.5 6.6 6.8 7A4.8 4.8 0 0 0 8 18v4m0-3c-3 .9-3-1.5-4-2"/></svg>
GitHub
</a>
</div>
</div>
</header>
<main>
<section class="hero-visual section-tight">
<div class="container">
<figure class="paper-panel hero-approach reveal">
<img src="static/images/fig1_approach.svg" alt="Plan2Pose approach overview from observed point cloud and symbolic plan to predicted geometric states and executable SE(3) targets">
<figcaption>Plan2Pose refines a complete symbolic plan into target object poses, one forward pass per step, before a motion planner or policy realizes each target.</figcaption>
</figure>
</div>
</section>
<section class="section abstract" id="abstract">
<div class="container container-narrow reveal">
<h2>Abstract</h2>
<p>Task and motion planning uses symbolic abstractions to determine a high-level, symbolic action plan, and then refines the actions into continuous robot configurations. This downward refinement typically makes use of hand-designed samplers, constraint solvers, or predicate-specific learned models, which require hand-crafted predicate-specific supervision or losses. Learned models are myopic and only consider the current configuration’s requirements, producing refinements that may cause subsequent actions to be unrealizable. This is particularly problematic in long-horizon tasks that require many sequential refinements. We propose a refinement learning method that is prescient where needed, i.e., it learns the relevant aspects of the remaining plan instead of myopically focusing on the current action. In contrast to previous approaches, we use a predicate-agnostic approach that avoids hand-crafted predicate-specific losses.</p>
<p>Given a PDDL plan and an object-segmented point cloud of the initial scene, we use a transformer architecture to embed the whole action sequence. For a given state, we learn placement fields that are composed to produce an SE(3) pose that satisfies all object relations described in the state. These placement fields are learned without predicate-specific losses, evaluators, or hand-coded geometric semantics, but rather purely from learning to replicate the distributions seen in demonstrations. Our approach, <strong>Plan2Pose</strong>, is trained on synthetic rearrangement plans with at most nine states and outperforms existing baselines on target-atom satisfaction, especially on OOD plan lengths. Under autoregressive rollout on 200 plans with 17–20 states, Plan2Pose satisfies 95.5% of state atoms in the final state, compared to 58.4–76.8% achieved by the evaluated learned baselines, demonstrating substantially more robust long-horizon refinement. Furthermore, Plan2Pose transfers zero-shot to real point clouds; fine-tuning on the limited real-capture set instead reduces performance.</p>
</div>
</section>
<section class="section placements-section">
<div class="container">
<div class="section-heading reveal">
<h2>Learned Placements and Composition</h2>
<p>Plan2Pose represents downward refinement with one transformer that encodes the observed scene and the complete symbolic plan. It is rolled out autoregressively to predict the target poses of the objects moved at each step.</p>
</div>
<div class="placement-explanation">
<article class="placement-topic reveal">
<h3>Scene and plan encoding</h3>
<p>Each object is represented by a segmented point cloud sampled to 512 points. A shared PointNet encodes the centered cloud, while pose and identity embeddings retain its initial position and link the same object across the plan. The initial state and the ADD and DEL events of subsequent actions are represented as predicate and argument tokens. Query tokens identify the objects whose poses must be predicted.</p>
</article>
<article class="placement-topic reveal">
<h3>Per-atom placement fields</h3>
<p>For a moved object, the prediction head considers atoms added or deleted by the current action and carried atoms that remain true in the current state. A shared FiLM-conditioned decoder translates each relevant atom into a placement field in the frame of its relatum. The decoder is shared across predicates, so relational geometry is represented by the transformer context rather than predicate-specific weights.</p>
</article>
<article class="placement-topic reveal">
<h3>Pose composition</h3>
<p>The fields are registered onto a common object-centered canvas and combined as a weighted product of experts. A location retains high probability only when every relevant factor supports it, so additional atoms narrow the admissible region. Height is predicted over 160 bins and rotation is pooled separately, producing the complete SE(3) target pose in one forward pass.</p>
</article>
</div>
<figure class="paper-panel placement-fields-figure reveal">
<img src="static/images/fig2_placement_fields.svg" alt="Decoded Plan2Pose placement fields for left, south, on, in, and the fine-tuned near predicate">
<figcaption>Decoded placement fields learned from demonstrations. The shape of the relatum affects the fields for <em>on</em> and <em>in</em>; <em>near</em> is acquired through fine-tuning.</figcaption>
</figure>
<div class="composition-detail reveal">
<div>
<h3>Weighted Product of Experts</h3>
<p>Each relevant symbolic atom produces its own spatial distribution. For example, one field may describe positions that are left of one object, while another describes positions that are south of a different object. Because these fields are initially expressed relative to their respective reference objects, Plan2Pose first shifts them onto a common canvas centered on the object being placed.</p>
<p>The aligned fields are then combined multiplicatively. This has an intersection-like effect: a position receives high probability only when it is supported by all relevant relations, while positions that violate any strongly weighted relation are suppressed. Each field also receives a learned confidence that determines how strongly it contributes to the combined prediction. A base factor is included so that the composition remains defined even when no symbolic atom directly constrains an object.</p>
<p>This differs from averaging or mixing the fields, which can retain positions that satisfy only one of several requirements. Product composition instead narrows the admissible region as more atoms are introduced. The resulting planar distribution is combined with the separately predicted height and orientation to form the final SE(3) target pose.</p>
</div>
</div>
<p class="training-note reveal"><strong>Training.</strong> The model is trained end-to-end on successful refinements with one loss on the composed prediction. Predicate geometry is learned only through its contribution to the demonstrated pose; no per-predicate loss or hand-coded geometric constraint is used.</p>
</div>
</section>
<section class="section results-table-section" aria-label="Main autoregressive refinement results">
<div class="container">
<div class="result-table-wrap reveal">
<table class="results-table">
<caption><strong>Main results.</strong> Autoregressive refinement on 200 plans in each state-count range. Values are means over three seeds. Task progress is measured over all states; collateral violation is lower-is-better.</caption>
<thead>
<tr>
<th rowspan="2" scope="col">Method</th>
<th colspan="3" scope="colgroup">Task progress (1)</th>
<th colspan="3" scope="colgroup">Final state correct (3)</th>
<th colspan="3" scope="colgroup">Mover-relevant satisfaction (5)</th>
<th colspan="3" scope="colgroup">Collateral violation (6) ↓</th>
</tr>
<tr>
<th scope="col">9–12</th><th scope="col">13–16</th><th scope="col">17–20</th>
<th scope="col">9–12</th><th scope="col">13–16</th><th scope="col">17–20</th>
<th scope="col">9–12</th><th scope="col">13–16</th><th scope="col">17–20</th>
<th scope="col">9–12</th><th scope="col">13–16</th><th scope="col">17–20</th>
</tr>
</thead>
<tbody>
<tr><th scope="row">B0 identity</th><td>0.630</td><td>0.610</td><td>0.575</td><td>0.005</td><td>0.000</td><td>0.005</td><td>0.524</td><td>0.519</td><td>0.501</td><td>0.299</td><td>0.326</td><td>0.374</td></tr>
<tr><th scope="row">B2 random</th><td>0.594</td><td>0.577</td><td>0.559</td><td>0.003</td><td>0.003</td><td>0.002</td><td>0.494</td><td>0.496</td><td>0.490</td><td>0.339</td><td>0.366</td><td>0.393</td></tr>
<tr><th scope="row">Rejection sampling</th><td>0.780</td><td>0.770</td><td>0.753</td><td>0.058</td><td>0.038</td><td>0.065</td><td>0.723</td><td>0.724</td><td>0.717</td><td>0.182</td><td>0.198</td><td>0.223</td></tr>
<tr><th scope="row"><a href="https://lingo-space.github.io/" target="_blank" rel="noopener" aria-label="LINGO-Space project website (opens in a new tab)">LINGO-Space<sup>z,s</sup></a></th><td>0.733</td><td>0.701</td><td>0.672</td><td>0.027</td><td>0.013</td><td>0.010</td><td>0.655</td><td>0.629</td><td>0.613</td><td>0.213</td><td>0.248</td><td>0.287</td></tr>
<tr><th scope="row"><a href="https://diffusion-ccsp.github.io/" target="_blank" rel="noopener" aria-label="Diffusion-CCSP project website (opens in a new tab)">Diffusion-CCSP</a></th><td>0.839</td><td>0.838</td><td>0.827</td><td>0.100</td><td>0.115</td><td>0.105</td><td>0.795</td><td>0.803</td><td>0.797</td><td>0.130</td><td>0.137</td><td>0.152</td></tr>
<tr><th scope="row"><a href="https://structdiffusion.github.io/" target="_blank" rel="noopener" aria-label="StructDiffusion project website (opens in a new tab)">StructDiffusion</a></th><td>0.817</td><td>0.806</td><td>0.804</td><td>0.075</td><td>0.097</td><td>0.075</td><td>0.767</td><td>0.760</td><td>0.767</td><td>0.149</td><td>0.162</td><td>0.170</td></tr>
<tr class="ours"><th scope="row">Plan2Pose (ours)</th><td>0.972</td><td>0.971</td><td>0.966</td><td>0.557</td><td>0.585</td><td>0.542</td><td>0.962</td><td>0.963</td><td>0.958</td><td>0.022</td><td>0.022</td><td>0.029</td></tr>
</tbody>
</table>
</div>
</div>
</section>
<section class="section selected-results-section" aria-labelledby="selected-results-title">
<div class="container">
<div class="section-heading reveal">
<h2 id="selected-results-title">Quantitative Results</h2>
<p>Key comparisons from the long-horizon, compounding, foresight, and deployment evaluations.</p>
</div>
<div class="metric-grid">
<article class="metric-card reveal">
<div class="metric-value">96.6%</div>
<h3>Task progress</h3>
<p>Across all states of plans with 17–20 states. The strongest evaluated baseline reaches 82.7%.</p>
<span>Higher is better</span>
</article>
<article class="metric-card reveal">
<div class="metric-value">54.2%</div>
<h3>Final states entirely correct</h3>
<p>Every atom in the final state is satisfied. The strongest evaluated baseline reaches 10.5%.</p>
<span>Higher is better</span>
</article>
<article class="metric-card reveal">
<div class="metric-value">94.6%</div>
<h3>Late-step satisfaction</h3>
<p>Autoregressive satisfaction at states 17–20. The strongest evaluated baseline reaches 78.7%.</p>
<span>Higher is better</span>
</article>
<article class="metric-card reveal">
<div class="metric-value">41.5 <small>mm</small></div>
<h3>Late-step xy deviation</h3>
<p>Median deviation written into the next reference frame at states 17–20, versus 98.1 mm for the strongest baseline.</p>
<span>Lower is better</span>
</article>
<article class="metric-card reveal">
<div class="metric-value">95.4%</div>
<h3>Platform completion</h3>
<p>In-distribution completion with whole-plan context, compared with 54.6% when the future is masked.</p>
<span>Higher is better</span>
</article>
<article class="metric-card reveal">
<div class="metric-value">29.9 <small>ms</small></div>
<h3>Latency per placement</h3>
<p>One forward pass on an NVIDIA L4, corresponding to an approximately 18.6–29.2× speedup over the diffusion-based refiners.</p>
<span>Lower is better</span>
</article>
</div>
</div>
</section>
<section class="section motivation">
<div class="container">
<div class="split-heading reveal">
<div>
<h2>Prescient Refinement</h2>
</div>
<p>Downward refinement is not only about satisfying the current symbolic state. Early choices must leave enough geometric room for the remainder of the plan.</p>
</div>
<div class="foresight-copy reveal">
<p>The foresight evaluation tests whether the first placement satisfies its current symbolic requirements while also preserving a valid refinement of the remaining plan. Each scene is paired with two possible futures whose extendable regions are disjoint. The current action and observation are otherwise unchanged, so the first placement can respond correctly only if the model uses information from later plan steps.</p>
<p><strong>Platform task.</strong> The first block is placed on a platform before later blocks introduce directional relations. The paired futures differ in which side later objects must occupy. A placement is counted as extendable only when it satisfies the current atoms and leaves the remainder of the plan geometrically satisfiable. With the future masked, the model repeatedly chooses the same upper-right region; with full plan context, Plan2Pose changes the first placement according to the later north–south requirements.</p>
<p><strong>Bridge task.</strong> The first two placements form the bridge pillars, while the final plan step names either a short or a long lintel. The short lintel requires the pillars to remain close together, whereas the long lintel requires greater separation. Plan2Pose conditions the pillar placement on the future lintel geometry; the future-masked variant receives identical input under both futures and therefore never changes the pillar spacing.</p>
<p><strong>Results and scope.</strong> On the in-distribution tests, Plan2Pose reaches 95.4% completion on the platform task and 89.3% on the bridge task, with a 36.5 mm average shift between paired platform futures. The result is limited to foresight patterns represented in the demonstrations: on the out-of-distribution six-block platform task, the predicted field favors the relevant region, but the decoded placement often misses it.</p>
</div>
<figure class="paper-panel foresight-figure reveal">
<img src="static/images/fig4_foresight.svg" alt="Platform and bridge evaluations comparing Plan2Pose with future context against future-masked predictions">
<figcaption>Foresight tests. Whole-plan context changes the first placement to preserve later platform relations and accommodate the geometry of a future bridge lintel.</figcaption>
</figure>
</div>
</section>
<section class="section method" id="method">
<div class="container">
<div class="section-heading reveal"> <h2>Model Architecture</h2>
<p>One shared architecture learns object geometry, plan context, and compositional relational constraints.</p>
</div>
<figure class="paper-panel architecture-figure reveal">
<img src="static/images/fig3_architecture.svg" alt="Plan2Pose architecture with scene and plan encoding, transformer tokens, FiLM-conditioned placement fields, product-of-experts composition, and pose heads">
<figcaption>The scene and complete plan become one token sequence. Per-atom fields are FiLM-conditioned and composed before decoding position, height, and rotation.</figcaption>
</figure>
<div class="method-steps">
<article class="method-step reveal"><span>01</span><h3>Encode shape</h3><p>A PointNet independently encodes each centered object point cloud, separating shape from world position.</p></article>
<article class="method-step reveal"><span>02</span><h3>Read the full plan</h3><p>Objects, predicates, arguments, and query tokens exchange information across the complete action sequence.</p></article>
<article class="method-step reveal"><span>03</span><h3>Compose constraints</h3><p>Each atom yields a normalized placement field in its reference object’s frame. A weighted product of experts intersects their support.</p></article>
<article class="method-step reveal"><span>04</span><h3>Decode and roll out</h3><p>Planar position, height, and orientation form an SE(3) pose. Refinement continues autoregressively, with no backtracking.</p></article>
</div>
</div>
</section>
<section class="section contributions">
<div class="container container-narrow reveal">
<h2>Contributions</h2>
<ol class="contribution-list">
<li><span>01</span><p>We formulate downward refinement as a conditional distribution over object poses given the whole symbolic plan, whose support should be restricted to refinements from which the remaining plan stays refinable.</p></li>
<li><span>02</span><p>We present Plan2Pose, a single model that encodes the whole plan with a transformer and composes per-atom placement fields as a product of experts. It requires no backtracking, manually defined constraints, samplers, object models, or predicate-specific losses, and acquires new predicates by fine-tuning.</p></li>
<li><span>03</span><p>We show that Plan2Pose, trained on plans with at most nine states, satisfies 95.5% of final-state atoms on plans with 17–20 states, compared to 58.4–76.8% for the evaluated learned baselines. In distribution, only Plan2Pose achieves an extendable first-placement rate above the 0.5 future-blind bound.</p></li>
<li><span>04</span><p>We show that Plan2Pose transfers zero-shot to real point clouds, acquires new predicates without forgetting, and refines each step in 29.9 ms, an approximately 18.6–29.2× speedup over diffusion-based refiners.</p></li>
</ol>
</div>
</section>
<section class="section citation" id="citation">
<div class="container container-narrow reveal">
<h2>Citation</h2>
<div class="citation-placeholder" role="note">
<span>BibTeX citation</span>
<p>To be added after the double-blind review period.</p>
</div>
</div>
</section>
</main>
<footer>
<div class="container footer-inner">
<a href="#" class="footer-mark" aria-label="Back to top">P<span>2</span>P</a>
<p>Plan2Pose · ICLR 2027 submission</p>
<p>Design reference: <a href="https://videomimic.net/" target="_blank" rel="noopener">VideoMimic</a>.</p>
</div>
</footer>
<script src="static/js/site.js"></script>
</body>
</html>