HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL

JunHyeok Oh1, Zian Jang1, Byung-Jun Lee1,†

1Korea University

Preprint † Corresponding author

Two route plans on Maze2D-Large growing from only a start and a goal: tokens are inserted until sigma reaches 1, then they only move into place. The near goal ends with 10 tokens, the far goal with 38.

HorizonFlow treats plan length as an output of generation. Starting from only the start and goal, tokens are inserted while the generation clock σ is below 1 and refined by flow matching; the near goal settles at 10 tokens, the far one at 38 — no horizon was given.

Abstract

Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is too short can force infeasible transitions, whereas one that is too long can introduce redundant motion. We introduce HorizonFlow, a hierarchical planner that treats plan length as an output of generation rather than a prescribed input. Its subgoal route planner guides its action-prefix controller through a sequence of latent subgoals. Both components combine insertion-based generation with flow matching to jointly generate continuous plan content and length, using the partially generated plan to guide token insertion. HorizonFlow reuses the resulting length information to select candidates and steer generation toward shorter plans without a separate learned value model. Across Maze2D, Multi2D, and OGBench navigation and visual manipulation benchmarks, HorizonFlow achieves the highest average performance among the compared methods.

A prescribed horizon is a guess about travel time

Inpainting planners pin the start and the goal and fill in the plan between them, so the plan length is an assumed travel time chosen before any content exists. Too short, and the plan is squeezed through walls; too long, and it wanders before it arrives.

Plans being generated on Maze2D-Large for prescribed horizons 64, 128, 256 and 384 steps. Diffuser and a fixed-length HorizonFlow planner produce wall-crossing plans for the far goal at short horizons and detours for the near goal at long ones; HorizonFlow with generated length reaches both goals the same way at every horizon.
Generation under a prescribed horizon. For each H = 64 → 128 → 256 → 384 steps, three planners each generate one plan from noise: Diffuser (256 denoising steps, every 8th shown), a fixed-length HorizonFlow planner with the same Transformer but trained without insertion, and HorizonFlow with insertion (20 flow steps). Given H = 64, both fixed-length planners reach the far goal (N* = 320) only by cutting through walls; given H = 384, they detour around the near goal (N* = 64). HorizonFlow ignores H and picks its own number of waypoints, 13 steps apart (3 for the near goal, 27 for the far one). Plans are shown before any controller runs.
Heatmaps of score, reach rate and steps to goal over nominal distance. Top row: a fixed-length planner at horizons 64 to 384 loses reach beyond its horizon and takes about H steps below it. Bottom row: HorizonFlow with 1 to 256 candidates reaches 99 to 100 percent at every distance.
Across 600 start–goal pairs. Maze2D-Large, 100 pairs per distance bin, every plan executed by the same waypoint controller. Top: a fixed-length planner given horizon H (outlined where H matches the distance) fails beyond its own length and wastes time below it. Bottom: HorizonFlow with generated length reaches nearly every goal, and more candidates bring its time to goal down. Darker is better.

Insertion and flow matching on one clock

A plan is a set of continuous tokens between two fixed anchors, each with an order coordinate. One Transformer predicts, for every gap between neighbouring tokens, how many tokens are still missing (a hurdle zero-truncated-Poisson head), and for every token a flow-matching velocity. Sampling advances a single clock σ from 0 to 2.

0 ≤ σ < 1gaps give birth to new noise tokens while existing ones are refined
1 ≤ σ ≤ 2the count is fixed; every token finishes its own refinement

At a fixed temporal resolution, the token count is the plan's duration. That makes length a free ranking signal: sample K candidates and keep the one with the fewest tokens, with no learned value function. During generation, Feynman–Kac steering uses the same count heads to move samples toward shorter plans.

Sixteen of 64 route candidates for one start and goal grow at the same time along different corridors; at the end all but the one with the fewest tokens fade out.
Best-of-K by count. One sampler call draws 64 route candidates for a start–goal pair with two corridors of equal length; the 16 shown were chosen to cover the different routes in the pool (upper corridor, lower corridor, loops). The candidate with the fewest tokens (n = 14) is kept.
Sixteen colour-coded chains generate plans in one maze for a far goal; at two checkpoints a panel shows each chain's predicted length and resampling weight, chains without offspring fade out and survivors are copied, so fewer colours remain; at the end the shortest chain is highlighted.
Feynman–Kac steering. Sixteen chains for a far goal (N* = 382), one colour per original chain. At σ = 0.3 and 0.7 each chain is scored by its predicted final length n̂ — tokens present plus the tokens the count heads still expect — and resampled with weight ∝ exp(−0.5 n̂). Chains without offspring are dropped and survivors are copied in their place, so the colours thin out and compute moves to short routes before they are finished. Faint strokes are tokens still being refined.

A route planner for where, a prefix controller for how

Writing a whole route at action resolution means very long sequences, so HorizonFlow applies the same insertion–flow model twice. The route planner generates a variable-length sequence of latent subgoals toward the final goal and is refreshed every half training stride. The action-prefix controller generates a variable-length action prefix toward the first subgoal; it replans every step and executes only the first action. Both pick their candidate by count.

AntMaze-Giant task 2: rendered episode next to a top-down map with the ant's trail and the current route plan, which is replaced every 8 steps and gets shorter as the ant nears the goal. The goal is reached at step 742.
AntMaze-Giant, task 2 — goal reached at step 742. Right: the agent's trail, the route the planner committed to at its latest update, the subgoal the controller is steering to (ring), and the route length chosen at every update, which shrinks as the goal gets closer.
HumanoidMaze-Giant task 3: rendered episode next to the top-down map; the route is re-planned every 32 steps and the goal is reached at step 3,254.
HumanoidMaze-Giant, task 3 — goal reached at step 3,254 of 4,000.
PointMaze-Giant task 4: rendered episode next to the top-down map; the goal is reached at step 619.
PointMaze-Giant, task 4 — goal reached at step 619.

Benchmark recipe (K = 16 routes × 4 prefixes, FK at both levels). Each clip plays a whole episode in 10 seconds. Latent subgoals are drawn at the mean position of their 8 nearest dataset states in the planner's latent space; counts exclude the start and goal anchors. Clips are successful episodes picked from 16 rendered episodes, 14 of which succeeded.

Results

HorizonFlow attains the highest average in every benchmark group, with the largest margins in long-horizon environments — 95.5% on HumanoidMaze-Giant against 35% for the best prior method.

BenchmarkMetricBest prior methodHorizonFlow
Maze2D (single goal), 3 mazesnormalized score154.2SSD165.8
Multi2D (goal resampled), 3 mazesnormalized score168.8SSD177.4
OGBench navigation, 9 mazessuccess %75.4SAW94.2
OGBench visual manipulation, 4 taskssuccess %49.5HIQL55.0

Averages over each group's environments. HorizonFlow uses 5, 8 and 4 training seeds for Maze2D/Multi2D, navigation and visual manipulation; 250 problems per OGBench environment (50 per task). Per-environment numbers and all baselines are in the paper's Table 1.

Does the generated count know the distance?

Length signal (2,400 Maze2D-Large pairs)Spearman ρ with N*ρ on pairs with N* ≥ 300
VH-Diffuser predicted length0.6530.070
HIQL value-derived length0.9240.371
HorizonFlow generated count0.9540.617

Both levels matter

Architecture (OGBench navigation)Average success %
Flat planner (no route planner)42.9
Route planner + behavior-cloning controller67.8
Route planner + prefix controller (HorizonFlow)94.2

About 8.5 ms of model computation per environment step at the default settings (RTX 5090).

BibTeX

The BibTeX entry will be posted here with the arXiv preprint.