sable-r/maze-trajectories
Reinforcement learning120k rows300 MBparquetLicense: mit#offline-rl#trajectories
Updated Sep 7, 2026
Part of the Open Pace seed catalog. Viewer rows below are illustrative samples built from the schema; the files stream at launch.
Summary
Recorded episodes from grid mazes for offline reinforcement learning.
About 120k rows, 300 MB on disk, stored as parquet and released under mit.
Schema
episode— intobservation— listaction— labelreward— float
Intended use
- Training and evaluating reinforcement learning systems.
- Benchmarks, as long as the validation split stays out of training.
Considerations
- Check the license of the dataset and of any linked source before redistribution.
- Report problems with a signed note on this page so the authors can see them.