sable-r/gridrunner-ppo
Reinforcement learning2M paramssafetensorsLicense: mit#ppo#grid-world
Updated Sep 9, 2026
Part of the Open Pace seed catalog. Weights are listed for the preview and become downloadable at launch.
Overview
A PPO agent that solves procedurally generated grid mazes.
sable-r/gridrunner-ppo is a 2M-parameter reinforcement learning model, published in safetensors under the mit license.
Intended use
- Reinforcement learning in the domains described above.
- Research, prototypes and products that keep a person in the loop.
- Fine-tuning as a starting point for a narrower task.
How to use
# Command-line client (planned; the shape may change)
pace pull sable-r/gridrunner-ppo
# Python (planned)
from openpace import load
model = load("sable-r/gridrunner-ppo")Limitations
- The agent only knows the environments it was trained in. Small rule changes can break it.
License
Released under mit. Read the license file before using the weights commercially.