OpenPace

sable-r/gridrunner-ppo

Reinforcement learning2M paramssafetensorsLicense: mit#ppo#grid-world
Updated Sep 9, 2026
Part of the Open Pace seed catalog. Weights are listed for the preview and become downloadable at launch.

Overview

A PPO agent that solves procedurally generated grid mazes.

sable-r/gridrunner-ppo is a 2M-parameter reinforcement learning model, published in safetensors under the mit license.

Intended use

  • Reinforcement learning in the domains described above.
  • Research, prototypes and products that keep a person in the loop.
  • Fine-tuning as a starting point for a narrower task.

How to use

# Command-line client (planned; the shape may change)
pace pull sable-r/gridrunner-ppo

# Python (planned)
from openpace import load
model = load("sable-r/gridrunner-ppo")

Limitations

  • The agent only knows the environments it was trained in. Small rule changes can break it.

License

Released under mit. Read the license file before using the weights commercially.

More reinforcement learning