Part of the Open Pace seed catalog. Viewer rows below are illustrative samples built from the schema; the files stream at launch.
Summary
Support questions grouped by intent, for testing clustering and near-repeat detection.
About 44k rows, 18 MB on disk, stored as parquet and released under mit.
Schema
text— stringcluster— label
Intended use
- Training and evaluating embeddings systems.
- Benchmarks, as long as the validation split stays out of training.
Considerations
- Check the license of the dataset and of any linked source before redistribution.
- Report problems with a signed note on this page so the authors can see them.