Agent Priors-guided Policy Learning, Puming Jiang★†, Tianrun Hu★†, Haozhe Du★, Yibo Li, Zhiwei Xue★, Xinhu Li★, and Harold Soh★, arXiv preprint
Links:
† Equal contribution. ★ CLeAR group member.

Robots trained on a handful of demonstrations must handle moved objects and new combinations of familiar actions. These challenges are connected: a planner can choose a sensible sequence, yet fail if a learned skill cannot execute from the state left by the previous one. A skill name alone says little about these limits.
Agent Priors-guided Policy Learning (APPL) makes the assumptions behind a policy available to the agent that uses it. For example, learning a grasp relative to an object can support transfer when that object moves. Documenting this same assumption helps the runtime agent decide whether to use the policy in a new scene.
A construction agent divides complete demonstrations into skills, proposes alternative priors, and trains and checks a policy for each prior. Each policy carries an interface describing its assumptions, handoff conditions, training support, and verification evidence. At deployment, a runtime agent uses these interfaces to select policies, set their arguments and stopping conditions, and compose them toward the requested goal. The policy library stays frozen during execution.
Results
- Few-demonstration skill learning: across six MetaWorld tasks, APPL reaches 89.6% out-of-distribution success with two demonstrations, compared with 37.9% for the fixed relational prior baseline.
- Long-horizon execution: on five ManiSkill tasks with twelve demonstrations per task, APPL achieves 50.0% success with shifted objects, versus 10.0% for full-task Diffusion Policy. It reaches 92.5% on variants that resume a task partway through or request an early stop.
- New compositions: APPL solves 8 of 16 unseen compositions. Hiding the interface information while keeping the same policy library reduces this to 2 of 16.
These simulation results highlight the value of sharing training assumptions with the agent that selects and combines skills.
Resources
Visit the project page for rollout videos and detailed results. The code repository provides experiment code, policy interfaces, evaluation records, and instructions for obtaining demonstrations and trained checkpoints. The project website source is available separately.
Citation
Please consider citing our paper if you build upon our results and ideas.
Puming Jiang★†, Tianrun Hu★†, Haozhe Du★, Yibo Li, Zhiwei Xue★, Xinhu Li★, and Harold Soh★, “Agent Priors-guided Policy Learning”, arXiv preprint
@article{jiang2026appl, title={Agent Priors-guided Policy Learning}, author={Jiang, Puming and Hu, Tianrun and Du, Haozhe and Li, Yibo and Xue, Zhiwei and Li, Xinhu and Soh, Harold}, journal={arXiv preprint arXiv:2609.35690}, year={2026} }
Contact
If you have questions or comments, please contact Puming or Harold.