Multi-Turn Interaction RL and Credit Assignment (Split)
New Locations
This page remains as an entry point for old links. Its original content has been split into two sections by topic:
- Formalization (multi-turn MDPs, POMDPs, trajectory structure, and action masks) → 19.2 Multi-Turn Reinforcement Learning
- Credit assignment (ORM/PRM/SALT, turn-level discounting, and the Mini Agent Loop experiment) → 19.3 Trajectory Credit Assignment
Use the two links above to open the corresponding material.