A Decentralized Partially Observable Team Decision Methodology with Delayed Information Sharing
A paper gives decentralized agents a way to learn team policies from delayed shared observations.
The paper studies decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models. Each member uses private information and delayed common information to learn an approximate low-rank Markov decision process and compute a policy with least-squares value iteration. The algorithm requires neither a central coordinator nor centralized training. The authors show that member-side solutions approximate the centralized team policy and derive finite-sample guarantees plus a sample-complexity bound.
- Agents learn approximate low-rank MDPs from local and delayed common information.
- Each member plans with least-squares value iteration and no centralized training.
- Member policies are shown to approximate the centralized team-optimal solution.
- The paper provides finite-sample guarantees and a sample-complexity bound.
Full article150 words · extracted from arxiv.org · click to collapse
We study decentralized partially observable team decision problems with low-rank latent dynamics and unknown system models. The proposed framework combines team-theoretic equivalence with low-rank model representations to address cooperative decision-making in partially observable Markov decision processes without prior knowledge of the transition model. Each team member makes decisions based on local private information and delayed common information shared across the team. Using only this available information, each member learns an approximate low-rank Markov decision process and applies least-squares value iteration to compute its policy. This yields a fully decentralized learning and planning algorithm that requires neither a centralized coordinator nor centralized training. We show that the resulting member-side solutions approximate the centralized team solution: despite partial observability, unknown dynamics, and delayed common information, each member recovers the corresponding component of an approximate team-optimal policy. We further establish finite-sample performance guarantees and derive a corresponding sample-complexity bound for the proposed algorithm.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.26783