Suboptimal policy determination for large-scale Markov decision processes. II: Implementation and numerical evaluation (Q5896459)
From MaRDI portal
!
This is the item page for this Wikibase entity, intended for internal use and editing purposes. Please use the normal view instead:
scientific article; zbMATH DE number 3874993
| Language | Label | Description | Also known as |
|---|---|---|---|
| default for all languages | No label defined |
||
| English | Suboptimal policy determination for large-scale Markov decision processes. II: Implementation and numerical evaluation |
scientific article; zbMATH DE number 3874993 |
Statements
Suboptimal policy determination for large-scale Markov decision processes. II: Implementation and numerical evaluation (English)
0 references
1985
0 references
We present an implementation of the procedure for determining a suboptimal policy for a large-scale Markov decision process (MDP) presented in Part 1 [see the preceding review]. An operation count analysis illuminates the significant computational benefits of this procedure for determining an optimal policy relative to a procedure for determining a suboptimal policy based on state and action space aggregation. Results of a preliminary numerical study indicate that the quality of the suboptimal policy produced by the 3MDP approach shows promise.
0 references
suboptimal policy
0 references
large-scale Markov decision process
0 references
0 references
0.9152982831001282
0 references
0.7883049249649048
0 references
0.7755498886108398
0 references
0.7645446062088013
0 references
0.757560670375824
0 references