DUFOUR, François; GENADOT, Alexandre

doi:10.1007/s00245-018-9533-6

Metadatos

Licencia de uso del documento

DUFOUR, François
Institut Polytechnique de Bordeaux [Bordeaux INP]
Quality control and dynamic reliability [CQFD]
Institut de Mathématiques de Bordeaux [IMB]

GENADOT, Alexandre
Institut de Mathématiques de Bordeaux [IMB]
Quality control and dynamic reliability [CQFD]

Idioma

Article de revue

Este ítem está publicado en

Applied Mathematics and Optimization. 2020, vol. 82, n° 2, p. 433-450

Springer Verlag (Germany)

Resumen en inglés

We consider a discrete-time Markov decision process with Borel state and action spaces. The performance criterion is to maximize a total expected utility determined by unbounded return function. It is shown the existence of optimal strategies under general conditions allowing the reward function to be unbounded both from above and below and the action sets available at each step to the decision maker to be not necessarily compact. To deal with unbounded reward functions, a new characterization for the weak convergence of probability measures is derived. Our results are illustrated by examples.< Leer menos

Palabras clave en inglés

Markov decision processes

Expected total reward

Unbounded return

Weak convergence of measure

Metadatos

Compartir este ítem

Licencia de uso del documento

On the Expected Total Reward with Unbounded Returns for Markov Decision Processes

Idioma

Este ítem está publicado en

Resumen en inglés

Palabras clave en inglés

URI

DOI

Orígen

Centros de investigación