Counterfactuals

Results

The 200-state study evaluates both kinds of change:

Change

True shift

Policy error

Value loss

Increase the first reward parameter by 1.0

0.0829

0.0064

0.0030

Slow engine deterioration

0.0454

0.0067

0.0018

Policy distance ranges from zero to one. Zero means the two policies choose each action with the same probability in every state. For both changes, the fitted-model policy is within 0.0067 of the true-parameter policy.

Expected-value loss compares the fitted counterfactual policy with the policy computed from the true parameters. It is 0.0030 for the reward change and 0.0018 for the transition change. See the Simulation Study for the corresponding estimation and inference results.