Counterfactuals
Read this page as reward-transfer evidence inside the supplied MDP. MCE-IRL does not expose the same one-call wrapper as NFXP, but the recovered reward can be re-solved in the simulation environment.
The public MCEIRL wrapper exposes the recovered policy, reward matrix, and
value function. It does not yet provide a one-call counterfactual method like
the structural likelihood wrappers.
The counterfactual results come from the simulation harness. The harness reruns the dynamic program under controlled changes and compares the recovered-reward policy with the oracle policy. See the simulation study page for the generator script, table source, and JSON results file.
Counterfactual Families
Type |
Intervention |
Purpose |
|---|---|---|
Type A |
Shift rewards and hold transitions fixed. |
Payoff counterfactual. |
Type B |
Change transitions and hold rewards fixed. |
State-dynamics counterfactual. |
Type C |
Disable one non-anchor action. |
Action-set or design counterfactual. |
Reported Results
These rows are from the primary mce_low_high_reward simulation results file.
Counterfactual |
Policy TV |
Policy KL |
Value RMSE |
Regret |
|---|---|---|---|---|
Type A |
0.006456 |
0.000157 |
0.000742 |
0.000433 |
Type B |
0.006284 |
0.000142 |
0.000523 |
0.000410 |
Type C |
0.004211 |
5.98e-5 |
0.000145 |
0.000094 |
Regret reports how the policy induced by the recovered reward compares with the oracle counterfactual policy.
API Boundary
For package users, the stable public objects are the fitted reward, policy, and value arrays. For controlled payoff, transition, or action-set interventions, use the simulation and evaluation utilities with an explicit problem and transition environment.