Quick Start
This page shows the feature-based MCE-IRL workflow. The reward that comes out is only as interpretable as the supplied feature matrix, transition tensor, and normalization.
The wrapper follows the sklearn-style pattern: build an estimator, call fit,
then read fitted attributes. For multi-action MCE-IRL, provide reward features
explicitly.
import numpy as np
from econirl.datasets import load_rust_bus
from econirl.estimators import MCEIRL
n_states = 90
n_actions = 2
features = np.zeros((n_states, n_actions, 2))
features[:, 0, 0] = -np.arange(n_states) / 100.0
features[:, 1, 1] = -1.0
df = load_rust_bus()
model = MCEIRL(
n_states=n_states,
n_actions=n_actions,
discount=0.99,
feature_matrix=features,
feature_names=["keep_mileage_cost", "replace_cost"],
)
model.fit(df, state="mileage_bin", action="replaced", id="bus_id")
print(model.params_)
print(model.policy_.shape)
For a problem other than the Rust bus, pass the dynamics explicitly. Supply a
transition tensor of shape (n_actions, n_states, n_states) and the observed
next-state column.
model.fit(
df,
state="state",
action="action",
id="id",
next_state="next_state",
transitions=transitions,
)
transitions=None estimates only the two-action Rust-bus keep/replace kernel. A
model with more than two actions requires an explicit tensor. Build one from
observed transitions with estimate_empirical_transitions(panel, n_actions, n_states) from econirl.estimators.
Estimates depend on the supplied reward features, transition specification, and inference settings, so no canonical output is shown here. The fitted estimator exposes reward parameters, standard errors when requested, the recovered reward, the policy, the value function, and feature-matching diagnostics.
Attribute |
Meaning |
|---|---|
|
Estimated reward parameters. |
|
Standard errors for the reward parameters when available. |
|
Structural reward matrix by state and action. |
|
Estimated action probabilities by state. |
|
Estimated value function by state. |
|
Log likelihood of the demonstrations under the recovered policy. |
Simulation Rerun
To reproduce the simulation, run the validation script:
PYTHONPATH=src:. python validation/estimators/mce_irl/run.py --quiet-progress --enforce-gates
The command writes the results file and reports the pass/fail summary for the two simulation cells.
Use econirl.estimation.mce_irl.MCEIRLEstimator when you need direct control
over Panel objects, utility objects, DDCProblem, transition tensors, the
root feature-matching optimizer, or standard-error computation.