References
This page lists the papers that the public estimator pages draw from. The estimator pages link here so the methodological source is visible before the usage examples.
Use this page to trace a method back to its source paper. It is a bibliography, not a claim that every source paper has a full numerical replication in the package.
Theory Surveys and Lecture Notes
Kang, E. H. (2026). “A Lecture Note on Offline RL and IRL: Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models.” arXiv preprint arXiv:2605.30843.
Rawat, P., and Rust, J. (2026). “Dynamic Discrete Choice and Inverse Reinforcement Learning: Inferring Preferences and Beliefs from Human Behavior.” Manuscript.
Dynamic Discrete Choice
Rust, J. (1987). “Optimal Replacement of GMC Bus Engines: An Empirical Model of Harold Zurcher.” Econometrica, 55(5), 999-1033.
Rust, J. (1994). “Structural Estimation of Markov Decision Processes.” In Handbook of Econometrics, Vol. 4, ch. 51, 3081-3143. North-Holland.
Hotz, V. J., and Miller, R. A. (1993). “Conditional Choice Probabilities and the Estimation of Dynamic Models.” Review of Economic Studies, 60(3), 497-529.
Aguirregabiria, V., and Mira, P. (2002). “Swapping the Nested Fixed Point Algorithm: A Class of Estimators for Discrete Markov Decision Models.” Econometrica, 70(4), 1519-1543.
Magnac, T., and Thesmar, D. (2002). “Identifying Dynamic Discrete Decision Processes.” Econometrica, 70(2), 801-816.
Hotz, V. J., Miller, R. A., Sanders, S., and Smith, J. (1994). “A Simulation Estimator for Dynamic Models of Discrete Choice.” Review of Economic Studies, 61(2), 265-289.
Kasahara, H., and Shimotsu, K. (2008). “Pseudo-Likelihood Estimation and Bootstrap-Based Inference for Structural Discrete Markov Decision Models.” Journal of Econometrics, 146(1), 92-106.
Kasahara, H., and Shimotsu, K. (2009). “Nonparametric Identification of Finite Mixture Models of Dynamic Discrete Choices.” Econometrica, 77(1), 135-175.
Aguirregabiria, V., and Mira, P. (2010). “Dynamic Discrete Choice Structural Models: A Survey.” Journal of Econometrics, 156(1), 38-67.
Arcidiacono, P., and Miller, R. A. (2011). “Conditional Choice Probability Estimation of Dynamic Discrete Choice Models With Unobserved Heterogeneity.” Econometrica, 79(6), 1823-1867.
Cameron, A. C., and Miller, D. L. (2015). “A Practitioner’s Guide to Cluster-Robust Inference.” Journal of Human Resources, 50(2), 317-372.
Su, C.-L., and Judd, K. L. (2012). “Constrained Optimization Approaches to Estimation of Structural Models.” Econometrica, 80(5), 2213-2230.
Iskhakov, F., Lee, J., Rust, J., Schjerning, B., and Seo, K. (2016). “Comment on ‘Constrained Optimization Approaches to Estimation of Structural Models’.” Econometrica, 84(1), 365-370.
Shapiro, A., and Xu, H. (2005). “Stochastic Mathematical Programs with Equilibrium Constraints, Modeling and Sample Average Approximation.” Published in Optimization, 57(3), 395-418 (2008).
Koiso, S., and Otani, S. (2024). “An MPEC Estimator for the Sequential Search Model.” arXiv:2409.04378.
Approximate Structural Estimation
Luo, Y., and Sang, P. (2024). “Efficient Estimation of Structural Models via Sieves.” Working paper, University of Toronto.
Nguyen, H. (2025). “Neural Networks for Efficient Estimation of High-Dimensional Dynamic Discrete Choice Models.” Working paper, Georgetown University.
Adusumilli, K., and Eckardt, D. (2025). “Temporal-Difference Estimation of Dynamic Discrete Choice Models.” Working paper.
Maximum-Entropy and Adversarial IRL
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K. (2008). “Maximum Entropy Inverse Reinforcement Learning.” Proceedings of the 23rd AAAI Conference on Artificial Intelligence, 1433-1438.
Ziebart, B. D. (2010). Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy. PhD thesis, Carnegie Mellon University.
Wulfmeier, M., Ondruska, P., and Posner, I. (2015). “Maximum Entropy Deep Inverse Reinforcement Learning.” NIPS Deep Reinforcement Learning Workshop.
Fu, J., Luo, K., and Levine, S. (2018). “Learning Robust Rewards with Adversarial Inverse Reinforcement Learning.” International Conference on Learning Representations.
Barnes, M., Abueg, M., Lange, O. F., Deeds, M., Trader, J., Molitor, D., Wulfmeier, M., and O’Banion, S. (2024). “Massively Scalable Inverse Reinforcement Learning in Google Maps.” International Conference on Learning Representations. arXiv:2305.11290.
Lee, P. S., Sudhir, K., and Wang, T. (2026). “Consumer Engagement with Sequential Content: A Content-Aware Dynamic Choice Model.” SSRN working paper, abstract 6331041.
Q-Function and Divergence IRL
Ni, T., Sikchi, H., Wang, Y., Gupta, T., Lee, L., and Eysenbach, B. (2020). “f-IRL: Inverse Reinforcement Learning via State Marginal Matching.” Proceedings of the 4th Conference on Robot Learning.
Garg, D., Chakraborty, S., Cundy, C., Song, J., and Ermon, S. (2021). “IQ-Learn: Inverse Soft-Q Learning for Imitation.” Advances in Neural Information Processing Systems.
Kang, E. H., Yoganarasimhan, H., and Jain, L. (2025). “An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model.” arXiv preprint arXiv:2502.14131.