Research

Job Market Paper

Econometrics with Pre-Trained Embeddings for Unstructured Data

presentation: Econometric Society Interdisciplinary Frontiers Conference on Economics and AI+ML (Ithaca), Chicago Booth AI and Economics Summer Conference, Midwest Econometrics Group (Cincinnati, scheduled), Canadian Econometrics Study Group (Vancouver, scheduled), Southern Economic Association (Houston, scheduled)

Abstract

Unstructured data, such as images and text, are increasingly used in empirical economics. Since training machine-learning models on unstructured data is costly, economists often use off-the-shelf pre-trained deep learning models developed by computer scientists to extract embeddings, which are then used as covariates in target economic analyses. Despite the popularity of this practice, its theoretical foundations remain limited. There are two main difficulties. First, pre-trained models are typically trained on different datasets and for different tasks, making it unclear when they can be used reliably for the target task. Second, the embedding function is subject to an identification problem, complicating the analysis of its estimation error and the effect of that error on the target task. We provide sufficient conditions to overcome these difficulties. A key condition, which we call transferability, governs the convergence rate we derive. This rate depends on three components: target estimation error with the pre-trained embeddings treated as standard covariates, source estimation error adjusted for the strength of transferability, and approximation error of the target model class using embeddings as inputs. We also extend this analysis for high-dimensional embeddings. To assess transferability, we develop a computationally feasible bootstrap test that does not require re-estimating the embeddings and nuisance functions. Our theory applies to a wide range of double machine learning applications, including partially linear regression with unstructured controls, price elasticity estimation in demand models accounting for product quality measured by images and text, missing-data imputation using unstructured data, and average treatment effect estimation with unstructured confounders. As an empirical application, we estimate the labor supply elasticity on Amazon Mechanical Turk, an online labor market platform, using job-description embeddings as controls.



Working Papers

Design-Based and Network Sampling-Based Uncertainties in Network Experiments

with Kensuke Sakamoto.
Revision Requested at Review of Economics and Statistics

Abstract

Ordinary least squares (OLS) estimators are widely used in network experiments to estimate spillover effects. We study the causal interpretation of, and inference for the OLS estimator under both design-based uncertainty from random treatment assignment and sampling-based uncertainty in network links. We show that correlations among regressors that capture the exposure to neighbors' treatments can induce contamination bias, preventing the OLS from aggregating heterogeneous spillover effects for clear causal interpretation. We derive the OLS estimator's asymptotic distribution and propose a network-robust variance estimator. Simulations and an empirical application demonstrate that contamination bias can be substantial, leading to inflated spillover estimates.


Optimal Testing in a Class of Nonregular Models

with Taisuke Otsu.
Revision Requested at Econometric Theory

Abstract

This paper studies optimal hypothesis testing for nonregular econometric models with parameter-dependent support. We consider both one-sided and two-sided hypothesis testing and develop asymptotically uniformly most powerful tests based on a limit experiment. Our two-sided test becomes asymptotically uniformly most powerful without imposing further restrictions such as unbiasedness, and can be inverted to construct a confidence set for the nonregular parameter. Simulation results illustrate desirable finite sample properties of the proposed tests.


Testing Inequalities Linear in Nuisance Parameters

With Gregory Cox and Xiaoxia Shi.

Abstract

This paper proposes a new test for inequalities that are linear in possibly partially identified nuisance parameters. This type of hypothesis arises in a broad set of problems, including subvector inference for linear unconditional moment (in)equality models, specification testing of such models, and inference for parameters bounded by linear programs. The new test uses a two-step test statistic and a chi-squared critical value with data-dependent degrees of freedom that can be calculated by an elementary formula. Its simple structure and tuning-parameter-free implementation make it attractive for practical use. We establish uniform asymptotic validity of the test, demonstrate its finite-sample size and power in simulations, and illustrate its use in an empirical application that analyzes women's labor supply in response to a welfare policy reform.



Publication

Nonparametric Regression under Cluster Sampling

Journal of Econometrics (2025)
Award: Kanematsu Prize 2023
[arXiv | R code]

Abstract

This paper develops a general asymptotic theory for nonparametric kernel regression in the presence of cluster dependence. We examine nonparametric density estimation, Nadaraya-Watson kernel regression, and local linear estimation. Our theory accommodates growing and heterogeneous cluster sizes. We derive asymptotic conditional bias and variance, establish uniform consistency, and prove asymptotic normality. Our findings reveal that under heterogeneous cluster sizes, the asymptotic variance includes a new term reflecting within-cluster dependence, which is overlooked when cluster sizes are presumed to be bounded. We propose valid approaches for bandwidth selection and inference, introduce estimators of the asymptotic variance, and demonstrate their consistency. In simulations, we verify the effectiveness of the cluster-robust bandwidth selection and show that the derived cluster-robust confidence interval improves the coverage ratio. We illustrate the application of these methods using a policy-targeting dataset in development economics.



Pre-Ph.D. Publication

Doubly Robust-type Estimation of Population Moments and Parameters in Biased Sampling

With Takahiro Hoshino.
Stat (2019)


Translation Work

Imbens, G. W. and D. B. Rubin, “Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction,” (translation into Japanese; responsible for Chapters 15 and 16)

Asakura Publishing (2023)