Probit model for panel data with heterogeneity and endogenous explanatory variables

In many cases, there is an unobservable heterogeneity in the probit model. For instance, when modelling the consumption choice of a certain brand, consumers’ personal preference is unobserved but needs to be considered in the model. Owing to omitted variable or measurement error, endogeneity issue also could arise. A probit model including both of these two issues can be represented as:

yit = 1[yit*>0]

yit* = xit(1)β + zitδ + ci + uit

zit = xit(1)γ1 + xit(2)γ2 + vit

where ci is the unobservable heterogeneity effect and uit ∣ xi N(0,1), vit|xi N(0,σ2). If vit and uit are independent, this model will degenerate to a probit model with unobservable heterogeneity. In this case, we can just integrate P(yiT,…,yi0xi,ci) against the density of ci conditional on xi, then P(yiT,…,yi0|xi) can be obtained and the objective for the conditional Maximum Likelihood Estimation is

$\sum_{i=1}^N \log [P(y_{iT},\ldots,y_{i0} |x_i)]$

If vit and uit are correlated, under the normality assumption, it can be assumed that vit =ρuit + ϵit, where ϵitiidN(0,σ2ρ2) and ϵi is independent with vi and ui. Then the model can be rewritten as:

yit = 1[xit(1)(β+δγ1)+xit(2)δγ2+ci+ωit>0]

where ωit = (1+ρδ)uit + δϵit, ωit/simN(0,(1+ρδ)2+δ2(σ2ρ2)) and $corr(\omega_{it} , \omega _{(it-s)} ) = \frac {(1 + \rho\delta)^2 corr (u_{it}, u_{(it-s)} )} { ((1 + \rho\delta)^2 + \delta^2 (\sigma^2 - \rho^2 ) )}.$

Based on this, following the same Maximum Likelihood Estimation procedure and the scaled parameter $(\beta + \delta\gamma_1 , \delta\gamma_2) / \sqrt {(1+\rho\delta)^2 + \delta^2 (\sigma^2 - \rho^2 )}$ can be consistently estimated, then the APE can be consistently estimated correspondingly.