--- title: "Case study: judge leniency and pretrial detention (Miami-Dade)" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Case study: judge leniency and pretrial detention (Miami-Dade)} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- This vignette runs the package on the empirical application of Frandsen, Leslie and McIntyre (2025): weekend bail hearings in Miami-Dade County, where quasi-randomly assigned bail judges differ in leniency, judge identities instrument for pretrial release, and errors are clustered by courtroom shift. It is a large-scale worked example: 91,421 defendants, 146 judges, and several thousand shift clusters. The companion vignette `vignette("queens-workflow")` introduces the same workflow on a smaller design. This vignette is precomputed: the data file cannot be redistributed inside an R package, so the code below was executed by the maintainer against a local copy and the outputs are stored. Every displayed result comes from the displayed code. ## Data The analysis file `clean_Miami_data.dta` is the cleaned weekend-arraignment file from the published replication materials of Frandsen, Leslie and McIntyre (2025, *Review of Economics and Statistics*); the underlying case records are those of Dobbie, Goldin and Yang (2018). Obtain it from the journal's replication archive and set `data_path` to your local copy: ```r data_path <- "clean_Miami_data.dta" ``` The file is one row per defendant. Following the paper, the sample is restricted to judges with at least 200 cases, and a cluster is a courtroom shift (court by hearing date): ```r library(clusterIV) df <- as.data.frame(haven::read_dta(data_path)) df <- df[df$n_cases >= 200, , drop = FALSE] # judges with >= 200 cases shift <- as.integer(factor(paste(df$court, df$clean_baildate))) y <- as.numeric(df$any_guilty) # outcome x <- as.numeric(df$bail_met) # endogenous: met bail Z <- as.matrix(df[, startsWith(names(df), "jfe_")]) # judge dummies W <- as.matrix(df[, startsWith(names(df), "fe_")]) # time-by-place dummies storage.mode(Z) <- "double"; storage.mode(W) <- "double" c(n = nrow(df), judges = length(unique(df$judgeid)), shifts = length(unique(shift))) #> n judges shifts #> 91421 146 1831 ``` Judge-dummy instrument sets need one hygiene step in any software: after the exogenous controls are partialled out, dummies for judges who never appear in the estimation sample (or are absorbed by the controls) carry no variation, and if every in-sample judge keeps a dummy their sum is collinear with the intercept. The package deliberately reports rank-deficient instrument matrices as errors instead of silently dropping columns, so the drop is made explicit: ```r qC <- qr(cbind(1, W), tol = 1e-10, LAPACK = FALSE) Zres <- qr.resid(qC, Z) keep <- sqrt(colSums(Zres^2)) > 1e-8 * max(sqrt(colSums(Zres^2))) Z <- Z[, keep, drop = FALSE] Zres <- Zres[, keep, drop = FALSE] qZ <- qr(Zres, tol = 1e-10, LAPACK = FALSE) if (qZ$rank < ncol(Zres)) { keep_idx <- sort(qZ$pivot[seq_len(qZ$rank)]) Z <- Z[, keep_idx, drop = FALSE] } ncol(Z) # instrument columns entering the analysis #> [1] 145 ``` ## Estimator comparison `iv_compare()` reports OLS, 2SLS, the improved jackknife IV (IJIVE row, historically labelled JIVE), and CJIVE on the identical design. The time-by-place dummies enter as dense controls, exactly as in the paper: ```r cmp <- iv_compare(y, x, Z, cluster = shift, # courtroom-shift clusters controls = W) # time-by-place fixed effects print(cmp, digits = 3) #> estimator coefficient se statistic p.value conf.low conf.high #> 1 OLS -0.233 0.00658 -35.33 2.25e-273 -0.245 -0.2196 #> 2 2SLS -0.274 0.06567 -4.18 2.96e-05 -0.403 -0.1456 #> 3 JIVE -0.298 0.10671 -2.79 5.22e-03 -0.507 -0.0889 #> 4 CJIVE -0.438 0.20206 -2.17 3.01e-02 -0.834 -0.0422 ``` The published Table 1 values are OLS −0.232 (0.007), 2SLS −0.275 (0.066), IJIVE −0.299 (0.107), and CJIVE −0.444 (0.206). The OLS, 2SLS and IJIVE rows above reproduce the published values to three decimals. The CJIVE estimate is validated differently: run on this deposited sample, the authors' own released Stata implementation (`cjive.ado`) produces the same constructed leave-cluster-out instrument pointwise and the same coefficient to the ado's single-precision limit (about 1e−7), which is the strongest available check that the package computes the estimator the paper defines. ## The CJIVE fit and its diagnostics ```r fit <- cjive(y, x, Z, cluster = shift, controls = W) summary(fit) #> Cluster-jackknife IV (CJIVE) #> Call: cjive.default(y = y, x = x, z = Z, cluster = shift, controls = W) #> #> coefficient = -0.4383 cluster-robust SE = 0.2021 #> z = -2.169 p = 0.03008 95% CI = [-0.8343, -0.04225] #> n = 91421 G = 1831 clusters k = 145 instruments path = dense #> max within-cluster leverage = 0.341 #> #> Coefficients: #> Estimate Std. Error z value Pr(>|z|) #> x -0.4383 0.2021 -2.169 0.0301 * #> --- #> Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1 #> #> Instrument strength: #> F_CJ requires a cjar() fit on the same design; iv_infer() reports the #> estimate, both strength statistics and the robust confidence set in one call. #> F_eff = 1.69 vs critical value 12.41 (Montiel Olea-Pflueger effective F, simplified TSLS, tau = 10%, alpha = 5%, K_eff = 59.61) #> (Confidence-set topology is read from the endpoint matrix printed #> above; no reporting branch infers it from F_eff.) ``` Two diagnostics matter in a design of this size. `maxlev` is the maximum within-cluster leverage of the first stage; values near 1 mean some shift nearly spans the instrument space and the leave-cluster-out fit is at the conditioning frontier. `F_eff` is the clustered Montiel Olea–Pflueger effective first-stage F with its simplified-TSLS critical value. ## Weak-instrument-robust inference Frandsen, Leslie and McIntyre publish no weak-instrument-robust interval for this application. With 100+ instruments, the CJIVE t-interval's validity rests on instrument strength; the cluster-jackknife Anderson–Rubin and score tests of Ligtenberg (2025) stay valid when instruments are weak or many. `iv_infer()` returns the estimate and both tests from one shared preprocessing pass: ```r panel <- iv_infer(y, x, Z, cluster = shift, controls = W) panel #> Cluster IV inference panel (CJIVE + CJAR + CJS) #> Call: iv_infer.default(y = y, x = x, z = Z, cluster = shift, controls = W) #> #> CJIVE/Wald (H0: beta = 0): coefficient = -0.4383 cluster-robust SE = 0.2021 z = -2.169 p = 0.03008 #> 95% Wald interval = [-0.8343, -0.04225] #> CJAR (H0: beta = 0): T = 3.681 one-sided p = 0.0004972 #> 95% confidence set = {} <- empty: no beta is accepted at this level #> CJS (H0: beta = 0): LM = 4.759 p = 0.02914 #> 95% confidence set = [-0.9866, -0.05018] #> F_CJ = 4.237 critical value = 1.709 #> F_CJS^2 = 13.9 critical value = 3.841 #> effective F (Montiel Olea-Pflueger) = 1.69 critical value (tau = 10%, alpha = 5%) = 12.41 K_eff = 59.61 #> n = 91421 G = 1831 clusters k = 145 instruments #> max within-cluster leverage = 0.341 ``` This panel is a lesson in reading weak-instrument-robust output rather than a clean confirmation. The effective F (1.69) is far below its critical value: with 145 instrument columns the first stage is weak by the Montiel Olea–Pflueger standard, so the Wald interval's nominal coverage is not guaranteed — precisely the regime the robust tests are for. The two robust results then differ in kind. The CJS test gives a bounded 95% set that excludes zero and comfortably contains the CJIVE estimate. The CJAR set is *empty*: no coefficient value is accepted at the 5% level. An empty AR set is a valid outcome, not a numerical failure — the AR statistic aggregates all 145 moment conditions, so it also has power against violations of the exclusion restrictions themselves, and with this much overidentification it can reject every candidate coefficient. It should be read as evidence of tension in the full instrument set rather than as an interval estimate (see the FAQ in the README). The sets are computed by analytic polynomial inversion, not a parameter grid, so empty, disjoint, and unbounded outcomes are exact statements. The p-value curves make all of this visible: ```r plot(panel) ``` ![plot of chunk pcurve](miami-pcurve-1.png) ## Notes on scale At this size (n = 91,421, k > 100 instruments, dense controls) the whole panel above runs in minutes on a laptop. The package keeps `Imports:` limited to `stats`; the loading step above uses `haven` only to read the Stata file, which is not a package dependency. For designs with many more fixed-effect levels, pass them as `fixed_effects =` factors instead of dense dummy columns: the matrix-free absorption never forms the dummy matrix. Note the caveat in `?cjive`: with many absorbed controls relative to the sample, the plug-in CJIVE standard error can over-reject (Kolesár, Min, Wang and Zhang 2026); the time-by-place controls here are few relative to n. ## References Dobbie, W., Goldin, J. and Yang, C. S. (2018). The effects of pretrial detention on conviction, future crime, and employment. *American Economic Review*, 108(2), 201–240. The underlying Miami-Dade case records. Frandsen, B., Leslie, E. and McIntyre, S. (2025). Cluster jackknife instrumental variables estimation. *Review of Economics and Statistics*. doi:10.1162/rest.a.263. The estimator, the application, and the replication materials used here. Ligtenberg, J. W. (2025). Inference in clustered IV models with many and weak instruments. arXiv:2306.08559v3. The CJAR and CJS tests reported by `iv_infer()`.