NYU

New York University · Stern School of Business

MS in Business Analytics & AI · Capstone

April 2026

California's Grid Has a Hidden Problem

Electric vehicles and data centers are reshaping the grid. The question nobody is asking: who pays when the grid can't predict itself?

Arguelles · Graff · Harrison · Jessup · Lindley

↓ Scroll to explore · Charts are interactive · Methodology at the bottom

Key Findings

Sign Reversal. EV adoption causally increases electricity price volatility — the opposite of what naïve regressions suggest.

Income Gradient. Lower-income counties bear a disproportionate share of that burden. The gap widens as adoption scales.

Data Centers: Null. Stable baseload demand does not destabilize grids. The concern is cost allocation, not volatility.

The Convergence

California promised an equitable energy transition. The data suggests otherwise.

Three forces are converging: AI data centers demanding continuous power, a state mandate electrifying millions of vehicles, and climate extremes straining generation capacity. The grid is bending. But the critical question isn't whether it can handle the volume — it's whether the unpredictability that follows falls equally on everyone. It doesn't.

0

EVs on CA Roads

18.7 GW

DC Capacity Requested

36 mo

Price Data Analyzed

Background

How California Prices Electricity

California's wholesale market uses locational marginal pricing — the price at a specific grid node reflects the real-time cost of delivering one more megawatt there. When transmission lines congest, prices at different nodes diverge. A node behind a bottleneck becomes cut off from cheaper generation. These differentials reflect decades of infrastructure investment decisions.

Key mechanism: Wholesale volatility doesn't hit bills directly. It flows through utilities' procurement costs, hedging, and infrastructure planning — passed to ratepayers over annual or multi-year cycles. The burden is real, but delayed and opaque.

The Trend

The EV Boom

California's EV fleet has more than doubled since April 2022, but charging infrastructure hasn't kept pace — reaching only 1.65× while adoption hit 2.3×. The stress isn't about total energy; it's about timing. Home charging clusters in the evening, precisely when solar ramps down.

EV Adoption vs. Tesla Store Expansion (Indexed, Base = 100)

Data Landscape

Exploring California's Grid Data

Multiple data sources mapped across 52 counties. Toggle the base layer to change county coloring. Check point layers to overlay infrastructure.

County Color

⏱ VARIES OVER TIME

▦ STATIC (COUNTY-LEVEL)

Point Layers

Time Period

All Months (Avg)

Price Volatility (MAD) by County

0.033
0.060

The Key Insight

Volatility ≠ Higher Prices

Average prices peaked in late 2022 (heat waves, fuel spikes) then declined. That story is over. But volatility — price unpredictability — is episodic, appearing in clusters throughout. The forces that drove the 2022 peak are not the forces creating recurring instability.

Monthly Avg. Wholesale LMP ($/MWh)

Monthly Volatility (MAD)

The forces that drove prices to their 2022 peak are not the same forces creating recurring instability. It's the unpredictability that affects utilities' planning — and flows through to household costs.

The Causal Finding

What Naïve Analysis Gets Wrong

OLS says more EVs = less volatility. But that's confounded — wealthier counties have both more EVs and better grids. Using Tesla store openings as an instrument (methodology ↓) (corporate decisions exogenous to local electricity prices), the sign flips:

OLS (Naïve)

−0.000347OLS coefficient — biased by confounders (wealth, grid infrastructure, urban density) that correlate with both EV adoption and lower volatility.

p = 0.016

2SLS (IV)

+0.0005932SLS coefficient from IV regression using Tesla store IDW exposure as instrument. County + month FE, cluster-robust SE. N=1,976.

p = 0.001

Each additional EV per 1,000 residents causally increases monthly price volatility. For +50 EVs/1,000, that's 0.7 standard deviations — roughly half the interquartile range (~48%) of observed volatility. A substantial shift in grid unpredictability.

The Equity Question

Who Bears the Burden?

Each $10,000 increase in county median income reduces the EV-induced volatility effect by ~$0.0003/MWh per EV/1,000 (p < 0.01). Lower-income counties absorb meaningfully more instability for the same adoption increase.

Explore: County Median Household Income

Income: $70,000

+0.029

$/MWh volatility increase per 50 EVs/1,000

Model specification (interaction 2SLS)

mad_ratec,t  =  α  +  β1·EVsc,t  +  β2·(EVsc,t × income10kc)  +  γ′Xc,t  +  μc  +  λt  +  εc,t
where EVsc,t and (EVs × income10k)c,t are both instrumented by Zc,t = [Tesla IDW exposure, Tesla IDW × income10k]

β1 = +0.003554 (main EV effect, p < 0.01)  ·  β2 = −0.000303 (income interaction, p < 0.01). Income is scaled in $10,000 units. The slider above evaluates the marginal effect (β1 + β2·income10k) × 50 EVs. Controls X: max temperature, cumulative data center power, generator capacity, lagged solar capacity. μc = county FE; λt = month FE. SEs clustered by county.

County Clusters: Income vs. Volatility

Point shade reflects county income: low income (higher burden) high income

Key: DAC designation (CalEnviroScreen) does not capture where this burden concentrates. Income is the more precise policy lever.

The Other Driver

What About Data Centers?

No statistically significant causal effect on volatility or spikes. A 1% increase in DC power associates with a $0.088/MWh decrease in average prices — predictable baseload demand smoothing operations. The equity concern isn't volatility; it's cost allocation — see the full results table ↓.

IV Coefficient Estimates: Data Center Power by Outcome

Whiskers show 95% confidence intervals. Values of zero (dashed line) indicate no effect.

Case Study

Policy Timing Matters

Two California policies shifted inside our study window from opposite sides of the grid. Exploiting their timing gives us a second identification strategy and a sharper picture of how short-run shocks propagate to volatility.

Case 1 — CVRP Expiration (November 2023)

The Clean Vehicle Rebate Project paid income-scaled rebates for EV purchases. California announced its closure in August 2023 and the program terminated in November. That three-month window produced a "run on the bank" — a burst of subsidy-anchored purchases followed by a sustained decline in adoption as the price signal disappeared.

We use the closure timing interacted with each county's pre-period CVRP reliance weight (share of applications that received low-income boosted rebates) as a shift-share instrument. The second-stage coefficient on EV adoption is −0.0007 (p = 0.11) — marginally negative but below conventional significance. We report this transparently: the CVRP-identified effect does not clear the 5% bar, and we interpret it as identifying a different complier population than the Tesla IV (rebate-sensitive rather than access-driven).

Monthly EV Adoption (New EVs per 1,000, All Counties)

Case 2 — NEM 2.0 Interconnection Backlog (January 2024)

California's Net Energy Metering framework transitioned from NEM 2.0 (retail-rate compensation) to NEM 3.0 (avoided-cost compensation) in April 2023. A large backlog of NEM 2.0 applications — systems grandfathered into the legacy payment structure — physically connected to the grid in the months that followed, peaking in early 2024. These were mostly battery-less solar-only installations that rushed to secure favorable rates before the rule changed.

We include a post-January-2024 structural break control (solar_x_nem3) in the final 2SLS specification to isolate this solar shock from the concurrent EV dynamics. The coefficient is +0.0177 (p < 0.06): the rapid wave of solar-only interconnection temporarily raised volatility, consistent with the duck-curve intuition — new capacity deepens midday troughs and steepens evening ramps without storage to buffer it.

A supply-side mirror of the EV problem. EVs add load when there isn't enough generation capacity; NEM 2.0 solar added generation when there isn't enough storage. Both are physical grid-timing problems rather than aggregate-capacity problems, and both point at the same policy priority: flexible infrastructure that buffers short-run ramps.

NEM 2.0 Rush: Solar Capacity (Cumulative MW per 1,000)

+0.0177**

The wave of battery-less NEM 2.0 systems connecting to the grid increased volatility

Summary

Causal Results Across Workstreams

AnalysisOutcomeOLS2SLS (IV)Significance
EV → VolatilityMAD rate−0.000347 **+0.000593 ***Sign reversal ✓
EV × IncomeMAD rate−0.000080 *−0.000303 ***Consistent ✓
EV × DACMAD ratePositive n.s.Positive n.s.Null result
EV → Price levelAvg LMPSmall, unstableSmall, unstableNo robust effect
DC → VolatilityMAD rateNot significantNot significantRobust null ✓
DC → Avg priceAvg LMPNot significant−8.8¢/MWh ***Significant ✓
CVRP (EV alt. IV)MAD rate−0.0007 (p=0.11)Not Significant
NEM 2.0 backlogMAD rate+0.0177 **Policy timing

*** p<0.01 | ** p<0.05 | * p<0.10 | n.s. = not significant

What This Means

Implications for Policy

Transportation electrification is a state mandate. But its grid impacts are not equity-neutral.

1. Target infrastructure investment by income, not just DAC status. CalEnviroScreen alone doesn't capture where volatility burden concentrates.

2. Prioritize distribution grid upgrades in lower-income counties as EV adoption scales.

3. Data center policy should focus on cost allocation, not volatility. Stable loads don't destabilize grids.

Under the Hood

Methodology & Data

All public data. BigQuery + dbt + Python. 1,976 county-month observations, 52 counties, Apr 2022 – May 2025.

County × month panel. ~36 months day-ahead LMP from CAISO OASIS (~9GB). CalEnviroScreen 4.0. CEC EV registrations. NOAA weather. EIA-860. 197 data centers (web-scraped). ACS census.
Median Absolute Deviation (MAD): median distance from median price per county-month. Limits extreme outlier influence.
IDW Tesla store exposure (log). 66 total California Tesla stores in the IDW kernel; the 29 new openings during the Apr 2022 – May 2025 window generate the identifying temporal variation. County + month FE. Cluster-robust SE. Permutation p = 0.014. Lead falsification: clean. Lag: signal through 2 months.
2014 FCC fiber: provider count + speeds. First-stage F: 60–77. Exclusion validated (direct regression insignificant).
500-permutation test. Lead falsification (3/6/12mo). Lag robustness (0–6). Tract-level. PSM for DCs. CVRP + NEM 2.0 discontinuity.

Full regression ladder · first-stage diagnostics · interaction spec · robustness grid

Looking Ahead

Limitations & Next Steps

Limitations & Caveats

Data center power is proxied. Operational load is estimated from web-scraped nameplate capacity using the CEC YoY ramp and JLARC seasonality profile, not observed consumption. DC price results are directional.

Two-month causal window. The Tesla IV coefficient retains ≈60% of its magnitude through a two-month lag before dropping sharply. EVs are a stock variable, so a purely cumulative-load mechanism would be expected to persist longer. This ambiguity between a genuine short-lived effect and a transient correlated shock is reported transparently.

Tract-level re-estimation does not replicate. At census-tract granularity the coefficient drops to near zero. We attribute this to a scale mismatch between LMP node geography and tracts, not to evidence against the county-level result, but it is a constraint on generalization.

CVRP IV is marginally significant. The policy-discontinuity IV produces a coefficient of −0.0007 (p=0.11) after full controls, below conventional thresholds. We interpret the Tesla and CVRP instruments as identifying different complier populations rather than as contradictory evidence — see the statistical appendix.

Next Steps

LLM-powered interactive dashboard. AI-driven application allowing stakeholders to query grid equity data in natural language.

Utility-level consumption data. Partner with CAISO for metered load profiles to validate data center estimates.

Sub-monthly pricing + charging. Hourly LMP data matched with EV charging event data to isolate the temporal mechanism.

For the Defense

Anticipated Questions

Questions we expect from reviewers, with concise answers. Click any question to expand.

Standard-deviation-based volatility measures are dominated by extreme outliers. California's wholesale prices exhibit heavy tails with occasional extreme spikes during grid emergencies — in our sample, monthly mean LMP reached ≈$270/MWh during the late-2022 heat wave. Including those spikes in a variance calculation would make volatility look like a story about rare crisis moments rather than the chronic planning-uncertainty problem utilities actually face.

MAD captures the typical dispersion households and utilities experience while limiting outlier influence. It also aligns more closely with the kind of volatility that translates into utility hedging costs and procurement reserves — the channel through which wholesale dispersion eventually reaches retail rates.

Three reasons. First, coverage: LMP nodes don't map cleanly to utility territories, and many tracts contain no pricing nodes at all. County aggregation achieves near-complete statewide coverage. Second, consistency: every other variable in our panel (EV registrations, CalEnviroScreen tracts, census demographics, weather) is available or clean to aggregate at the county level. Third, policy relevance: California's environmental justice framework, disadvantaged community designations, and many infrastructure funding programs operate at the county level.

We tested tract-level re-estimation and report the null result transparently (Section H of the statistical appendix) — we attribute it to scale mismatch, not to evidence against the county result.

Both instruments are reported. The Tesla IV is the preferred main specification because (a) its first stage is strong and clean, (b) it passes the permutation test at p = 0.014, and (c) store-opening timing is plausibly exogenous on multi-year corporate real-estate horizons.

The CVRP shift-share IV is reported as a complementary specification. The two instruments identify different compliers — the Tesla instrument captures access-driven adoption (higher-income, urban), the CVRP instrument captures rebate-sensitive adoption (lower-income, higher-reliance). That the two produce different second-stage signs is informative about heterogeneity, not contradictory evidence.

This is the single most-asked question about the paper and is fully addressed in Section B of the statistical appendix. Briefly: without fixed effects, cross-county variation dominates and the sign is positive — counties with more EVs have higher volatility on average. Adding county FE strips the cross-sectional signal and reveals the within-county pattern: months where a given county has higher-than-usual EV adoption see lower volatility. Adding month FE on its own does not flip the sign because it only absorbs statewide time shocks.

The TWFE-OLS gives −0.000347 (p=0.015), but this is still biased — it confounds EV adoption with time-varying county investments in grid infrastructure (battery rollouts, TOU rates, distribution upgrades) that are negatively correlated with volatility. Once the Tesla IV removes that confounded variation, the coefficient returns to positive: +0.000593 (p=0.001). The sign flips are each driven by a specific source of bias being removed.

Intentionally. The dominant source of uncertainty in the forecast is structural, not statistical — it depends on whether the causal coefficients β₁ and β₂ stay approximately constant over the next decade, which they may not. Adding parametric 95% CIs from the 2SLS would imply a false precision that our honest uncertainty doesn't support.

The parametric CI on the $40k / 2030 point would be roughly ±40% of the central estimate. If the real mechanism attenuates as utilities deploy storage, V2G, or dynamic rates, the actual 2030 burden could be substantially lower. We frame the forecast as a "no-intervention projection" that motivates action, not as a point forecast.

It is a fair concern and we report it transparently. EVs are a stock variable; a mechanism operating purely through cumulative charging load would be expected to persist longer than two months. The permutation test (p = 0.014) confirms the effect at contemporaneous timing is statistically real, but cannot distinguish between a genuine short-lived causal effect and a transient correlated shock.

Two non-exclusive interpretations: (a) utilities and distribution networks respond to new EV load within a quarter (demand response, rate adjustments, feeder upgrades), transiently reducing the marginal volatility contribution of any new cohort; (b) the signal is real at contemporaneous timing but attenuates as charging habits stabilize. Distinguishing these requires sub-monthly charging event data, which is in the paper's "Next Steps."

DAC designation at the tract level, when aggregated to a county share, is a noisy and coarse measure of economic vulnerability compared to continuous county median household income. CalEnviroScreen weights environmental burdens, pollution exposure, health outcomes, and socioeconomic factors — income is only one of many inputs. Two counties can have similar DAC shares but very different income distributions.

Our result — significant income gradient, null DAC gradient — is consistent with income being the precise policy lever for EV-induced volatility burden, not the broader cumulative-burden framework DAC captures. This is actually a policy-relevant finding in its own right: DAC designation is useful for many things but does not predict volatility exposure here.

We treat it as directional, not precise. The collapsed IV coefficient (−0.174, p=0.001, translating to −$0.088/MWh per 1% DC power) is statistically significant and directionally consistent with PSM. The PSM robustness check produces −0.087 on arcsinh-transformed prices, in the same direction.

The caveat: our DC operational load is proxied from imputed nameplate capacity using the CEC YoY ramp methodology, not observed consumption. Variation is primarily cross-sectional across a limited set of host counties, constraining precision. Our strongest DC claim is the robust null on volatility; the negative price effect is reported as suggestive.

Sources

References

California Air Resources Board. (2022). 2022 Scoping Plan for Achieving Carbon Neutrality.

California Public Utilities Commission. (2025). How will data center growth impact California ratepayers?

Elmallah, S., Brockway, A. M., & Callaway, D. (2022). Can distribution grid infrastructure accommodate residential electrification and EV adoption in Northern California? Environmental Research: Infrastructure and Sustainability, 2, 045005.

Jenn, A., & Highwayman, J. (2021). Distribution grid impacts of electric vehicles: A California case study. iScience, 25(1), 103686.

Li, Y., & Jenn, A. (2024). Impact of EV charging demand on power distribution grid congestion. Proceedings of the National Academy of Sciences, 121(18).

Padilla, S. (2025). Senate Bill 57: Ratepayer and Technological Innovation Protection Act. California Legislature.

Ren, S. (2025). An assessment of California data centers' environmental and public health impacts. Next 10.

Schweppe, F. C., Caramanis, M. C., Tabors, R. D., & Bohn, R. E. (1988). Spot pricing of electricity. Kluwer Academic Publishers.

Shehabi, A., et al. (2024). 2024 United States data center energy usage report. Lawrence Berkeley National Laboratory.

Talkington, S., West, A., & Haider, R. (2024). Locational marginal burden: Quantifying the equity of optimal power flow solutions. E-Energy '24.

U.S. Department of Energy, EERE. (n.d.). Energy accessibility.

How to Cite

Arguelles, S., Graff, V., Harrison, E., Jessup, R., & Lindley, S. (2026). Electricity Price Volatility and Grid Pressure in California, 2022–2025: A Data Infrastructure for Analyzing Household Impacts. NYU Stern School of Business, MS in Business Analytics & AI Capstone Project.

BibTeX available on request · Data pipeline: BigQuery + dbt + Python · All sources publicly available

Dashboard-only · Not in the Capstone Report

Additional: Forward Projections & Source Code

The content in this section sits outside the scope of the formal capstone report. It is included here because the dashboard serves a different purpose than the report, and a few things make sense in one venue but not the other.

Why these live in the dashboard but not the report. The report is the empirical deliverable — it defends claims about 2022–2025, the window the data actually covers, and every coefficient reported there is bounded by what the panel can identify. Extrapolating the two interaction coefficients β1 and β2 forward to 2035 requires assumptions the data cannot test: that utilities do not deploy managed charging at scale, that V2G and dynamic rates remain limited, that the EV fleet's charging behavior stays similar to today's, and that the causal mechanism itself stays stable over a decade. Those are assumptions a stakeholder-facing decision aid can legitimately make (“if nothing changes, here is where we land”), but that a peer-reviewable empirical paper should not. The forecast therefore belongs in a policy-monitoring tool, not in the paper's results section. The same logic applies to the raw source code: it's essential for reproducibility and anyone who wants to audit the pipeline, but it would consume half the report page count without adding to the argument. Keeping both here keeps the report focused and the dashboard useful.

Forward projection · 2025 → 2035

California targets 5 million EVs by 2030 and 100% ZEV sales by 2035. Under the no-intervention assumption — β1 and β2 held constant at their Apr 2022 – May 2025 values — the slider below projects how the volatility burden would shift across income levels over the next decade. Treat as a directional scenario, not a point forecast.

Projected EV Adoption & Volatility Impact by County Income

Projection Year

2025

Est. EVs Statewide

1.5M

Volatility Burden ($40k County)

0.03%

of electricity spend

Volatility Burden ($120k County)

~0%

of electricity spend

Projection formula

Burden%y,i  =  max(0,   [ (β1 + β2·income10ki) × EVsPerKy × 7 MWh ]  /  ElecSpendi  ) × 100
where EVsPerKy = EVsy / CA population × 1,000  ·  ElecSpendi = annual $ household electricity spend at income tier i

Uses the same β1 = +0.003554 and β2 = −0.000303 from the interaction 2SLS. EV trajectory anchors: 1.5M (2025), 5.0M (2030, state target), 12.2M (2035, 100% ZEV sales). Household consumption fixed at 7 MWh/yr (CA residential average). Spend benchmarks by income tier: $2,200 ($40k county), $2,700 ($70k), $3,800 ($120k) — from the energy-burden literature. Floored at 0. No parametric CIs are shaded — the dominant uncertainty here is structural (coefficient stability over a decade), not sampling variance, and narrow parametric bands would imply a precision this extrapolation does not have.

Source code · full Python notebooks

Every cell from the seven production notebooks is included below, grouped by workstream. Collapsed by default — click to expand. These are the complete notebooks as run; nothing has been edited for brevity aside from trimming one very long scraped-data literal for readability (flagged inline). Long; use Ctrl/Cmd+F once a notebook is expanded.

EV Ecosystem · causal identification

Two complementary identification strategies for EV adoption: Tesla-store IDW exposure as the preferred IV, CVRP closure as a shift-share falsification check.

Tesla IV – ev_iv_tesla_2sls.ipynb46 cells
Cell [1]
import pandas as _hex_pandas
import datetime as _hex_datetime
import json as _hex_json
Cell [2]
pip install linearmodels
Cell [4]
from google.colab import drive
drive.mount('/content/drive')
Cell [5]
from pathlib import Path

## YOU WILL NEED TO UPDATE MYDRIVE TO SHAREDWITHME IF YOU ARE TRYING TO RUN THIS AND AREN'T RACHEL
DATA_DIR = Path(
    "/content/drive/MyDrive/Capstone: Watts the Problem/Datasets - Colab"
)

DATA_DIR.exists()
Cell [6]
import pandas as pd

panel_df = pd.read_csv(DATA_DIR / "panel_df.csv")
county_shape_df = pd.read_csv(DATA_DIR / "county_shapefile_df.csv")
tesla_loc_df = pd.read_csv(DATA_DIR / "tesla_loc_df.csv")
census_df = pd.read_csv(DATA_DIR / "census_df.csv")
solar_df = pd.read_csv(DATA_DIR / "solar_control_variable.csv")
tract_shape_df = pd.read_csv(DATA_DIR / "tract_shapefile_df.csv")
tract_panel_df = pd.read_csv(DATA_DIR / "tract_panel_df.csv")
tract_census_df = pd.read_csv(DATA_DIR / "tract_census_df.csv")
Cell [10]
import numpy as np
import pandas as pd

# -------------------------
# Helpers
# -------------------------
def haversine_km(lat1, lon1, lat2, lon2):
    R = 6371.0
    lat1, lon1, lat2, lon2 = map(np.radians, [lat1, lon1, lat2, lon2])
    dlat = lat2 - lat1
    dlon = lon2 - lon1
    a = np.sin(dlat / 2) ** 2 + np.cos(lat1) * np.cos(lat2) * np.sin(dlon / 2) ** 2
    return 2 * R * np.arcsin(np.sqrt(a))

def to_month_start(x):
    return pd.to_datetime(x, errors="coerce").dt.to_period("M").dt.to_timestamp()

# -------------------------
# Prepare inputs
# -------------------------
df = panel_df.copy()
df["month"] = to_month_start(df["month"])

county_points_df = (
    county_shape_df[["county_geoid", "internal_point_latitude", "internal_point_longitude"]]
    .drop_duplicates("county_geoid")
    .rename(columns={"internal_point_latitude": "county_lat", "internal_point_longitude": "county_lon"})
    .copy()
)
county_points_df["county_lat"] = pd.to_numeric(county_points_df["county_lat"], errors="coerce")
county_points_df["county_lon"] = pd.to_numeric(county_points_df["county_lon"], errors="coerce")
county_points_df = county_points_df.dropna(subset=["county_lat", "county_lon"])

tesla_stores_df = tesla_loc_df[["latitude", "longitude", "open_month"]].copy()
tesla_stores_df["latitude"] = pd.to_numeric(tesla_stores_df["latitude"], errors="coerce")
tesla_stores_df["longitude"] = pd.to_numeric(tesla_stores_df["longitude"], errors="coerce")
tesla_stores_df["open_month"] = to_month_start(tesla_stores_df["open_month"])
tesla_stores_df = tesla_stores_df.dropna(subset=["latitude", "longitude", "open_month"])

# -------------------------
# Compute tesla_store_exposure_log across all months
# This measures inverse-distance-weighted exposure to open Tesla stores
# -------------------------
months = pd.DatetimeIndex(sorted(df["month"].dropna().unique()))
counties = county_points_df[["county_geoid", "county_lat", "county_lon"]].copy()

eps_km = 1.0  # minimum distance cap to avoid division by zero
iv_rows = []

for m in months:
    open_stores_m = tesla_stores_df.loc[tesla_stores_df["open_month"] <= m]
    if open_stores_m.empty:
        continue

    tmp_m = counties.assign(_k=1).merge(open_stores_m.assign(_k=1), on="_k").drop(columns="_k")
    tmp_m["dist_km"] = haversine_km(tmp_m["county_lat"], tmp_m["county_lon"], tmp_m["latitude"], tmp_m["longitude"])
    tmp_m["dist_km_capped"] = np.maximum(tmp_m["dist_km"], eps_km)

    iv_m = tmp_m.groupby("county_geoid", as_index=False).agg(
        idw_exposure=("dist_km_capped", lambda s: float(np.sum(1.0 / s))),
    )
    iv_m["month"] = m
    iv_rows.append(iv_m)

iv_table = pd.concat(iv_rows, ignore_index=True)
iv_table["tesla_store_exposure"] = iv_table["idw_exposure"]
iv_table["iv_tesla_store_exposure_log"] = np.log1p(iv_table["idw_exposure"])
iv_table = iv_table[["county_geoid", "month", "tesla_store_exposure", "iv_tesla_store_exposure_log"]]

# Merge baseline IV onto the panel
df_iv = df.merge(iv_table, on=["county_geoid", "month"], how="left", validate="m:1")

df_iv
Cell [12]
import pandas as pd
import numpy as np
import geopandas as gpd
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
from matplotlib.lines import Line2D

# Get the maximum month in the sample
max_month = pd.to_datetime(df_iv['month'].max())
print(f"Mapping Instrument for: {max_month.strftime('%B %Y')}")

# Filter data for max month
iv_max_month = df_iv[df_iv['month'] == df_iv['month'].max()].copy()

# Get stores open by max month
tesla_open_dates = pd.to_datetime(tesla_loc_df['open_month']).dt.tz_localize(None)
stores_open = tesla_loc_df[tesla_open_dates <= max_month].copy()
print(f"Tesla stores open: {len(stores_open)}")

# Create GeoDataFrame with correct geometry column
county_gdf = gpd.GeoDataFrame(
    county_shape_df,
    geometry=gpd.GeoSeries.from_wkt(county_shape_df['county_geometry']),
    crs='EPSG:4326'
)

# Merge with IV values
county_map_data = county_gdf.merge(
    iv_max_month[['county_geoid', 'iv_tesla_store_exposure_log', 'tesla_store_exposure']],
    on='county_geoid',
    how='left'
)

# Get county centroids
county_centroids = county_points_df.copy()

# Create figure
fig, ax = plt.subplots(1, 1, figsize=(16, 14))

# Plot county boundaries colored by log exposure
county_map_data.plot(
    column='iv_tesla_store_exposure_log',
    ax=ax,
    cmap='YlOrRd',
    edgecolor='black',
    linewidth=0.3,
    legend=True,
    legend_kwds={'label': 'Tesla Store Exposure (log)\nlog(1 + Σ(1/dist))', 'shrink': 0.6},
    missing_kwds={'color': 'lightgrey', 'edgecolor': 'black', 'linewidth': 0.3}
)

# Plot county centroids
ax.scatter(
    county_centroids['county_lon'],
    county_centroids['county_lat'],
    c='blue',
    s=20,
    alpha=0.7,
    marker='o',
    edgecolors='darkblue',
    linewidths=0.5,
    label='County Centroids',
    zorder=3
)

# Plot Tesla stores
ax.scatter(
    stores_open['longitude'],
    stores_open['latitude'],
    c='red',
    s=180,
    alpha=0.95,
    marker='*',
    edgecolors='darkred',
    linewidths=1.5,
    label=f'Tesla Stores (n={len(stores_open)})',
    zorder=4
)

# Add title
ax.set_title(
    f'Instrument Construction: Inverse-Distance Weighted Tesla Store Exposure\n{max_month.strftime("%B %Y")}',
    fontsize=16, fontweight='bold', pad=20
)
ax.set_xlabel('Longitude', fontsize=12, fontweight='bold')
ax.set_ylabel('Latitude', fontsize=12, fontweight='bold')

# Customize legend
ax.legend(loc='lower left', fontsize=11, framealpha=0.95, edgecolor='black')

# Clean axes
ax.set_xticks([])
ax.set_yticks([])
for spine in ax.spines.values():
    spine.set_visible(False)

plt.tight_layout()
plt.show()

# Print summary
print("\n" + "=" * 80)
print(f"TESLA STORE EXPOSURE SUMMARY ({max_month.strftime('%B %Y')})")
print("=" * 80)
print(f"\nRaw Exposure (Σ 1/distance_km):")
print(iv_max_month['tesla_store_exposure'].describe().round(3))
print(f"\nLog Exposure (used in regressions):")
print(iv_max_month['iv_tesla_store_exposure_log'].describe().round(3))

# Top counties
print("\n" + "=" * 80)
print("TOP 10 COUNTIES BY TESLA STORE EXPOSURE")
print("=" * 80)
top_counties = iv_max_month.nlargest(10, 'tesla_store_exposure')[
    ['county_name', 'tesla_store_exposure', 'iv_tesla_store_exposure_log', 'evs_per_1000']
].round(3)
display(top_counties)

print("\n" + "=" * 80)
print("MAP INTERPRETATION")
print("=" * 80)
print("• DARK RED counties = High exposure (near multiple Tesla stores)")
print("• YELLOW counties = Low exposure (rural, far from stores)")
print("• BLUE DOTS = County centroids (distances calculated from this point)")
print("• RED STARS = Tesla stores (70 locations statewide)")
print("\nGEOGRAPHIC PATTERNS:")
print("• Bay Area: Highest exposure (San Francisco, Santa Clara, Alameda)")
print("• Southern CA: Concentrated around Los Angeles, Orange, San Diego")
print("• Central Valley & Rural: Minimal exposure (far from all stores)")
print("\nCALCULATION: For each county, exposure = log(1 + Σ(1/distance_km to each store))")
print("• Closer stores contribute MORE (1/10km = 0.10 vs 1/100km = 0.01)")
print("• Multiple nearby stores compound the exposure")
print("• This creates continuous spatial variation exploited by the instrument")
print("=" * 80)
Cell [15]
import pandas as pd
import numpy as np

# Convert both to string for merge
df_iv_temp = df_iv.copy()
df_iv_temp['county_geoid'] = df_iv_temp['county_geoid'].astype(str)

# Drop old census columns if they exist (from previous merge attempts)
census_cols = ['percent_dac_tracts', 'median_household_income',
               'percent_bachelor_or_higher', 'percent_no_vehicle_available']
for col in census_cols:
    if col in df_iv_temp.columns:
        df_iv_temp = df_iv_temp.drop(columns=[col])

census_temp = census_df.copy()
census_temp['county_geoid'] = census_temp['county_geoid'].astype(str)

# Merge census demographics onto IV panel
df_with_census = df_iv_temp.merge(
    census_temp[['county_geoid'] + census_cols],
    on='county_geoid',
    how='left',
    validate='m:1'
)

print("=" * 80)
print("CENSUS DATA MERGE")
print("=" * 80)
print(f"\nOriginal df_iv: {len(df_iv)} rows, {df_iv['county_geoid'].nunique()} counties")
print(f"After merge: {len(df_with_census)} rows, {df_with_census['county_geoid'].nunique()} counties")

print(f"\nMissing census data by variable:")
for var in census_cols:
    missing = df_with_census[var].isna().sum()
    pct_missing = missing/len(df_with_census)*100
    print(f"  {var}: {missing} missing ({pct_missing:.1f}%)")

# -------------------------
# Create Interaction Terms for Heterogeneity Analysis
# -------------------------
print("\n" + "=" * 80)
print("CREATING INTERACTION TERMS")
print("=" * 80)

# Standardize income to avoid numerical issues (convert to $10k units)
df_with_census["median_household_income_10k"] = df_with_census["median_household_income"] / 10000

# Create EV × Demographics interactions (for 2nd stage)
df_with_census["ev_x_dac"] = df_with_census["evs_per_1000"] * df_with_census["percent_dac_tracts"]
df_with_census["ev_x_income"] = df_with_census["evs_per_1000"] * df_with_census["median_household_income_10k"]

# Create IV × Demographics interactions (for 1st stage)
df_with_census["iv_x_dac"] = df_with_census["iv_tesla_store_exposure_log"] * df_with_census["percent_dac_tracts"]
df_with_census["iv_x_income"] = df_with_census["iv_tesla_store_exposure_log"] * df_with_census["median_household_income_10k"]

print("\nInteraction terms created:")
print("  • ev_x_dac          (evs_per_1000 × percent_dac_tracts)")
print("  • ev_x_income       (evs_per_1000 × median_household_income_10k)")
print("  • iv_x_dac          (iv_tesla_store_exposure_log × percent_dac_tracts)")
print("  • iv_x_income       (iv_tesla_store_exposure_log × median_household_income_10k)")

print("\n" + "=" * 80)
print("Census data and interaction terms ready for heterogeneity analysis")
print("=" * 80)

df_with_census
Cell [18]
import pandas as pd
import numpy as np

# -------------------------
# Create Interaction Terms for Heterogeneity Analysis
# -------------------------
df_interactions = df_with_census.copy()

print("=" * 80)
print("CREATING INTERACTION TERMS FOR HETEROGENEITY ANALYSIS")
print("=" * 80)

# Standardize income to avoid numerical issues (convert to $10k units)
df_interactions["median_household_income_10k"] = df_interactions["median_household_income"] / 10000

print("\n1. Income standardized to $10k units")
print(f"   Original range: ${df_interactions['median_household_income'].min():.0f} - ${df_interactions['median_household_income'].max():.0f}")
print(f"   Standardized range: {df_interactions['median_household_income_10k'].min():.1f} - {df_interactions['median_household_income_10k'].max():.1f}")

# Create EV × Demographics interactions (for 2nd stage)
df_interactions["ev_x_dac"] = df_interactions["evs_per_1000"] * df_interactions["percent_dac_tracts"]
df_interactions["ev_x_income"] = df_interactions["evs_per_1000"] * df_interactions["median_household_income_10k"]

# Create IV × Demographics interactions (for 1st stage instruments)
df_interactions["iv_x_dac"] = df_interactions["iv_tesla_store_exposure_log"] * df_interactions["percent_dac_tracts"]
df_interactions["iv_x_income"] = df_interactions["iv_tesla_store_exposure_log"] * df_interactions["median_household_income_10k"]

print("\n2. Interaction terms created:")
print("   For 2nd stage (endogenous interactions):")
print("     • ev_x_dac          = evs_per_1000 × percent_dac_tracts")
print("     • ev_x_income       = evs_per_1000 × median_household_income_10k")
print("\n   For 1st stage (instruments for interactions):")
print("     • iv_x_dac          = iv_tesla_store_exposure_log × percent_dac_tracts")
print("     • iv_x_income       = iv_tesla_store_exposure_log × median_household_income_10k")

# Check for missing values in interaction terms
print("\n3. Missing values check:")
interaction_cols = ["ev_x_dac", "ev_x_income", "iv_x_dac", "iv_x_income"]
for col in interaction_cols:
    missing = df_interactions[col].isna().sum()
    pct = missing / len(df_interactions) * 100
    print(f"   {col:<20} {missing:>6} missing ({pct:>5.1f}%)")

print("\n" + "=" * 80)
print("Interaction terms ready for 2SLS heterogeneity analysis")
print("=" * 80)
print("\nNote: Time-invariant moderators (DAC %, income) will be absorbed by")
print("      county fixed effects. Only the interaction terms are identified,")
print("      allowing us to test how EV effects vary across the demographic gradient.")
print("=" * 80)

df_interactions
Cell [21]
import pandas as pd

# -------------------------
# Load & merge solar control variable
# -------------------------
#solar_df = pd.read_csv("solar_control_variable.csv")
solar_df["month"] = pd.to_datetime(solar_df["month"])

# -------------------------
# Validate solar lookup coverage
# -------------------------
solar_keys = set(zip(solar_df["county_name"], solar_df["month"]))
panel_keys = set(zip(df_interactions["county_name"], pd.to_datetime(df_interactions["month"])))
solar_only = solar_keys - panel_keys
panel_only_missing = panel_keys - solar_keys

print("SOLAR MERGE VALIDATION")
print("=" * 80)
print(f"  Solar rows:              {len(solar_df):,}")
print(f"  Panel rows:              {len(df_interactions):,}")
print(f"  Solar keys matched:      {len(solar_keys & panel_keys):,} / {len(solar_keys):,}")
if solar_only:
    print(f"\n  ⚠ {len(solar_only)} solar rows have NO match in df_interactions:")
    for cn, m in sorted(solar_only)[:20]:
        print(f"      {cn} | {m.strftime('%Y-%m-%d')}")
    if len(solar_only) > 20:
        print(f"      ... and {len(solar_only) - 20} more")
else:
    print("\n  ✓ Every solar row matched a panel row")

if panel_only_missing:
    print(f"\n  ⚠ {len(panel_only_missing)} panel rows will have NULL solar values (no solar data):")
    for cn, m in sorted(panel_only_missing)[:20]:
        print(f"      {cn} | {m.strftime('%Y-%m-%d')}")
    if len(panel_only_missing) > 20:
        print(f"      ... and {len(panel_only_missing) - 20} more")
else:
    print("\n  ✓ Every panel row has a solar match")
print("=" * 80 + "\n")

# -------------------------
# Create Final Analysis Dataset
# -------------------------
print("=" * 80)
print("CREATING FINAL ANALYSIS DATASET")
print("=" * 80)

# Select only the columns needed for 2SLS analysis
analysis_cols = [
    # Identifiers
    "county_geoid",
    "county_name",
    "month",

    # Outcomes (dependent variables)
    "avg_lmp_weighted",
    "mad_rate",
    "median_of_medians",

    # Endogenous variable
    "evs_per_1000",

    # Instrument
    "iv_tesla_store_exposure_log",

    # Controls
    "tmax_c",
    "cumulative_data_center_power_mw",
    "total_generator_active_capacity_mw",
    "solar_mw_per_1000_lag1",

    # Demographics (moderators)
    "percent_dac_tracts",
    "median_household_income_10k",

    # Interaction terms (2nd stage)
    "ev_x_dac",
    "ev_x_income",

    # Interaction instruments (1st stage)
    "iv_x_dac",
    "iv_x_income"
]

df_final = df_interactions.merge(
    solar_df, on=["county_name", "month"], how="left"
)[analysis_cols].copy()

print("\n1. DATASET STRUCTURE")
print("=" * 80)
print(f"Rows: {len(df_final):,}")
print(f"Columns: {len(df_final.columns)}")
print(f"Counties: {df_final['county_geoid'].nunique()}")
print(f"Time periods: {df_final['month'].nunique()}")
print(f"Date range: {df_final['month'].min()} to {df_final['month'].max()}")

print("\n2. COLUMN GROUPS")
print("=" * 80)
print("\n  IDENTIFIERS (3):")
print("    • county_geoid, county_name, month")

print("\n  OUTCOMES (3):")
print("    • avg_lmp_weighted       - Average electricity price")
print("    • mad_rate               - Price volatility (MAD)")
print("    • median_of_medians      - Median electricity price")

print("\n  TREATMENT & INSTRUMENT (2):")
print("    • evs_per_1000                  - EV adoption (endogenous)")
print("    • iv_tesla_store_exposure_log   - Tesla store exposure (instrument)")

print("\n  CONTROLS (4):")
print("    • tmax_c                                - Max temperature")
print("    • cumulative_data_center_power_mw       - Data center demand")
print("    • total_generator_active_capacity_mw    - Generation capacity")
print("    • solar_mw_per_1000_lag1                - Solar capacity (lagged)")

print("\n  DEMOGRAPHICS (2):")
print("    • percent_dac_tracts              - % disadvantaged tracts")
print("    • median_household_income_10k     - Median income ($10k units)")

print("\n  INTERACTIONS (4):")
print("    • ev_x_dac       - EV × DAC (2nd stage)")
print("    • ev_x_income    - EV × Income (2nd stage)")
print("    • iv_x_dac       - IV × DAC (1st stage instrument)")
print("    • iv_x_income    - IV × Income (1st stage instrument)")

print("\n3. MISSING DATA")
print("=" * 80)
total_missing = df_final.isna().sum()
if total_missing.sum() > 0:
    print("\nColumns with missing values:")
    for col, missing in total_missing[total_missing > 0].items():
        pct = missing / len(df_final) * 100
        print(f"  {col:<40} {missing:>6} ({pct:>5.1f}%)")
else:
    print("\n  ✓ No missing values in any column")

print("\n" + "=" * 80)
print("Final analysis dataset ready for 2SLS regressions")
print("=" * 80)

df_final
Cell [22]
df_final.columns.tolist()
Cell [23]
import scipy.stats as stats

# County-level: mean solar capacity and income
county_level = df_final.groupby('county_name').agg(
    mean_solar=('solar_mw_per_1000_lag1', 'mean'),
    income=('median_household_income_10k', 'first')
).reset_index()

# Correlation
r, p = stats.pearsonr(county_level['mean_solar'], county_level['income'])
print(f"Pearson r: {r:.3f}, p-value: {p:.4f}")

# Scatter plot
import matplotlib.pyplot as plt
plt.scatter(county_level['income'], county_level['mean_solar'])
for _, row in county_level.iterrows():
    plt.annotate(row['county_name'], (row['income'], row['mean_solar']), fontsize=6)
plt.xlabel('Median Household Income ($10k)')
plt.ylabel('Mean Solar Capacity per 1,000')
plt.title('Solar Capacity vs Income by County')
plt.tight_layout()
plt.show()
Cell [26]
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

# -------------------------------------------------------------------
# First-stage diagnostics: testing instrument relevance
# Including interaction instruments for heterogeneity analysis
# with cluster-robust standard errors (clustered by county_geoid)
# -------------------------------------------------------------------

df = df_final.copy()

# Ensure proper types
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)

num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce").astype(float)

# Controls (matching second-stage specification)
controls = [
    "tmax_c",
    "cumulative_data_center_power_mw",
    "total_generator_active_capacity_mw",
    "solar_mw_per_1000_lag1"
]

print("=" * 80)
print("FIRST-STAGE DIAGNOSTICS")
print("=" * 80)
print(f"\nControls: {', '.join(controls)}")
print(f"Fixed Effects: County FE + Month FE")
print(f"Standard Errors: Clustered by county")
print("\n" + "=" * 80)

first_stage_results = []

# ==========================================
# 1. Main Effect: evs_per_1000 ~ iv_tesla_store_exposure_log
# ==========================================
print("\n1. MAIN EFFECT INSTRUMENT")
print("-" * 80)

needed_cols = ["county_geoid", "month", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls
data_main = df[needed_cols].dropna().copy()

formula_main = (
    "evs_per_1000 ~ iv_tesla_store_exposure_log"
    + " + " + " + ".join(controls)
    + " + C(county_geoid) + C(month)"
)

model_main_classic = smf.ols(formula_main, data=data_main).fit()
model_main_cluster = smf.ols(formula_main, data=data_main).fit(
    cov_type="cluster",
    cov_kwds={"groups": data_main["county_geoid"]},
)

classic_test_main = model_main_classic.f_test("iv_tesla_store_exposure_log = 0")
classic_f_main = float(np.asarray(classic_test_main.fvalue).squeeze())

robust_test_main = model_main_cluster.wald_test("iv_tesla_store_exposure_log = 0", scalar=True)
robust_stat_main = float(robust_test_main.statistic)

first_stage_results.append({
    "Endogenous Variable": "evs_per_1000",
    "Instrument": "iv_tesla_store_exposure_log",
    "Coefficient": float(model_main_cluster.params["iv_tesla_store_exposure_log"]),
    "SE (cluster)": float(model_main_cluster.bse["iv_tesla_store_exposure_log"]),
    "P-value": float(model_main_cluster.pvalues["iv_tesla_store_exposure_log"]),
    "F-stat (classic)": classic_f_main,
    "Wald (cluster)": robust_stat_main,
    "N": int(model_main_cluster.nobs),
})

print(f"Dependent: evs_per_1000")
print(f"Instrument: iv_tesla_store_exposure_log")
print(f"F-stat (classic): {classic_f_main:.2f}  |  Wald (cluster): {robust_stat_main:.2f}")
status_main = '✓ STRONG' if classic_f_main > 10 else '✗ WEAK'
print(f"Status: {status_main} (F > 10)")

# ==========================================
# 2. DAC Interaction: ev_x_dac ~ iv_x_dac
# ==========================================
print("\n2. DAC INTERACTION INSTRUMENT")
print("-" * 80)

needed_cols_dac = ["county_geoid", "month", "ev_x_dac", "iv_x_dac"] + controls
data_dac = df[needed_cols_dac].dropna().copy()

formula_dac = (
    "ev_x_dac ~ iv_x_dac"
    + " + " + " + ".join(controls)
    + " + C(county_geoid) + C(month)"
)

model_dac_classic = smf.ols(formula_dac, data=data_dac).fit()
model_dac_cluster = smf.ols(formula_dac, data=data_dac).fit(
    cov_type="cluster",
    cov_kwds={"groups": data_dac["county_geoid"]},
)

classic_test_dac = model_dac_classic.f_test("iv_x_dac = 0")
classic_f_dac = float(np.asarray(classic_test_dac.fvalue).squeeze())

robust_test_dac = model_dac_cluster.wald_test("iv_x_dac = 0", scalar=True)
robust_stat_dac = float(robust_test_dac.statistic)

first_stage_results.append({
    "Endogenous Variable": "ev_x_dac",
    "Instrument": "iv_x_dac",
    "Coefficient": float(model_dac_cluster.params["iv_x_dac"]),
    "SE (cluster)": float(model_dac_cluster.bse["iv_x_dac"]),
    "P-value": float(model_dac_cluster.pvalues["iv_x_dac"]),
    "F-stat (classic)": classic_f_dac,
    "Wald (cluster)": robust_stat_dac,
    "N": int(model_dac_cluster.nobs),
})

print(f"Dependent: ev_x_dac (EV × percent_dac_tracts)")
print(f"Instrument: iv_x_dac (IV × percent_dac_tracts)")
print(f"F-stat (classic): {classic_f_dac:.2f}  |  Wald (cluster): {robust_stat_dac:.2f}")
status_dac = '✓ STRONG' if classic_f_dac > 10 else '✗ WEAK'
print(f"Status: {status_dac} (F > 10)")

# ==========================================
# 3. Income Interaction: ev_x_income ~ iv_x_income
# ==========================================
print("\n3. INCOME INTERACTION INSTRUMENT")
print("-" * 80)

needed_cols_income = ["county_geoid", "month", "ev_x_income", "iv_x_income"] + controls
data_income = df[needed_cols_income].dropna().copy()

formula_income = (
    "ev_x_income ~ iv_x_income"
    + " + " + " + ".join(controls)
    + " + C(county_geoid) + C(month)"
)

model_income_classic = smf.ols(formula_income, data=data_income).fit()
model_income_cluster = smf.ols(formula_income, data=data_income).fit(
    cov_type="cluster",
    cov_kwds={"groups": data_income["county_geoid"]},
)

classic_test_income = model_income_classic.f_test("iv_x_income = 0")
classic_f_income = float(np.asarray(classic_test_income.fvalue).squeeze())

robust_test_income = model_income_cluster.wald_test("iv_x_income = 0", scalar=True)
robust_stat_income = float(robust_test_income.statistic)

first_stage_results.append({
    "Endogenous Variable": "ev_x_income",
    "Instrument": "iv_x_income",
    "Coefficient": float(model_income_cluster.params["iv_x_income"]),
    "SE (cluster)": float(model_income_cluster.bse["iv_x_income"]),
    "P-value": float(model_income_cluster.pvalues["iv_x_income"]),
    "F-stat (classic)": classic_f_income,
    "Wald (cluster)": robust_stat_income,
    "N": int(model_income_cluster.nobs),
})

print(f"Dependent: ev_x_income (EV × median_household_income_10k)")
print(f"Instrument: iv_x_income (IV × median_household_income_10k)")
print(f"F-stat (classic): {classic_f_income:.2f}  |  Wald (cluster): {robust_stat_income:.2f}")
status_income = '✓ STRONG' if classic_f_income > 10 else '✗ WEAK'
print(f"Status: {status_income} (F > 10)")

# ==========================================
# Summary Table
# ==========================================
print("\n" + "=" * 80)
print("SUMMARY: ALL FIRST-STAGE INSTRUMENTS")
print("=" * 80)

summary_df = pd.DataFrame(first_stage_results)
display(summary_df)

print("\n" + "=" * 80)
print("INTERPRETATION")
print("=" * 80)
all_strong = all(r["F-stat (classic)"] > 10 for r in first_stage_results)
if all_strong:
    print("\n✓ ALL instruments are STRONG (F > 10)")
    print("✓ Main effect AND interaction effects are well-identified")
    print("✓ Can proceed with heterogeneity analysis with confidence")
else:
    print("\n⚠️ Some instruments are WEAK (F < 10)")
    print("⚠️ Weak instrument bias may affect coefficient estimates")

print("=" * 80)
Cell [29]
# ==========================================
# NAIVE OLS: mad_rate ~ evs_per_1000 (no FEs, no IV)
# ==========================================

data_ols_nofe = df[
    ["county_geoid", "mad_rate", "evs_per_1000"] + controls
].dropna().copy()

formula_ols_nofe = "mad_rate ~ evs_per_1000 + " + " + ".join(controls)

model_ols_nofe = smf.ols(formula_ols_nofe, data=data_ols_nofe).fit(
    cov_type="cluster",
    cov_kwds={"groups": data_ols_nofe["county_geoid"]},
)
print(model_ols_nofe.summary())
Cell [31]
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

# ==========================================
# FULL FIRST-STAGE REGRESSION TABLES
# ==========================================
# This cell shows detailed regression output for all three first-stage regressions
# (Main effect, DAC interaction, Income interaction)

df = df_final.copy()

# Ensure proper types
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)

num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce").astype(float)

controls = [
    "tmax_c",
    "cumulative_data_center_power_mw",
    "total_generator_active_capacity_mw"
]

print("=" * 80)
print("DETAILED FIRST-STAGE REGRESSION OUTPUTS")
print("=" * 80)
print("\nShowing full regression tables with cluster-robust standard errors")
print("=" * 80)

# ==========================================
# 1. Main Effect
# ==========================================
print("\n\n1. MAIN EFFECT: evs_per_1000 ~ iv_tesla_store_exposure_log")
print("=" * 80)

needed_cols = ["county_geoid", "month", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls
data_main = df[needed_cols].dropna().copy()

formula_main = (
    "evs_per_1000 ~ iv_tesla_store_exposure_log"
    + " + " + " + ".join(controls)
    + " + C(county_geoid) + C(month)"
)

model_main = smf.ols(formula_main, data=data_main).fit(
    cov_type="cluster",
    cov_kwds={"groups": data_main["county_geoid"]},
)

print(model_main.summary())

# ==========================================
# 2. DAC Interaction
# ==========================================
print("\n\n2. DAC INTERACTION: ev_x_dac ~ iv_x_dac")
print("=" * 80)

needed_cols_dac = ["county_geoid", "month", "ev_x_dac", "iv_x_dac"] + controls
data_dac = df[needed_cols_dac].dropna().copy()

formula_dac = (
    "ev_x_dac ~ iv_x_dac"
    + " + " + " + ".join(controls)
    + " + C(county_geoid) + C(month)"
)

model_dac = smf.ols(formula_dac, data=data_dac).fit(
    cov_type="cluster",
    cov_kwds={"groups": data_dac["county_geoid"]},
)

print(model_dac.summary())

# ==========================================
# 3. Income Interaction
# ==========================================
print("\n\n3. INCOME INTERACTION: ev_x_income ~ iv_x_income")
print("=" * 80)

needed_cols_income = ["county_geoid", "month", "ev_x_income", "iv_x_income"] + controls
data_income = df[needed_cols_income].dropna().copy()

formula_income = (
    "ev_x_income ~ iv_x_income"
    + " + " + " + ".join(controls)
    + " + C(county_geoid) + C(month)"
)

model_income = smf.ols(formula_income, data=data_income).fit(
    cov_type="cluster",
    cov_kwds={"groups": data_income["county_geoid"]},
)

print(model_income.summary())

print("\n" + "=" * 80)
print("END OF FIRST-STAGE REGRESSION TABLES")
print("=" * 80)
Cell [35]
import numpy as np
import pandas as pd
from linearmodels.iv import IV2SLS

df = df_final.copy()

# Type prep
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)

# Convert numeric columns (excluding identifiers)
num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")

# EXPANDED CONTROLS (matching first-stage)
# Adding generator capacity to control for supply-side factors
controls = [
    "tmax_c",
    "cumulative_data_center_power_mw",
    "total_generator_active_capacity_mw",
    "solar_mw_per_1000_lag1"
]

control_str = " + ".join(controls)

results = {}

print("=" * 80)
print("TWO-STAGE LEAST SQUARES REGRESSION")
print("=" * 80)
print(f"\nExpanded controls ({len(controls)}):")
for ctrl in controls:
    print(f"  • {ctrl}")
print(f"\nFixed Effects: County FE + Month FE")
print(f"Standard Errors: Clustered by county")
print("\n" + "=" * 80)

# ==========================================
# Spec 1: avg_lmp_weighted
# ==========================================
iv_cols = ["avg_lmp_weighted", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]

iv_data = df[iv_cols].dropna().copy()
for c in iv_data.columns:
    if pd.api.types.is_numeric_dtype(iv_data[c]):
        iv_data[c] = iv_data[c].astype(float)

iv_formula = f"""
avg_lmp_weighted ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""

iv_model = IV2SLS.from_formula(iv_formula, data=iv_data).fit(
    cov_type="clustered",
    clusters=iv_data["county_geoid"],
)

results["avg_lmp_weighted"] = {
    "outcome": "avg_lmp_weighted",
    "coef": iv_model.params["evs_per_1000"],
    "se": iv_model.std_errors["evs_per_1000"],
    "pval": iv_model.pvalues["evs_per_1000"],
    "n": int(iv_model.nobs),
}

print("\nSpec 1: EV Adoption → Average Price (avg_lmp_weighted)")
print("=" * 80)
print(iv_model.summary.tables[1])

# ==========================================
# Spec 2: mad_rate
# ==========================================
iv_cols_mad = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]

iv_data_mad = df[iv_cols_mad].dropna().copy()
for c in iv_data_mad.columns:
    if pd.api.types.is_numeric_dtype(iv_data_mad[c]):
        iv_data_mad[c] = iv_data_mad[c].astype(float)

iv_formula_mad = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""

iv_model_mad = IV2SLS.from_formula(iv_formula_mad, data=iv_data_mad).fit(
    cov_type="clustered",
    clusters=iv_data_mad["county_geoid"],
)

results["mad_rate"] = {
    "outcome": "mad_rate",
    "coef": iv_model_mad.params["evs_per_1000"],
    "se": iv_model_mad.std_errors["evs_per_1000"],
    "pval": iv_model_mad.pvalues["evs_per_1000"],
    "n": int(iv_model_mad.nobs),
}

print("\n" + "=" * 80)
print("Spec 2: EV Adoption → Price Volatility (mad_rate)")
print("=" * 80)
print(iv_model_mad.summary.tables[1])

# ==========================================
# Spec 3: median_of_medians
# ==========================================
iv_cols_median = ["median_of_medians", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]

iv_data_median = df[iv_cols_median].dropna().copy()
for c in iv_data_median.columns:
    if pd.api.types.is_numeric_dtype(iv_data_median[c]):
        iv_data_median[c] = iv_data_median[c].astype(float)

iv_formula_median = f"""
median_of_medians ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""

iv_model_median = IV2SLS.from_formula(iv_formula_median, data=iv_data_median).fit(
    cov_type="clustered",
    clusters=iv_data_median["county_geoid"],
)

results["median_of_medians"] = {
    "outcome": "median_of_medians",
    "coef": iv_model_median.params["evs_per_1000"],
    "se": iv_model_median.std_errors["evs_per_1000"],
    "pval": iv_model_median.pvalues["evs_per_1000"],
    "n": int(iv_model_median.nobs),
}

print("\n" + "=" * 80)
print("Spec 3: EV Adoption → Median Price (median_of_medians)")
print("=" * 80)
print(iv_model_median.summary.tables[1])

# ==========================================
# Spec 4: mad_rate with DAC Interaction
# ==========================================
print("\n" + "=" * 80)
print("Spec 4: EV × DAC Interaction → Price Volatility (mad_rate)")
print("=" * 80)

iv_cols_mad_dac = ["mad_rate", "evs_per_1000", "ev_x_dac",
                    "iv_tesla_store_exposure_log", "iv_x_dac"] + controls + ["county_geoid", "month"]

iv_data_mad_dac = df[iv_cols_mad_dac].dropna().copy()
for c in iv_data_mad_dac.columns:
    if pd.api.types.is_numeric_dtype(iv_data_mad_dac[c]):
        iv_data_mad_dac[c] = iv_data_mad_dac[c].astype(float)

iv_formula_mad_dac = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 + ev_x_dac ~ iv_tesla_store_exposure_log + iv_x_dac]
"""

iv_model_mad_dac = IV2SLS.from_formula(iv_formula_mad_dac, data=iv_data_mad_dac).fit(
    cov_type="clustered",
    clusters=iv_data_mad_dac["county_geoid"],
)

results["mad_rate_dac_interaction"] = {
    "outcome": "mad_rate (DAC interaction)",
    "main_coef": iv_model_mad_dac.params["evs_per_1000"],
    "main_se": iv_model_mad_dac.std_errors["evs_per_1000"],
    "main_pval": iv_model_mad_dac.pvalues["evs_per_1000"],
    "interaction_coef": iv_model_mad_dac.params["ev_x_dac"],
    "interaction_se": iv_model_mad_dac.std_errors["ev_x_dac"],
    "interaction_pval": iv_model_mad_dac.pvalues["ev_x_dac"],
    "n": int(iv_model_mad_dac.nobs),
}

print(iv_model_mad_dac.summary.tables[1])

# ==========================================
# Spec 5: mad_rate with Income Interaction
# ==========================================
print("\n" + "=" * 80)
print("Spec 5: EV × Income Interaction → Price Volatility (mad_rate)")
print("=" * 80)

iv_cols_mad_income = ["mad_rate", "evs_per_1000", "ev_x_income",
                       "iv_tesla_store_exposure_log", "iv_x_income"] + controls + ["county_geoid", "month"]

iv_data_mad_income = df[iv_cols_mad_income].dropna().copy()
for c in iv_data_mad_income.columns:
    if pd.api.types.is_numeric_dtype(iv_data_mad_income[c]):
        iv_data_mad_income[c] = iv_data_mad_income[c].astype(float)

iv_formula_mad_income = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 + ev_x_income ~ iv_tesla_store_exposure_log + iv_x_income]
"""

iv_model_mad_income = IV2SLS.from_formula(iv_formula_mad_income, data=iv_data_mad_income).fit(
    cov_type="clustered",
    clusters=iv_data_mad_income["county_geoid"],
)

results["mad_rate_income_interaction"] = {
    "outcome": "mad_rate (Income interaction)",
    "main_coef": iv_model_mad_income.params["evs_per_1000"],
    "main_se": iv_model_mad_income.std_errors["evs_per_1000"],
    "main_pval": iv_model_mad_income.pvalues["evs_per_1000"],
    "interaction_coef": iv_model_mad_income.params["ev_x_income"],
    "interaction_se": iv_model_mad_income.std_errors["ev_x_income"],
    "interaction_pval": iv_model_mad_income.pvalues["ev_x_income"],
    "n": int(iv_model_mad_income.nobs),
}

print(iv_model_mad_income.summary.tables[1])

results
Cell [36]

import pandas as pd

models = {
    "mad_rate": iv_model_mad,
    "mad_rate (DAC int.)": iv_model_mad_dac,
    "mad_rate (Income int.)": iv_model_mad_income,
}

controls_to_report = [
    "tmax_c",
    "cumulative_data_center_power_mw",
    "total_generator_active_capacity_mw",
    "solar_mw_per_1000_lag1",
]

rows = []
for spec_name, mod in models.items():
    for var in controls_to_report:
        if var in mod.params.index:
            pval = mod.pvalues[var]
            sig = "***" if pval < 0.01 else "**" if pval < 0.05 else "*" if pval < 0.10 else ""
            rows.append({
                "Specification": spec_name,
                "Variable": var,
                "Coefficient": mod.params[var],
                "P-value": pval,
                "Sig.": sig,
            })

control_summary = pd.DataFrame(rows)
control_summary["Coefficient"] = control_summary["Coefficient"].map("{:.6f}".format)
control_summary["P-value"] = control_summary["P-value"].map("{:.4f}".format)

#print("=" * 70)
#print("SECOND-STAGE 2SLS: CONTROL VARIABLE COEFFICIENTS")
#print("=" * 70)

control_summary
Cell [38]
import pandas as pd
import statsmodels.formula.api as smf

print("=" * 80)
print("OLS vs 2SLS COMPARISON: mad_rate (Price Volatility) ")
print("=" * 80)
print("\nAll specifications include:")
print("  • Controls: Temperature, Data Centers, Generation Capacity")
print("  • Fixed Effects: County FE + Month FE")
print("  • Standard Errors: Clustered by county")
print("\n" + "=" * 80)
print("IMPORTANT NOTES ON INTERPRETATION:")
print("=" * 80)
print("\n1. SIGN CHANGES:")
print("   'YES' means OLS and 2SLS have opposite signs → severe endogeneity bias")
print("\n2. INTERACTION MODELS (Sections 2 & 3):")
print("   • Main effects are conditional on moderator = 0 (theoretical/extrapolated)")
print("   • DAC model: evs_per_1000 coef = effect when percent_dac_tracts = 0")
print("   • Income model: evs_per_1000 coef = effect when income = $0")
print("   • These are absorbed by county FE and not directly interpretable")
print("\n3. WHAT MATTERS:")
print("   ✓ Baseline spec: evs_per_1000 coefficient = average causal effect")
print("   ✓ Interaction specs: INTERACTION coefficient = how effect varies by demographic")
print("   ✗ DO NOT compare main effects across specifications with/without interactions")
print("\n4. INTERPRETING INTERACTIONS:")
print("   Interaction coef = change in EV effect per unit change in moderator")
print("   • If negative: EV effect decreases as moderator increases")
print("   • If positive: EV effect increases as moderator increases")
print("\n" + "=" * 80)

# ==========================================
# RUN OLS REGRESSIONS
# ==========================================
df = df_final.copy()
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)

num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")

controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
control_str = " + ".join(controls)

# Baseline OLS
ols_cols_baseline = ["mad_rate", "evs_per_1000"] + controls + ["county_geoid", "month"]
ols_data_baseline = df[ols_cols_baseline].dropna().copy()
ols_formula_baseline = f"mad_rate ~ evs_per_1000 + {control_str} + C(county_geoid) + C(month)"
ols_baseline = smf.ols(ols_formula_baseline, data=ols_data_baseline).fit(
    cov_type="cluster", cov_kwds={"groups": ols_data_baseline["county_geoid"]})

# DAC Interaction OLS
ols_cols_dac = ["mad_rate", "evs_per_1000", "ev_x_dac"] + controls + ["county_geoid", "month"]
ols_data_dac = df[ols_cols_dac].dropna().copy()
ols_formula_dac = f"mad_rate ~ evs_per_1000 + ev_x_dac + {control_str} + C(county_geoid) + C(month)"
ols_dac = smf.ols(ols_formula_dac, data=ols_data_dac).fit(
    cov_type="cluster", cov_kwds={"groups": ols_data_dac["county_geoid"]})

# Income Interaction OLS
ols_cols_income = ["mad_rate", "evs_per_1000", "ev_x_income"] + controls + ["county_geoid", "month"]
ols_data_income = df[ols_cols_income].dropna().copy()
ols_formula_income = f"mad_rate ~ evs_per_1000 + ev_x_income + {control_str} + C(county_geoid) + C(month)"
ols_income = smf.ols(ols_formula_income, data=ols_data_income).fit(
    cov_type="cluster", cov_kwds={"groups": ols_data_income["county_geoid"]})

# ==========================================
# SECTION 1: BASELINE SPECIFICATION
# ==========================================
print("\n\n" + "=" * 80)
print("SECTION 1: BASELINE SPECIFICATION")
print("=" * 80)
print("\nModel: mad_rate ~ evs_per_1000 + controls + FE")

tsls_base = results["mad_rate"]

baseline_table = pd.DataFrame([{
    "Method": "OLS",
    "Coefficient": f"{ols_baseline.params['evs_per_1000']:.6f}",
    "Std. Error": f"{ols_baseline.bse['evs_per_1000']:.6f}",
    "P-value": f"{ols_baseline.pvalues['evs_per_1000']:.4f}",
    "Sig.": "**" if ols_baseline.pvalues['evs_per_1000'] < 0.05 else "*" if ols_baseline.pvalues['evs_per_1000'] < 0.10 else "",
    "N": int(ols_baseline.nobs)
}, {
    "Method": "2SLS (IV)",
    "Coefficient": f"{tsls_base['coef']:.6f}",
    "Std. Error": f"{tsls_base['se']:.6f}",
    "P-value": f"{tsls_base['pval']:.4f}",
    "Sig.": "***" if tsls_base['pval'] < 0.01 else "**" if tsls_base['pval'] < 0.05 else "*" if tsls_base['pval'] < 0.10 else "",
    "N": tsls_base['n']
}])

print("\n")
display(baseline_table)

sign_change_baseline = (ols_baseline.params['evs_per_1000'] * tsls_base['coef']) < 0
print(f"\nSign Change: {'YES' if sign_change_baseline else 'NO'}")

# ==========================================
# SECTION 2: DAC INTERACTION
# ==========================================
print("\n\n" + "=" * 80)
print("SECTION 2: DAC INTERACTION SPECIFICATION")
print("=" * 80)
print("\nModel: mad_rate ~ evs_per_1000 + (evs_per_1000 × percent_dac_tracts) + controls + FE")

tsls_dac = results["mad_rate_dac_interaction"]

dac_table = pd.DataFrame([{
    "Variable": "evs_per_1000 (main)",
    "OLS Coef": f"{ols_dac.params['evs_per_1000']:.6f}",
    "OLS P-val": f"{ols_dac.pvalues['evs_per_1000']:.4f}",
    "2SLS Coef": f"{tsls_dac['main_coef']:.6f}",
    "2SLS P-val": f"{tsls_dac['main_pval']:.4f}",
}, {
    "Variable": "ev_x_dac (interaction)",
    "OLS Coef": f"{ols_dac.params['ev_x_dac']:.6f}",
    "OLS P-val": f"{ols_dac.pvalues['ev_x_dac']:.4f}",
    "2SLS Coef": f"{tsls_dac['interaction_coef']:.6f}",
    "2SLS P-val": f"{tsls_dac['interaction_pval']:.4f}",
}])

print("\n")
display(dac_table)

sign_change_dac_main = (ols_dac.params['evs_per_1000'] * tsls_dac['main_coef']) < 0
sign_change_dac_int = (ols_dac.params['ev_x_dac'] * tsls_dac['interaction_coef']) < 0
print(f"\nSign Changes: Main = {'YES' if sign_change_dac_main else 'NO'}, Interaction = {'YES' if sign_change_dac_int else 'NO'}")

# ==========================================
# SECTION 3: INCOME INTERACTION
# ==========================================
print("\n\n" + "=" * 80)
print("SECTION 3: INCOME INTERACTION SPECIFICATION")
print("=" * 80)
print("\nModel: mad_rate ~ evs_per_1000 + (evs_per_1000 × median_income_10k) + controls + FE")

tsls_income = results["mad_rate_income_interaction"]

income_table = pd.DataFrame([{
    "Variable": "evs_per_1000 (main)",
    "OLS Coef": f"{ols_income.params['evs_per_1000']:.6f}",
    "OLS P-val": f"{ols_income.pvalues['evs_per_1000']:.4f}",
    "2SLS Coef": f"{tsls_income['main_coef']:.6f}",
    "2SLS P-val": f"{tsls_income['main_pval']:.4f}",
}, {
    "Variable": "ev_x_income (interaction)",
    "OLS Coef": f"{ols_income.params['ev_x_income']:.6f}",
    "OLS P-val": f"{ols_income.pvalues['ev_x_income']:.4f}",
    "2SLS Coef": f"{tsls_income['interaction_coef']:.6f}",
    "2SLS P-val": f"{tsls_income['interaction_pval']:.4f}",
}])

print("\n")
display(income_table)

sign_change_income_main = (ols_income.params['evs_per_1000'] * tsls_income['main_coef']) < 0
sign_change_income_int = (ols_income.params['ev_x_income'] * tsls_income['interaction_coef']) < 0
print(f"\nSign Changes: Main = {'YES' if sign_change_income_main else 'NO'}, Interaction = {'YES' if sign_change_income_int else 'NO'}")

print("\n" + "=" * 80)
print("Significance: *** p<0.01, ** p<0.05, * p<0.10")
print("=" * 80)

# Store results for markdown interpretation
ols_vs_2sls_results = {
    "baseline": {
        "ols_coef": ols_baseline.params['evs_per_1000'],
        "ols_pval": ols_baseline.pvalues['evs_per_1000'],
        "tsls_coef": tsls_base['coef'],
        "tsls_pval": tsls_base['pval'],
    },
    "dac_main": {
        "ols_coef": ols_dac.params['evs_per_1000'],
        "ols_pval": ols_dac.pvalues['evs_per_1000'],
        "tsls_coef": tsls_dac['main_coef'],
        "tsls_pval": tsls_dac['main_pval'],
    },
    "dac_interaction": {
        "ols_coef": ols_dac.params['ev_x_dac'],
        "ols_pval": ols_dac.pvalues['ev_x_dac'],
        "tsls_coef": tsls_dac['interaction_coef'],
        "tsls_pval": tsls_dac['interaction_pval'],
    },
    "income_main": {
        "ols_coef": ols_income.params['evs_per_1000'],
        "ols_pval": ols_income.pvalues['evs_per_1000'],
        "tsls_coef": tsls_income['main_coef'],
        "tsls_pval": tsls_income['main_pval'],
    },
    "income_interaction": {
        "ols_coef": ols_income.params['ev_x_income'],
        "ols_pval": ols_income.pvalues['ev_x_income'],
        "tsls_coef": tsls_income['interaction_coef'],
        "tsls_pval": tsls_income['interaction_pval'],
    }
}

ols_vs_2sls_results
Cell [39]
import pandas as pd
import matplotlib.pyplot as plt
import re

def parse_and_filter_summary(model, spec_name):
    """Extract parameter table, drop FE rows, return as DataFrame."""
    table_str = str(model.summary.tables[1])
    lines = table_str.split('\n')

    rows = []
    for line in lines:
        # Skip county and month FE rows
        if 'C(county_geoid)' in line or 'C(month)' in line:
            continue
        # Skip separator lines
        if set(line.strip()) <= set('=-'):
            continue
        # Skip empty lines
        if not line.strip():
            continue
        # Skip header line
        if 'Parameter' in line and 'Std. Err' in line:
            continue
        rows.append(line)

    # Parse into columns
    data = []
    for row in rows:
        parts = row.split()
        if len(parts) >= 6:
            # Handle multi-word row names
            try:
                # Last 6 values are numeric
                nums = parts[-6:]
                name = ' '.join(parts[:-6])
                data.append([name] + nums)
            except:
                pass

    df = pd.DataFrame(data, columns=['Variable', 'Coef', 'Std. Err.', 'T-stat', 'P-value', 'Lower CI', 'Upper CI'])
    return df

def save_summary_png(model, spec_name, filename):
    # Header info
    h = model.summary.tables[0]
    header_str = str(h)

    # Parameter table
    df = parse_and_filter_summary(model, spec_name)

    fig = plt.figure(figsize=(10, len(df) * 0.35 + 3))

    # Header text
    ax_header = fig.add_axes([0, 0.85, 1, 0.15])
    ax_header.axis('off')
    ax_header.text(0.01, 0.8, f"Appendix: Full Regression Output — {spec_name}",
                   fontsize=10, fontweight='bold', va='top')
    ax_header.text(0.01, 0.4,
                   f"N={int(model.nobs):,}   R²={model.rsquared:.4f}   "
                   f"Clustered SEs by county   County & Month FE included (coefficients omitted)",
                   fontsize=8, color='#555555', va='top')

    # Parameter table
    ax_table = fig.add_axes([0, 0, 1, 0.85])
    ax_table.axis('off')

    tbl = ax_table.table(
        cellText=df.values,
        colLabels=df.columns,
        cellLoc='center',
        loc='center',
    )
    tbl.auto_set_font_size(False)
    tbl.set_fontsize(8.5)
    tbl.auto_set_column_width(col=list(range(len(df.columns))))

    for (row, col), cell in tbl.get_celld().items():
        cell.set_edgecolor('#dddddd')
        cell.set_linewidth(0.5)
        if row == 0:
            cell.set_facecolor('#f0f0f0')
            cell.set_text_props(fontweight='bold')
        else:
            cell.set_facecolor('white')
        if col == 0:
            cell.set_text_props(ha='left')
        cell.set_height(0.06)

    plt.savefig(filename, dpi=150, bbox_inches='tight')
    plt.close()
    print(f"Saved: {filename}")

# ── Save all three ─────────────────────────────────────────────────────────────
save_summary_png(iv_model_mad,        "Baseline",          "appendix_baseline.png")
save_summary_png(iv_model_mad_dac,    "DAC Interaction",   "appendix_dac.png")
save_summary_png(iv_model_mad_income, "Income Interaction","appendix_income.png")
Cell [40]
import pandas as pd
import matplotlib.pyplot as plt

def parse_and_filter_summary(model, spec_name):
    """Extract parameter table, drop FE rows, return as DataFrame."""
    table_str = str(model.summary.tables[1])
    lines = table_str.split('\n')
    rows = []
    for line in lines:
        if 'C(county_geoid)' in line or 'C(month)' in line:
            continue
        if set(line.strip()) <= set('=-'):
            continue
        if not line.strip():
            continue
        if 'Parameter' in line and 'Std. Err' in line:
            continue
        rows.append(line)

    data = []
    for row in rows:
        parts = row.split()
        if len(parts) >= 6:
            try:
                nums = parts[-6:]
                name = ' '.join(parts[:-6])
                data.append([name] + nums)
            except:
                pass

    df = pd.DataFrame(data, columns=['Variable', 'Coef', 'Std. Err.', 'T-stat', 'P-value', 'Lower CI', 'Upper CI'])
    return df

def save_summary_png(model, spec_name, filename):
    df = parse_and_filter_summary(model, spec_name)

    fig = plt.figure(figsize=(10, len(df) * 0.35 + 3))

    ax_header = fig.add_axes([0, 0.85, 1, 0.15])
    ax_header.axis('off')
    ax_header.text(0.01, 0.8, f"Appendix: Full Regression Output — {spec_name}",
                   fontsize=10, fontweight='bold', va='top')
    ax_header.text(0.01, 0.4,
                   f"N={int(model.nobs):,}   R²={model.rsquared:.4f}   "
                   f"Clustered SEs by county   County & Month FE included (coefficients omitted)",
                   fontsize=8, color='#555555', va='top')

    ax_table = fig.add_axes([0, 0, 1, 0.85])
    ax_table.axis('off')
    tbl = ax_table.table(
        cellText=df.values,
        colLabels=df.columns,
        cellLoc='center',
        loc='center',
    )
    tbl.auto_set_font_size(False)
    tbl.set_fontsize(8.5)
    tbl.auto_set_column_width(col=list(range(len(df.columns))))
    for (row, col), cell in tbl.get_celld().items():
        cell.set_edgecolor('#dddddd')
        cell.set_linewidth(0.5)
        if row == 0:
            cell.set_facecolor('#f0f0f0')
            cell.set_text_props(fontweight='bold')
        else:
            cell.set_facecolor('white')
        if col == 0:
            cell.set_text_props(ha='left')
        cell.set_height(0.06)

    plt.savefig(filename, dpi=150, bbox_inches='tight')
    plt.close()
    print(f"Saved: {filename}")

# ── Save all five specs ────────────────────────────────────────────────────────
save_summary_png(iv_model_mad,        "Baseline (mad_rate)",              "appendix_baseline_mad.png")
save_summary_png(iv_model_mad_dac,    "DAC Interaction (mad_rate)",       "appendix_dac_mad.png")
save_summary_png(iv_model_mad_income, "Income Interaction (mad_rate)",    "appendix_income_mad.png")
save_summary_png(iv_model,            "Baseline (avg_lmp_weighted)",      "appendix_baseline_lmp.png")
save_summary_png(iv_model_median,     "Baseline (median_of_medians)",     "appendix_baseline_median.png")
Cell [41]
import matplotlib.pyplot as plt
from collections import OrderedDict

# ── Extract F-stats ───────────────────────────────────────────────────────────
wald_main   = float(iv_model_mad.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
wald_dac    = float(iv_model_mad_dac.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
wald_income = float(iv_model_mad_income.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])

classical_main   = float(model_main.fvalue)
classical_dac    = float(model_dac.fvalue)
classical_income = float(model_income.fvalue)

import matplotlib.pyplot as plt

def get_stats(model):
    is_iv = hasattr(model, 'std_errors')
    if is_iv:
        params = model.params
        bse    = model.std_errors
        pvals  = model.pvalues
        nobs   = int(model.nobs)
        rsq    = float(model.rsquared)
    else:
        params = model.params
        bse    = model.bse
        pvals  = model.pvalues
        nobs   = int(model.nobs)
        rsq    = float(model.rsquared)
    mask = ~(
        params.index.str.startswith('C(county_geoid)') |
        params.index.str.startswith('C(month)')
    )
    return params[mask], bse[mask], pvals[mask], nobs, rsq

def sig_stars(p):
    if p < 0.01: return '***'
    if p < 0.05: return '**'
    return ''

def make_reg_table(
    models,
    col_labels,
    var_labels,
    county_fe,
    month_fe,
    filename,
    dep_var_label='Dependent Variable: Price Volatility (mad_rate)',
    note='Clustered standard errors in parentheses (county level).  *** p<0.01,  ** p<0.05',
    fstats=None,
    row_h=0.30,
    top_h=0.90,
    bot_h=0.55,
    var_col_w=2.8,
    data_col_w=1.55,
):
    all_stats = [get_stats(m) for m in models]

    seen, var_order = set(), []
    for params, *_ in all_stats:
        for v in params.index:
            if v not in seen:
                var_order.append(v)
                seen.add(v)

    data_rows = []
    for var in var_order:
        display = var_labels.get(var, var)
        coef_cells, se_cells = [], []
        for params, bse, pvals, *_ in all_stats:
            if var in params.index:
                c, s, p = params[var], bse[var], pvals[var]
                coef_cells.append(f"{c:.4f}{sig_stars(p)}")
                se_cells.append(f"({s:.4f})")
            else:
                coef_cells.append('')
                se_cells.append('')
        data_rows.append((True,  display, coef_cells))
        data_rows.append((False, display, se_cells))

    footer = []
    if fstats:
        wald_row, classical_row = [], []
        for col in col_labels:
            if col in fstats:
                wald, classical = fstats[col]
                wald_row.append(f"{wald:.2f}")
                classical_row.append(f"{classical:.2f}")
            else:
                wald_row.append('—')
                classical_row.append('—')
        footer.append(('Cluster-Robust Wald F', wald_row))
        footer.append(('Classical F-stat',      classical_row))

    footer += [
        ('County FEs',   [('Yes' if county_fe[i] else 'No') for i in range(len(models))]),
        ('Month FEs',    [('Yes' if month_fe[i]  else 'No') for i in range(len(models))]),
        ('Observations', [f"{s[3]:,}"                       for s in all_stats]),
        ('R-squared',    [f"{s[4]:.4f}"                     for s in all_stats]),
    ]

    n_cols  = len(models)
    n_data  = len(data_rows)
    n_foot  = len(footer)
    n_total = n_data + n_foot

    fig_w = var_col_w + n_cols * data_col_w
    fig_h = top_h + n_total * row_h + bot_h

    fig, ax = plt.subplots(figsize=(fig_w, fig_h))
    ax.set_xlim(0, fig_w)
    ax.set_ylim(0, fig_h)
    ax.axis('off')

    BLACK = '#000000'
    GRAY  = '#444444'
    FONT  = 'DejaVu Sans'

    col_x = [var_col_w + (i + 0.5) * data_col_w for i in range(n_cols)]

    def row_y(i):
        return fig_h - top_h - (i + 0.5) * row_h

    # dep var label
    ax.text(fig_w / 2, fig_h - 0.12, dep_var_label,
            ha='center', va='center', fontsize=8, fontfamily=FONT, color=GRAY)

    # column headers
    for cx, label in zip(col_x, col_labels):
        ax.text(cx, fig_h - 0.52, label,
                ha='center', va='center', fontsize=8.5,
                fontfamily=FONT, color=BLACK)

    # thick rule above column headers
    ax.plot([0.05, fig_w - 0.05], [fig_h - 0.28, fig_h - 0.28],
            color=BLACK, lw=1.4)

    # thick rule below column headers
    ax.plot([0.05, fig_w - 0.05], [fig_h - top_h, fig_h - top_h],
            color=BLACK, lw=1.4)

    # data rows
    for ri, (is_coef, var_display, cells) in enumerate(data_rows):
        y = row_y(ri)
        if is_coef:
            ax.text(0.10, y, var_display,
                    ha='left', va='center', fontsize=8,
                    fontfamily=FONT, color=BLACK, fontweight='bold')
        for cx, val in zip(col_x, cells):
            ax.text(cx, y, val,
                    ha='center', va='center', fontsize=8,
                    fontfamily=FONT, color=BLACK)

    # footer rows — no rule, flows directly from data rows
    for fi, (label, vals) in enumerate(footer):
        y = row_y(n_data + fi)
        ax.text(0.10, y, label,
                ha='left', va='center', fontsize=8,
                fontfamily=FONT, color=BLACK)
        for cx, val in zip(col_x, vals):
            ax.text(cx, y, val,
                    ha='center', va='center', fontsize=8,
                    fontfamily=FONT, color=BLACK)

    # thick bottom rule
    bot_rule_y = fig_h - top_h - n_total * row_h
    ax.plot([0.05, fig_w - 0.05], [bot_rule_y, bot_rule_y],
            color=BLACK, lw=1.4)

    # notes
    ax.text(0.10, bot_rule_y - 0.10, note,
            ha='left', va='top', fontsize=7,
            fontfamily=FONT, color=GRAY, style='italic')

    plt.savefig(filename, dpi=200, bbox_inches='tight', facecolor='white')
    plt.close()
    print(f"Saved: {filename}")

VAR_LABELS = {
    'evs_per_1000':                       'EVs per 1,000 Residents',
    'ev_x_dac':                           'EVs × DAC Share',
    'ev_x_income':                        'EVs × Median Income ($10k)',
    'percent_dac_tracts':                 'DAC Share',
    'median_household_income_10k':        'Median Income ($10k)',
    'tmax_c':                             'Max Temperature (°C)',
    'cumulative_data_center_power_mw':    'Cumulative Data Center Power (MW)',
    'total_generator_active_capacity_mw': 'Total Generator Capacity (MW)',
    'solar_mw_per_1000_lag1':             'Solar Capacity per 1,000 (lagged)',
    'Intercept':                          'Constant',
}

print("Functions defined.")
Cell [42]
wald_main   = float(iv_model_mad.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
wald_dac    = float(iv_model_mad_dac.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
wald_income = float(iv_model_mad_income.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])

classical_main   = float(model_main.fvalue)
classical_dac    = float(model_dac.fvalue)
classical_income = float(model_income.fvalue)

make_reg_table(
    models     = [model_ols_nofe, ols_baseline, iv_model_mad],
    col_labels = ['(1) Naive OLS', '(2) OLS', '(3) 2SLS'],
    var_labels = VAR_LABELS,
    county_fe  = [False, True, True],
    month_fe   = [False, True, True],
    dep_var_label = 'Dependent Variable: Price Volatility (mad_rate)',
    filename   = 'table_A_baseline.png',
    fstats     = {'(3) 2SLS': (wald_main, classical_main)},
)

make_reg_table(
    models     = [ols_dac, iv_model_mad_dac],
    col_labels = ['(1) OLS', '(2) 2SLS'],
    var_labels = VAR_LABELS,
    county_fe  = [True, True],
    month_fe   = [True, True],
    dep_var_label = 'Dependent Variable: Price Volatility (mad_rate)',
    filename   = 'table_B_dac.png',
    fstats     = {'(2) 2SLS': (wald_dac, classical_dac)},
)

make_reg_table(
    models     = [ols_income, iv_model_mad_income],
    col_labels = ['(1) OLS', '(2) 2SLS'],
    var_labels = VAR_LABELS,
    county_fe  = [True, True],
    month_fe   = [True, True],
    dep_var_label = 'Dependent Variable: Price Volatility (mad_rate)',
    filename   = 'table_C_income.png',
    fstats     = {'(2) 2SLS': (wald_income, classical_income)},
)
Cell [46]
import pandas as pd
import statsmodels.formula.api as smf
from linearmodels.iv import IV2SLS

# ==========================================
# MANUAL 2SLS: Showing the Two Stages Explicitly
# ==========================================
print("=" * 80)
print("MANUAL 2SLS VERIFICATION")
print("=" * 80)
print("\nThis cell manually runs the two stages to verify the IV2SLS package")
print("is using first-stage predictions correctly in the second stage.")
print("\n" + "=" * 80)

df = df_final.copy()
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)

num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")

controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw"]
control_str = " + ".join(controls)

# Prepare data
needed_cols = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]
data = df[needed_cols].dropna().copy()

# ==========================================
# STAGE 1: Predict evs_per_1000 using IV
# ==========================================
print("\nSTAGE 1: First-Stage Regression")
print("-" * 80)
print("Regressing: evs_per_1000 ~ iv_tesla_store_exposure_log + controls + FE")

first_stage_formula = (
    f"evs_per_1000 ~ iv_tesla_store_exposure_log + {control_str} + C(county_geoid) + C(month)"
)

first_stage_model = smf.ols(first_stage_formula, data=data).fit()

# Get predicted values from first stage
data["evs_per_1000_predicted"] = first_stage_model.fittedvalues

print(f"\nFirst-stage predictions created:")
print(f"  Mean predicted EV adoption: {data['evs_per_1000_predicted'].mean():.4f}")
print(f"  Std dev: {data['evs_per_1000_predicted'].std():.4f}")
print(f"  Min: {data['evs_per_1000_predicted'].min():.4f}")
print(f"  Max: {data['evs_per_1000_predicted'].max():.4f}")

# ==========================================
# STAGE 2: Use predictions in second stage
# ==========================================
print("\n\nSTAGE 2: Second-Stage Regression")
print("-" * 80)
print("Regressing: mad_rate ~ evs_per_1000_predicted + controls + FE")
print("\nNOTE: Standard errors from naive OLS would be WRONG here.")
print("      We're just showing the coefficient for comparison.")

second_stage_formula = (
    f"mad_rate ~ evs_per_1000_predicted + {control_str} + C(county_geoid) + C(month)"
)

second_stage_model = smf.ols(second_stage_formula, data=data).fit()

manual_2sls_coef = second_stage_model.params["evs_per_1000_predicted"]
manual_2sls_se_naive = second_stage_model.bse["evs_per_1000_predicted"]  # WRONG SE!

print(f"\nManual 2SLS coefficient: {manual_2sls_coef:.6f}")
print(f"Naive SE (INCORRECT): {manual_2sls_se_naive:.6f}")

# ==========================================
# COMPARE TO IV2SLS PACKAGE
# ==========================================
print("\n\n" + "=" * 80)
print("COMPARISON: Manual vs IV2SLS Package")
print("=" * 80)

# Run IV2SLS package version
iv_formula = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""

iv_model = IV2SLS.from_formula(iv_formula, data=data).fit(
    cov_type="clustered",
    clusters=data["county_geoid"],
)

package_coef = iv_model.params["evs_per_1000"]
package_se = iv_model.std_errors["evs_per_1000"]
package_pval = iv_model.pvalues["evs_per_1000"]

print("\nCoefficient Comparison:")
print(f"  Manual 2SLS:      {manual_2sls_coef:.6f}")
print(f"  IV2SLS Package:   {package_coef:.6f}")
print(f"  Difference:       {abs(manual_2sls_coef - package_coef):.10f}")

if abs(manual_2sls_coef - package_coef) < 0.000001:
    print("\n  ✓ COEFFICIENTS MATCH (within rounding error)")
    print("  ✓ The package IS using first-stage predictions correctly")
else:
    print("\n  ✗ Coefficients differ - something is wrong")

print("\nStandard Error Comparison:")
print(f"  Manual 2SLS (naive OLS SE):  {manual_2sls_se_naive:.6f}  ← WRONG!")
print(f"  IV2SLS Package (correct SE): {package_se:.6f}  ← CORRECT")
print(f"\n  The package computes proper 2SLS standard errors that account for")
print(f"  the two-stage estimation process. Naive OLS SE underestimates uncertainty.")

print("\nFinal Result:")
print(f"  Coefficient: {package_coef:.6f}")
print(f"  Std. Error:  {package_se:.6f}")
print(f"  P-value:     {package_pval:.4f}")

print("\n" + "=" * 80)
print("CONCLUSION: We are doing 2SLS correctly!")
print("=" * 80)
print("\nThe IV2SLS package:")
print("  ✓ Runs first stage internally")
print("  ✓ Uses predicted values in second stage")
print("  ✓ Computes correct standard errors (accounting for two-stage uncertainty)")
print("  ✓ Handles fixed effects and clustering properly")
print("\nYou can trust the previous 2SLS results - they're methodologically sound.")
print("=" * 80)
Cell [47]

# -------------------------
# Falsification IVs: Lag and Lead variants recomputed at source
# -------------------------
variants = {
    "lag2":   lambda m: m - pd.DateOffset(months=2),
    "lag3":   lambda m: m - pd.DateOffset(months=3),
    "lag6":   lambda m: m - pd.DateOffset(months=6),
    "lead3":  lambda m: m + pd.DateOffset(months=3),
    "lead6":  lambda m: m + pd.DateOffset(months=6),
    "lead12": lambda m: m + pd.DateOffset(months=12),
}

iv_falsification_tables = {}

for label, shift_fn in variants.items():
    rows = []
    is_lead = label.startswith("lead")

    for m in months:
        ref_month = shift_fn(m)

        if is_lead:
            open_stores_m = tesla_stores_df.loc[
                (tesla_stores_df["open_month"] > m) & (tesla_stores_df["open_month"] <= ref_month)
            ]
        else:
            open_stores_m = tesla_stores_df.loc[tesla_stores_df["open_month"] <= ref_month]

        if open_stores_m.empty:
            iv_m = counties[["county_geoid"]].copy()
            iv_m["idw_exposure"] = 0.0
            iv_m["month"] = m
            rows.append(iv_m)
            continue

        tmp_m = counties.assign(_k=1).merge(open_stores_m.assign(_k=1), on="_k").drop(columns="_k")
        tmp_m["dist_km"] = haversine_km(
            tmp_m["county_lat"], tmp_m["county_lon"],
            tmp_m["latitude"], tmp_m["longitude"]
        )
        tmp_m["dist_km_capped"] = np.maximum(tmp_m["dist_km"], eps_km)

        iv_m = tmp_m.groupby("county_geoid", as_index=False).agg(
            idw_exposure=("dist_km_capped", lambda s: float(np.sum(1.0 / s))),
        )
        iv_m["month"] = m
        rows.append(iv_m)

    if rows:
        tbl = pd.concat(rows, ignore_index=True)
        col_name = f"iv_tesla_store_exposure_log_{label}"
        tbl[col_name] = np.log1p(tbl["idw_exposure"])
        iv_falsification_tables[label] = tbl[["county_geoid", "month", col_name]]

# Merge all falsification IVs onto df_iv
df_iv_falsification = df_iv.copy()
for label, tbl in iv_falsification_tables.items():
    df_iv_falsification = df_iv_falsification.merge(
        tbl, on=["county_geoid", "month"], how="left", validate="m:1"
    )

falsification_cols = [c for c in df_iv_falsification.columns if "lag" in c or "lead" in c]
print(f"Falsification IVs merged onto df_iv_falsification ({len(df_iv_falsification):,} rows)")
for c in falsification_cols:
    nn = df_iv_falsification[c].notna().sum()
    print(f"  {c:<45} {nn:>6} non-null")

df_iv_falsification
Cell [48]

import pandas as pd

# -------------------------
# Robustness Dataset: lag/lead IV variants recomputed at source
# -------------------------
robustness_iv_cols = [
    "iv_tesla_store_exposure_log_lag2",
    "iv_tesla_store_exposure_log_lag3",
    "iv_tesla_store_exposure_log_lag6",
    "iv_tesla_store_exposure_log_lead3",
    "iv_tesla_store_exposure_log_lead6",
    "iv_tesla_store_exposure_log_lead12",
]

# Start from df_final columns + lag/lead IVs from df_iv_falsification
base_cols = list(df_final.columns)
extra_cols = [c for c in robustness_iv_cols if c in df_iv_falsification.columns]

# Merge the falsification IVs onto df_final via county_geoid + month
falsification_merge = df_iv_falsification[["county_geoid", "month"] + extra_cols].copy()
falsification_merge["county_geoid"] = falsification_merge["county_geoid"].astype(str)

df_final_falsification = df_final.merge(
    falsification_merge,
    on=["county_geoid", "month"],
    how="left",
    validate="m:1",
)

# Create interaction instruments for each lag/lead variant
for iv_col in extra_cols:
    suffix = iv_col.replace("iv_tesla_store_exposure_log", "")  # e.g. "_lag6"
    df_final_falsification[f"iv_x_dac{suffix}"] = (
        df_final_falsification[iv_col] * df_final_falsification["percent_dac_tracts"]
    )
    df_final_falsification[f"iv_x_income{suffix}"] = (
        df_final_falsification[iv_col] * df_final_falsification["median_household_income_10k"]
    )

print(f"df_final_falsification: {len(df_final_falsification):,} rows × {len(df_final_falsification.columns)} cols")
print(f"\nLag/lead IV columns carried through:")
for c in extra_cols:
    non_null = df_final_falsification[c].notna().sum()
    print(f"  {c:<45} {non_null:>6} non-null")

print(f"\nNew interaction instruments:")
for c in df_final_falsification.columns:
    if c.startswith("iv_x_") and c not in ["iv_x_dac", "iv_x_income"]:
        non_null = df_final_falsification[c].notna().sum()
        print(f"  {c:<45} {non_null:>6} non-null")

df_final_falsification
Cell [50]
ts = tesla_loc_df.copy()
ts["open_month"] = pd.to_datetime(ts["open_month"], errors="coerce").dt.to_period("M").dt.to_timestamp()
ts = ts.dropna(subset=["open_month"])

openings = ts.groupby("open_month").size().rename("new_stores")
cumulative = openings.cumsum().rename("cumulative_stores")

summary = pd.concat([openings, cumulative], axis=1).reset_index()
summary.rename(columns={"open_month": "month"}, inplace=True)

# Panel time range for reference
panel_min = pd.to_datetime(panel_df["month"], errors="coerce").min()
panel_max = pd.to_datetime(panel_df["month"], errors="coerce").max()

print(f"Panel range: {panel_min.strftime('%Y-%m')} to {panel_max.strftime('%Y-%m')}")
print(f"Total store openings: {openings.sum()}")
print(f"Openings within panel range: {openings.loc[(openings.index >= panel_min) & (openings.index <= panel_max)].sum()}")
print()
print(summary.to_string(index=False))
Cell [51]
from linearmodels.iv import IV2SLS
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

# -------------------------------------------------------------------
# LAG DECAY: Run 2SLS for lag-0 through lag-6 in 1-month increments
# Plot coefficient path to see exactly where the signal breaks down
# -------------------------------------------------------------------
controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
control_str = " + ".join(controls)

# Prep base panel (same as main spec)
df_base = df_final.copy()
df_base["month"] = pd.to_datetime(df_base["month"], errors="coerce")
df_base["month_str"] = df_base["month"].dt.to_period("M").astype(str)
df_base["county_geoid"] = df_base["county_geoid"].astype(str)
for c in ["mad_rate", "evs_per_1000"] + controls:
    df_base[c] = pd.to_numeric(df_base[c], errors="coerce")

# Precompute inverse distances (reuse county-level setup)
county_arr = counties[["county_geoid", "county_lat", "county_lon"]].reset_index(drop=True)
store_arr = tesla_stores_df[["latitude", "longitude", "open_month"]].reset_index(drop=True)
clat = county_arr["county_lat"].values[:, None]
clon = county_arr["county_lon"].values[:, None]
slat = store_arr["latitude"].values[None, :]
slon = store_arr["longitude"].values[None, :]
R = 6371.0
la1, lo1, la2, lo2 = map(np.radians, [clat, clon, slat, slon])
a = np.sin((la2 - la1) / 2)**2 + np.cos(la1) * np.cos(la2) * np.sin((lo2 - lo1) / 2)**2
inv_dist = 1.0 / np.maximum(2 * R * np.arcsin(np.sqrt(a)), 1.0)
inv_dist_county = inv_dist.copy()  # ← save county-level before anything overwrites inv_dist

county_ids = county_arr["county_geoid"].values
open_dates = pd.to_datetime(store_arr["open_month"]).values
panel_months = pd.DatetimeIndex(sorted(df_base["month"].dropna().unique()))

# Run lag-0 through lag-6
lag_results = []
for lag_k in range(7):
    # Compute IV with lag
    shifted_months = panel_months - pd.DateOffset(months=lag_k)
    is_open = (open_dates[:, None] <= shifted_months.values[None, :]).astype(float)
    exposure = inv_dist_county @ is_open  # ← use county-level matrix
    iv_log = np.log1p(exposure)

    # Build IV table
    cids = np.repeat(county_ids, len(panel_months))
    mths = np.tile(panel_months.values, len(county_ids))
    iv_df = pd.DataFrame({
        "county_geoid": cids.astype(str),
        "month": mths,
        "iv_lag": iv_log.ravel(),
    })

    # Merge and run 2SLS
    d = df_base[["county_geoid", "month", "month_str", "mad_rate", "evs_per_1000"] + controls].dropna().copy()
    d = d.merge(iv_df, on=["county_geoid", "month"], how="inner").dropna()
    for c in d.columns:
        if pd.api.types.is_numeric_dtype(d[c]):
            d[c] = d[c].astype(float)

    formula = f"mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month_str) [evs_per_1000 ~ iv_lag]"
    model = IV2SLS.from_formula(formula, data=d).fit(
        cov_type="clustered", clusters=d["county_geoid"]
    )
    coef = model.params["evs_per_1000"]
    se = model.std_errors["evs_per_1000"]
    pval = model.pvalues["evs_per_1000"]
    fstat = float(model.first_stage.diagnostics.loc["evs_per_1000", "f.stat"])
    lag_results.append({"Lag": lag_k, "Coef": coef, "SE": se, "P-value": pval, "F-stat": fstat, "N": int(model.nobs)})
    print(f"Lag-{lag_k}: coef={coef:.6f}  SE={se:.6f}  p={pval:.4f}  F={fstat:.1f}")

results_df = pd.DataFrame(lag_results)

Cell [52]
fig, ax = plt.subplots(figsize=(8, 5))

ax.errorbar(results_df["Lag"], results_df["Coef"],
            yerr=1.96 * results_df["SE"], fmt="o-", color="steelblue",
            capsize=4, linewidth=2, markersize=6)

ax.axhline(0, color="gray", linewidth=0.8, linestyle="-", alpha=0.4)

ax.set_xlabel("Lag (months)", fontsize=11)
ax.set_ylabel("2SLS Coefficient (evs_per_1000)", fontsize=11)
ax.set_title("Lag Decay: 2SLS Coefficient by IV Timing",
             fontsize=12, fontweight="bold")
ax.set_xticks(range(7))
ax.set_xticklabels([f"Lag-{k}" for k in range(7)])

ax.spines["top"].set_visible(False)
ax.spines["right"].set_visible(False)
ax.spines["left"].set_color("#666666")
ax.spines["bottom"].set_color("#666666")
ax.tick_params(colors="#666666")

plt.tight_layout()
plt.savefig("lag_decay.png", dpi=150, bbox_inches="tight")
plt.show()
display(results_df)
Cell [53]
from linearmodels.iv import IV2SLS
import numpy as np
import pandas as pd

df_rob = df_final_falsification.copy()
df_rob["month"] = pd.to_datetime(df_rob["month"], errors="coerce").dt.to_period("M").astype(str)
df_rob["county_geoid"] = df_rob["county_geoid"].astype(str)
num_cols = [c for c in df_rob.columns if c not in ["county_geoid", "county_name", "month"]]
df_rob[num_cols] = df_rob[num_cols].apply(pd.to_numeric, errors="coerce").astype(float)

controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
control_str = " + ".join(controls)

# --- Run each spec ---
specs = {
    "Main (contemporaneous IV)": "iv_tesla_store_exposure_log",
    "Lag-2 IV (exposure 2mo prior)": "iv_tesla_store_exposure_log_lag2",
    "Lag-3 IV (exposure 3mo prior)": "iv_tesla_store_exposure_log_lag3",
    "Lag-6 IV (exposure 6mo prior)": "iv_tesla_store_exposure_log_lag6",
}

rows = []
for label, iv_col in specs.items():
    cols = ["mad_rate", "evs_per_1000", iv_col] + controls + ["county_geoid", "month"]
    d = df_rob[cols].dropna().copy()
    for c in d.columns:
        if pd.api.types.is_numeric_dtype(d[c]):
            d[c] = d[c].astype(float)

    formula = f"""
    mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
    [evs_per_1000 ~ {iv_col}]
    """
    model = IV2SLS.from_formula(formula, data=d).fit(
        cov_type="clustered", clusters=d["county_geoid"]
    )

    rows.append({
        "Specification": label,
        "2SLS Coef": round(model.params["evs_per_1000"], 4),
        "SE (clustered)": round(model.std_errors["evs_per_1000"], 4),
        "P-value": round(model.pvalues["evs_per_1000"], 4),
        "First-stage F": round(float(model.first_stage.diagnostics.loc["evs_per_1000", "f.stat"]), 2),
        "N": int(model.nobs),
    })

summary = pd.DataFrame(rows)

print("=" * 80)
print("LAG ROBUSTNESS: mad_rate ~ evs_per_1000 (instrumented)")
print("=" * 80)
print(f"Controls: {', '.join(controls)}")
print(f"FEs: county + month | SEs: clustered by county")
print()
display(summary)
Cell [56]
import statsmodels.formula.api as smf

# -------------------------------------------------------------------
# LEAD FALSIFICATION: Reduced-form OLS (NOT 2SLS)
#
# Purpose: Test whether proximity to stores that don't exist yet
# predicts current electricity market outcomes.
#
# Construction: Each lead variable measures IDW exposure ONLY to
# stores opening strictly after month m but within k months.
# These stores are not yet open — they cannot causally affect
# current EV adoption or electricity prices.
#
# Why OLS, not 2SLS: We are not instrumenting anything. We are
# directly testing whether future store openings (which shouldn't
# matter yet) correlate with current outcomes. This is a pure
# reduced-form check of the exclusion restriction.
#
# Interpretation:
#   - Insignificant lead coefficients → future stores don't predict
#     current outcomes → supports instrument validity
#   - Significant lead coefficients → counties that will later get
#     stores were already trending differently → confounding concern
#
# Caveat: With only 29 store openings in the panel, these tests
# are underpowered. A null result is necessary but not sufficient
# evidence of instrument validity.
# -------------------------------------------------------------------

lead_ivs = [
    "iv_tesla_store_exposure_log_lead3",   # stores opening in next 3 months
    "iv_tesla_store_exposure_log_lead6",   # stores opening in next 6 months
    "iv_tesla_store_exposure_log_lead12",  # stores opening in next 12 months
]

print("=" * 80)
print("LEAD IVs: REDUCED-FORM FALSIFICATION (should be insignificant)")
print("=" * 80)

falsification_results = []

outcomes = ["mad_rate"]

print(f"\ndf_rob shape: {df_rob.shape}")
print(f"Controls: {controls}")
print(f"control_str: {control_str}")
for lc in lead_ivs:
    present = lc in df_rob.columns
    nn = df_rob[lc].notna().sum() if present else 0
    print(f"  {lc}: present={present}, non-null={nn}")

for outcome in outcomes:
    for lead_col in lead_ivs:
        cols = [outcome, lead_col] + controls + ["county_geoid", "month"]
        missing_cols = [c for c in cols if c not in df_rob.columns]
        if missing_cols:
            print(f"\n⚠ {lead_col}: missing columns in df_rob: {missing_cols}")
            continue

        d = df_rob[cols].dropna().copy()
        print(f"\n{lead_col}: {len(d)} rows after dropna (from {len(df_rob)})")

        if len(d) == 0:
            print(f"  ⚠ No rows remaining — skipping")
            continue

        for c in d.columns:
            if pd.api.types.is_numeric_dtype(d[c]):
                d[c] = d[c].astype(float)

        formula = f"{outcome} ~ {lead_col} + {control_str} + C(county_geoid) + C(month)"
        model = smf.ols(formula, data=d).fit(
            cov_type="cluster", cov_kwds={"groups": d["county_geoid"]}
        )

        coef = model.params[lead_col]
        se = model.bse[lead_col]
        pval = model.pvalues[lead_col]
        sig = "*" if pval < 0.1 else ""

        falsification_results.append({
            "Outcome": outcome,
            "Lead IV": lead_col.replace("iv_tesla_store_exposure_log_", ""),
            "Coef": round(coef, 4),
            "SE": round(se, 4),
            "P-value": round(pval, 4),
            "Sig": sig,
            "N": int(model.nobs),
        })

falsification_df = pd.DataFrame(falsification_results)
print("\n")
print(falsification_df.to_string(index=False))
print("\n* = p < 0.10. Significant leads suggest confounding / pre-trends.")
Cell [59]
import numpy as np
import pandas as pd

np.random.seed(42)

# Original store opening dates
store_arr = tesla_stores_df[["latitude", "longitude", "open_month"]].copy()
store_arr["open_month"] = pd.to_datetime(store_arr["open_month"], errors="coerce")
store_arr = store_arr.dropna(subset=["open_month"]).reset_index(drop=True)

real_open = store_arr["open_month"].values
print(f"Stores: {len(real_open)}")
print(f"Date range: {pd.Timestamp(real_open.min()).strftime('%Y-%m')} to {pd.Timestamp(real_open.max()).strftime('%Y-%m')}")

# One example shuffle — dates reassigned across locations
shuffled_open = np.random.permutation(real_open)

# Show a few rows to confirm
check = pd.DataFrame({
    "lat": store_arr["latitude"].values[:10],
    "lon": store_arr["longitude"].values[:10],
    "real_open": pd.to_datetime(real_open[:10]),
    "shuffled_open": pd.to_datetime(shuffled_open[:10]),
})
print("\nSample (first 10 stores):")
print(check.to_string(index=False))
Cell [60]
# Precompute pairwise inverse distances (fixed — locations don't change)
county_arr = counties[["county_geoid", "county_lat", "county_lon"]].reset_index(drop=True)

clat = county_arr["county_lat"].values[:, None]
clon = county_arr["county_lon"].values[:, None]
slat = store_arr["latitude"].values[None, :]
slon = store_arr["longitude"].values[None, :]

R = 6371.0
la1, lo1, la2, lo2 = map(np.radians, [clat, clon, slat, slon])
a = np.sin((la2 - la1) / 2)**2 + np.cos(la1) * np.cos(la2) * np.sin((lo2 - lo1) / 2)**2
inv_dist = 1.0 / np.maximum(2 * R * np.arcsin(np.sqrt(a)), 1.0)
inv_dist_county = inv_dist.copy()

print(f"inv_dist shape: {inv_dist.shape}  (counties × stores)")

# Compute IV for one set of opening dates
panel_months = pd.DatetimeIndex(sorted(df_final["month"].dropna().unique()))

def compute_iv_table(open_dates):
    """opening dates → DataFrame of (county_geoid, month, iv_perm_log)"""
    is_open = (open_dates[:, None] <= panel_months.values[None, :]).astype(float)
    exposure = inv_dist_county @ is_open  # was inv_dist
    iv_log = np.log1p(exposure)

    # Reshape to long DataFrame
    cids = np.repeat(county_arr["county_geoid"].values, len(panel_months))
    mths = np.tile(panel_months.values, len(county_arr))
    vals = iv_log.ravel()

    return pd.DataFrame({"county_geoid": cids, "month": mths, "iv_perm_log": vals})

# Test with real dates
iv_real = compute_iv_table(real_open)
print(f"\nIV table: {len(iv_real)} rows")
print(iv_real.head())
Cell [61]
import statsmodels.formula.api as smf

# -------------------------------------------------------------------
# REDUCED-FORM REGRESSION: REAL STORE OPENING DATES
#
# This is the baseline for the permutation test.
# We regress mad_rate directly on the IV (Tesla store exposure)
# with controls and two-way FEs. The coefficient tells us whether
# the IV predicts price volatility in the actual data.
#
# The permutation test will then ask: does this coefficient remain
# significant when we randomly reassign opening dates to stores?
# -------------------------------------------------------------------

df_reg = df_final[["county_geoid", "month", "mad_rate",
                    "tmax_c", "cumulative_data_center_power_mw",
                    "total_generator_active_capacity_mw",
                    "solar_mw_per_1000_lag1"]].copy()
df_reg["county_geoid"] = df_reg["county_geoid"].astype(str)
df_reg["month"] = pd.to_datetime(df_reg["month"], errors="coerce")

iv_real["county_geoid"] = iv_real["county_geoid"].astype(str)

df_reg = df_reg.merge(iv_real, on=["county_geoid", "month"], how="inner")
df_reg = df_reg.dropna().reset_index(drop=True)
df_reg["month_str"] = df_reg["month"].dt.to_period("M").astype(str)

controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
formula = ("mad_rate ~ iv_perm_log + tmax_c + cumulative_data_center_power_mw"
           " + total_generator_active_capacity_mw + solar_mw_per_1000_lag1 + C(county_geoid) + C(month_str)")

model = smf.ols(formula, data=df_reg).fit(
    cov_type="cluster", cov_kwds={"groups": df_reg["county_geoid"]}
)

real_coef = model.params["iv_perm_log"]
real_se = model.bse["iv_perm_log"]
real_pval = model.pvalues["iv_perm_log"]

print("=" * 60)
print("BASELINE: REDUCED-FORM WITH REAL OPENING DATES")
print("=" * 60)
print(f"Outcome:          mad_rate (price volatility)")
print(f"Key regressor:    iv_perm_log (log IDW Tesla store exposure)")
print(f"Controls:         {', '.join(controls)}")
print(f"Fixed effects:    county + month")
print(f"Standard errors:  clustered by county")
print(f"N:                {int(model.nobs)}")
print(f"")
print(f"Coefficient:      {real_coef:.6f}")
print(f"SE (clustered):   {real_se:.6f}")
print(f"P-value:          {real_pval:.4f}")
print(f"Significant:      {'Yes (p < 0.05)' if real_pval < 0.05 else 'No'}")
print(f"")
print(f"Interpretation:   A one-unit increase in log Tesla store")
print(f"                  exposure is associated with a {real_coef:.4f}")
print(f"                  increase in mad_rate, after absorbing")
print(f"                  county and month fixed effects.")
print("=" * 60)
Cell [62]
N_PERMS = 500
perm_coefs = []

print(f"Running {N_PERMS} permutations...")

for i in range(N_PERMS):
    shuffled = np.random.permutation(real_open)
    iv_perm = compute_iv_table(shuffled)
    iv_perm["county_geoid"] = iv_perm["county_geoid"].astype(str)

    df_p = df_final[["county_geoid", "month", "mad_rate",
                      "tmax_c", "cumulative_data_center_power_mw",
                      "total_generator_active_capacity_mw",
                      "solar_mw_per_1000_lag1"]].copy()
    df_p["county_geoid"] = df_p["county_geoid"].astype(str)
    df_p["month"] = pd.to_datetime(df_p["month"], errors="coerce")
    df_p = df_p.merge(iv_perm, on=["county_geoid", "month"], how="inner").dropna()
    df_p["month_str"] = df_p["month"].dt.to_period("M").astype(str)

    try:
        m = smf.ols(formula, data=df_p).fit(
            cov_type="cluster", cov_kwds={"groups": df_p["county_geoid"]}
        )
        perm_coefs.append(m.params["iv_perm_log"])
    except Exception as e:
        if i == 0:  # only print error on first failure to avoid spam
            print(f"Model error (iteration {i}): {e}")

    if (i + 1) % 100 == 0:
        print(f"  {i + 1}/{N_PERMS} done ({len(perm_coefs)} valid)")

perm_coefs = np.array(perm_coefs)
p_two_sided = np.mean(np.abs(perm_coefs) >= np.abs(real_coef))

Cell [63]
print(f"\nReal coef:          {real_coef:.6f}")
print(f"Valid permutations: {len(perm_coefs)}/{N_PERMS}")
print(f"Permutation p-value (two-sided): {p_two_sided:.4f}")

fig, ax = plt.subplots(figsize=(8, 4))

ax.hist(perm_coefs, bins=40, color="steelblue", alpha=0.7, edgecolor="white")
ax.axvline(real_coef, color="red", linewidth=2, label=f"Real = {real_coef:.3f}")

ax.set_xlabel("Reduced-form coefficient (mad_rate ~ IV)")
ax.set_ylabel("Count")
ax.set_title(f"Permutation Test (n={len(perm_coefs)}, p={p_two_sided:.3f})")
ax.legend()

ax.spines["top"].set_visible(False)
ax.spines["right"].set_visible(False)
ax.spines["left"].set_color("#666666")
ax.spines["bottom"].set_color("#666666")
ax.tick_params(colors="#666666")

plt.tight_layout()
plt.savefig("permutation_test.png", dpi=150, bbox_inches="tight")
plt.show()
Cell [65]
from linearmodels.iv import IV2SLS
import numpy as np
import pandas as pd

# -------------------------------------------------------------------
# COMPARISON: Dummy Variable 2SLS vs Demeaned 2SLS
#
# Both estimate:
#   mad_rate ~ evs_per_1000 (instrumented) + controls + county FE + month FE
#
# Dummy: includes FEs as explicit regressors
# Demeaned: iteratively absorbs county + month means from all
#   variables, then runs IV2SLS on residualized data
# -------------------------------------------------------------------

df = df_final.copy()
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)

num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")

controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
control_str = " + ".join(controls)

iv_cols = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]
data = df[iv_cols].dropna().copy()
for c in data.columns:
    if pd.api.types.is_numeric_dtype(data[c]):
        data[c] = data[c].astype(float)

# ==========================================
# Method 1: Dummy Variable 2SLS (current)
# ==========================================
formula_dummy = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""

model_dummy = IV2SLS.from_formula(formula_dummy, data=data).fit(
    cov_type="clustered", clusters=data["county_geoid"]
)

# ==========================================
# Method 2: Demeaned 2SLS
# ==========================================

# Iterative two-way demeaning to absorb county + month FEs
def demean_twoway(vals, entity, time, max_iter=100, tol=1e-10):
    s = vals.copy().astype(float)
    ent = pd.Series(entity)
    tm = pd.Series(time)
    for _ in range(max_iter):
        old = s.copy()
        s -= pd.Series(s).groupby(ent).transform("mean").values
        s -= pd.Series(s).groupby(tm).transform("mean").values
        if np.max(np.abs(s - old)) < tol:
            break
    return s

entity = data["county_geoid"].values
time = data["month"].values

# Demean all variables
vars_to_demean = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls
data_dm = pd.DataFrame()
for v in vars_to_demean:
    data_dm[v] = demean_twoway(data[v].values, entity, time)

# Run IV2SLS on demeaned data (no FEs needed)
formula_dm = f"""
mad_rate ~ {control_str}
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""

# Clustered SEs still need the original county IDs
model_dm = IV2SLS.from_formula(formula_dm, data=data_dm).fit(
    cov_type="clustered", clusters=data["county_geoid"]
)

# ==========================================
# Comparison Table
# ==========================================
n_fe = data["county_geoid"].nunique() + data["month"].nunique()

summary = pd.DataFrame([
    {
        "Method": "Dummy Variable 2SLS",
        "Coef": round(model_dummy.params["evs_per_1000"], 6),
        "SE (clustered)": round(model_dummy.std_errors["evs_per_1000"], 6),
        "P-value": round(model_dummy.pvalues["evs_per_1000"], 4),
        "First-stage F": round(float(model_dummy.first_stage.diagnostics.loc["evs_per_1000", "f.stat"]), 2),
        "N": int(model_dummy.nobs),
    },
    {
        "Method": "Demeaned 2SLS",
        "Coef": round(model_dm.params["evs_per_1000"], 6),
        "SE (clustered)": round(model_dm.std_errors["evs_per_1000"], 6),
        "P-value": round(model_dm.pvalues["evs_per_1000"], 4),
        "First-stage F": round(float(model_dm.first_stage.diagnostics.loc["evs_per_1000", "f.stat"]), 2),
        "N": int(model_dm.nobs),
    },
])

print("=" * 70)
print("COMPARISON: DUMMY VARIABLE vs DEMEANED 2SLS")
print("=" * 70)
print(f"Outcome:     mad_rate")
print(f"Endogenous:  evs_per_1000")
print(f"Instrument:  iv_tesla_store_exposure_log")
print(f"Controls:    {', '.join(controls)}")
print(f"FEs:         county + month (~{n_fe} parameters)")
print(f"SEs:         clustered by county")
print()
display(summary)

se_dummy = model_dummy.std_errors["evs_per_1000"]
se_dm = model_dm.std_errors["evs_per_1000"]
pct_diff = 100 * (se_dm - se_dummy) / se_dummy

print(f"\nCoefficient difference: {abs(model_dummy.params['evs_per_1000'] - model_dm.params['evs_per_1000']):.8f}")
print(f"SE difference: {pct_diff:+.1f}%")
if abs(pct_diff) < 5:
    print("Negligible — both methods yield effectively the same inference.")
else:
    print(f"Demeaned SEs are {'tighter' if pct_diff < 0 else 'wider'} — "
          f"degrees-of-freedom correction matters with {len(data)} obs "
          f"and ~{n_fe} FE parameters.")
Cell [67]
import geopandas as gpd
from shapely import wkt

tract_geo = gpd.GeoDataFrame(
    tract_shape_df,
    geometry=tract_shape_df["geometry"].apply(wkt.loads),
    crs="EPSG:4326",
)

tract_geo = tract_geo.to_crs("EPSG:3310")
centroids = tract_geo.geometry.centroid.to_crs("EPSG:4326")
tract_geo["tract_lat"] = centroids.y
tract_geo["tract_lon"] = centroids.x

tract_points = tract_geo[["census_tract", "tract_lat", "tract_lon"]].copy()
print(f"Tract centroids: {len(tract_points)}")
print(tract_points.head())
Cell [68]
import numpy as np

# Prep Tesla stores
tesla_stores = tesla_loc_df[["latitude", "longitude", "open_month"]].copy()
tesla_stores["latitude"] = pd.to_numeric(tesla_stores["latitude"], errors="coerce")
tesla_stores["longitude"] = pd.to_numeric(tesla_stores["longitude"], errors="coerce")
tesla_stores["open_month"] = pd.to_datetime(tesla_stores["open_month"], errors="coerce").dt.to_period("M").dt.to_timestamp()
tesla_stores = tesla_stores.dropna().reset_index(drop=True)

# Pairwise inverse distances: (n_tracts, n_stores)
tlat = tract_points["tract_lat"].values[:, None]
tlon = tract_points["tract_lon"].values[:, None]
slat = tesla_stores["latitude"].values[None, :]
slon = tesla_stores["longitude"].values[None, :]

R = 6371.0
la1, lo1, la2, lo2 = map(np.radians, [tlat, tlon, slat, slon])
a = np.sin((la2 - la1) / 2)**2 + np.cos(la1) * np.cos(la2) * np.sin((lo2 - lo1) / 2)**2
dist = 2 * R * np.arcsin(np.sqrt(a))
inv_dist = 1.0 / np.maximum(dist, 1.0)

print(f"inv_dist shape: {inv_dist.shape}  (tracts × stores)")

# Compute IV for each tract-month
panel_months = pd.DatetimeIndex(sorted(panel_df["month"].dropna().unique()))
open_dates = tesla_stores["open_month"].values

is_open = (open_dates[:, None] <= panel_months.values[None, :]).astype(float)  # (stores, months)
exposure = inv_dist @ is_open  # (tracts, months)

# Reshape to long DataFrame
tract_ids = np.repeat(tract_points["census_tract"].values, len(panel_months))
months_tile = np.tile(panel_months.values, len(tract_points))

iv_tract = pd.DataFrame({
    "census_tract": tract_ids,
    "month": months_tile,
    "tesla_store_exposure": exposure.ravel(),
    "iv_tesla_store_exposure_log": np.log1p(exposure.ravel()),
})

print(f"IV table: {len(iv_tract):,} rows")
print(iv_tract.head())
Cell [69]
print(tract_points.columns.tolist())
print(tract_points.head(3))
Cell [70]
from linearmodels.iv import IV2SLS

df = tract_panel_df.copy()
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["census_tract"] = df["census_tract"].astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)

num_cols = [c for c in df.columns if c not in ["census_tract", "county_geoid", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")

# Controls (dropping tmax_c and data center — not available at tract level)
controls = ["total_active_capacity_mw"]
control_str = " + ".join(controls)

cols = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["census_tract", "county_geoid", "month"]
data = df[cols].dropna().copy()
for c in data.columns:
    if pd.api.types.is_numeric_dtype(data[c]):
        data[c] = data[c].astype(float)

# --- Tract FE + Month FE, clustered by county ---
formula = f"""
mad_rate ~ 1 + {control_str} + C(census_tract) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""

model = IV2SLS.from_formula(formula, data=data).fit(
    cov_type="clustered", clusters=data["county_geoid"]
)

print("=" * 70)
print("TRACT-LEVEL 2SLS: mad_rate ~ evs_per_1000 (instrumented)")
print("=" * 70)
print(f"FEs:         tract + month")
print(f"Controls:    {', '.join(controls)}")
print(f"SEs:         clustered by county ({data['county_geoid'].nunique()} clusters)")
print(f"N:           {int(model.nobs)}")
print(f"Tracts:      {data['census_tract'].nunique()}")
print(f"")
print(f"Coef:        {model.params['evs_per_1000']:.6f}")
print(f"SE:          {model.std_errors['evs_per_1000']:.6f}")
print(f"P-value:     {model.pvalues['evs_per_1000']:.4f}")
print(f"First-stg F: {float(model.first_stage.diagnostics.loc['evs_per_1000', 'f.stat']):.2f}")
print("=" * 70)
#print(model.summary.tables[1])
Cell [71]
print([c for c in tract_panel_df.columns if "iv" in c.lower() or "tesla" in c.lower() or "exposure" in c.lower()])
Cell [74]

# ── mad_rate distribution stats ───────────────────────────────────────────────
mad_stats  = df_final['mad_rate'].describe(percentiles=[0.25, 0.5, 0.75])
mad_std    = df_final['mad_rate'].std()
mad_iqr    = df_final['mad_rate'].quantile(0.75) - df_final['mad_rate'].quantile(0.25)
mad_median = df_final['mad_rate'].median()

# ── Income percentiles (county-level, one obs per county) ─────────────────────
income_by_county = df_final.groupby('county_geoid')['median_household_income_10k'].first()
p10_income = income_by_county.quantile(0.10)
p25_income = income_by_county.quantile(0.25)
p75_income = income_by_county.quantile(0.75)
p90_income = income_by_county.quantile(0.90)

# ── EV adoption distribution ──────────────────────────────────────────────────
ev_stats = df_final['evs_per_1000'].describe(percentiles=[0.25, 0.5, 0.75, 0.90, 0.95])

# ── Pull coefficients directly from model object ───────────────────────────────
coef_ev    = iv_model_mad_income.params['evs_per_1000']
coef_int   = iv_model_mad_income.params['ev_x_income']
breakeven  = -coef_ev / coef_int

print("=== Income-interaction 2SLS coefficients ===")
print(f"Main EV coef:        {coef_ev:.6f}")
print(f"Interaction coef:    {coef_int:.6f}")
print(f"Break-even income:   ${breakeven * 10:.0f}k")

# ── Marginal effect function ───────────────────────────────────────────────────
def marginal_effect(income_10k):
    return coef_ev + coef_int * income_10k

# ── Marginal effects at income percentiles ────────────────────────────────────
print("\n=== Marginal EV effect on mad_rate by income percentile ===")
income_anchors = [
    ("P10", p10_income),
    ("P25", p25_income),
    ("P75", p75_income),
    ("P90", p90_income),
]
for label, inc in income_anchors:
    me = marginal_effect(inc)
    scaled = me * 50
    print(f"  {label} (${inc*10:.0f}k): marginal = {me:.6f}, scaled (50 EVs) = {scaled:.4f} $/MWh  ({scaled/mad_std:.2f} std devs)")

# ── P25 vs P75 ratio ──────────────────────────────────────────────────────────
me_p25 = marginal_effect(p25_income)
me_p75 = marginal_effect(p75_income)
me_p10 = marginal_effect(p10_income)
me_p90 = marginal_effect(p90_income)

print(f"\nP25/P75 effect ratio: {me_p25/me_p75:.2f}x")
print(f"P10/P90 effect ratio: {me_p10/me_p90:.2f}x")

# ── Scaled effect context (50 EV scenario at median income) ───────────────────
me_median = marginal_effect(income_by_county.median())
scaled_median = me_median * 50

print(f"\n=== 50-EV scenario at median county income (${income_by_county.median()*10:.0f}k) ===")
print(f"Implied mad_rate increase: {scaled_median:.4f} $/MWh")
print(f"As % of std dev:           {scaled_median/mad_std:.2%}")
print(f"As % of IQR:               {scaled_median/mad_iqr:.2%}")
print(f"As % of median:            {scaled_median/mad_median:.2%}")

# ── EV adoption context (justifying the 50-EV scenario) ───────────────────────
print("\n=== evs_per_1000 distribution (justifying 50-EV scenario) ===")
print(ev_stats)
Cell [75]
coef_baseline = iv_model_mad.params['evs_per_1000']
delta_ev = 50
scaled = coef_baseline * delta_ev

print(f"Baseline coef:        {coef_baseline:.6f}")
print(f"Scaled (50 EVs):      {scaled:.4f} $/MWh")
print(f"As % of std dev:      {scaled/mad_std:.2%}")
print(f"As % of IQR:          {scaled/mad_iqr:.2%}")
print(f"As % of median:       {scaled/mad_median:.2%}")
print(f"\n90th pctile adoption: {ev_stats['90%']:.1f} EVs/1,000")
print(f"75th pctile adoption: {ev_stats['75%']:.1f} EVs/1,000")
Cell [76]
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import numpy as np

# ── Pull coefficients from model ───────────────────────────────────────────────
coef_ev  = iv_model_mad_income.params['evs_per_1000']
coef_int = iv_model_mad_income.params['ev_x_income']
se_ev    = iv_model_mad_income.std_errors['evs_per_1000']
se_int   = iv_model_mad_income.std_errors['ev_x_income']
cov_ev_int = iv_model_mad_income.cov.loc['evs_per_1000', 'ev_x_income']

# ── Income range for plot ──────────────────────────────────────────────────────
income_min = income_by_county.min()
income_max = income_by_county.max()
income_grid = np.linspace(income_min, income_max, 200)

# ── Marginal effect and 95% CI ─────────────────────────────────────────────────
me     = coef_ev + coef_int * income_grid
me_var = se_ev**2 + income_grid**2 * se_int**2 + 2 * income_grid * cov_ev_int
me_se  = np.sqrt(me_var)
me_upper = me + 1.96 * me_se
me_lower = me - 1.96 * me_se

breakeven_k = (-coef_ev / coef_int) * 10

# ── Plot ───────────────────────────────────────────────────────────────────────
fig, ax = plt.subplots(figsize=(8, 4.5))

# CI band
ax.fill_between(income_grid * 10, me_lower, me_upper,
                alpha=0.15, color='steelblue', label='95% Confidence Interval')

# Main line
ax.plot(income_grid * 10, me, color='steelblue', linewidth=2,
        label='Marginal Effect of EV Adoption')

# Zero line (solid gray)
ax.axhline(0, color='gray', linewidth=0.8, linestyle='-', alpha=0.4)

# Break-even line (red dotted)
ax.axvline(breakeven_k, color='firebrick', linewidth=1.2, linestyle=':',
           label=f'Break-even income (${breakeven_k:.0f}k)')

# Income percentile markers (solid gray, low alpha)
markers = {
    'P25 ($57k)': p25_income,
    'P75 ($85k)': p75_income,
}
for label, val in markers.items():
    ax.axvline(val * 10, color='gray', linewidth=0.8, linestyle='-', alpha=0.4)
    ax.text(val * 10, me_upper.max() * 1.02, label,
            ha='center', va='bottom', fontsize=7.5, color='gray')

# Rug plot
county_incomes_k = income_by_county.values * 10
ax.plot(county_incomes_k,
        np.full_like(county_incomes_k, me_lower.min() * 0.85),
        '|', color='steelblue', alpha=0.5, markersize=8, markeredgewidth=1.2,
        label='California counties')

# Notable county labels
county_income_df = df_final.groupby('county_name')['median_household_income_10k'].first()
label_counties = ['Los Angeles', 'San Francisco', 'Fresno', 'Santa Clara', 'Tulare']

for county in label_counties:
    if county in county_income_df.index:
        inc_k = county_income_df[county] * 10
        me_at_inc = coef_ev + coef_int * (inc_k / 10)
        ax.annotate(county, xy=(inc_k, me_at_inc),
                    xytext=(inc_k, me_at_inc + 0.0004),
                    fontsize=7, color='dimgray', ha='center',
                    arrowprops=dict(arrowstyle='-', color='lightgray', lw=0.8))

# Labels and formatting
ax.set_xlabel('Median Household Income ($k)', fontsize=11)
ax.set_ylabel('Marginal Effect on mad_rate\n($/MWh per EV per 1,000)', fontsize=10)
ax.set_title('Marginal Effect of EV Adoption on Price Volatility\nby County Income Level (2SLS Estimates)',
             fontsize=12, fontweight='bold')
ax.legend(fontsize=9, framealpha=0.9)

xticks = np.arange(np.floor(income_min * 10 / 10) * 10,
                   np.ceil(income_max * 10 / 10) * 10 + 1, 10)
ax.set_xticks(xticks)
ax.set_xticklabels([f'${int(x)}k' for x in xticks])

ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
plt.tight_layout()
plt.savefig('ev_income_marginal_effect.png', dpi=150, bbox_inches='tight')
plt.show()
Cell [77]
# Filter out zero or near-zero counties for cleaner plot
dac_plot = (df_final.groupby('county_name')['percent_dac_tracts']
            .first()
            .sort_values(ascending=True))

# Drop counties with 0% DAC share, add a note instead
dac_plot_filtered = dac_plot[dac_plot > 0]
n_zero = (dac_plot == 0).sum()

fig, ax = plt.subplots(figsize=(8, 7))  # shorter now

colors = ['firebrick' if v >= 50 else 'steelblue' if v >= 25 else 'lightsteelblue'
          for v in dac_plot_filtered.values]

bars = ax.barh(dac_plot_filtered.index, dac_plot_filtered.values,
               color=colors, edgecolor='none', height=0.7)

# Reference lines
ax.axvline(25, color='gray', linewidth=0.8, linestyle='--', alpha=0.5)
ax.axvline(50, color='gray', linewidth=0.8, linestyle='--', alpha=0.5)
ax.text(25.5, -1.2, '25%', fontsize=7.5, color='gray', va='top')
ax.text(50.5, -1.2, '50%', fontsize=7.5, color='gray', va='top')

# Value labels
for bar, val in zip(bars, dac_plot_filtered.values):
    ax.text(val + 0.5, bar.get_y() + bar.get_height()/2,
            f'{val:.0f}%', va='center', ha='left', fontsize=7.5, color='dimgray')

# Note about zero counties
ax.text(0, -2.5, f'Note: {n_zero} counties have 0% DAC-designated tracts and are excluded.',
        fontsize=7.5, color='gray', style='italic')

ax.set_xlabel('Share of Census Tracts Designated as Disadvantaged (%)', fontsize=10)
ax.set_title('DAC Intensity by California County\n(Share of Disadvantaged Community Tracts)',
             fontsize=12, fontweight='bold')
ax.set_xlim(0, dac_plot_filtered.max() * 1.15)

import matplotlib.patches as mpatches
legend_elements = [
    mpatches.Patch(color='firebrick', label='≥50% DAC tracts'),
    mpatches.Patch(color='steelblue', label='25–50% DAC tracts'),
    mpatches.Patch(color='lightsteelblue', label='<25% DAC tracts'),
]
ax.legend(handles=legend_elements, fontsize=8, loc='lower right')

ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
ax.spines['left'].set_visible(False)
ax.tick_params(left=False)
plt.tight_layout()
plt.savefig('dac_share_by_county.png', dpi=150, bbox_inches='tight')
plt.show()
Cell [78]
county_summary = (df_final.groupby('county_name')
                  .agg(
                      dac_share=('percent_dac_tracts', 'first'),
                      median_income_10k=('median_household_income_10k', 'first'),
                      mean_evs_per_1000=('evs_per_1000', 'mean'),
                      mean_mad_rate=('mad_rate', 'mean')
                  )
                  .reset_index()
                  .sort_values('dac_share', ascending=False))

# High DAC + low income counties
print("=== High DAC share + low income ===")
print(county_summary[county_summary['dac_share'] > 20]
      [['county_name', 'dac_share', 'median_income_10k', 'mean_evs_per_1000']]
      .sort_values('median_income_10k').to_string())
CVRP shift-share – ev_iv_cvrp_shift_share.ipynb46 cells
Cell [2]
from google.colab import drive
drive.mount('/content/drive')

!pip install linearmodels
!pip install pandas numpy

# Setup & Imports
import os
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
import statsmodels.api as sm
from linearmodels.iv import IV2SLS
import warnings

warnings.filterwarnings('ignore', category=pd.errors.SettingWithCopyWarning)

# Define DATA_DIR - Modify this path as needed for your Colab environment
from pathlib import Path

DATA_DIR = Path(
    "/content/drive/MyDrive/Capstone_Datasets"
)

DATA_DIR.exists()

# Initialize a list to store all regression results for the Step 5 summary table
all_model_results = []

def print_sanity_check(df, df_name="Dataset"):
    if 'county_name' in df.columns:
        print(f"[{df_name}] Sanity Check - Current N: {len(df)}, Unique Counties: {df['county_name'].nunique()}")
    else:
        print(f"[{df_name}] Sanity Check - Current N: {len(df)}")
Cell [4]
# STEP 0A — COMPUTE RELIANCE WEIGHT
print("--- STEP 0A: COMPUTE RELIANCE WEIGHT ---")

try:
    cvrp_path = os.path.join(DATA_DIR, "CVRPStats_data.csv")
    cvrp_df = pd.read_csv(cvrp_path)

    # Clean Air District (actually county)
    cvrp_df['county_name'] = cvrp_df['Air District'].astype(str).str.replace(' County', '', case=False).str.title().str.strip()

    # Parse Application Date
    cvrp_df['Application Date'] = pd.to_datetime(cvrp_df['Application Date'], format='%m/%d/%Y', errors='coerce')

    # Filter to pre-treatment baseline (2019-2021)
    cvrp_baseline = cvrp_df[(cvrp_df['Application Date'].dt.year >= 2019) &
                            (cvrp_df['Application Date'].dt.year <= 2021)]

    # Identify low-income boosted applications (Assuming column name 'Low-/Moderate-Income Increased Rebate' or similar; adjust if needed)
    li_col = 'Low-/Moderate-Income Increased Rebate'
    if li_col not in cvrp_baseline.columns:
        # Fallback if exact column name differs slightly in your local file
        li_col = [c for c in cvrp_baseline.columns if 'Low' in c and 'Income' in c][0]

    cvrp_baseline['is_low_income'] = (cvrp_baseline[li_col] > 0).astype(int)

    # Aggregate to county level
    county_reliance = cvrp_baseline.groupby('county_name').agg(
        total_applications=('ID', 'count'),
        li_applications=('is_low_income', 'sum')
    ).reset_index()

    county_reliance['reliance_weight'] = county_reliance['li_applications'] / county_reliance['total_applications']

    # Apply minimum threshold
    threshold_mask = county_reliance['total_applications'] < 10
    excluded_counties = county_reliance[threshold_mask]

    county_reliance.loc[threshold_mask, 'reliance_weight'] = np.nan

    print(f"Counties below the 10-application threshold (assigned NaN reliance):")
    for _, row in excluded_counties.iterrows():
        print(f" - {row['county_name']}: {row['total_applications']} applications")

    print("\nCounty reliance dataframe created.")

except Exception as e:
    print(f"ERROR loading or processing CVRP data: {e}")
Cell [6]
# STEP 0B — LOCK THE ANALYSIS SAMPLE
print("--- STEP 0B: LOCK THE ANALYSIS SAMPLE ---")

try:
    # Load Main Panel
    panel_path = os.path.join(DATA_DIR, "panel_df.csv")
    panel_df = pd.read_csv(panel_path)
    panel_df['month'] = pd.to_datetime(panel_df['month'])

    # Load Solar Control
    solar_path = os.path.join(DATA_DIR, "solar_control_variable.csv")
    solar_df = pd.read_csv(solar_path)
    solar_df['month'] = pd.to_datetime(solar_df['month'])

    # Load Voter Registration (Political Data)
    voter_path = os.path.join(DATA_DIR, "voter_registration.csv")
    voter_df = pd.read_csv(voter_path)
    # Date Parsing
    voter_df['report_date'] = pd.to_datetime(voter_df[['year', 'month', 'day']])
    # CRITICAL: Filter for Pre-April 2022 Data (Strict Exogeneity)
    study_start_date = '2022-04-01'
    history_mask = voter_df['report_date'] < study_start_date
    baseline_voter = voter_df[history_mask].copy()
    # Calculate Dem Share
    baseline_voter['dem_share'] = baseline_voter['democratic'] / baseline_voter['registered']
    # Clean County Names
    baseline_voter['county_clean'] = baseline_voter['county'].astype(str).str.strip().str.title()
    # Collapse to Static County Level
    # Result has columns: ['county_clean', 'dem_share']
    county_politics = baseline_voter.groupby('county_clean')['dem_share'].mean().reset_index()

    # --- ERROR FIX IS HERE ---
    county_politics = county_politics.rename(columns={
        'county_clean': 'county_name'
    })
    voter_df = county_politics.copy()

    # Merge datasets
    analysis_df = panel_df.merge(county_reliance[['county_name', 'reliance_weight']], on='county_name', how='left')
    analysis_df = analysis_df.merge(solar_df[['county_name', 'month', 'solar_mw_per_1000_lag1']], on=['county_name', 'month'], how='left')
    analysis_df = analysis_df.merge(voter_df[['county_name', 'dem_share']], on='county_name', how='left')

    # Check completeness per county
    county_completeness = analysis_df.groupby('county_name').agg(
        has_reliance=('reliance_weight', lambda x: x.notnull().any()),
        has_solar=('solar_mw_per_1000_lag1', lambda x: x.notnull().any()),
        has_voter=('dem_share', lambda x: x.notnull().any())
    ).reset_index()

    # Lock Sample
    valid_counties = county_completeness[
        county_completeness['has_reliance'] &
        county_completeness['has_solar'] &
        county_completeness['has_voter']
    ]['county_name'].tolist()

    excluded_counties = county_completeness[~county_completeness['county_name'].isin(valid_counties)]

    analysis_df = analysis_df[analysis_df['county_name'].isin(valid_counties)].copy()

    print(f"Included Counties ({len(valid_counties)}): {', '.join(valid_counties)}\n")
    print(f"Excluded Counties ({len(excluded_counties)}):")
    for _, row in excluded_counties.iterrows():
        reasons = []
        if not row['has_reliance']: reasons.append("Missing Reliance Weight")
        if not row['has_solar']: reasons.append("Missing Solar Data")
        if not row['has_voter']: reasons.append("Missing Voter Data")
        print(f" - {row['county_name']}: {', '.join(reasons)}")

    # Ensure Categorical Data for Clustering & Demeaning
    analysis_df['county_cat'] = pd.Categorical(analysis_df['county_name'])
    analysis_df['month_cat'] = pd.Categorical(analysis_df['month'])

    print(f"\nLOCKED SAMPLE SANITY CHECK: N = {len(analysis_df)}, Unique Counties = {analysis_df['county_cat'].nunique()}")

except Exception as e:
    print(f"ERROR merging datasets: {e}")
Cell [7]
print(panel_df.columns)
Cell [9]
# STEP 1 — DESCRIPTIVE STATISTICS & BALANCE TABLE
print("--- STEP 1: DESCRIPTIVE STATISTICS & BALANCE TABLE ---")

# Sanity Check
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")

# Plot over time
time_agg = analysis_df.groupby('month')[['evs_per_1000', 'monthly_new_evs']].mean().reset_index()

fig, ax1 = plt.subplots(figsize=(10, 5))
ax2 = ax1.twinx()

ax1.plot(time_agg['month'], time_agg['evs_per_1000'], color='blue', label='EVs per 1000 (Cumulative)')
ax2.plot(time_agg['month'], time_agg['monthly_new_evs'], color='red', label='Monthly New EVs', linestyle='--')

ax1.set_xlabel('Month')
ax1.set_ylabel('EVs per 1000', color='blue')
ax2.set_ylabel('Monthly New EVs', color='red')
plt.title('Average EV Adoption Over Time (Locked Sample)')
fig.legend(loc="upper left", bbox_to_anchor=(0.1,0.9))
plt.show()

# Balance Table (Pre-treatment implies prior to Sept 2023)
# We reconstruct panel_df to get excluded counties for comparison
pre_treat = panel_df[panel_df['month'] < '2023-09-01'].copy()
pre_treat['included'] = pre_treat['county_name'].isin(valid_counties)

# We also merge CVRP total_applications for the balance table
pre_treat = pre_treat.merge(county_reliance[['county_name', 'total_applications']], on='county_name', how='left')

balance_cols = ['evs_per_1000', 'mad_rate', 'tmax_c', 'total_applications']
balance_table = pre_treat.groupby('included')[balance_cols].mean().T
balance_table.columns = ['Excluded Counties', 'Included Counties']
print("\nPRE-TREATMENT BALANCE TABLE (Means):")
print(balance_table.round(3))
Cell [11]
# STEP 2 — BUILD THE HIGH/LOW RELIANCE DUMMIES
print("--- STEP 2: BUILD HIGH/LOW RELIANCE DUMMIES ---")

# Sanity Check
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")

# Median Split
median_reliance = analysis_df[['county_name', 'reliance_weight']].drop_duplicates()['reliance_weight'].median()
analysis_df['is_high_reliance'] = (analysis_df['reliance_weight'] > median_reliance).astype(int)

# Tercile Split
terciles = pd.qcut(analysis_df[['county_name', 'reliance_weight']].drop_duplicates()['reliance_weight'], 3, labels=[0, 1, 2])
county_tercile_map = dict(zip(analysis_df[['county_name']].drop_duplicates()['county_name'], terciles))

# Map to analysis_df: 0 = Bottom, 2 = Top (we rename to 1), 1 = Middle (NaN)
def map_tercile(county):
    t = county_tercile_map[county]
    if t == 2: return 1
    elif t == 0: return 0
    else: return np.nan

analysis_df['reliance_tercile'] = analysis_df['county_name'].apply(map_tercile)

# Print lists
high_rel_counties = analysis_df[analysis_df['is_high_reliance']==1]['county_name'].unique()
low_rel_counties = analysis_df[analysis_df['is_high_reliance']==0]['county_name'].unique()
top_tercile = analysis_df[analysis_df['reliance_tercile']==1]['county_name'].unique()
bottom_tercile = analysis_df[analysis_df['reliance_tercile']==0]['county_name'].unique()

print(f"Median Threshold: {median_reliance:.4f}")
print(f"High Reliance Counties (Median Split): {', '.join(high_rel_counties)}")
print(f"Low Reliance Counties (Median Split): {', '.join(low_rel_counties)}\n")
print(f"Top Tercile Counties: {', '.join(top_tercile)}")
print(f"Bottom Tercile Counties: {', '.join(bottom_tercile)}\n")

# Plot Histogram
rel_unique = analysis_df[['county_name', 'reliance_weight']].drop_duplicates()
plt.figure(figsize=(8,4))
plt.hist(rel_unique['reliance_weight'], bins=15, color='gray', edgecolor='black')
plt.axvline(median_reliance, color='red', linestyle='dashed', linewidth=2, label='Median')
tercile_bounds = np.percentile(rel_unique['reliance_weight'].dropna(), [33.33, 66.67])
plt.axvline(tercile_bounds[0], color='blue', linestyle='dotted', linewidth=2, label='Tercile 1/2')
plt.axvline(tercile_bounds[1], color='blue', linestyle='dotted', linewidth=2, label='Tercile 2/3')
plt.title('Distribution of CVRP Reliance Weight')
plt.xlabel('Share of Low-Income Boosted Applications')
plt.ylabel('Count of Counties')
plt.legend()
plt.show()
Cell [13]
# STEP 3 — BUILD INSTRUMENTS AND POLICY VARIABLES
print("--- STEP 3: BUILD INSTRUMENTS & POLICY VARIABLES ---")

# Sanity Check
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")

# Create temporal logic
analysis_df['post_cvrp'] = (analysis_df['month'] >= '2023-09-01').astype(int)

# Trend break: months elapsed since Nov 2023
analysis_df['iv_trend_break'] = ((analysis_df['month'].dt.year - 2023) * 12 + (analysis_df['month'].dt.month - 11)).clip(lower=0)

# Anticipation stock shift (Sep 2023 - Nov 2023)
analysis_df['iv_anticipation_stock_shift'] = ((analysis_df['month'] > '2023-08-01') &
                                              (analysis_df['month'] <= '2023-11-30')).astype(int)

# NEM 3.0 control
analysis_df['post_nem3'] = (analysis_df['month'] >= '2024-01-01').astype(int)
analysis_df['solar_x_nem3'] = analysis_df['solar_mw_per_1000_lag1'] * analysis_df['post_nem3']

# Interactions
analysis_df['ev_x_high_reliance'] = analysis_df['evs_per_1000'] * analysis_df['is_high_reliance']
analysis_df['iv_trend_x_high_reliance'] = analysis_df['iv_trend_break'] * analysis_df['is_high_reliance']
analysis_df['iv_stock_x_high_reliance'] = analysis_df['iv_anticipation_stock_shift'] * analysis_df['is_high_reliance']

# Add continuous reliance interactions for later
analysis_df['ev_x_reliance_weight'] = analysis_df['evs_per_1000'] * analysis_df['reliance_weight']
analysis_df['iv_trend_x_reliance_weight'] = analysis_df['iv_trend_break'] * analysis_df['reliance_weight']
analysis_df['iv_stock_x_reliance_weight'] = analysis_df['iv_anticipation_stock_shift'] * analysis_df['reliance_weight']

# Print Phase Counts
phase_counts = pd.Series({
    'Pre-Announcement (Before Sept 2023)': len(analysis_df[analysis_df['month'] < '2023-09-01']),
    'Anticipation (Sept - Nov 2023)': len(analysis_df[(analysis_df['month'] >= '2023-09-01') & (analysis_df['month'] <= '2023-11-30')]),
    'Post-Program (Dec 2023 Onward)': len(analysis_df[analysis_df['month'] > '2023-11-30'])
})
print(phase_counts)
Cell [15]
# STEP 4 — DEMEANING FUNCTIONS
print("--- STEP 4: DEMEANING FUNCTIONS ---")

def iterative_demean(df, value_cols, cat1='county_cat', cat2='month_cat',
                     tol=1e-10, max_iter=1000, verbose=True):
    """
    Iterative (Gauss-Seidel) two-way demeaning implementing Frisch-Waugh-Lovell.

    Returns df with new columns `{col}_dm` for each value_col.

    IMPORTANT CAVEAT ON DEGREES OF FREEDOM
    --------------------------------------
    When you run OLS/IV on the demeaned variables, the software does NOT know
    that `N_cat1 + N_cat2 - 1` parameters have been absorbed. Residual degrees
    of freedom and (non-cluster) standard errors will be too small. Under
    cluster-robust SEs with enough clusters (~50+), the distortion is modest
    because the cluster count dominates the adjustment — but for exact
    inference, use linearmodels.panel.PanelOLS with entity_effects=True /
    time_effects=True, or the dof-corrected IV helper defined below.
    """
    df_out = df.copy()
    subset = [cat1, cat2] + value_cols
    valid_idx = df_out[subset].dropna().index
    n_valid = len(valid_idx)
    n_total = len(df_out)

    if verbose:
        print(f"  Demeaning {len(value_cols)} columns on {n_valid} rows "
              f"({n_total - n_valid} dropped due to NaN)")

    max_iter_used = 0
    convergence_issues = []

    for col in value_cols:
        demeaned = df_out.loc[valid_idx, col].values.astype(float).copy()
        c1 = df_out.loc[valid_idx, cat1].values
        c2 = df_out.loc[valid_idx, cat2].values

        diff = np.inf
        iteration = 0

        while diff > tol and iteration < max_iter:
            old_demeaned = demeaned.copy()

            # Sweep cat1
            means1 = pd.Series(demeaned).groupby(c1, observed=True).transform('mean').values
            demeaned = demeaned - means1

            # Sweep cat2
            means2 = pd.Series(demeaned).groupby(c2, observed=True).transform('mean').values
            demeaned = demeaned - means2

            diff = np.max(np.abs(demeaned - old_demeaned))
            iteration += 1

        max_iter_used = max(max_iter_used, iteration)
        if iteration >= max_iter:
            convergence_issues.append(col)

        df_out.loc[valid_idx, f"{col}_dm"] = demeaned

    if verbose:
        print(f"  Max iterations across columns: {max_iter_used} "
              f"(tol={tol:.0e}, max_iter={max_iter})")
        if convergence_issues:
            print(f"  WARNING: these columns hit max_iter without converging: {convergence_issues}")

    return df_out


def county_demean(df, cols, verbose=False):
    """One-way demeaning by county. Single pass (exact)."""
    df_out = df.copy()
    valid_idx = df_out[['county_cat'] + cols].dropna().index
    for c in cols:
        means = df_out.loc[valid_idx, c].groupby(
            df_out.loc[valid_idx, 'county_cat'], observed=True
        ).transform('mean')
        df_out.loc[valid_idx, f"{c}_cdm"] = df_out.loc[valid_idx, c] - means
    if verbose:
        print(f"  County-demeaned {len(cols)} columns on {len(valid_idx)} rows")
    return df_out


def k_absorbed_twfe(df, entity='county_cat', time='month_cat'):
    """
    Number of linearly independent FE parameters absorbed by TWFE demeaning:
    N_entity + N_time - 1 (the -1 because one common constant is shared).
    Used for dof correction when running IV2SLS on demeaned data.
    """
    n_e = df[entity].nunique()
    n_t = df[time].nunique()
    return n_e + n_t - 1


print("Demeaning functions defined (iterative_demean, county_demean, k_absorbed_twfe).")
Cell [17]
# MODEL 1: Naive OLS (no FE, no IV, no controls)
print("--- MODEL 1: NAIVE OLS ---")

# Sanity Check
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")

# Formula String
print("Formula: mad_rate ~ evs_per_1000\n")

# Model
df_clean = analysis_df[['mad_rate', 'evs_per_1000', 'county_cat']].dropna()
exog = sm.add_constant(df_clean[['evs_per_1000']])

mod1 = IV2SLS(dependent=df_clean['mad_rate'],
              exog=exog,
              endog=None,
              instruments=None)

res1 = mod1.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res1.summary)
Cell [19]
# MODEL 2: OLS with Controls
print("--- MODEL 2: OLS WITH CONTROLS ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate ~ evs_per_1000 + tmax_c + cumulative_data_center_power_mw + solar_mw_per_1000_lag1 + total_generator_active_capacity_mw\n")

cols = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw', 'county_cat']
df_clean = analysis_df[cols].dropna()

exog = sm.add_constant(df_clean[['evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw']])

mod2 = IV2SLS(dependent=df_clean['mad_rate'],
              exog=exog,
              endog=None,
              instruments=None)

res2 = mod2.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res2.summary)
Cell [21]
# MODEL 2 FE VARIANTS (2a: county FE, 2b: month FE, 2c: TWFE)
from linearmodels.panel import PanelOLS

print("--- MODEL 2 FE VARIANTS ---")

cols = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw']
df_m2fe = analysis_df[cols + ['county_name', 'month']].dropna().copy()

# PanelOLS requires a MultiIndex: (entity, time)
df_m2fe = df_m2fe.set_index(['county_name', 'month'])

y = df_m2fe['mad_rate']
X = df_m2fe[['evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw']]
X_const = sm.add_constant(X)  # constant absorbed when FE included, but PanelOLS handles it

print(f"N = {len(df_m2fe)}, Unique counties = {df_m2fe.index.get_level_values(0).nunique()}, "
      f"Unique months = {df_m2fe.index.get_level_values(1).nunique()}\n")

# -------- MODEL 2a: County FE only --------
print("=" * 70)
print("MODEL 2a: OLS + Controls + County FE")
print("Formula: mad_rate ~ evs_per_1000 + tmax_c + data_center + solar_lag1 + total_generator_active_capacity_mw + county_FE")
print("=" * 70)
mod2a = PanelOLS(y, X_const, entity_effects=True, drop_absorbed=True)
res2a = mod2a.fit(cov_type='clustered', cluster_entity=True)
print(res2a.summary)
print()

# -------- MODEL 2b: Month FE only --------
print("=" * 70)
print("MODEL 2b: OLS + Controls + Month FE")
print("Formula: mad_rate ~ evs_per_1000 + tmax_c + data_center + solar_lag1 + total_generator_active_capacity_mw + month_FE")
print("=" * 70)
mod2b = PanelOLS(y, X_const, time_effects=True, drop_absorbed=True)
res2b = mod2b.fit(cov_type='clustered', cluster_entity=True)
print(res2b.summary)
print()

# -------- MODEL 2c: County + Month FE (TWFE) --------
print("=" * 70)
print("MODEL 2c: OLS + Controls + County FE + Month FE (TWFE)")
print("Formula: mad_rate ~ evs_per_1000 + tmax_c + data_center + solar_lag1 + total_generator_active_capacity_mw + county_FE + month_FE")
print("=" * 70)
mod2c = PanelOLS(y, X_const, entity_effects=True, time_effects=True, drop_absorbed=True)
res2c = mod2c.fit(cov_type='clustered', cluster_entity=True)
print(res2c.summary)
print()

# -------- Side-by-side comparison of evs_per_1000 coefficient --------
print("=" * 70)
print("COMPARISON: evs_per_1000 coefficient across Model 2 variants")
print("=" * 70)
comparison = pd.DataFrame({
    'Model': ['2 (no FE)', '2a (county FE)', '2b (month FE)', '2c (TWFE)'],
    'evs_per_1000 coef': [
        res2e.params['evs_per_1000'],
        res2a.params['evs_per_1000'],
        res2b.params['evs_per_1000'],
        res2c.params['evs_per_1000'],
    ],
    'Std. Err': [
        res2e.std_errors['evs_per_1000'],
        res2a.std_errors['evs_per_1000'],
        res2b.std_errors['evs_per_1000'],
        res2c.std_errors['evs_per_1000'],
    ],
    'p-value': [
        res2e.pvalues['evs_per_1000'],
        res2a.pvalues['evs_per_1000'],
        res2b.pvalues['evs_per_1000'],
        res2c.pvalues['evs_per_1000'],
    ],
    'Within R2': [
        res2e.rsquared,
        res2a.rsquared_within,
        res2b.rsquared_within,
        res2c.rsquared_within,
    ],
})
print(comparison.to_string(index=False))
Cell [23]
# MODEL 3: TWFE OLS
print("--- MODEL 3: TWFE OLS ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ evs_per_1000_dm + tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + total_generator_active_capacity_mw_dm \n")

cols_to_demean = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw']
df_dm = iterative_demean(analysis_df, cols_to_demean)

df_clean = df_dm[[c+"_dm" for c in cols_to_demean] + ['county_cat']].dropna()

exog = df_clean[['evs_per_1000_dm', 'tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'solar_mw_per_1000_lag1_dm','total_generator_active_capacity_mw_dm']]

mod3 = IV2SLS(dependent=df_clean['mad_rate_dm'],
              exog=exog,
              endog=None,
              instruments=None)

res3 = mod3.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res3.summary)
Cell [25]
# EVENT STUDY PLOT — PRE/POST PARALLEL TRENDS
print("--- EVENT STUDY PLOT ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")

event_df = analysis_df.dropna(subset=['mad_rate', 'is_high_reliance'])
agg_trends = event_df.groupby(['month', 'is_high_reliance'])['mad_rate'].mean().unstack()

plt.figure(figsize=(12, 6))
plt.plot(agg_trends.index, agg_trends[1], label='High Reliance Counties', color='darkred', linewidth=2)
plt.plot(agg_trends.index, agg_trends[0], label='Low Reliance Counties', color='darkblue', linewidth=2)

# Policy timeline
plt.axvline(pd.to_datetime('2023-09-01'), color='black', linestyle='--', label='CVRP Closure Announced (Sep 2023)')
plt.axvline(pd.to_datetime('2023-11-01'), color='black', linestyle='--', label='CVRP Effectively Ended (Nov 2023)')

# Shade Anticipation Window
plt.axvspan(pd.to_datetime('2023-09-01'), pd.to_datetime('2023-11-01'), color='gray', alpha=0.3, label='Anticipation Window')

plt.title('Parallel Trends: Mean Absolute Deviation (MAD) of LMP')
plt.ylabel('MAD Rate')
plt.xlabel('Month')
plt.legend()
plt.tight_layout()
plt.show()
Cell [27]
# MODEL 4a: Weak IV (post_cvrp as Standalone Instrument)
print("--- MODEL 4a: WEAK IV (STRAW MAN) ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + total_generator_active_capacity_mw_dm + [evs_per_1000_dm ~ post_cvrp_dm]\n")

cols_to_demean = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'post_cvrp','total_generator_active_capacity_mw']
df_dm = iterative_demean(analysis_df, cols_to_demean)

dm_cols = [c+"_dm" for c in cols_to_demean]
df_clean = df_dm[dm_cols + ['county_cat']].dropna()

exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'solar_mw_per_1000_lag1_dm','total_generator_active_capacity_mw_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['post_cvrp_dm']]

try:
    mod4a = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
    res4a = mod4a.fit(cov_type='clustered', clusters=df_clean['county_cat'])
    print(res4a.summary)
except ValueError as e:
    print(f"ECONOMETRIC RANK FAILURE CAUGHT:\nValueError: {e}\n")

print("""
DIAGNOSTIC — WHY THIS INSTRUMENT FAILS (AND THROWS AN ERROR):
1. Perfect Collinearity: 'post_cvrp' is a universal time indicator. After two-way demeaning (subtracting month fixed effects), it has strictly zero cross-sectional variation. The demeaned column is perfectly zeros, causing the design matrix to drop rank, hence the ValueError.
2. Identification Failure: This proves you cannot use a universal macro-shock as an instrument in a TWFE model. The instrument is completely absorbed by the time fixed effects.
3. This specification is presented deliberately as a failed straw man to motivate the preferred shift-share instruments (iv_trend_break × reliance), which introduce necessary cross-sectional variation.
""")
Cell [29]
# MODEL 4b: Simple DiD (county FE only)
print("--- MODEL 4b: SIMPLE DiD ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate ~ evs_per_1000 + post_cvrp + post_cvrp × is_high_reliance + tmax_c + cumulative_data_center_power_mw + solar_mw_per_1000_lag1 + total_generator_active_capacity_mw + county_FE\n")

analysis_df['post_x_reliance'] = analysis_df['post_cvrp'] * analysis_df['is_high_reliance']
cols = ['mad_rate', 'evs_per_1000', 'post_cvrp', 'post_x_reliance', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'total_generator_active_capacity_mw']

df_cdm = county_demean(analysis_df, cols)
df_clean = df_cdm[[c+"_cdm" for c in cols] + ['county_cat']].dropna()

exog = df_clean[[c+"_cdm" for c in cols if c != 'mad_rate']]

mod4b = IV2SLS(dependent=df_clean['mad_rate_cdm'], exog=exog, endog=None, instruments=None)
res4b = mod4b.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res4b.summary)

print("\nNOTE: Time fixed effects are intentionally excluded from this model. With full TWFE, the post_cvrp dummy would be collinear with the month fixed effects since the CVRP closure is a universal California event. This model is a deliberate straw man demonstrating the identification problem that motivates the reliance-based IV design.")
Cell [30]
# MODEL 4c: Simple DiD (county + month FE )
print("--- MODEL 4C: SIMPLE DiD County + Month FE ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ evs_per_1000_dm + (post_cvrp × is_high_reliance)_dm  + tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + total_generator_active_capacity_mw_dm")

analysis_df['post_x_reliance'] = analysis_df['post_cvrp'] * analysis_df['is_high_reliance']
cols = ['mad_rate', 'evs_per_1000', 'post_x_reliance', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'total_generator_active_capacity_mw']

df_dm = iterative_demean(analysis_df, cols)
df_clean = df_dm[[c+"_dm" for c in cols] + ['county_cat']].dropna()

exog = df_clean[[c+"_dm" for c in cols if c != 'mad_rate']]

mod4c = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=None, instruments=None)
res4c = mod4c.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res4c.summary)
Cell [31]
# MODEL 4d: Simple DiD (county + month FE)
print("--- MODEL 4D: SIMPLE DiD County + Month FE ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ evs_per_1000_dm + (post_cvrp × reliance_weight)_dm  + tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + total_generator_active_capacity_mw_dm")

analysis_df['post_x_reliance_weight'] = analysis_df['post_cvrp'] * analysis_df['reliance_weight']
cols = ['mad_rate', 'evs_per_1000', 'post_x_reliance_weight', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'total_generator_active_capacity_mw']

df_dm = iterative_demean(analysis_df, cols)
df_clean = df_dm[[c+"_dm" for c in cols] + ['county_cat']].dropna()

exog = df_clean[[c+"_dm" for c in cols if c != 'mad_rate']]

mod4d = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=None, instruments=None)
res4d = mod4d.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res4d.summary)
Cell [33]
# === MODIFICATION 1 (REVISED): MODEL 5 REDUCED FORM ===
# MODEL 5: Reduced Form (TWFE)
# The base instruments iv_trend_break and iv_anticipation_stock_shift are purely
# time-varying (identical across counties within each month). After subtracting
# month means in TWFE they collapse to zero and cause rank failure. The correct
# instruments for a TWFE design are the shift-share versions that interact the
# time shock with the pre-treatment county reliance weight, preserving
# cross-sectional variation after month demeaning.
print("--- MODEL 5: REDUCED FORM ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm\n")

cols_to_demean = [
    'mad_rate', 'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
    'tmax_c', 'cumulative_data_center_power_mw','total_generator_active_capacity_mw','solar_mw_per_1000_lag1', 'solar_x_nem3'
]

df_dm = iterative_demean(analysis_df, cols_to_demean)

dm_cols = [c + "_dm" for c in cols_to_demean]
df_clean = df_dm[dm_cols + ['county_cat']].dropna()

exog = df_clean[[c for c in dm_cols if c != 'mad_rate_dm']]

mod5 = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=None, instruments=None)
res5 = mod5.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res5.summary)
Cell [34]
print("--- MODEL 5A: REDUCED FORM ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ evs_per_1000_dm + iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm\n")

cols_to_demean = [
    'mad_rate', 'evs_per_1000', 'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
    'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
    'total_generator_active_capacity_mw'
]

df_dm = iterative_demean(analysis_df, cols_to_demean)

dm_cols = [c + "_dm" for c in cols_to_demean]
df_clean = df_dm[dm_cols + ['county_cat']].dropna()

exog = df_clean[[c for c in dm_cols if c != 'mad_rate_dm']]

mod5 = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=None, instruments=None)
res5 = mod5.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res5.summary)
Cell [36]
# === MODIFICATION 2 (REVISED): MODEL 6 IV-2SLS BASELINE ===
# MODEL 6: IV-2SLS Baseline (TWFE)
print("--- MODEL 6: IV-2SLS BASELINE ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1 + [evs_per_1000_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm]\n")

cols_to_demean = [
    'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1',
    'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight', 'total_generator_active_capacity_mw'
]

df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]

mod6 = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res6 = mod6.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res6.summary)

f_stat = res6.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat:.4f}")
if f_stat < 10:
    print("WARNING: First-stage F-statistic is below 10. Instrument may be weak.")

# SARGAN-HANSEN OVERIDENTIFICATION TEST — MODEL 6 (manual computation)
# S = n * R² from regressing 2SLS residuals on the full instrument matrix
# Distributed as chi2(L - K) where L = excluded instruments, K = endogenous vars
# Model 6: L=2, K=1 -> df=1
from scipy import stats as scipy_stats

try:
    e      = res6.resids.values
    Z_full = np.column_stack([exog.values, instr.values])  # exog + excluded instruments
    n      = len(e)

    # Project residuals onto instrument space
    ZtZ_inv   = np.linalg.pinv(Z_full.T @ Z_full)
    e_fitted  = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
    sargan_stat = float(n * (e_fitted @ e_fitted) / (e @ e))
    df_sargan   = instr.shape[1] - endog.shape[1]   # 2 - 1 = 1
    sargan_pval = 1 - scipy_stats.chi2.cdf(sargan_stat, df=df_sargan)

    print("\n--- SARGAN-HANSEN OVERIDENTIFICATION TEST ---")
    print(f"Statistic:    {sargan_stat:.4f}")
    print(f"P-value:      {sargan_pval:.4f}")
    print(f"Distribution: chi2({df_sargan})  "
          f"[degrees of freedom = {instr.shape[1]} instruments - {endog.shape[1]} endogenous var]")
    if sargan_pval > 0.10:
        print("RESULT: Fail to reject H0 (p > 0.10). Instruments pass the overidentification test.")
        print("        Both instruments are jointly consistent with instrument validity.")
    elif sargan_pval > 0.05:
        print("RESULT: Marginal rejection at 10% level. Interpret instrument validity with caution.")
    else:
        print("RESULT: Reject H0 (p < 0.05). At least one instrument may be invalid.")
        print("        Consider whether exclusion restriction holds for both instruments separately.")

except Exception as e_err:
    print(f"Sargan test could not be computed: {e_err}")
Cell [38]
df_clean.columns
Cell [39]
# === MODIFICATION 3 (REVISED): MODEL 7 IV-2SLS FINAL ===
# MODEL 7: IV-2SLS Final (TWFE)
print("--- MODEL 7: IV-2SLS FINAL ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + total_generator_active_capacity_mw_dm + [evs_per_1000_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm]\n")

cols_to_demean = [
    'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
    'solar_mw_per_1000_lag1', 'solar_x_nem3', 'total_generator_active_capacity_mw' ,
    'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight'
]

df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm', 'total_generator_active_capacity_mw_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]

mod7 = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res7 = mod7.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res7.summary)

f_stat = res7.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat:.4f}")
if f_stat < 10:
    print("WARNING: First-stage F-statistic is below 10. Instrument may be weak.")

# SARGAN-HANSEN OVERIDENTIFICATION TEST — MODEL 7 (manual computation)
# S = n * R² from regressing 2SLS residuals on the full instrument matrix
# Distributed as chi2(L - K) where L = excluded instruments, K = endogenous vars
# Model 7: L=2, K=1 -> df=1
from scipy import stats as scipy_stats

try:
    e      = res7.resids.values
    Z_full = np.column_stack([exog.values, instr.values])  # exog + excluded instruments
    n      = len(e)

    # Project residuals onto instrument space
    ZtZ_inv   = np.linalg.pinv(Z_full.T @ Z_full)
    e_fitted  = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
    sargan_stat = float(n * (e_fitted @ e_fitted) / (e @ e))
    df_sargan   = instr.shape[1] - endog.shape[1]   # 2 - 1 = 1
    sargan_pval = 1 - scipy_stats.chi2.cdf(sargan_stat, df=df_sargan)

    print("\n--- SARGAN-HANSEN OVERIDENTIFICATION TEST ---")
    print(f"Statistic:    {sargan_stat:.4f}")
    print(f"P-value:      {sargan_pval:.4f}")
    print(f"Distribution: chi2({df_sargan})  "
          f"[degrees of freedom = {instr.shape[1]} instruments - {endog.shape[1]} endogenous var]")
    if sargan_pval > 0.10:
        print("RESULT: Fail to reject H0 (p > 0.10). Instruments pass the overidentification test.")
        print("        Both instruments are jointly consistent with instrument validity.")
    elif sargan_pval > 0.05:
        print("RESULT: Marginal rejection at 10% level. Interpret instrument validity with caution.")
    else:
        print("RESULT: Reject H0 (p < 0.05). At least one instrument may be invalid.")
        print("        Consider whether exclusion restriction holds for both instruments separately.")

except Exception as e_err:
    print(f"Sargan test could not be computed: {e_err}")
Cell [40]
# === MODEL 7R IV-2SLS FINAL ===
# MODEL 7R: IV-2SLS Final - No FE
print("--- MODEL 7R: IV-2SLS NO FE ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate ~ tmax_c + cumulative_data_center_power_mw + solar_mw_per_1000_lag1 + solar_x_nem3 + total_generator_active_capacity_mw + [evs_per_1000 ~ iv_trend_x_reliance_weight + iv_stock_x_reliance_weight]\n")

exog = sm.add_constant(analysis_df[['tmax_c', 'cumulative_data_center_power_mw',
                                     'solar_mw_per_1000_lag1', 'solar_x_nem3',
                                     'total_generator_active_capacity_mw']])
endog = analysis_df[['evs_per_1000']]
instr = analysis_df[['iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']]

mod7r = IV2SLS(dependent=analysis_df['mad_rate'], exog=exog, endog=endog, instruments=instr)
res7r = mod7r.fit(cov_type='clustered', clusters=analysis_df['county_cat'])
print(res7r.summary)

f_stat = res7r.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat:.4f}")
if f_stat < 10:
    print("WARNING: First-stage F-statistic is below 10. Instrument may be weak.")

# SARGAN-HANSEN OVERIDENTIFICATION TEST — MODEL 7 (manual computation)
# S = n * R² from regressing 2SLS residuals on the full instrument matrix
# Distributed as chi2(L - K) where L = excluded instruments, K = endogenous vars
# Model 7: L=2, K=1 -> df=1
from scipy import stats as scipy_stats

try:
    e      = res7r.resids.values
    Z_full = np.column_stack([exog.values, instr.values])  # exog + excluded instruments
    n      = len(e)

    # Project residuals onto instrument space
    ZtZ_inv   = np.linalg.pinv(Z_full.T @ Z_full)
    e_fitted  = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
    sargan_stat = float(n * (e_fitted @ e_fitted) / (e @ e))
    df_sargan   = instr.shape[1] - endog.shape[1]   # 2 - 1 = 1
    sargan_pval = 1 - scipy_stats.chi2.cdf(sargan_stat, df=df_sargan)

    print("\n--- SARGAN-HANSEN OVERIDENTIFICATION TEST ---")
    print(f"Statistic:    {sargan_stat:.4f}")
    print(f"P-value:      {sargan_pval:.4f}")
    print(f"Distribution: chi2({df_sargan})  "
          f"[degrees of freedom = {instr.shape[1]} instruments - {endog.shape[1]} endogenous var]")
    if sargan_pval > 0.10:
        print("RESULT: Fail to reject H0 (p > 0.10). Instruments pass the overidentification test.")
        print("        Both instruments are jointly consistent with instrument validity.")
    elif sargan_pval > 0.05:
        print("RESULT: Marginal rejection at 10% level. Interpret instrument validity with caution.")
    else:
        print("RESULT: Reject H0 (p < 0.05). At least one instrument may be invalid.")
        print("        Consider whether exclusion restriction holds for both instruments separately.")

except Exception as e_err:
    print(f"Sargan test could not be computed: {e_err}")
Cell [41]
# === MODEL 7S IV-2SLS FINAL ===
# MODEL 7S: IV-2SLS Final - No FE
print("--- MODEL 7S: IV-2SLS NO FE ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate ~ tmax_c + cumulative_data_center_power_mw + solar_mw_per_1000_lag1 + total_generator_active_capacity_mw + [evs_per_1000 ~ iv_trend_x_reliance_weight + iv_stock_x_reliance_weight]\n")

exog = sm.add_constant(analysis_df[['tmax_c', 'cumulative_data_center_power_mw',
                                     'solar_mw_per_1000_lag1',
                                     'total_generator_active_capacity_mw']])
endog = analysis_df[['evs_per_1000']]
instr = analysis_df[['iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']]

mod7s = IV2SLS(dependent=analysis_df['mad_rate'], exog=exog, endog=endog, instruments=instr)
res7s = mod7s.fit(cov_type='clustered', clusters=analysis_df['county_cat'])
print(res7s.summary)

f_stat = res7s.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat:.4f}")
if f_stat < 10:
    print("WARNING: First-stage F-statistic is below 10. Instrument may be weak.")

# SARGAN-HANSEN OVERIDENTIFICATION TEST — MODEL 7 (manual computation)
# S = n * R² from regressing 2SLS residuals on the full instrument matrix
# Distributed as chi2(L - K) where L = excluded instruments, K = endogenous vars
# Model 7: L=2, K=1 -> df=1
from scipy import stats as scipy_stats

try:
    e      = res7s.resids.values
    Z_full = np.column_stack([exog.values, instr.values])  # exog + excluded instruments
    n      = len(e)

    # Project residuals onto instrument space
    ZtZ_inv   = np.linalg.pinv(Z_full.T @ Z_full)
    e_fitted  = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
    sargan_stat = float(n * (e_fitted @ e_fitted) / (e @ e))
    df_sargan   = instr.shape[1] - endog.shape[1]   # 2 - 1 = 1
    sargan_pval = 1 - scipy_stats.chi2.cdf(sargan_stat, df=df_sargan)

    print("\n--- SARGAN-HANSEN OVERIDENTIFICATION TEST ---")
    print(f"Statistic:    {sargan_stat:.4f}")
    print(f"P-value:      {sargan_pval:.4f}")
    print(f"Distribution: chi2({df_sargan})  "
          f"[degrees of freedom = {instr.shape[1]} instruments - {endog.shape[1]} endogenous var]")
    if sargan_pval > 0.10:
        print("RESULT: Fail to reject H0 (p > 0.10). Instruments pass the overidentification test.")
        print("        Both instruments are jointly consistent with instrument validity.")
    elif sargan_pval > 0.05:
        print("RESULT: Marginal rejection at 10% level. Interpret instrument validity with caution.")
    else:
        print("RESULT: Reject H0 (p < 0.05). At least one instrument may be invalid.")
        print("        Consider whether exclusion restriction holds for both instruments separately.")

except Exception as e_err:
    print(f"Sargan test could not be computed: {e_err}")
Cell [42]
# MODEL 7a: IV-2SLS Trend Only (TWFE)
print("--- MODEL 7a: IV-2SLS TREND INSTRUMENT ONLY ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm ~ iv_trend_x_reliance_weight_dm]\n")
print("NOTE: Exactly identified (1 instrument, 1 endogenous variable). Sargan test not available.\n")

cols_to_demean = [
    'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'total_generator_active_capacity_mw',
    'solar_mw_per_1000_lag1', 'solar_x_nem3',
    'iv_trend_x_reliance_weight'
]

df_dm    = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm',
                  'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm']]

mod7a = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res7a = mod7a.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res7a.summary)

f_stat_7a = res7a.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat_7a:.4f}")
if f_stat_7a < 10:
    print("WARNING: First-stage F-statistic is below 10. Weak instrument bias is a concern.")
else:
    print("First-stage F-statistic is above 10. Instrument passes relevance threshold.")

print(f"\nEV Coefficient:  {res7a.params['evs_per_1000_dm']:.6f}")
print(f"P-value:         {res7a.pvalues['evs_per_1000_dm']:.4f}")
Cell [43]
# MODEL 7b: IV-2SLS Anticipation Only (TWFE)
print("--- MODEL 7b: IV-2SLS ANTICIPATION INSTRUMENT ONLY ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm ~ iv_stock_x_reliance_weight_dm]\n")
print("NOTE: Exactly identified (1 instrument, 1 endogenous variable). Sargan test not available.\n")

cols_to_demean = [
    'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'total_generator_active_capacity_mw',
    'solar_mw_per_1000_lag1', 'solar_x_nem3',
    'iv_stock_x_reliance_weight'
]

df_dm    = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm',
                  'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_stock_x_reliance_weight_dm']]

mod7b = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res7b = mod7b.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res7b.summary)

f_stat_7b = res7b.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat_7b:.4f}")
if f_stat_7b < 10:
    print("WARNING: First-stage F-statistic is below 10. Weak instrument bias is a concern.")
else:
    print("First-stage F-statistic is above 10. Instrument passes relevance threshold.")

print(f"\nEV Coefficient:  {res7b.params['evs_per_1000_dm']:.6f}")
print(f"P-value:         {res7b.pvalues['evs_per_1000_dm']:.4f}")
Cell [45]
# ROBUSTNESS: Native-FE re-estimation of Models 3, 4b/4c/4d (OLS) and 5, 6, 7 (IV)
from linearmodels.panel import PanelOLS
from linearmodels.iv import IV2SLS

print("=" * 78)
print("ROBUSTNESS — Native FE Re-estimation")
print("=" * 78)

# Build a panel-indexed dataframe once
panel_df = analysis_df.copy()
if 'post_x_reliance' not in panel_df.columns:
    panel_df['post_x_reliance'] = panel_df['post_cvrp'] * panel_df['is_high_reliance']
if 'post_x_reliance_weight' not in panel_df.columns:
    panel_df['post_x_reliance_weight'] = panel_df['post_cvrp'] * panel_df['reliance_weight']

panel_df = panel_df.set_index(['county_name', 'month']).sort_index()

def run_panel_ols(y_col, x_cols, entity=True, time=True, label=''):
    """Fit PanelOLS with requested FE and clustered SEs by entity."""
    sub = panel_df[[y_col] + x_cols].dropna()
    y = sub[y_col]
    X = sub[x_cols]
    mod = PanelOLS(y, X, entity_effects=entity, time_effects=time, drop_absorbed=True)
    res = mod.fit(cov_type='clustered', cluster_entity=True)
    print(f"\n{'=' * 78}\n{label}\n{'=' * 78}")
    print(f"N={int(res.nobs)}, entities={sub.index.get_level_values(0).nunique()}, "
          f"FE: entity={entity}, time={time}")
    print(res.summary)
    return res

# ---- Model 3 native (TWFE) ----
res3_native = run_panel_ols(
    'mad_rate',
    ['evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw'],
    entity=True, time=True,
    label='MODEL 3 (native TWFE via PanelOLS)'
)

# ---- Model 4b native (county FE only, deliberate straw man) ----
res4b_native = run_panel_ols(
    'mad_rate',
    ['evs_per_1000', 'post_cvrp', 'post_x_reliance', 'tmax_c',
     'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw'],
    entity=True, time=False,
    label='MODEL 4b (native county FE only — straw man)'
)

# ---- Model 4c native (TWFE + binary interaction) ----
res4c_native = run_panel_ols(
    'mad_rate',
    ['evs_per_1000', 'post_x_reliance', 'tmax_c',
     'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw'],
    entity=True, time=True,
    label='MODEL 4c (native TWFE, binary post×reliance)'
)

# ---- Model 4d native (TWFE + continuous interaction) ----
res4d_native = run_panel_ols(
    'mad_rate',
    ['evs_per_1000', 'post_x_reliance_weight', 'tmax_c',
     'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw'],
    entity=True, time=True,
    label='MODEL 4d (native TWFE, continuous post×reliance_weight)'
)
Cell [46]
# ROBUSTNESS (cont'd) — dof-corrected IV2SLS for Models 5, 6, 7
# ------------------------------------------------------------------
# Strategy: demean via FWL (which we've already validated gives exact
# coefficients), then fit IV2SLS on the demeaned data, but rescale the
# cluster-robust variance by (N-1)/(N - k - k_fe) vs. linearmodels' default
# (N-1)/(N - k), where k_fe = N_county + N_month - 1.
#
# This mirrors what Stata's ivreghdfe and R's fixest do internally.
# ------------------------------------------------------------------
from scipy import stats as scipy_stats

def fit_iv_dof_corrected(df_dm_clean, dep_col, exog_cols, endog_cols, instr_cols,
                         k_fe, cluster_col='county_cat', label=''):
    """
    Fit IV2SLS on pre-demeaned data and apply FE dof correction to the
    cluster-robust standard errors.

    linearmodels reports SEs using a small-sample factor that assumes
    residual dof = N - k (where k = # non-FE regressors). The true residual
    dof after absorbing FEs is N - k - k_fe. We rescale the variance matrix
    by (N - k) / (N - k - k_fe) and recompute t-stats, p-values, CIs.
    """
    dep   = df_dm_clean[dep_col]
    exog  = df_dm_clean[exog_cols] if exog_cols else None
    endog = df_dm_clean[endog_cols] if endog_cols else None
    instr = df_dm_clean[instr_cols] if instr_cols else None

    mod = IV2SLS(dependent=dep, exog=exog, endog=endog, instruments=instr)
    res = mod.fit(cov_type='clustered', clusters=df_dm_clean[cluster_col], debiased=True)

    # Rescale
    N = int(res.nobs)
    k_total = len(res.params)  # non-FE regressors including constant if present
    naive_dof   = N - k_total
    corrected_dof = N - k_total - k_fe
    scale = naive_dof / corrected_dof  # variance inflation factor

    V_corrected = res.cov * scale
    se_corrected = np.sqrt(np.diag(V_corrected))
    tstats = res.params.values / se_corrected
    pvals = 2 * (1 - scipy_stats.t.cdf(np.abs(tstats), df=corrected_dof))

    print(f"\n{'=' * 78}\n{label}\n{'=' * 78}")
    print(f"N={N}, non-FE params k={k_total}, absorbed FE params k_fe={k_fe}")
    print(f"Variance inflation factor (dof correction): {scale:.4f}")
    print(f"\n{'Parameter':<45s} {'Coef':>12s} {'Naive SE':>12s} {'Corr. SE':>12s} {'Corr. p':>10s}")
    print('-' * 93)
    for i, name in enumerate(res.params.index):
        print(f"{name:<45s} {res.params.values[i]:>12.4e} "
              f"{res.std_errors.values[i]:>12.4e} "
              f"{se_corrected[i]:>12.4e} {pvals[i]:>10.4f}")

    return {
        'res': res,
        'N': N,
        'k_fe': k_fe,
        'scale': scale,
        'params': res.params,
        'se_naive': res.std_errors,
        'se_corrected': pd.Series(se_corrected, index=res.params.index),
        'pvals_corrected': pd.Series(pvals, index=res.params.index),
    }

k_fe = k_absorbed_twfe(analysis_df)
print(f"\nAbsorbed TWFE parameters (N_county + N_month - 1): {k_fe}")

# ---- Model 5 (Reduced Form) ----
cols5 = ['mad_rate', 'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
         'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3','total_generator_active_capacity_mw']
df5 = iterative_demean(analysis_df, cols5, verbose=False)
df5_clean = df5[[c + '_dm' for c in cols5] + ['county_cat']].dropna()
r5_corr = fit_iv_dof_corrected(
    df5_clean, 'mad_rate_dm',
    exog_cols=[c + '_dm' for c in cols5[1:]],
    endog_cols=None, instr_cols=None,
    k_fe=k_fe, label='MODEL 5 (Reduced Form) — dof-corrected'
)

# ---- Model 6 (IV Baseline) ----
cols6 = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
         'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight','total_generator_active_capacity_mw']
df6 = iterative_demean(analysis_df, cols6, verbose=False)
df6_clean = df6[[c + '_dm' for c in cols6] + ['county_cat']].dropna()
r6_corr = fit_iv_dof_corrected(
    df6_clean, 'mad_rate_dm',
    exog_cols=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm'],
    endog_cols=['evs_per_1000_dm'],
    instr_cols=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
    k_fe=k_fe, label='MODEL 6 (IV Baseline) — dof-corrected'
)

# ---- Model 7 (IV Final — preferred specification) ----
cols7 = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
         'solar_mw_per_1000_lag1', 'solar_x_nem3', 'total_generator_active_capacity_mw',
         'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']
df7 = iterative_demean(analysis_df, cols7, verbose=False)
df7_clean = df7[[c + '_dm' for c in cols7] + ['county_cat']].dropna()
r7_corr = fit_iv_dof_corrected(
    df7_clean, 'mad_rate_dm',
    exog_cols=['tmax_c_dm', 'cumulative_data_center_power_mw_dm',
               'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm','total_generator_active_capacity_mw_dm'],
    endog_cols=['evs_per_1000_dm'],
    instr_cols=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
    k_fe=k_fe, label='MODEL 7 (IV Final — PREFERRED) — dof-corrected'
)

# ---- Summary comparison of evs_per_1000 across Models 3, 6, 7 (original vs corrected) ----
print(f"\n{'=' * 78}\nSUMMARY: evs_per_1000 across demeaned vs native/corrected\n{'=' * 78}")
summary_rows = []
summary_rows.append({
    'Model': '3 (TWFE OLS)',
    'Source': 'Demeaned (original)',
    'Coef': res3.params['evs_per_1000_dm'],
    'SE':   res3.std_errors['evs_per_1000_dm'],
    'p':    res3.pvalues['evs_per_1000_dm'],
})
summary_rows.append({
    'Model': '3 (TWFE OLS)',
    'Source': 'PanelOLS native',
    'Coef': res3_native.params['evs_per_1000'],
    'SE':   res3_native.std_errors['evs_per_1000'],
    'p':    res3_native.pvalues['evs_per_1000'],
})
summary_rows.append({
    'Model': '6 (IV baseline)',
    'Source': 'Demeaned (original)',
    'Coef': res6.params['evs_per_1000_dm'],
    'SE':   res6.std_errors['evs_per_1000_dm'],
    'p':    res6.pvalues['evs_per_1000_dm'],
})
summary_rows.append({
    'Model': '6 (IV baseline)',
    'Source': 'dof-corrected',
    'Coef': r6_corr['params']['evs_per_1000_dm'],
    'SE':   r6_corr['se_corrected']['evs_per_1000_dm'],
    'p':    r6_corr['pvals_corrected']['evs_per_1000_dm'],
})
summary_rows.append({
    'Model': '7 (IV final)',
    'Source': 'Demeaned (original)',
    'Coef': res7.params['evs_per_1000_dm'],
    'SE':   res7.std_errors['evs_per_1000_dm'],
    'p':    res7.pvalues['evs_per_1000_dm'],
})
summary_rows.append({
    'Model': '7 (IV final)',
    'Source': 'dof-corrected',
    'Coef': r7_corr['params']['evs_per_1000_dm'],
    'SE':   r7_corr['se_corrected']['evs_per_1000_dm'],
    'p':    r7_corr['pvals_corrected']['evs_per_1000_dm'],
})
print(pd.DataFrame(summary_rows).to_string(index=False))
print()
print("Interpretation: if Coef matches across Source rows within a Model, the")
print("demeaning is correct. SE differences reveal the magnitude of the dof bias.")
Cell [47]
fs6 = res6.first_stage.individual['evs_per_1000_dm']
instruments = ['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']
first_stage_table = pd.DataFrame({
    'Coef':    [fs6.params[i]      for i in instruments],
    'SE':      [fs6.std_errors[i]  for i in instruments],
    't-stat':  [fs6.tstats[i]      for i in instruments],
    'p-value': [fs6.pvalues[i]     for i in instruments],
}, index=['Trend × Reliance', 'Anticipation × Reliance'])
print("\nModel 6 First Stage (dep var: evs_per_1000_dm):")
print(first_stage_table.round(4))
print(f"\nFirst-stage F: {res6.first_stage.diagnostics['f.stat'].iloc[0]:.2f}")
print(f"Partial R²:    {res6.first_stage.diagnostics['partial.rsquared'].iloc[0]:.4f}")

fs7 = res7.first_stage.individual['evs_per_1000_dm']
instruments = ['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']
first_stage_table = pd.DataFrame({
    'Coef':    [fs7.params[i]      for i in instruments],
    'SE':      [fs7.std_errors[i]  for i in instruments],
    't-stat':  [fs7.tstats[i]      for i in instruments],
    'p-value': [fs7.pvalues[i]     for i in instruments],
}, index=['Trend × Reliance', 'Anticipation × Reliance'])
print("\nModel 7 First Stage (dep var: evs_per_1000_dm):")
print(first_stage_table.round(4))
print(f"\nFirst-stage F: {res7.first_stage.diagnostics['f.stat'].iloc[0]:.2f}")
print(f"Partial R²:    {res7.first_stage.diagnostics['partial.rsquared'].iloc[0]:.4f}")
Cell [49]
# === MODIFICATION 4 (REVISED): MODEL 8a IV HETEROGENEITY MEDIAN SPLIT ===
# MODEL 8a: IV Heterogeneity, Median Split (TWFE)
print("--- MODEL 8a: IV HETEROGENEITY (MEDIAN SPLIT) ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + total_generator_active_capacity_mw + [evs_per_1000_dm + ev_x_high_reliance_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + iv_trend_x_high_reliance_dm + iv_stock_x_high_reliance_dm]\n")

# iv_trend_break and iv_anticipation_stock_shift are purely time-varying and become
# zero columns after month demeaning — they cannot serve as instruments in TWFE.
# iv_trend_x_reliance_weight and iv_stock_x_reliance_weight vary cross-sectionally
# (by county reliance weight) and survive two-way demeaning; they instrument evs_per_1000_dm.
# iv_trend_x_high_reliance and iv_stock_x_high_reliance instrument ev_x_high_reliance_dm.
cols_to_demean = [
    'mad_rate', 'evs_per_1000', 'ev_x_high_reliance', 'tmax_c',
    'cumulative_data_center_power_mw','total_generator_active_capacity_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
    'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
    'iv_trend_x_high_reliance', 'iv_stock_x_high_reliance'
]

df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm', 'ev_x_high_reliance_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm',
                  'iv_trend_x_high_reliance_dm', 'iv_stock_x_high_reliance_dm']]

mod8a = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res8a = mod8a.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res8a.summary)

print("\n--- FIRST STAGE DIAGNOSTICS ---")
print(res8a.first_stage.diagnostics['f.stat'])
for var in endog.columns:
    f_stat = res8a.first_stage.diagnostics.loc[var, 'f.stat']
    if f_stat < 10:
        print(f"WARNING: First-stage F-statistic for {var} is {f_stat:.4f} — below 10. Instrument may be weak.")

base_coef  = res8a.params['evs_per_1000_dm']
int_coef   = res8a.params['ev_x_high_reliance_dm']
p_val      = res8a.pvalues['ev_x_high_reliance_dm']
tot_effect = base_coef + int_coef
conclusion = "significant at 10%" if p_val < 0.10 else "not significant"

print(f"""
Base Effect (Low Reliance Counties):    {base_coef:.6f}
Interaction Effect (Differential):      {int_coef:.6f}
Interaction P-Value:                     {p_val:.4f}
Total Effect (High Reliance Counties):  {tot_effect:.6f}
Conclusion: {conclusion}
""")
Cell [51]
# === MODIFICATION 5 (REVISED): MODEL 8b IV HETEROGENEITY CONTINUOUS RELIANCE ===
# MODEL 8b: IV Heterogeneity, Continuous Reliance Weight (TWFE)
print("--- MODEL 8b: IV HETEROGENEITY (CONTINUOUS) ---")

print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm + ev_x_reliance_weight_dm ~ iv_trend_x_high_reliance_dm + iv_stock_x_high_reliance_dm + iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm]\n")

# For the continuous model: iv_trend_x_high_reliance and iv_stock_x_high_reliance
# instrument evs_per_1000_dm (binary cross-sectional variation survives TWFE);
# iv_trend_x_reliance_weight and iv_stock_x_reliance_weight instrument
# ev_x_reliance_weight_dm (continuous cross-sectional variation survives TWFE).
cols_to_demean = [
    'mad_rate', 'evs_per_1000', 'ev_x_reliance_weight', 'tmax_c',
    'cumulative_data_center_power_mw','total_generator_active_capacity_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
    'iv_trend_x_high_reliance', 'iv_stock_x_high_reliance',
    'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight'
]

df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm','solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm', 'ev_x_reliance_weight_dm']]
instr = df_clean[['iv_trend_x_high_reliance_dm', 'iv_stock_x_high_reliance_dm',
                  'iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]

mod8b = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res8b = mod8b.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res8b.summary)

print("\n--- FIRST STAGE DIAGNOSTICS ---")
print(res8b.first_stage.diagnostics['f.stat'])
for var in endog.columns:
    f_stat = res8b.first_stage.diagnostics.loc[var, 'f.stat']
    if f_stat < 10:
        print(f"WARNING: First-stage F-statistic for {var} is {f_stat:.4f} — below 10. Instrument may be weak.")
Cell [53]
# === MODIFICATION 6 (REVISED): MODEL 8c IV HETEROGENEITY TERCILE SPLIT ===
# MODEL 8c: IV Heterogeneity, Tercile Split (TWFE)
print("--- MODEL 8c: IV HETEROGENEITY (TERCILE) ---")

tercile_df = analysis_df.dropna(subset=['reliance_tercile']).copy()
print(f"Current N (Excluding Middle Tercile): {len(tercile_df)}, Unique counties: {tercile_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm + ev_x_tercile_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + iv_trend_x_tercile_dm + iv_stock_x_tercile_dm]\n")

tercile_df['ev_x_tercile']       = tercile_df['evs_per_1000']                 * tercile_df['reliance_tercile']
tercile_df['iv_trend_x_tercile'] = tercile_df['iv_trend_break']               * tercile_df['reliance_tercile']
tercile_df['iv_stock_x_tercile'] = tercile_df['iv_anticipation_stock_shift']  * tercile_df['reliance_tercile']

# iv_trend_x_reliance_weight and iv_stock_x_reliance_weight instrument evs_per_1000_dm;
# iv_trend_x_tercile and iv_stock_x_tercile instrument ev_x_tercile_dm.
cols_to_demean = [
    'mad_rate', 'evs_per_1000', 'ev_x_tercile', 'tmax_c',
    'cumulative_data_center_power_mw','total_generator_active_capacity_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
    'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
    'iv_trend_x_tercile', 'iv_stock_x_tercile'
]

df_dm = iterative_demean(tercile_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm', 'ev_x_tercile_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm',
                  'iv_trend_x_tercile_dm', 'iv_stock_x_tercile_dm']]

mod8c = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res8c = mod8c.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res8c.summary)

print("\n--- FIRST STAGE DIAGNOSTICS ---")
print(res8c.first_stage.diagnostics['f.stat'])
for var in endog.columns:
    f_stat = res8c.first_stage.diagnostics.loc[var, 'f.stat']
    if f_stat < 10:
        print(f"WARNING: First-stage F-statistic for {var} is {f_stat:.4f} — below 10. Instrument may be weak.")
Cell [55]
# UPDATED COMBINED SUMMARY TABLE (includes Models 7a, 7b, and Sargan test)
def get_stars(pval):
    if pval < 0.01:   return '***'
    elif pval < 0.05: return '**'
    elif pval < 0.10: return '*'
    else:             return ''

def get_sargan(res, exog, instr, endog):
    """
    Manual Sargan-Hansen test: n * R² from regressing 2SLS residuals
    on the full instrument matrix (exog + excluded instruments).
    Only valid for overidentified models (n_instruments > n_endogenous).
    Returns (np.nan, np.nan) for exactly identified models.
    """
    from scipy import stats as scipy_stats
    try:
        df_sargan = instr.shape[1] - endog.shape[1]
        if df_sargan < 1:
            return np.nan, np.nan          # exactly identified — test undefined
        e         = res.resids.values
        Z_full    = np.column_stack([exog.values, instr.values])
        n         = len(e)
        ZtZ_inv   = np.linalg.pinv(Z_full.T @ Z_full)
        e_fitted  = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
        stat      = float(n * (e_fitted @ e_fitted) / (e @ e))
        pval      = 1 - scipy_stats.chi2.cdf(stat, df=df_sargan)
        return stat, pval
    except:
        return np.nan, np.nan

# Pre-compute Sargan results for overidentified models before the summary loop
sargan_results = {}

# Model 6 — re-run demeaning to recover exog/instr in scope
cols_m6 = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
            'solar_mw_per_1000_lag1','total_generator_active_capacity_mw',
            'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']
df_dm_m6   = iterative_demean(analysis_df, cols_m6)
df_c_m6    = df_dm_m6[[c+"_dm" for c in cols_m6] + ['county_cat']].dropna()
exog_m6    = df_c_m6[['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm']]
endog_m6   = df_c_m6[['evs_per_1000_dm']]
instr_m6   = df_c_m6[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]
sargan_results['Model 6: IV-2SLS Base'] = get_sargan(res6, exog_m6, instr_m6, endog_m6)

# Model 7 — re-run demeaning to recover exog/instr in scope
cols_m7 = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'total_generator_active_capacity_mw',
            'solar_mw_per_1000_lag1', 'solar_x_nem3',
            'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']
df_dm_m7   = iterative_demean(analysis_df, cols_m7)
df_c_m7    = df_dm_m7[[c+"_dm" for c in cols_m7] + ['county_cat']].dropna()
exog_m7    = df_c_m7[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm',
                       'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog_m7   = df_c_m7[['evs_per_1000_dm']]
instr_m7   = df_c_m7[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]
sargan_results['Model 7: IV-2SLS Final (Both)'] = get_sargan(res7, exog_m7, instr_m7, endog_m7)



models = [
    ('Model 1: Naive OLS',              res1,  'evs_per_1000',        'None',                             False),
    ('Model 2: OLS w/ Controls',        res2,  'evs_per_1000',        'None',                             False),
    ('Model 3: TWFE OLS',               res3,  'evs_per_1000_dm',     'None',                             False),
    ('Model 4b: Simple DiD + County FE',res4b, 'evs_per_1000_cdm',    'None',                             False),

    ('Model 6: IV-2SLS Base',           res6,  'evs_per_1000_dm',     'Trend + Anticip.',                 True),
    ('Model 7: IV-2SLS Final (Both)',   res7,  'evs_per_1000_dm',     'Trend + Anticip.',                 True),
    ('Model 7a: IV-2SLS Trend Only',    res7a, 'evs_per_1000_dm',     'Trend only',                       False),
    ('Model 7b: IV-2SLS Anticip. Only', res7b, 'evs_per_1000_dm',     'Anticip. only',                    False),
    ('Model 8a: IV Het (Base)',         res8a, 'evs_per_1000_dm',     'Trend + Anticip. (×reliance)',     True),
    ('Model 8a: IV Het (Interaction)',  res8a, 'ev_x_high_reliance_dm','Trend + Anticip. (×reliance)',    True),
]

# Note: Models 6 and 7 are overidentified (2 instruments, 1 endogenous var) -> Sargan available
# Models 7a and 7b are exactly identified (1 instrument, 1 endogenous var) -> Sargan undefined
# Model 8a has 4 instruments and 2 endogenous vars -> overidentified by 2 -> Sargan available
# OLS models have no instruments -> Sargan undefined

summary_data = []
for name, res, var, iv_label, is_iv in models:

    # Second-stage coefficient, SE, p-value
    coef  = res.params.get(var, np.nan)
    se    = res.std_errors.get(var, np.nan)
    pval  = res.pvalues.get(var, np.nan)
    stars = get_stars(pval) if not pd.isna(pval) else ''

    # First-stage F-statistic (IV models only)
    fstat = np.nan
    if is_iv:
        try:
            idx   = var if var in res.first_stage.diagnostics.index \
                    else res.first_stage.diagnostics.index[0]
            fstat = res.first_stage.diagnostics.loc[idx, 'f.stat']
        except:
            pass

    # Sargan-Hansen overidentification test
    # Only computed for overidentified models; exactly identified -> n_instr == n_endog
    # We check overidentification by catching the test result
    sargan_stat, sargan_pval = np.nan, np.nan
    if is_iv:
        sargan_stat, sargan_pval = sargan_results.get(name, (np.nan, np.nan))

    summary_data.append({
        'Model':            name,
        'N':                int(res.nobs),
        'EV Coef':          f"{coef:.5f}{stars}" if not pd.isna(coef)        else '-',
        'Std. Err.':        f"({se:.5f})"        if not pd.isna(se)          else '-',
        'P-value':          f"{pval:.3f}"         if not pd.isna(pval)        else '-',
        'F-stat (1st)':     f"{fstat:.2f}"        if not pd.isna(fstat)       else '-',
        'Sargan stat':      f"{sargan_stat:.4f}"  if not pd.isna(sargan_stat) else 'n/a',
        'Sargan p':         f"{sargan_pval:.4f}"  if not pd.isna(sargan_pval) else 'n/a',
        'Instruments':      iv_label,
    })

sum_df = pd.DataFrame(summary_data)

print("=" * 130)
print("COMBINED SUMMARY TABLE — DEPENDENT VARIABLE: mad_rate")
print("=" * 130)
print(sum_df.to_string(index=False))

print("\nSignificance: * p<0.10   ** p<0.05   *** p<0.01")
print("\nNotes:")
print("  - All shift-share instruments defined as: time_shock × reliance_weight")
print("    Trend IV   = months_since_CVRP_end × reliance_weight")
print("    Anticip IV = anticipation_phase_dummy × reliance_weight")
print("  - Sargan-Hansen test H0: instruments are valid (model not overidentified).")
print("    A p-value > 0.10 fails to reject H0 — instruments pass the validity test.")
print("    'n/a' indicates exactly identified model (test undefined) or OLS (no instruments).")
print("  - Model 4b uses county FE only by design — post_cvrp would be absorbed by time FE.")
Cell [58]
# === MODIFICATION 9 (REVISED): STEP 6 NEM 3.0 SENSITIVITY ANALYSIS ===
print("--- STEP 6: NEM 3.0 SENSITIVITY ANALYSIS ---")

candidate_dates = pd.date_range(start='2023-04-01', end='2024-04-01', freq='MS')
results = []

for d in candidate_dates:
    temp_df = analysis_df.copy()
    temp_df['post_nem3_sens']    = (temp_df['month'] >= d).astype(int)
    temp_df['solar_x_nem3_sens'] = temp_df['solar_mw_per_1000_lag1'] * temp_df['post_nem3_sens']

    cols_to_demean = [
        'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
        'solar_mw_per_1000_lag1', 'solar_x_nem3_sens',
        'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight'
    ]

    df_dm    = iterative_demean(temp_df, cols_to_demean)
    df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

    exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm',
                      'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_sens_dm']]
    endog = df_clean[['evs_per_1000_dm']]
    instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]

    mod = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
    res = mod.fit(cov_type='clustered', clusters=df_clean['county_cat'])

    results.append({
        'Date':  d,
        'Coef':  res.params['evs_per_1000_dm'],
        'SE':    res.std_errors['evs_per_1000_dm'],
        'P-val': res.pvalues['evs_per_1000_dm']
    })

sens_df = pd.DataFrame(results)

plt.figure(figsize=(10, 5))
plt.plot(sens_df['Date'], sens_df['Coef'], marker='o', color='blue', label='EV Coefficient')
plt.fill_between(sens_df['Date'],
                 sens_df['Coef'] - 1.645 * sens_df['SE'],
                 sens_df['Coef'] + 1.645 * sens_df['SE'],
                 color='blue', alpha=0.2, label='90% CI')
plt.axhline(0, color='black', linestyle='--')
plt.title('NEM 3.0 Structural Break Sensitivity Analysis')
plt.xlabel('Candidate Break Month')
plt.ylabel('EV Coefficient on MAD Rate')
plt.legend()
plt.show()

best_fit = sens_df.loc[sens_df['SE'].idxmin()]
print(f"Month with best model fit (lowest standard error): {best_fit['Date'].strftime('%Y-%m')}")
print("\nNOTE: Confirm visually from the plot whether the EV coefficient remains stable (confidence band excludes zero) across the tested break dates.")
Cell [60]
# STEP 7A: POLITICAL HETEROGENEITY DATA PREP
print("--- STEP 7A: POLITICAL DATA PROCESSING ---")

try:
    med_dem = analysis_df[['county_name', 'dem_share']].drop_duplicates()['dem_share'].median()
    analysis_df['is_high_dem'] = (analysis_df['dem_share'] > med_dem).astype(int)

    analysis_df['ev_x_high_dem'] = analysis_df['evs_per_1000'] * analysis_df['is_high_dem']
    analysis_df['iv_trend_x_high_dem'] = analysis_df['iv_trend_break'] * analysis_df['is_high_dem']
    analysis_df['iv_stock_x_high_dem'] = analysis_df['iv_anticipation_stock_shift'] * analysis_df['is_high_dem']

    print(f"Median Democratic Share: {med_dem:.4f}")
    print(f"Counties coded as High Dem: {len(analysis_df[analysis_df['is_high_dem']==1]['county_name'].unique())}")

except Exception as e:
    print(f"ERROR processing political data: {e}")
Cell [61]
# === MODIFICATION 7 (REVISED): STEP 7B POLITICAL HETEROGENEITY MODEL ===
# STEP 7B: POLITICAL HETEROGENEITY MODEL (TWFE)
print("--- STEP 7B: POLITICAL HETEROGENEITY MODEL ---")

try:
    print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm + ev_x_high_dem_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + iv_trend_x_high_dem_dm + iv_stock_x_high_dem_dm]\n")

    # iv_trend_x_reliance_weight and iv_stock_x_reliance_weight instrument evs_per_1000_dm;
    # iv_trend_x_high_dem and iv_stock_x_high_dem instrument ev_x_high_dem_dm.
    cols_to_demean = [
        'mad_rate', 'evs_per_1000', 'ev_x_high_dem', 'tmax_c',
        'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
        'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
        'iv_trend_x_high_dem', 'iv_stock_x_high_dem'
    ]

    df_dm = iterative_demean(analysis_df, cols_to_demean)
    df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()

    exog  = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
    endog = df_clean[['evs_per_1000_dm', 'ev_x_high_dem_dm']]
    instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm',
                      'iv_trend_x_high_dem_dm', 'iv_stock_x_high_dem_dm']]

    mod_pol = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
    res_pol = mod_pol.fit(cov_type='clustered', clusters=df_clean['county_cat'])
    print(res_pol.summary)

    print("\n--- FIRST STAGE DIAGNOSTICS ---")
    print(res_pol.first_stage.diagnostics['f.stat'])
    for var in endog.columns:
        f_stat = res_pol.first_stage.diagnostics.loc[var, 'f.stat']
        if f_stat < 10:
            print(f"WARNING: First-stage F-statistic for {var} is {f_stat:.4f} — below 10. Instrument may be weak.")

    base_coef  = res_pol.params['evs_per_1000_dm']
    int_coef   = res_pol.params['ev_x_high_dem_dm']
    p_val      = res_pol.pvalues['ev_x_high_dem_dm']
    tot_effect = base_coef + int_coef
    sig_label  = "significant at 10%" if p_val < 0.10 else "not significant"
    if p_val < 0.10:
        pol_conclusion = f"{sig_label} — EV grid impacts are concentrated in politically distinct county groups, suggesting the burden or benefit is not distributed uniformly across the political geography of California."
    else:
        pol_conclusion = f"{sig_label} — EV grid impacts do not appear to differ systematically by county-level political affiliation; the evidence does not support a politically concentrated distribution of grid volatility effects."

    print(f"""
Base Effect (Low Dem / Conservative Counties):   {base_coef:.6f}
Interaction Effect (Differential):               {int_coef:.6f}
Interaction P-Value:                              {p_val:.4f}
Total Effect (High Dem / Liberal Counties):       {tot_effect:.6f}
Conclusion: {pol_conclusion}
""")

except Exception as e:
    print(f"ERROR running political heterogeneity model: {e}")
Cell [63]
# ============================================================
# CVRP DAILY APPLICATIONS — EVENT CHART
# Visualization style: Sosulski (clean, annotated, purposeful)
# ============================================================
# ---- EASILY EDITABLE PARAMETERS ----
PLOT_START     = '2022-12-31'   # Start of date range to display
PLOT_END       = '2023-12-31'   # End of date range to display
ROLLING_WINDOW = 1              # Days for rolling average smoothing (set to 1 for raw daily)

# Key event dates
DATE_ANNOUNCE  = '2023-08-21'   # CVRP closure announced
DATE_STANDBY   = '2023-09-06'   # Applications placed on standby
DATE_CLOSURE   = '2023-11-08'   # Effective closure

# Colors — edit here to restyle the entire chart
COLOR_RAW      = '#C8D8E8'      # Raw daily bars (muted blue-grey)
COLOR_ROLLING  = '#1A5276'      # Rolling average line (dark navy)
COLOR_SHADE    = '#F5CBA7'      # Anticipation window shading (warm amber)
COLOR_ANNOUNCE = '#E74C3C'      # Announcement line (red)
COLOR_STANDBY  = '#E74C3C'      # Standby line (purple)
COLOR_CLOSURE  = '#E74C3C'      # Closure line (green)
# ---- END EDITABLE PARAMETERS ----

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import matplotlib.ticker as mticker
import warnings
warnings.filterwarnings('ignore')

# ---- BUILD DAILY APPLICATION COUNTS ----
# Ensure app_date column exists; re-parse if needed
cvrp_plot = cvrp_df.copy()
cvrp_plot['app_date'] = pd.to_datetime(cvrp_plot['Application Date'], errors='coerce')
cvrp_plot = cvrp_plot.dropna(subset=['app_date'])

# Aggregate to daily counts
daily = (
    cvrp_plot
    .groupby('app_date')
    .size()
    .reset_index(name='applications')
    .rename(columns={'app_date': 'date'})
)

# Filter to plot range
daily = daily[
    (daily['date'] >= PLOT_START) &
    (daily['date'] <= PLOT_END)
].copy()

# Fill missing dates with zero (no applications = 0, not missing)
full_range = pd.DataFrame({'date': pd.date_range(PLOT_START, PLOT_END)})
daily      = full_range.merge(daily, on='date', how='left').fillna(0)

# Rolling average
daily['rolling'] = daily['applications'].rolling(window=ROLLING_WINDOW, center=True).mean()

# Convert key dates
d_announce = pd.Timestamp(DATE_ANNOUNCE)
d_standby  = pd.Timestamp(DATE_STANDBY)
d_closure  = pd.Timestamp(DATE_CLOSURE)

# ---- CHART ----
fig, ax = plt.subplots(figsize=(14, 6))
fig.patch.set_facecolor('white')
ax.set_facecolor('white')

# Shade the anticipation window (announcement → standby)
ax.axvspan(d_announce, d_closure, color=COLOR_SHADE, alpha=0.35, zorder=1,
           label='Anticipation Window')

# Raw daily bars — thin, muted, background layer
ax.bar(daily['date'], daily['applications'],
       color=COLOR_RAW, width=1.0, alpha=0.6, zorder=2, label='Daily Applications')

# Rolling average — foreground signal
ax.plot(daily['date'], daily['rolling'],
        color=COLOR_ROLLING, linewidth=2.2, zorder=3,
        label=f'{ROLLING_WINDOW}-Day Rolling Average')

# ---- VERTICAL EVENT LINES ----
line_ymax = daily['applications'].max() * 1.05

for date, color, ls in [
    (d_announce, COLOR_ANNOUNCE, '--'),
    (d_standby,  COLOR_STANDBY,  '-'),
    (d_closure,  COLOR_CLOSURE,  ':'),
]:
    ax.axvline(date, color=color, linewidth=1.6, linestyle=ls, zorder=4, alpha=0.9)

# ---- ANNOTATIONS (Sosulski style: horizontal callout with leader line) ----
annotation_cfg = [
    dict(
        date  = d_announce,
        label = "Closure\nAnnounced\nAug 21, 2023",
        color = COLOR_ANNOUNCE,
        x_off = -2,    # days offset for text box (negative = left)
        y_pos = 0.7,  # fraction of y-axis height
        ha    = 'right',
    ),
    dict(
        date  = d_standby,
        label = "Application\nStandby\nSep 6, 2023",
        color = COLOR_STANDBY,
        x_off = 2,
        y_pos = 0.7,
        ha    = 'left',
    ),
    dict(
        date  = d_closure,
        label = "Effective\nClosure\nNov 8, 2023",
        color = COLOR_CLOSURE,
        x_off = 2,
        y_pos = 0.7,
        ha    = 'left',
    ),
]

y_top = daily['applications'].max()

for cfg in annotation_cfg:
    text_x = cfg['date'] + pd.Timedelta(days=cfg['x_off'])
    text_y = y_top * cfg['y_pos']
    ax.annotate(
        cfg['label'],
        xy        = (cfg['date'], text_y * 0.80),        # arrow tip on the line
        xytext    = (text_x, text_y),                     # text box position
        color     = cfg['color'],
        fontsize  = 12,
        fontweight= 'semibold',
        ha        = cfg['ha'],
        va        = 'top',
        arrowprops= dict(
            arrowstyle = '-',
            color      = cfg['color'],
            lw         = 1.2,
            linestyle  = 'dashed',
        ),
        bbox = dict(
            boxstyle    = 'round,pad=0.25',
            facecolor   = 'white',
            edgecolor   = cfg['color'],
            linewidth   = 1.0,
            alpha       = 0.9,
        ),
        zorder = 5,
    )

# ---- SPIKE CALLOUT ----
# Find the peak day and annotate it directly
#peak_row  = daily.loc[daily['applications'].idxmax()]
#peak_date = peak_row['date']
#peak_val  = peak_row['applications']

#ax.annotate(
#    f"Peak: {int(peak_val):,} applications\n{peak_date.strftime('%b %d, %Y')}",
#    xy        = (peak_date, peak_val),
#    xytext    = (peak_date - pd.Timedelta(days=18), peak_val * 0.88),
#    fontsize  = 12,
#    color     = '#1A1A1A',
#    ha        = 'right',
#    va        = 'top',
#    arrowprops= dict(
#       arrowstyle = '-|>',
#        color      = '#555555',
#        lw         = 1.2,
#    ),
#    bbox = dict(
#        boxstyle  = 'round,pad=0.3',
#        facecolor = '#FDFEFE',
#        edgecolor = '#AAAAAA',
#        linewidth = 0.8,
#        alpha     = 0.9,
#    ),
#    zorder = 6,
#)

# ---- AXES STYLING (Sosulski: remove top/right spines, minimal grid) ----
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
ax.spines['left'].set_color('#CCCCCC')
ax.spines['bottom'].set_color('#CCCCCC')

ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, _: f'{int(x):,}'))
import matplotlib.dates as mdates
ax.xaxis.set_major_locator(mdates.MonthLocator())
ax.xaxis.set_major_formatter(mdates.DateFormatter('%b-%y'))
ax.tick_params(axis='both', labelsize=9, color='#AAAAAA')
ax.yaxis.grid(True, color='#EEEEEE', linewidth=0.8, zorder=0)
ax.set_axisbelow(True)

# ---- LABELS ----
# ax.set_xlabel('Application Date', fontsize=10, color='#444444', labelpad=8)
# ax.set_ylabel('Daily Applications', fontsize=10, color='#444444', labelpad=8)
#ax.set_title(
#    'CVRP Daily Applications — Behavioral Surge Around Program Closure',
#    fontsize=13, fontweight='bold', color='#1A1A1A', pad=14, loc='left'
#)

# Subtitle
#fig.text(
#    0.013, 0.91,
#    f'Shaded region = anticipation window (Aug 21 – Sep 6, 2023)  |  '
#    f'{ROLLING_WINDOW}-day rolling average overlaid on raw daily counts',
#    fontsize=8.5, color='#777777'
#)

# ---- LEGEND ----
#handles = [
#    mpatches.Patch(color=COLOR_RAW,     alpha=0.6,  label='Daily Applications'),
#    plt.Line2D([0], [0], color=COLOR_ROLLING, lw=2.2, label=f'{ROLLING_WINDOW}-Day Rolling Avg'),
#    mpatches.Patch(color=COLOR_SHADE,   alpha=0.35, label='Anticipation Window'),
#    plt.Line2D([0], [0], color=COLOR_ANNOUNCE, lw=1.6, ls='--', label='Closure Announced'),
#    plt.Line2D([0], [0], color=COLOR_STANDBY,  lw=1.6, ls='-',  label='Applications on Standby'),
#    plt.Line2D([0], [0], color=COLOR_CLOSURE,  lw=1.6, ls=':',  label='Effective Closure'),
#]
#ax.legend(
#    handles   = handles,
#    loc       = 'upper left',
#    fontsize  = 8.5,
#    frameon   = True,
#    framealpha= 0.9,
#    edgecolor = '#DDDDDD',
#    ncol      = 2,
#)

ax.set_xlim(pd.Timestamp(PLOT_START), pd.Timestamp(PLOT_END))
plt.show()

print(f"\nDate range plotted: {PLOT_START} to {PLOT_END}")
# print(f"Peak day: {peak_date.strftime('%Y-%m-%d')} with {int(peak_val):,} applications")
print(f"Rolling window: {ROLLING_WINDOW} days")
Cell [64]
# ============================================================
# THE POLICY CLIFF — MONTHLY EV ADOPTION
# Visualization style: Sosulski (clean, annotated, purposeful)
# ============================================================
# ---- EASILY EDITABLE PARAMETERS ----
PLOT_START       = '2022-05-01'   # Start of date range
PLOT_END         = '2025-05-01'   # End of date range

# Key event date
DATE_ANNOUNCE    = '2023-08-01'   # CVRP closure announced (month-level)

# Colors
COLOR_LINE       = '#1A3A4A'      # Main line color (dark navy)
COLOR_FILL       = '#E8F4F8'      # Area fill under line (light blue)
COLOR_ANNOUNCE   = '#E74C3C'      # Announcement line (red)
COLOR_ANNOT_BG   = 'white'        # Annotation box background
# ---- END EDITABLE PARAMETERS ----

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
import matplotlib.ticker as mticker
import warnings
warnings.filterwarnings('ignore')

# ---- BUILD MONTHLY EV TOTALS ----
monthly_ev = (
    analysis_df
    .groupby('month')
    .agg(
        total_new_evs = ('monthly_new_evs', 'sum')
    )
    .reset_index()
)

# Filter to plot range
monthly_ev = monthly_ev[
    (monthly_ev['month'] >= PLOT_START) &
    (monthly_ev['month'] <= PLOT_END)
].copy()

# Key date
d_announce = pd.Timestamp(DATE_ANNOUNCE)

# ---- CHART ----
fig, ax = plt.subplots(figsize=(10, 6))
fig.patch.set_facecolor('white')
ax.set_facecolor('white')

# Area fill under the line
ax.fill_between(
    monthly_ev['month'],
    monthly_ev['total_new_evs'],
    color=COLOR_FILL,
    alpha=1.0,
    zorder=1
)

# Main line
ax.plot(
    monthly_ev['month'],
    monthly_ev['total_new_evs'],
    color=COLOR_LINE,
    linewidth=2.0,
    zorder=2
)

# ---- VERTICAL EVENT LINE ----
ax.axvline(
    d_announce,
    color     = COLOR_ANNOUNCE,
    linewidth = 1.6,
    linestyle = '--',
    zorder    = 3,
    alpha     = 0.9
)

# ---- ANNOTATION ----
y_top  = monthly_ev['total_new_evs'].max()
text_x = d_announce - pd.Timedelta(days=10)
text_y = y_top * 0.88

ax.annotate(
    "Closure\nAnnounced\nAug, 2023",
    xy        = (d_announce, text_y * 0.78),
    xytext    = (text_x, text_y),
    color     = COLOR_ANNOUNCE,
    fontsize  = 11,
    fontweight= 'semibold',
    ha        = 'right',
    va        = 'top',
    arrowprops= dict(
        arrowstyle = '-',
        color      = COLOR_ANNOUNCE,
        lw         = 1.2,
        linestyle  = 'dashed',
    ),
    bbox = dict(
        boxstyle  = 'round,pad=0.25',
        facecolor = COLOR_ANNOT_BG,
        edgecolor = COLOR_ANNOUNCE,
        linewidth = 1.0,
        alpha     = 0.9,
    ),
    zorder = 4,
)

# ---- AXES STYLING ----
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
ax.spines['left'].set_color('#CCCCCC')
ax.spines['bottom'].set_color('#CCCCCC')

ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, _: f'{int(x):,}'))
ax.xaxis.set_major_locator(mdates.MonthLocator(interval=4))
ax.xaxis.set_major_formatter(mdates.DateFormatter('%b %Y'))
ax.tick_params(axis='both', labelsize=9, color='#AAAAAA')
ax.yaxis.grid(True, color='#EEEEEE', linewidth=0.8, zorder=0)
ax.set_axisbelow(True)

# Force x-axis to match plot range exactly
ax.set_xlim(pd.Timestamp(PLOT_START), pd.Timestamp(PLOT_END))

# Floor y-axis at zero
ax.set_ylim(bottom=0)

# ---- LABELS ----
ax.set_xlabel('', labelpad=8)
ax.set_ylabel(
    'New EVs Registered (Monthly, All Counties)',
    fontsize  = 10,
    color     = '#444444',
    labelpad  = 8
)
#ax.set_title(
#    'The Policy Cliff: Monthly EV Adoption',
#    fontsize  = 13,
#    fontweight= 'bold',
#    color     = '#1A1A1A',
#    pad       = 14,
#    loc       = 'left'
#)

plt.tight_layout()
plt.show()

print(f"\nDate range plotted: {PLOT_START} to {PLOT_END}")
print(f"Peak month: {monthly_ev.loc[monthly_ev['total_new_evs'].idxmax(), 'month'].strftime('%Y-%m')} "
      f"with {monthly_ev['total_new_evs'].max():,.0f} new EVs")
print(f"Announcement month ({DATE_ANNOUNCE}): "
      f"{monthly_ev.loc[monthly_ev['month'] == DATE_ANNOUNCE, 'total_new_evs'].values[0]:,.0f} new EVs"
      if DATE_ANNOUNCE in monthly_ev['month'].astype(str).values
      else f"Announcement date {DATE_ANNOUNCE} not in plotted range")
Cell [65]
# ============================================================
# THE NEM 2.0 RUSH — MONTHLY SOLAR CAPACITY ADDITIONS
# Visualization style: Sosulski (clean, annotated, purposeful)
# ============================================================
# ---- EASILY EDITABLE PARAMETERS ----
PLOT_START        = '2022-05-01'   # Start of date range
PLOT_END          = '2025-05-01'   # End of date range

# Key NEM 3.0 event dates
DATE_NEM3_APPROVE = '2022-12-01'   # CPUC approves NEM 3.0
DATE_NEM3_DEADLINE= '2023-04-01'   # NEM 2.0 grandfathering deadline (Apr 14)
DATE_BACKLOG      = '2024-01-01'   # NEM 2.0 backlog physically connects to grid

# Colors
COLOR_LINE        = '#1A3A4A'      # Main line (dark navy)
COLOR_FILL        = '#E8F4F8'      # Area fill (light blue)
COLOR_NEM_APPROVE = '#E67E22'      # Approval line (orange)
COLOR_NEM_DEAD    = '#E74C3C'      # Deadline line (red)
COLOR_BACKLOG     = '#1E8449'      # Backlog line (green)
COLOR_ANNOT_BG    = 'white'
# ---- END EDITABLE PARAMETERS ----

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
import matplotlib.ticker as mticker
import warnings
warnings.filterwarnings('ignore')

# ---- BUILD MONTHLY NEW SOLAR CAPACITY (TOTAL MW, SUM ACROSS COUNTIES) ----

# Step 1: recover total MW from per-1000 variable by multiplying back by population
#         solar_mw_per_1000_lag1 × total_population / 1000 = total MW (lagged)
df_solar = analysis_df[['county_name', 'month',
                         'solar_mw_per_1000_lag1',
                         'total_population']].copy()

df_solar['solar_mw_total'] = (
    df_solar['solar_mw_per_1000_lag1'] * df_solar['total_population'] / 1000
)

# Step 2: sum across all counties to get state-level cumulative (still lagged)
solar_monthly = (
    df_solar
    .groupby('month')
    .agg(solar_cumulative_lagged = ('solar_mw_total', 'sum'))
    .reset_index()
    .sort_values('month')
)

# Step 3: undo the 1-month lag to recover the current-month cumulative
solar_monthly['solar_cumulative'] = solar_monthly['solar_cumulative_lagged'].shift(-1)

# Step 4: first difference → new MW installed each month
solar_monthly['new_solar_mw'] = solar_monthly['solar_cumulative'].diff()

# Step 5: clip negatives (data artifacts) and filter to plot range
solar_monthly['new_solar_mw'] = solar_monthly['new_solar_mw'].clip(lower=0)
solar_monthly = solar_monthly[
    (solar_monthly['month'] >= PLOT_START) &
    (solar_monthly['month'] <= PLOT_END)
].dropna(subset=['new_solar_mw']).copy()

# Sanity check
print(f"Months in series: {len(solar_monthly)}")
print(f"Peak month:       {solar_monthly.loc[solar_monthly['new_solar_mw'].idxmax(), 'month'].strftime('%Y-%m')} "
      f"— {solar_monthly['new_solar_mw'].max():,.1f} MW")
print(f"Mean new MW/month: {solar_monthly['new_solar_mw'].mean():,.1f} MW")

# Key dates
d_approve  = pd.Timestamp(DATE_NEM3_APPROVE)
d_deadline = pd.Timestamp(DATE_NEM3_DEADLINE)
d_backlog  = pd.Timestamp(DATE_BACKLOG)

# ---- CHART ----
fig, ax = plt.subplots(figsize=(10, 6))
fig.patch.set_facecolor('white')
ax.set_facecolor('white')

# Shade the NEM 2.0 rush window (approval → deadline)
ax.axvspan(
    d_deadline, d_backlog,
    color='#FDEBD0', alpha=0.5, zorder=1
)

# Area fill
ax.fill_between(
    solar_monthly['month'],
    solar_monthly['new_solar_mw'],
    color=COLOR_FILL,
    alpha=1.0,
    zorder=2
)

# Main line
ax.plot(
    solar_monthly['month'],
    solar_monthly['new_solar_mw'],
    color=COLOR_LINE,
    linewidth=2.0,
    zorder=3
)

# ---- VERTICAL EVENT LINES ----
for date, color, ls in [
#    (d_approve,  COLOR_NEM_APPROVE, '--'),
    (d_deadline, COLOR_NEM_DEAD,    '-'),
#    (d_backlog,  COLOR_BACKLOG,     ':'),
]:
    ax.axvline(date, color=color, linewidth=1.6, linestyle=ls, zorder=4, alpha=0.9)

# ---- ANNOTATIONS ----
y_top = solar_monthly['new_solar_mw'].max()

annotation_cfg = [
#    dict(
#        date   = d_approve,
#        label  = "NEM 3.0\nApproved\nDec 2022",
#        color  = COLOR_NEM_APPROVE,
#        x_off  = -15,
#        y_pos  = 0.98,
#        ha     = 'right',
#    ),
    dict(
        date   = d_deadline,
        label  = "Application\nDeadline\nApr 14, 2023",
        color  = COLOR_NEM_DEAD,
        x_off  = -10,
        y_pos  = 0.92,
        ha     = 'right',
    ),
#    dict(
#        date   = d_backlog,
#        label  = "Backlog Connects\nto Grid\nJan 2024",
#        color  = COLOR_BACKLOG,
#        x_off  = 2,
#        y_pos  = 0.98,
#        ha     = 'left',
#    ),
]

for cfg in annotation_cfg:
    text_x = cfg['date'] + pd.Timedelta(days=cfg['x_off'])
    text_y = y_top * cfg['y_pos']
    ax.annotate(
        cfg['label'],
        xy        = (cfg['date'], text_y * 0.80),
        xytext    = (text_x, text_y),
        color     = cfg['color'],
        fontsize  = 11,
        fontweight= 'semibold',
        ha        = cfg['ha'],
        va        = 'top',
        arrowprops= dict(
            arrowstyle = '-',
            color      = cfg['color'],
            lw         = 1.2,
            linestyle  = 'dashed',
        ),
        bbox = dict(
            boxstyle  = 'round,pad=0.25',
            facecolor = COLOR_ANNOT_BG,
            edgecolor = cfg['color'],
            linewidth = 1.0,
            alpha     = 0.9,
        ),
        zorder = 5,
    )

# ---- PEAK CALLOUT ----
peak_row  = solar_monthly.loc[solar_monthly['new_solar_mw'].idxmax()]
peak_date = peak_row['month']
peak_val  = peak_row['new_solar_mw']

ax.annotate(
    f"Peak: {peak_val:,.0f} MW\n{peak_date.strftime('%b %Y')}",
    xy        = (peak_date, peak_val),
    xytext    = (peak_date + pd.Timedelta(days=45), peak_val * 0.88),
    fontsize  = 12,
    color     = '#1A1A1A',
    ha        = 'left',
    va        = 'top',
    arrowprops= dict(
        arrowstyle = '-|>',
        color      = '#555555',
        lw         = 1.2,
    ),
    bbox = dict(
        boxstyle  = 'round,pad=0.3',
        facecolor = '#FDFEFE',
        edgecolor = '#AAAAAA',
        linewidth = 0.8,
        alpha     = 0.9,
    ),
    zorder = 6,
)

# ---- AXES STYLING ----
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
ax.spines['left'].set_color('#CCCCCC')
ax.spines['bottom'].set_color('#CCCCCC')

ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, _: f'{x:,.0f}'))
ax.xaxis.set_major_locator(mdates.MonthLocator(interval=4))
ax.xaxis.set_major_formatter(mdates.DateFormatter('%b %Y'))
ax.tick_params(axis='both', labelsize=9, color='#AAAAAA')
ax.yaxis.grid(True, color='#EEEEEE', linewidth=0.8, zorder=0)
ax.set_axisbelow(True)

ax.set_xlim(pd.Timestamp(PLOT_START), pd.Timestamp(PLOT_END))
ax.set_ylim(bottom=0)

# ---- LABELS ----
ax.set_xlabel('', labelpad=8)
ax.set_ylabel(
    'New Solar Capacity Installed — Total MW',
    fontsize  = 12,
    color     = '#444444',
    labelpad  = 8
)
#ax.set_title(
#    'The NEM 2.0 Rush: Monthly Solar Capacity Additions',
#    fontsize  = 13,
#    fontweight= 'bold',
#    color     = '#1A1A1A',
#    pad       = 14,
#    loc       = 'left'
#)

plt.tight_layout()
plt.show()
Cell [67]
# ============================================================
# CALIFORNIA COUNTY MAP — CVRP LEGACY RELIANCE RATE
# Visualization style: Sosulski (clean, purposeful, annotated)
# ============================================================
# ---- EASILY EDITABLE PARAMETERS ----
COLORMAP        = 'Blues'     # Matplotlib colormap — try 'Blues', 'RdYlGn_r', 'Oranges', 'YlOrRd'
FIGSIZE         = (10, 12)     # Figure dimensions
LABEL_COUNTIES  = True         # Show county name labels on map
LABEL_MIN_SIZE  = 6            # Minimum font size for labels
SHOW_EXCLUDED   = True         # Show excluded counties (NaN reliance) in gray
COLOR_EXCLUDED  = '#D5D8DC'    # Color for excluded/NaN counties
COLOR_BORDER    = 'white'      # County border color
BORDER_WIDTH    = 0.5          # County border line width
# ---- END EDITABLE PARAMETERS ----

import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import matplotlib.ticker as mticker
from matplotlib.colors import Normalize
from matplotlib.cm import ScalarMappable
import warnings
warnings.filterwarnings('ignore')

# ---- INSTALL AND IMPORT GEOPANDAS ----
try:
    import geopandas as gpd
except ImportError:
    import subprocess
    subprocess.run(['pip', 'install', 'geopandas', '-q'], check=True)
    import geopandas as gpd

# ---- DOWNLOAD CALIFORNIA COUNTY SHAPEFILE FROM CENSUS BUREAU ----
import urllib.request, zipfile, os

shp_url  = "https://www2.census.gov/geo/tiger/GENZ2022/shp/cb_2022_us_county_5m.zip"
shp_zip  = "/tmp/us_counties.zip"
shp_dir  = "/tmp/us_counties"

if not os.path.exists(shp_dir):
    print("Downloading US county shapefile from Census Bureau...")
    urllib.request.urlretrieve(shp_url, shp_zip)
    with zipfile.ZipFile(shp_zip, 'r') as z:
        z.extractall(shp_dir)
    print("Download complete.")

shp_file = [f for f in os.listdir(shp_dir) if f.endswith('.shp')][0]
counties  = gpd.read_file(os.path.join(shp_dir, shp_file))

# Filter to California (FIPS = 06)
ca_counties = counties[counties['STATEFP'] == '06'].copy()
ca_counties['county_name'] = ca_counties['NAME'].str.strip().str.title()

print(f"California counties in shapefile: {len(ca_counties)}")

# ---- BUILD RELIANCE DATA ----
# Get one row per county from analysis_df
reliance_data = (
    analysis_df[['county_name', 'reliance_weight', 'is_high_reliance']]
    .drop_duplicates(subset='county_name')
    .copy()
)

# Also pull from county_reliance to get excluded counties
all_reliance = county_reliance[['county_name', 'reliance_weight']].copy()

print(f"Counties in reliance data:        {len(reliance_data)}")
print(f"Counties in locked sample:        {reliance_data['reliance_weight'].notna().sum()}")

# ---- MERGE SHAPEFILE WITH RELIANCE DATA ----
# Merge locked sample first
ca_map = ca_counties.merge(
    reliance_data[['county_name', 'reliance_weight', 'is_high_reliance']],
    on='county_name',
    how='left'
)

# For counties excluded from locked sample but present in county_reliance,
# fill in their reliance weight so they still show data
ca_map = ca_map.merge(
    all_reliance.rename(columns={'reliance_weight': 'reliance_all'}),
    on='county_name',
    how='left'
)

# Use locked sample weight where available, otherwise use all_reliance
ca_map['plot_reliance'] = ca_map['reliance_weight'].fillna(ca_map['reliance_all'])

# Flag: in locked sample, excluded by threshold, or not in CVRP at all
ca_map['status'] = 'no_data'
ca_map.loc[ca_map['reliance_weight'].notna(), 'status']     = 'included'
ca_map.loc[
    ca_map['reliance_weight'].isna() & ca_map['reliance_all'].notna(),
    'status'
] = 'excluded'

print(f"\nMap county status breakdown:")
print(ca_map['status'].value_counts().to_string())

# ---- COMPUTE CENTROID FOR LABELS ----
ca_map = ca_map.copy()
ca_map['centroid_x'] = ca_map.geometry.centroid.x
ca_map['centroid_y'] = ca_map.geometry.centroid.y

# ---- CHART ----
fig, ax = plt.subplots(figsize=FIGSIZE)
fig.patch.set_facecolor('white')
ax.set_facecolor('white')

# Normalize colormap to reliance range (included counties only)
vmin = ca_map.loc[ca_map['status'] == 'included', 'plot_reliance'].min()
vmax = ca_map.loc[ca_map['status'] == 'included', 'plot_reliance'].max()
norm = Normalize(vmin=vmin, vmax=vmax)
cmap = plt.get_cmap(COLORMAP)

# ---- PLOT INCLUDED COUNTIES (colored by reliance) ----
included = ca_map[ca_map['status'] == 'included'].copy()
included.plot(
    column     = 'plot_reliance',
    ax         = ax,
    cmap       = COLORMAP,
    norm       = norm,
    linewidth  = BORDER_WIDTH,
    edgecolor  = COLOR_BORDER,
    zorder     = 2
)

# ---- PLOT EXCLUDED COUNTIES (gray, below threshold) ----
if SHOW_EXCLUDED:
    excluded = ca_map[ca_map['status'] == 'excluded'].copy()
    if len(excluded) > 0:
        excluded.plot(
            ax        = ax,
            color     = COLOR_EXCLUDED,
            linewidth = BORDER_WIDTH,
            edgecolor = COLOR_BORDER,
            zorder    = 2
        )

# ---- PLOT NO-DATA COUNTIES (light gray) ----
no_data = ca_map[ca_map['status'] == 'no_data'].copy()
if len(no_data) > 0:
    no_data.plot(
        ax        = ax,
        color     = '#F2F3F4',
        linewidth = BORDER_WIDTH,
        edgecolor = COLOR_BORDER,
        zorder    = 2
    )
'''
# ---- COUNTY LABELS ----
if LABEL_COUNTIES:
    for _, row in ca_map.iterrows():
        if pd.isna(row['plot_reliance']) and row['status'] == 'no_data':
            continue
        # Shorten long county names for readability
        name = row['county_name']
        short = (name
                 .replace(' County', '')
                 .replace('San ', 'S. ')
                 .replace('Santa ', 'Sta. ')
                 .replace('Los Angeles', 'L.A.')
                 .replace('San Francisco', 'S.F.'))

        # Scale font size by county area (larger counties get bigger labels)
        area  = row.geometry.area
        fsize = np.clip(6 + np.log10(area + 1) * 0.4, LABEL_MIN_SIZE, 8)

        # Bold label for high-reliance counties
        weight = 'bold' if row.get('is_high_reliance', 0) == 1 else 'normal'
        color  = 'white' if (
            pd.notna(row['plot_reliance']) and
            row['plot_reliance'] > vmin + 0.65 * (vmax - vmin)
        ) else '#1A1A1A'

        ax.text(
            row['centroid_x'], row['centroid_y'],
            short,
            fontsize  = fsize,
            fontweight= weight,
            ha        = 'center',
            va        = 'center',
            color     = color,
            zorder    = 3,
        )
'''
# ---- COLORBAR ----
sm = ScalarMappable(cmap=cmap, norm=norm)
sm.set_array([])
cbar = fig.colorbar(sm, ax=ax, fraction=0.03, pad=0.02, aspect=30)
'''
cbar.set_label(
    'Share of Low-Income Boosted CVRP Applications (2019–2021)',
    fontsize=20, color='#444444', labelpad=10
)
'''
cbar.ax.tick_params(labelsize=8)
cbar.outline.set_edgecolor('#CCCCCC')

'''
# ---- LEGEND FOR EXCLUDED / NO DATA ----
legend_handles = []
if SHOW_EXCLUDED:
    legend_handles.append(
        mpatches.Patch(color=COLOR_EXCLUDED, label='Excluded (<10 applications, NaN reliance)')
    )
legend_handles.append(
    mpatches.Patch(color='#F2F3F4', label='Not in analysis sample')
)
legend_handles.append(
    mpatches.Patch(
        facecolor='none', edgecolor='#1A1A1A',
        linewidth=1.5, label='Bold label = High Reliance (above median)'
    )
)
ax.legend(
    handles   = legend_handles,
    loc       = 'lower left',
    fontsize  = 8,
    frameon   = True,
    framealpha= 0.9,
    edgecolor = '#DDDDDD',
)

# ---- MEDIAN THRESHOLD ANNOTATION ----
median_rel = reliance_data['reliance_weight'].median()
ax.text(
    0.02, 0.12,
    f"Median reliance threshold: {median_rel:.3f}\n"
    f"High reliance: {(reliance_data['is_high_reliance']==1).sum()} counties\n"
    f"Low reliance:  {(reliance_data['is_high_reliance']==0).sum()} counties",
    transform = ax.transAxes,
    fontsize  = 8,
    color     = '#555555',
    va        = 'bottom',
    bbox      = dict(
        boxstyle  = 'round,pad=0.4',
        facecolor = 'white',
        edgecolor = '#CCCCCC',
        alpha     = 0.9
    )
)
'''
# ---- AXES STYLING ----
ax.set_axis_off()

'''
ax.set_title(
    'CVRP Legacy Reliance Rate by California County',
    fontsize  = 20,
    fontweight= 'bold',
    color     = '#1A1A1A',
    pad       = 16,
    loc       = 'left'
)
#fig.text(
#    0.02, 0.96,
#    'Share of CVRP applications receiving low-income boost, 2019–2021 baseline  |  '
#    'Bold county names = above-median reliance (high treatment intensity)',
#    fontsize = 8,
#    color    = '#777777'
#)
'''
plt.tight_layout()
plt.show()

# print(f"\nMedian reliance threshold: {median_rel:.4f}")
print(f"Min reliance (included):   {vmin:.4f}")
print(f"Max reliance (included):   {vmax:.4f}")
Cell [69]
import statsmodels.api as sm
# DIAGNOSTIC 1A: County-level means + Cook's distance / DFBETA on Model 2
print("--- DIAGNOSTIC 1A: COUNTY MEANS & INFLUENCE STATISTICS ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")

controls_pc = ['evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1']

# County-level means (for scatter and reference)
county_means = analysis_df.groupby('county_name').agg(
    mean_evs=('evs_per_1000', 'mean'),
    mean_mad=('mad_rate', 'mean'),
    mean_tmax=('tmax_c', 'mean'),
    mean_dc=('cumulative_data_center_power_mw', 'mean'),
    mean_solar=('solar_mw_per_1000_lag1', 'mean'),
    pop=('total_population', 'mean')
).reset_index()

print("Top 10 counties by EVs/1000:")
print(county_means.sort_values('mean_evs', ascending=False).head(10).to_string(index=False))
print("\nTop 10 counties by mad_rate:")
print(county_means.sort_values('mean_mad', ascending=False).head(10).to_string(index=False))

# Compute influence stats via statsmodels OLS (Model 2 specification)
df2 = analysis_df[['mad_rate'] + controls_pc + ['county_cat']].dropna()
ols_sm = sm.OLS(df2['mad_rate'], sm.add_constant(df2[controls_pc])).fit()
infl = ols_sm.get_influence()
df2 = df2.copy()
df2['cooks_d'] = infl.cooks_distance[0]
df2['dfbeta_evs'] = infl.dfbetas[:, 1]  # column 1 = evs_per_1000 (after const)
df2['leverage'] = infl.hat_matrix_diag

county_infl = df2.groupby('county_cat', observed=True).agg(
    max_cooks=('cooks_d', 'max'),
    sum_abs_dfbeta_evs=('dfbeta_evs', lambda s: s.abs().sum()),
    max_leverage=('leverage', 'max')
).reset_index().sort_values('sum_abs_dfbeta_evs', ascending=False)

n = len(df2)
print(f"\nInfluence thresholds: Cook's D > {4/n:.5f}, |DFBETA| > {2/np.sqrt(n):.5f}")
print("\nCounties ranked by total influence on evs_per_1000 coefficient:")
print(county_infl.head(15).to_string(index=False))
Cell [70]
# DIAGNOSTIC 1B: Sequential leave-out tests
# Re-run Model 2 dropping the top counties by (i) EV penetration, (ii) mad_rate, (iii) DFBETA influence
print("--- DIAGNOSTIC 1B: SEQUENTIAL LEAVE-OUT TESTS ---\n")

def fit_model2(sub_df):
    d = sub_df[['mad_rate'] + controls_pc + ['county_cat']].dropna()
    e = sm.add_constant(d[controls_pc])
    r = IV2SLS(d['mad_rate'], e, None, None).fit(cov_type='clustered', clusters=d['county_cat'])
    return {
        'N': len(d),
        'counties': d['county_cat'].nunique(),
        'evs_coef': r.params['evs_per_1000'],
        'evs_se': r.std_errors['evs_per_1000'],
        'evs_pval': r.pvalues['evs_per_1000'],
    }

def sequential_drop(rank_list, label):
    rows = []
    for k in [0, 1, 2, 3, 4, 5]:
        if k == 0:
            sub = analysis_df
            dropped = "none"
        else:
            sub = analysis_df[~analysis_df['county_name'].isin(rank_list[:k])]
            dropped = ", ".join(rank_list[:k])
        row = {'k': k, 'dropped': dropped, **fit_model2(sub)}
        rows.append(row)
    print(f"### Drop top counties by {label} ###")
    print(pd.DataFrame(rows).to_string(index=False))
    print()

drops_by_ev = county_means.sort_values('mean_evs', ascending=False)['county_name'].tolist()
drops_by_mad = county_means.sort_values('mean_mad', ascending=False)['county_name'].tolist()
drops_by_infl = county_infl['county_cat'].astype(str).tolist()

sequential_drop(drops_by_ev, "EV penetration")
sequential_drop(drops_by_mad, "mad_rate")
sequential_drop(drops_by_infl, "DFBETA influence on evs_per_1000")

# Winsorization check
print("### Winsorize mad_rate at 1st/99th percentile ###")
lo, hi = analysis_df['mad_rate'].quantile([0.01, 0.99])
df_w = analysis_df.copy()
df_w['mad_rate'] = df_w['mad_rate'].clip(lo, hi)
res = fit_model2(df_w)
print(f"Trimmed to [{lo:.4f}, {hi:.4f}]")
print(f"evs_coef={res['evs_coef']:.6e}, SE={res['evs_se']:.6e}, p={res['evs_pval']:.4f}")
Cell [71]
# DIAGNOSTIC 1C: Scatter plot of county means with Bay Area & high-volatility groups labeled
try:
    from adjustText import adjust_text
except ImportError:
    !pip install adjustText -q
    from adjustText import adjust_text

print("--- DIAGNOSTIC 1C: COUNTY MEANS SCATTER ---\n")

bay_area = ['Santa Clara', 'San Mateo', 'Marin', 'Alameda', 'Contra Costa',
            'San Francisco', 'Napa', 'Sonoma', 'Solano']
high_mad = ['Lake', 'Kings', 'Mendocino', 'Nevada', 'Sutter', 'Madera', 'Yuba', 'Colusa']

cm = county_means.copy()
cm['group'] = 'Other'
cm.loc[cm['county_name'].isin(bay_area), 'group'] = 'Bay Area'
cm.loc[cm['county_name'].isin(high_mad), 'group'] = 'High-volatility rural'

fig, ax = plt.subplots(figsize=(11, 7.5))
colors = {'Other': '#bdbdbd', 'Bay Area': '#2166ac', 'High-volatility rural': '#b2182b'}
sizes = {'Other': 60, 'Bay Area': 110, 'High-volatility rural': 110}

for grp, sub in cm.groupby('group'):
    ax.scatter(sub['mean_evs'], sub['mean_mad'],
               s=sizes[grp], c=colors[grp], alpha=0.85,
               edgecolors='black', linewidths=0.6, label=grp, zorder=3)

# OLS through county means (illustrative — between variation only)
slope, intercept = np.polyfit(cm['mean_evs'], cm['mean_mad'], 1)
xx = np.linspace(cm['mean_evs'].min(), cm['mean_evs'].max(), 100)
ax.plot(xx, intercept + slope * xx, color='#444', linestyle='--', linewidth=1.5,
        label=f'OLS through county means (slope={slope:.2e})', zorder=2)

labels_to_show = bay_area + high_mad + ['Los Angeles', 'Orange', 'San Diego']
texts = []
for _, row in cm.iterrows():
    if row['county_name'] in labels_to_show:
        texts.append(ax.text(row['mean_evs'], row['mean_mad'], row['county_name'], fontsize=8.5, zorder=4))
adjust_text(texts, ax=ax, arrowprops=dict(arrowstyle='-', color='#666', lw=0.5))

ax.set_xlabel('Mean EVs per 1,000 residents (county average)', fontsize=11)
ax.set_ylabel('Mean MAD rate (county average)', fontsize=11)
ax.set_title('Between-county relationship: EV penetration vs LMP volatility\n'
             f'(48 counties, {analysis_df["month"].min().strftime("%Y-%m")} to {analysis_df["month"].max().strftime("%Y-%m")})',
             fontsize=12.5, fontweight='bold')
ax.legend(loc='upper right', frameon=True, framealpha=0.95)
ax.grid(True, alpha=0.3)
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
plt.tight_layout()
plt.show()

print(f"\nSlope through county means: {slope:.4e}")
print("For reference, Model 2 panel coef on evs_per_1000 was ~+8.86e-05.")
print("The cross-sectional pattern is essentially flat — Bay Area counties sit on")
print("the line, and high-volatility rural counties sit *above* it (pulling slope DOWN).")
Cell [73]
# DIAGNOSTIC 2A: Re-run Models 2, 2a, 2c using cumulative_evs instead of evs_per_1000
print("--- DIAGNOSTIC 2A: OLS / FE MODELS WITH LEVELS ---\n")

controls_lvl = ['cumulative_evs', 'tmax_c', 'cumulative_data_center_power_mw','total_generator_active_capacity_mw', 'solar_mw_per_1000_lag1']

# Model 2 (no FE) — levels
df_l = analysis_df[['mad_rate'] + controls_lvl + ['county_cat']].dropna()
exog = sm.add_constant(df_l[controls_lvl])
r2_lvl = IV2SLS(df_l['mad_rate'], exog, None, None).fit(cov_type='clustered', clusters=df_l['county_cat'])
print("### Model 2 (no FE) — LEVELS ###")
#print(f"  cumulative_evs: coef={r2_lvl.params['cumulative_evs']:.6e}, SE={r2_lvl.std_errors['cumulative_evs']:.6e}, p={r2_lvl.pvalues['cumulative_evs']:.4f}")
#print(f"  R2={r2_lvl.rsquared:.4f}\n")

print(r2_lvl.summary)

# Models 2a and 2c via PanelOLS — levels
df_panel = analysis_df[['mad_rate'] + controls_lvl + ['county_name', 'month']].dropna().set_index(['county_name', 'month'])
y = df_panel['mad_rate']
X = sm.add_constant(df_panel[controls_lvl])

print("### Model 2a (county FE) — LEVELS ###")
m2a_lvl = PanelOLS(y, X, entity_effects=True, drop_absorbed=True).fit(cov_type='clustered', cluster_entity=True)
#print(f"  cumulative_evs: coef={m2a_lvl.params['cumulative_evs']:.6e}, SE={m2a_lvl.std_errors['cumulative_evs']:.6e}, p={m2a_lvl.pvalues['cumulative_evs']:.4f}")
#print(f"  Within R2={m2a_lvl.rsquared_within:.4f}\n")
print(m2a_lvl.summary)

print("### Model 2c (TWFE) — LEVELS ###")
m2c_lvl = PanelOLS(y, X, entity_effects=True, time_effects=True, drop_absorbed=True).fit(cov_type='clustered', cluster_entity=True)
#print(f"  cumulative_evs: coef={m2c_lvl.params['cumulative_evs']:.6e}, SE={m2c_lvl.std_errors['cumulative_evs']:.6e}, p={m2c_lvl.pvalues['cumulative_evs']:.4f}")
##print(f"  Within R2={m2c_lvl.rsquared_within:.4f}\n")
print(m2c_lvl.summary)

# Side-by-side: per-capita vs levels (TWFE)
df_pc = analysis_df[['mad_rate'] + controls_pc + ['county_name', 'month']].dropna().set_index(['county_name', 'month'])
y_pc = df_pc['mad_rate']
X_pc = sm.add_constant(df_pc[controls_pc])
m2c_pc = PanelOLS(y_pc, X_pc, entity_effects=True, time_effects=True, drop_absorbed=True).fit(cov_type='clustered', cluster_entity=True)

sd_pc = analysis_df['evs_per_1000'].std()
sd_lvl = analysis_df['cumulative_evs'].std()
comp = pd.DataFrame({
    'spec': ['per-capita (evs_per_1000)', 'levels (cumulative_evs)'],
    'coef': [m2c_pc.params['evs_per_1000'], m2c_lvl.params['cumulative_evs']],
    'SE': [m2c_pc.std_errors['evs_per_1000'], m2c_lvl.std_errors['cumulative_evs']],
    'pval': [m2c_pc.pvalues['evs_per_1000'], m2c_lvl.pvalues['cumulative_evs']],
    'sd_x': [sd_pc, sd_lvl],
    'std_effect_per_1sd': [m2c_pc.params['evs_per_1000'] * sd_pc, m2c_lvl.params['cumulative_evs'] * sd_lvl],
})
print("### TWFE side-by-side: per-capita vs levels ###")
print(comp.to_string(index=False))
print("\n(std_effect_per_1sd = change in mad_rate from a 1 SD increase in the EV variable)")
Cell [74]
# DIAGNOSTIC 2B: IV-2SLS Models 6 and 7 with cumulative_evs (LEVELS)
# This is the critical test — does instrument relevance survive when we drop per-capita?
print("--- DIAGNOSTIC 2B: IV-2SLS WITH LEVELS ---\n")

# >>> CHANGE 1: Demean ALL variables we'll need, in one call, into a new dataframe.
# This avoids relying on _dm columns that may or may not exist on analysis_df
# from earlier cells, and avoids mutating the global.
cols_needed = [
    'mad_rate', 'evs_per_1000', 'cumulative_evs',
    'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
    'tmax_c', 'cumulative_data_center_power_mw', 'total_generator_active_capacity_mw',
    'solar_mw_per_1000_lag1', 'solar_x_nem3',
]
df_diag = iterative_demean(analysis_df, cols_needed)

# >>> CHANGE 2: Function now takes `data` as an argument instead of
# reaching out to the global analysis_df. Also validates columns up front.
def run_iv_diagnostic(data, dep, endog, instruments, exog_controls, label):
    required = [dep, endog] + list(instruments) + list(exog_controls) + ['county_cat']
    missing = [c for c in required if c not in data.columns]
    if missing:
        available_dm = [c for c in data.columns if c.endswith('_dm')]
        raise ValueError(
            f"Missing columns: {missing}\nAvailable _dm columns: {available_dm}"
        )
    d = data[required].dropna()
    mod = IV2SLS(
        dependent=d[dep],
        exog=d[exog_controls],
        endog=d[[endog]],
        instruments=d[instruments]
    )
    res = mod.fit(cov_type='clustered', clusters=d['county_cat'])
    print(f"### {label} ###")
    print(f"  N = {len(d)}")
    print(f"  {endog}: coef={res.params[endog]:.6e}, SE={res.std_errors[endog]:.6e}, p={res.pvalues[endog]:.4f}")
    ci = res.conf_int().loc[endog]
    print(f"  95% CI: [{ci['lower']:.6e}, {ci['upper']:.6e}]")
    diag = res.first_stage.diagnostics
    f_col = [c for c in diag.columns if 'f.stat' in c.lower()][0]
    pr2_col = [c for c in diag.columns if 'partial' in c.lower()]
    print(f"  First-stage F: {diag.loc[endog, f_col]:.3f}")
    if pr2_col:
        print(f"  Partial R2:    {diag.loc[endog, pr2_col[0]]:.4f}")
    fs = res.first_stage.individual[endog]
    for inst in instruments:
        print(f"    {inst}: coef={fs.params[inst]:.4f}, t={fs.tstats[inst]:.3f}, p={fs.pvalues[inst]:.4f}")
    try:
        sh = res.sargan
        print(f"  Sargan-Hansen: stat={sh.stat:.4f}, p={sh.pval:.4f}")
    except Exception:
        pass
    print()
    return res

# Per-capita replication (sanity check — should match Models 6 and 7)
print("=" * 70)
print("PER-CAPITA REPLICATION (should match Models 6 and 7)")
print("=" * 70)
m6_pc = run_iv_diagnostic(
    data=df_diag,  # >>> CHANGE 3: pass df_diag explicitly
    dep='mad_rate_dm', endog='evs_per_1000_dm',
    instruments=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
    exog_controls=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm'],
    label="Model 6 (per-capita)"
)
m7_pc = run_iv_diagnostic(
    data=df_diag,
    dep='mad_rate_dm', endog='evs_per_1000_dm',
    instruments=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
    exog_controls=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm'],
    label="Model 7 (per-capita)"
)

# Levels versions
print("=" * 70)
print("LEVELS VERSIONS (cumulative_evs)")
print("=" * 70)
m6_lvl = run_iv_diagnostic(
    data=df_diag,
    dep='mad_rate_dm', endog='cumulative_evs_dm',
    instruments=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
    exog_controls=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm'],
    label="Model 6 (levels)"
)
m7_lvl = run_iv_diagnostic(
    data=df_diag,
    dep='mad_rate_dm', endog='cumulative_evs_dm',
    instruments=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
    exog_controls=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm'],
    label="Model 7 (levels)"
)

# >>> CHANGE 4: Compute sd_pc and sd_lvl here rather than assuming they
# already exist from an earlier cell (which is the next KeyError waiting to happen).
sd_pc  = df_diag['evs_per_1000'].std()
sd_lvl = df_diag['cumulative_evs'].std()

# Standardized comparison
print("=" * 70)
print("STANDARDIZED COMPARISON (effect of 1 SD change in EV variable)")
print("=" * 70)
f_col = [c for c in m6_pc.first_stage.diagnostics.columns if 'f.stat' in c.lower()][0]
summary = pd.DataFrame({
    'spec': ['M6 per-capita', 'M6 levels', 'M7 per-capita', 'M7 levels'],
    'coef': [m6_pc.params['evs_per_1000_dm'], m6_lvl.params['cumulative_evs_dm'],
             m7_pc.params['evs_per_1000_dm'], m7_lvl.params['cumulative_evs_dm']],
    'pval': [m6_pc.pvalues['evs_per_1000_dm'], m6_lvl.pvalues['cumulative_evs_dm'],
             m7_pc.pvalues['evs_per_1000_dm'], m7_lvl.pvalues['cumulative_evs_dm']],
    'first_stage_F': [
        m6_pc.first_stage.diagnostics.iloc[0][f_col],
        m6_lvl.first_stage.diagnostics.iloc[0][f_col],
        m7_pc.first_stage.diagnostics.iloc[0][f_col],
        m7_lvl.first_stage.diagnostics.iloc[0][f_col],
    ],
    'sd_x': [sd_pc, sd_lvl, sd_pc, sd_lvl],
})
summary['effect_per_1sd'] = summary['coef'] * summary['sd_x']
print(summary.to_string(index=False))
Cell [75]
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from scipy import stats

# --- 1. DEFINING TIME PERIODS FROM LOCKED SAMPLE ---
# Using your locked panel dataframe: 'analysis_df'

# Define first 6 months (April 2022 - Sept 2022)
start_period_mask = (analysis_df['month'] >= '2022-04-01') & (analysis_df['month'] <= '2022-09-30')
df_start = analysis_df[start_period_mask].groupby('county_name')[['evs_per_1000', 'mad_rate']].mean().reset_index()

# Define last 6 months (Dec 2024 - May 2025)
end_period_mask = (analysis_df['month'] >= '2024-12-01') & (analysis_df['month'] <= '2025-05-31')
df_end = analysis_df[end_period_mask].groupby('county_name')[['evs_per_1000', 'mad_rate']].mean().reset_index()


# --- 2. FIGURE 1: First 6 Months Cross-Section ---
plt.figure(figsize=(10, 6))
sns.scatterplot(data=df_start, x='evs_per_1000', y='mad_rate', s=80, alpha=0.7)

# Add OLS trendline
slope1, intercept1, r_value1, p_value1, std_err1 = stats.linregress(df_start['evs_per_1000'], df_start['mad_rate'])
plt.plot(df_start['evs_per_1000'], intercept1 + slope1 * df_start['evs_per_1000'], 'k--', label=f'OLS (slope={slope1:.2e})')

# Formatting
plt.title('Figure 1: EV Penetration vs LMP Volatility (First 6 Months: Apr-Sep 2022)', fontsize=14)
plt.xlabel('Mean EVs per 1,000 residents', fontsize=12)
plt.ylabel('Mean MAD rate', fontsize=12)
plt.legend()
#plt.grid(True, alpha=0.3)
plt.tight_layout()
plt.savefig('figure1_start_period.png', dpi=300)
plt.show()


# --- 3. FIGURE 2: Changes Over Time (End Period - Start Period) ---
# Merge start and end periods to calculate changes
df_changes = pd.merge(df_start, df_end, on='county_name', suffixes=('_start', '_end'))
df_changes['delta_evs'] = df_changes['evs_per_1000_end'] - df_changes['evs_per_1000_start']
df_changes['delta_mad'] = df_changes['mad_rate_end'] - df_changes['mad_rate_start']

plt.figure(figsize=(10, 6))
sns.scatterplot(data=df_changes, x='delta_evs', y='delta_mad', s=80, alpha=0.7, color='teal')

# Add OLS trendline for changes
slope2, intercept2, r_value2, p_value2, std_err2 = stats.linregress(df_changes['delta_evs'], df_changes['delta_mad'])
plt.plot(df_changes['delta_evs'], intercept2 + slope2 * df_changes['delta_evs'], 'k--', label=f'OLS (slope={slope2:.2e})')

# Formatting
plt.title('Figure 2: Change in LMP Volatility vs Change in EV Penetration\n(Apr-Sep 2022 to Dec 2024-May 2025)', fontsize=14)
plt.xlabel('Change in Mean EVs per 1,000 residents', fontsize=12)
plt.ylabel('Change in Mean MAD rate', fontsize=12)
plt.legend()
#plt.grid(True, alpha=0.3)

# Add zero lines to easily distinguish positive/negative changes
plt.axhline(0, color='gray', linewidth=1)
plt.axvline(0, color='gray', linewidth=1)

plt.tight_layout()
plt.savefig('figure2_changes.png', dpi=300)
plt.show()

Data Centers · Main causal analysis

The substantive IV regression producing the headline DC findings.

Causal analysis – dc_iv_fiber_2sls.ipynb10 cells · primary

Fiber-instrumented 2SLS producing the headline DC results reported above: the −0.174 log-DC-power coefficient on arcsinh-transformed LMP (collapsed IV, first-stage F ≈ 60), the consistent null on volatility and spikes, and the PSM robustness check on DAC-heavy counties.

Cell [5]
#Importing the dataset
import pandas as pd
df = pd.read_excel('/content/dmn_county_monthly_summary_final_with_DC.xlsx')
df.head()
Cell [6]
#Converting date to date format
df["date"] = pd.to_datetime(df["Year"].astype(str) + "-" + df["Month"].astype(str) + "-01")

#Rename Variables
df.columns = (
    df.columns.str.strip()
              .str.replace(r"[^\w]+", "_", regex=True)
              .str.replace(r"__+", "_", regex=True)
)
Cell [7]
# Import census demographics info
df_census = pd.read_csv('/content/dmn_county_census_demographics_info.csv')

# Rename 'County GEOID' to 'county_geoid' for merging, if different
# And ensure 'median household income' is clean
df_census.rename(columns={'County GEOID': 'county_geoid', 'median household income': 'median_household_income', 'percent_dac_tracts': 'percent_dac_tracts'}, inplace=True)

# Select only unique county_geoid and their median household income and percent_dac_tracts
df_census_unique = df_census[['county_geoid', 'median_household_income', 'percent_dac_tracts', 'percent_below_poverty']].drop_duplicates(subset=['county_geoid'])

# Merge into the main dataframe
df = pd.merge(df, df_census_unique, on='county_geoid', how='left')

print("Median Household Income and Percent DAC Tracts added to df. First 5 rows:")
print(df[['county_geoid', 'county_name', 'median_household_income', 'percent_dac_tracts']].head())
Cell [8]
#STEP #1 - OLS BASIC REGRESSION

import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

# ---------------------------
# 1) Create variables
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]

# Prices: asinh handles negatives
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])

# Create median_income_10k_units and then log-transform it
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])

# ---------------------------
# 2) Helper: clean coefficient table (hide FE rows)
# ---------------------------
def tidy_no_fe(res, drop_prefixes=("C(county_geoid)", "C(date)")):

    tab = res.summary2().tables[1].copy()  # includes Coef., Std.Err., P>|t|/P>|z|, [0.025, 0.975]

    # Identify the p-value column name (can be P>|t| or P>|z|)
    pcol = "P>|t|" if "P>|t|" in tab.columns else ("P>|z|" if "P>|z|" in tab.columns else None)
    if pcol is None:
        raise ValueError("Could not find p-value column in summary2 output.")

    # Drop fixed effects rows
    mask_fe = np.zeros(len(tab), dtype=bool)
    for pref in drop_prefixes:
        mask_fe |= tab.index.to_series().str.startswith(pref)
    tab = tab[~mask_fe]

    # Keep and rename columns nicely
    out = tab.rename(columns={
        "Coef.": "Coef",
        "Std.Err.": "Std_err",
        pcol: "P_value",
        "[0.025": "C.I.95_lo",
        "0.975]": "C.I.95_hi",
    })[["Coef", "Std_err", "P_value", "C.I.95_lo", "C.I.95_hi"]]

    return out

# ---------------------------
# 3) Run regressions (county FE + month FE, clustered by county)
# ---------------------------

# Spikes regression dataset
cols_spikes = ["log_spikes", "log_dc_power", "log_pop", "temp_f", "log_median_income_10k_units", "county_geoid", "date"]
d_spikes = df[cols_spikes].dropna().copy()

m_spikes = smf.ols(
    "log_spikes ~ log_dc_power + log_pop + temp_f + log_median_income_10k_units + C(county_geoid) + C(date)",
    data=d_spikes
).fit(cov_type="cluster", cov_kwds={"groups": d_spikes["county_geoid"]})

# Prices regression dataset
cols_prices = ["asinh_lmp", "log_dc_power", "log_pop", "temp_f", "log_median_income_10k_units", "county_geoid", "date"]
d_prices = df[cols_prices].dropna().copy()

m_prices = smf.ols(
    "asinh_lmp ~ log_dc_power + log_pop + temp_f + log_median_income_10k_units + C(county_geoid) + C(date)",
    data=d_prices
).fit(cov_type="cluster", cov_kwds={"groups": d_prices["county_geoid"]})

# ---------------------------
# 4) Print clean tables (no FE rows)
# ---------------------------

print("\n=== Spikes model (log1p) — FE included, FE rows hidden ===")
print(tidy_no_fe(m_spikes))

print("\n=== Prices model (asinh) — FE included, FE rows hidden ===")
print(tidy_no_fe(m_prices))
Cell [13]
# If needed in Colab (run once):
!pip -q install linearmodels
from linearmodels.iv import IV2SLS
Cell [14]
# STEP #2 — IV REGRESSION (2SLS) in the PANEL (county × month)
# Spec: Month FE only, clustered SE by county
# Outcomes: log_spikes and asinh_lmp
# Endogenous regressor: log_dc_power
# Instrument: fiber provider counts, average max upstream speed, average max downstream speed
# ============================================================

import numpy as np
import pandas as pd

# If needed in Colab (run once):
!pip -q install linearmodels
from linearmodels.iv import IV2SLS

# ---------------------------
# 1) Create variables
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]

# Prices: asinh handles negatives
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])

# Instrument (fiber providers)
df["fiber"] = df["Fiber_Provider_Counts"]
df["log_upstream"] = np.log1p(df["Average_Max_Upstream"])
df["log_downstream"] = np.log1p(df["Average_Max_Downstream"])

# Create median_income_10k_units and then log-transform it
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])

# ---------------------------
# 2) Helpers: tidy IV output + drop FE rows from display
# ---------------------------
def tidy_iv(res):
    """Return coef, std err, p-values, and 95% CI from linearmodels IV results."""
    return pd.DataFrame({
        "Coef": res.params,
        "Std_err": res.std_errors,
        "P_value": res.pvalues,
        "C.I.95_lo": res.conf_int().iloc[:, 0],
        "C.I.95_hi": res.conf_int().iloc[:, 1],
    })

def drop_fe_rows(tab, prefixes=("C(date)",)):
    """Remove FE dummy rows from the printed table (FE stay in the model)."""
    idx = tab.index.astype(str)
    mask_fe = np.zeros(len(tab), dtype=bool)
    for pref in prefixes:
        mask_fe |= pd.Series(idx).str.startswith(pref).values
    return tab[~mask_fe]

# ---------------------------
# 3) Build estimation datasets (drop NaNs consistently)
# ---------------------------
cols_spikes = ["county_geoid", "date", "log_spikes", "log_dc_power", "fiber", "log_pop", "temp_f", "log_median_income_10k_units","log_upstream", "log_downstream"]
d_spikes = df[cols_spikes].dropna().copy()

cols_prices = ["county_geoid", "date", "asinh_lmp", "log_dc_power", "fiber", "log_pop", "temp_f", "log_median_income_10k_units","log_upstream", "log_downstream"]
d_prices = df[cols_prices].dropna().copy()

# ---------------------------
# 4) Run IV regressions (2SLS) — Month FE only
# ---------------------------

# Outcome: spikes
iv_spikes_monthfe = IV2SLS.from_formula(
    "log_spikes ~ 1 + log_pop + temp_f + log_median_income_10k_units + C(date) + [log_dc_power ~ fiber + log_upstream + log_downstream]",
    data=d_spikes
).fit(cov_type="clustered", clusters=d_spikes["county_geoid"])

# Outcome: prices
iv_prices_monthfe = IV2SLS.from_formula(
    "asinh_lmp ~ 1 + log_pop + temp_f + log_median_income_10k_units + C(date) + [log_dc_power ~ fiber + log_upstream + log_downstream]",
    data=d_prices
).fit(cov_type="clustered", clusters=d_prices["county_geoid"])

# ---------------------------
# 5) Print clean coefficient tables (hide month FE dummies)
# ---------------------------
print("\n=== IV (2SLS) — Spikes | Month FE | Clustered by county ===")
t_spikes = drop_fe_rows(tidy_iv(iv_spikes_monthfe), prefixes=("C(date)",))
print(t_spikes)

print("\n=== IV (2SLS) — Prices | Month FE | Clustered by county ===")
t_prices = drop_fe_rows(tidy_iv(iv_prices_monthfe), prefixes=("C(date)",))
print(t_prices)

# ---------------------------
# 6) First-stage diagnostics (print for BOTH outcomes)
# ---------------------------
print("\n=== First-stage diagnostics (Spikes sample) ===")
print(iv_spikes_monthfe.first_stage)

print("\n=== First-stage diagnostics (Prices sample) ===")
print(iv_prices_monthfe.first_stage)
Cell [17]
# STEP #3 — COLLAPSED REGRESSION (cross-county / between design)
# Collapse county × month -> county-level means
# Outcomes: log_spikes (mean) and asinh_lmp (mean)
# Regressor: log_dc_power (mean)
# Controls: log_pop (mean), temp_f (mean)
# Robust SE (HC3). No fixed effects needed after collapsing.

import numpy as np
import pandas as pd
import statsmodels.formula.api as smf

# ---------------------------
# 1) Create variables (same as before)
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])

# Create median_income_10k_units and then log-transform it
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])

# ---------------------------
# 2) Helpers: tidy OLS output (no FE rows to drop here, but keep consistent style)
# ---------------------------
def tidy_ols(res):
    tab = res.summary2().tables[1].copy()
    pcol = "P>|t|" if "P>|t|" in tab.columns else ("P>|z|" if "P>|z|" in tab.columns else None)
    if pcol is None:
        raise ValueError("Could not find p-value column in summary2 output.")

    out = tab.rename(columns={
        "Coef.": "Coef",
        "Std.Err.": "Std_err",
        pcol: "P_value",
        "[0.025": "C.I.95_lo",
        "0.975]": "C.I.95_hi",
    })[["Coef", "Std_err", "P_value", "C.I.95_lo", "C.I.95_hi"]]
    return out

# ---------------------------
# 3) Collapse to cross-county dataset (means by county)
# ---------------------------
cols_collapse = ["county_geoid", "log_spikes", "asinh_lmp", "log_dc_power", "log_pop", "temp_f", "log_median_income_10k_units"]
g = (df[cols_collapse]
     .dropna()
     .groupby("county_geoid", as_index=False)
     .mean()
)

# ---------------------------
# 4) Run collapsed OLS regressions (cross-county)
# ---------------------------

# Outcome: spikes (mean)
collapsed_spikes = smf.ols(
    "log_spikes ~ log_dc_power + log_pop + temp_f + log_median_income_10k_units",
    data=g
).fit(cov_type="HC3")

# Outcome: prices (mean)
collapsed_prices = smf.ols(
    "asinh_lmp ~ log_dc_power + log_pop + temp_f + log_median_income_10k_units",
    data=g
).fit(cov_type="HC3")

# ---------------------------
# 5) Print clean coefficient tables
# ---------------------------
print("\n=== Collapsed OLS — Spikes (county means) | Robust SE (HC3) ===")
print(tidy_ols(collapsed_spikes))

print("\n=== Collapsed OLS — Prices (county means) | Robust SE (HC3) ===")
print(tidy_ols(collapsed_prices))
Cell [20]
# STEP #3 — COLLAPSED REGRESSION (cross-county / between design)
# Collapse county × month -> county-level means
# Outcomes: log_spikes (mean) and asinh_lmp (mean)
# Regressor: log_dc_power (mean)
# Instrument: fiber (mean), log upstream + downstream speeds
# Controls: log_pop (mean), temp_f (mean)
# Robust SE. No fixed effects needed after collapsing.

import numpy as np
import pandas as pd
from linearmodels.iv import IV2SLS

# ---------------------------
# 1) Create variables (same as before)
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])
df["fiber"] = df["Fiber_Provider_Counts"]
df["log_upstream"] = np.log1p(df["Average_Max_Upstream"])
df["log_downstream"] = np.log1p(df["Average_Max_Downstream"])

# Create median_income_10k_units and then log-transform it
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])

# ---------------------------
# 2) Helpers: tidy IV output
# ---------------------------
def tidy_iv(res):
    """Return coef, std err, p-values, and 95% CI from linearmodels IV results."""
    return pd.DataFrame({
        "Coef": res.params,
        "Std_err": res.std_errors,
        "P_value": res.pvalues,
        "C.I.95_lo": res.conf_int().iloc[:, 0],
        "C.I.95_hi": res.conf_int().iloc[:, 1],
    })

# ---------------------------
# 3) Collapse to cross-county dataset (means by county)
# ---------------------------
cols_collapse = [
    "county_geoid", "log_spikes", "asinh_lmp",
    "log_dc_power", "log_pop", "temp_f", "fiber","log_median_income_10k_units",
    "log_upstream", "log_downstream"
]
g = (
    df[cols_collapse]
    .dropna()
    .groupby("county_geoid", as_index=False)
    .mean()
)

# ---------------------------
# 4) Run collapsed IV regressions (2SLS) — cross-county
# ---------------------------

# Outcome: spikes (mean)
iv_collapsed_spikes = IV2SLS.from_formula(
    "log_spikes ~ log_pop + temp_f + log_median_income_10k_units + [log_dc_power ~ fiber + log_upstream + log_downstream]",
    data=g
).fit(cov_type="robust")

# Outcome: prices (mean)
iv_collapsed_prices = IV2SLS.from_formula(
    "asinh_lmp ~ log_pop + temp_f + log_median_income_10k_units + [log_dc_power ~ fiber + log_upstream + log_downstream]",
    data=g
).fit(cov_type="robust")

# ---------------------------
# 5) Print clean coefficient tables
# ---------------------------
print("\n=== Collapsed IV (2SLS) — Spikes (county means) | Robust SE ===")
print(tidy_iv(iv_collapsed_spikes))

print("\n=== Collapsed IV (2SLS) — Prices (county means) | Robust SE ===")
print(tidy_iv(iv_collapsed_prices))

# ---------------------------
# 6) First-stage diagnostics (print for BOTH outcomes)
# ---------------------------
print("\n=== First-stage diagnostics (Spikes sample - collapsed) ===")
print(iv_collapsed_spikes.first_stage)

print("\n=== First-stage diagnostics (Prices sample - collapsed) ===")
print(iv_collapsed_prices.first_stage)
Cell [29]
import statsmodels.formula.api as smf

# Ensure all necessary variables are in the 'g' DataFrame
# The 'g' DataFrame was prepared in cell 'fo_kWVJVYhjR' with these columns.

print("\n=== OLS Robustness Check for Exclusion Restriction (Prices) ===")
print("\nOutcome: asinh_lmp | Explanatory: log_dc_power + Instruments + Controls")

m_prices_exclusion_check = smf.ols(
    "asinh_lmp ~ log_dc_power + fiber + log_upstream + log_downstream + log_pop + temp_f + log_median_income_10k_units",
    data=g
).fit(cov_type="HC3")

# Helper function to print clean tables (adapted from earlier helper)
def tidy_ols(res):
    tab = res.summary2().tables[1].copy()
    pcol = "P>|t|" if "P>|t|" in tab.columns else ("P>|z|" if "P>|z|" in tab.columns else None)
    if pcol is None:
        raise ValueError("Could not find p-value column in summary2 output.")

    out = tab.rename(columns={
        "Coef.": "Coef",
        "Std.Err.": "Std_err",
        pcol: "P_value",
        "[0.025": "C.I.95_lo",
        "0.975]": "C.I.95_hi",
    })[["Coef", "Std_err", "P_value", "C.I.95_lo", "C.I.95_hi"]]
    return out

print(tidy_ols(m_prices_exclusion_check))

print("\n--- Interpretation Hint ---")
print("If the instruments (fiber, log_upstream, log_downstream) were *valid*, their coefficients in this OLS regression should be statistically *insignificant* when log_dc_power is already in the model. A significant coefficient suggests a direct effect on prices, violating the exclusion restriction.")
Cell [33]
# ============================================================
# STEP #4 — PROPENSITY SCORE MATCHING (PSM) [Corrected v2] - CALIPER 0.15
# Treatment: above median percent_dac_tracts vs below median percent_dac_tracts
# Unit: County (collapsed to county-level means)
# Outcomes: log_spikes (mean) and asinh_lmp (mean)
# Matching: 1-to-1 nearest neighbor on propensity score + optional caliper
# ============================================================

import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import NearestNeighbors

# --- Re-import and merge census data to ensure necessary columns are present ---
df_census = pd.read_csv('/content/dmn_county_census_demographics_info.csv')
df_census.rename(columns={'County GEOID': 'county_geoid', 'median household income': 'median_household_income', 'percent_dac_tracts': 'percent_dac_tracts'}, inplace=True)
df_census_unique = df_census[['county_geoid', 'median_household_income', 'percent_dac_tracts']].drop_duplicates(subset=['county_geoid'])

# Identify columns to be updated/added from df_census_unique
census_cols_to_merge = ['median_household_income', 'percent_dac_tracts']

# Drop existing versions of these columns from df before merging to avoid conflicts
for col in census_cols_to_merge:
    if col in df.columns:
        df = df.drop(columns=[col])
    if col + '_orig' in df.columns: # Also drop any _orig versions that might clash if present
        df = df.drop(columns=[col + '_orig'])

# Perform the merge. Since we dropped potential conflicts, no suffixes are needed.
# This ensures the columns from df_census_unique are added cleanly.
df = pd.merge(df, df_census_unique, on='county_geoid', how='left')
# ------------------------------------------------------------------------------------

# ---------------------------
# 1) Create variables
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])
df["fiber"] = df["Fiber_Provider_Counts"]
df["log_upstream"] = np.log1p(df["Average_Max_Upstream"])
df["log_downstream"] = np.log1p(df["Average_Max_Downstream"])

df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])

# ---------------------------
# 2) Collapse to county-level means
# ---------------------------
cols_psm = ["county_geoid", "log_spikes", "asinh_lmp", "log_dc_power", "log_pop", "temp_f", "fiber", "log_upstream", "log_downstream", "log_median_income_10k_units", "percent_dac_tracts"]

g = (
    df[cols_psm]
    .dropna() # Drop NaNs after adding all columns needed for PSM
    .groupby("county_geoid", as_index=False)
    .mean()
)

# --- NEW: Filter to counties WITH Data Centers only ---
g_filtered_dc = g[g["log_dc_power"] > 0].copy()

# ---------------------------
# 3) Define treatment: above median percent_dac_tracts vs below median percent_dac_tracts
# ---------------------------
# Calculate median for the filtered group (counties with DC > 0)
median_dac_percent = g_filtered_dc["percent_dac_tracts"].median()

g_psm = g_filtered_dc.copy()
g_psm["treat"] = (g_psm["percent_dac_tracts"] >= median_dac_percent).astype(int)

print("=== Treatment counts (must include 0 and 1) ===")
print(g_psm["treat"].value_counts())
print("\nMedian Percent DAC Tracts (for counties with DC > 0):", median_dac_percent)
print("Counties in PSM sample:", g_psm.shape[0])

# ---------------------------
# 4) Estimate propensity scores
# ---------------------------
covariates = ["log_dc_power", "log_pop", "temp_f", "fiber", "log_upstream", "log_downstream", "log_median_income_10k_units"]
g_psm = g_psm.dropna(subset=covariates + ["treat", "log_spikes", "asinh_lmp", "percent_dac_tracts"]).copy() # Ensure all relevant columns are dropped consistently

X = g_psm[covariates].values
t = g_psm["treat"].values

scaler = StandardScaler()
Xz = scaler.fit_transform(X)

logit = LogisticRegression(max_iter=1000, solver="lbfgs")
logit.fit(Xz, t)

g_psm["pscore"] = logit.predict_proba(Xz)[:, 1]

# ---------------------------
# 5) Matching (reset indices to avoid index alignment errors)
# ---------------------------
treated = g_psm[g_psm["treat"] == 1].reset_index(drop=True).copy()
control = g_psm[g_psm["treat"] == 0].reset_index(drop=True).copy()

nn = NearestNeighbors(n_neighbors=1, metric="euclidean")
nn.fit(control[["pscore"]].values)

dist, idx = nn.kneighbors(treated[["pscore"]].values)

# idx is position in CONTROL (0..len(control)-1), aligned to treated rows (0..len(treated)-1)
matched_control = control.iloc[idx.flatten()].reset_index(drop=True).copy()
matched_treated = treated.reset_index(drop=True).copy()
matched_treated["match_dist"] = dist.flatten()

# Optional caliper
caliper = 0.3
keep = matched_treated["match_dist"] <= caliper

matched_treated = matched_treated.loc[keep].reset_index(drop=True)
matched_control = matched_control.loc[keep].reset_index(drop=True)

print("\nMatched pairs after caliper:", len(matched_treated))

# ---------------------------
# 6) ATT (Average Treatment effect on the Treated)
# ---------------------------
att_spikes = (matched_treated["log_spikes"] - matched_control["log_spikes"]).mean()
att_prices = (matched_treated["asinh_lmp"] - matched_control["asinh_lmp"]).mean()

# ---------------------------
# 7) Results tables
# ---------------------------
psm_results = pd.DataFrame({
    "Outcome": ["log_spikes", "asinh_lmp"],
    "ATT (Above Median % DACs - Below Median % DACs)": [att_spikes, att_prices],
    "Matched pairs": [len(matched_treated), len(matched_treated)],
    "Caliper": [caliper, caliper],
})

print("\n=== Propensity Score Matching Results (ATT) ===")
print(psm_results)

balance_post = pd.DataFrame({
    "Mean Treated (matched)": matched_treated[covariates].mean(),
    "Mean Control (matched)": matched_control[covariates].mean(),
    "Diff (T - C)": matched_treated[covariates].mean() - matched_control[covariates].mean()
})

Data Centers · Supporting pipeline · data engineering

The four notebooks that build the DC dataset — imputation of missing fields, web-scraping facility attributes and founding dates, and constructing the fiber-broadband instrument.

Imputation — Total Power / Usable Space – dc_imputation_power_space.ipynb23 cells · supporting

Random Forest / XGBoost / KNN / Linear imputation for missing Total Power and Usable Space fields in the 197-facility dataset, with 5-fold cross-validation.

Cell [2]
import pandas as pd


# Load the data
df = pd.read_excel('datacenter_list_11.6.25.xlsx', engine='openpyxl')


df.info()


Cell [4]
# df_train: rows where both 'usable_space' and 'total_power' are not null
df_train = df[df['Usable Space'].notna() & df['Total Power (MW)'].notna()]

df_train.info()
Cell [6]
from sklearn.model_selection import KFold, cross_val_score

# Define k-fold cross-validation
kf = KFold(n_splits=5, shuffle=True, random_state=42)
Cell [8]
from sklearn.model_selection import cross_validate
from sklearn.metrics import make_scorer, r2_score, mean_absolute_error, mean_squared_error
import numpy as np

# If sklearn >= 0.24, you can import this directly:
from sklearn.metrics import mean_absolute_percentage_error

def mape(y_true, y_pred):
    return mean_absolute_percentage_error(y_true, y_pred) * 100  # percentage

# Define scoring dictionary
scoring = {
    "R2": make_scorer(r2_score),
    "MAE": make_scorer(mean_absolute_error),
    "RMSE": make_scorer(lambda y_true, y_pred: np.sqrt(mean_squared_error(y_true, y_pred))),
    "MAPE": make_scorer(mape)
}

# Initialize results list
results = []
Cell [10]
from sklearn.preprocessing import StandardScaler

# List of numeric features to scale
numeric_features = ["Usable Space", "Total Power (MW)", "Year", "Latitude", "Longitude"]

# Initialize scaler
scaler = StandardScaler()

# Fit and transform numeric features
scaled_values = scaler.fit_transform(df_train[numeric_features])

# Create a DataFrame with new column names for scaled features
scaled_df = pd.DataFrame(scaled_values,
                         columns=[f"{col}_scaled" for col in numeric_features],
                         index=df_train.index)

# Add the scaled features to df_train without replacing the originals
df_train = pd.concat([df_train, scaled_df], axis=1)


Cell [12]

from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_validate
from sklearn.metrics import make_scorer, r2_score, mean_absolute_error, mean_squared_error
import statsmodels.api as sm
import numpy as np




# -------- Predicting Total Power --------
X_total_power = df_train[["Usable Space_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]]  # exclude Total Power
y_total_power = df_train["Total Power (MW)"]


model_lr_tp = LinearRegression()
cv_lr_tp = cross_validate(model_lr_tp, X_total_power, y_total_power, cv=kf, scoring=scoring)

# Fit statsmodels OLS for summary
X_sm_tp = sm.add_constant(X_total_power)
ols_tp = sm.OLS(y_total_power, X_sm_tp).fit()

print("\n=== OLS Regression Summary: Predicting Total Power ===")
print(ols_tp.summary())

# Append metrics to results
results.append({
    "Target": "Total Power",
    "Model": "Linear Regression (OLS)",
    "R2": np.mean(cv_lr_tp["test_R2"]),
    "MAE": np.mean(cv_lr_tp["test_MAE"]),
    "RMSE": np.mean(cv_lr_tp["test_RMSE"]),
    "MAPE": np.mean(cv_lr_tp["test_MAPE"])
})

# -------- Predicting Usable Space --------
X_usable_space = df_train[["Total Power (MW)_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]]
y_usable_space = df_train["Usable Space"]

model_lr_us = LinearRegression()
cv_lr_us = cross_validate(model_lr_us, X_usable_space, y_usable_space, cv=kf, scoring=scoring)

# Fit statsmodels OLS for summary
X_sm_us = sm.add_constant(X_usable_space)
ols_us = sm.OLS(y_usable_space, X_sm_us).fit()

print("\n=== OLS Regression Summary: Predicting Usable Space ===")
print(ols_us.summary())

# Append metrics to results
results.append({
    "Target": "Usable Space",
    "Model": "Linear Regression (OLS)",
    "R2": np.mean(cv_lr_us["test_R2"]),
    "MAE": np.mean(cv_lr_us["test_MAE"]),
    "RMSE": np.mean(cv_lr_us["test_RMSE"]),
    "MAPE": np.mean(cv_lr_us["test_MAPE"])
})

# Convert results to DataFrame
results_df = pd.DataFrame(results)
print("\n=== Cross-Validation Results ===")
print(results_df)
Cell [14]
from sklearn.neighbors import KNeighborsRegressor
from sklearn.model_selection import GridSearchCV
from sklearn.preprocessing import StandardScaler
from sklearn.inspection import permutation_importance


# -------- Predicting Total Power --------
X_total_power = df_train[["Usable Space_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]]  # exclude Total Power
y_total_power = df_train["Total Power (MW)"]


# Grid search for best k
param_grid = {'n_neighbors': list(range(1, 10))}
knn = KNeighborsRegressor()
grid_total = GridSearchCV(knn, param_grid, cv=kf, scoring='r2')
grid_total.fit(X_total_power, y_total_power)
best_knn_total = grid_total.best_estimator_

# Cross-validation
cv_knn_tp = cross_validate(best_knn_total, X_total_power, y_total_power, cv=kf, scoring=scoring)


# Permutation importance
perm_imp_total = permutation_importance(best_knn_total, X_total_power, y_total_power, n_repeats=10, random_state=42)
feature_importance_total = pd.Series(perm_imp_total.importances_mean, index=X_total_power.columns)

# Append results
results.append({
    "Target": "Total Power",
    "Model": "KNN",
    "R2": np.mean(cv_knn_tp['test_R2']),
    "MAE": np.mean(cv_knn_tp['test_MAE']),
    "RMSE": np.mean(cv_knn_tp['test_RMSE']),
    "MAPE": np.mean(cv_knn_tp['test_MAPE'])
})

print("Best k for Total Power:", grid_total.best_params_['n_neighbors'])
print("Feature importance (Total Power):")
print(feature_importance_total.sort_values(ascending=False))

# ---------------------------
# Model 2: Predict Usable Space
# ---------------------------
X_usable_space = df_train[["Total Power (MW)_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]]
y_usable_space = df_train["Usable Space"]


# Grid search for best k
grid_usable = GridSearchCV(knn, param_grid, cv=kf, scoring="r2")
grid_usable.fit(X_usable_space, y_usable_space)
best_knn_us = grid_usable.best_estimator_

# Cross-validation
cv_knn_us = cross_validate(best_knn_us, X_usable_space, y_usable_space, cv=kf, scoring=scoring)


# Permutation importance
perm_imp_us = permutation_importance(best_knn_us, X_usable_space, y_usable_space, n_repeats=10, random_state=42)
feature_importance_us = pd.Series(perm_imp_us.importances_mean, index=X_usable_space.columns)

# Append results
results.append({
    "Target": "Usable Space",
    "Model": "KNN",
    "R2": np.mean(cv_knn_us['test_R2']),
    "MAE": np.mean(cv_knn_us['test_MAE']),
    "RMSE": np.mean(cv_knn_us['test_RMSE']),
    "MAPE": np.mean(cv_knn_us['test_MAPE'])
})

print("Best k for Usable Space:", grid_usable.best_params_['n_neighbors'])
print("Feature importance (Usable Space):")
print(feature_importance_us.sort_values(ascending=False))

# Convert results to DataFrame
results_df = pd.DataFrame(results)
print(results_df)
Cell [16]
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import GridSearchCV, cross_validate
import pandas as pd
import numpy as np

#Using zipcode instead of lat/long that had a slightly worse result

# -------- Predicting Total Power --------
X_total_power = df_train[["Usable Space", "Colocation Flag", "Year", "Zipcode"]]
y_total_power = df_train["Total Power (MW)"]

param_grid = {
    'n_estimators': [50, 100, 200],
    'max_depth': [None, 5, 10],
    'random_state': [42]
}

rf = RandomForestRegressor()
grid_total = GridSearchCV(rf, param_grid, cv=kf, scoring='r2')
grid_total.fit(X_total_power, y_total_power)
best_rf_total = grid_total.best_estimator_

# Cross-validation
cv_rf_tp = cross_validate(best_rf_total, X_total_power, y_total_power, cv=kf, scoring=scoring)

# Feature importance (computed from the model trained on the same unscaled features)
feature_importance_total = pd.Series(best_rf_total.feature_importances_, index=X_total_power.columns)

# Append results
results.append({
    "Target": "Total Power",
    "Model": "Random Forest",
    "R2": np.mean(cv_rf_tp['test_R2']),
    "MAE": np.mean(cv_rf_tp['test_MAE']),
    "RMSE": np.mean(cv_rf_tp['test_RMSE']),
    "MAPE": np.mean(cv_rf_tp['test_MAPE'])
})

print("Best params for Total Power:", grid_total.best_params_)
print("Feature importance (Total Power):")
print(feature_importance_total.sort_values(ascending=False))

# ---------------------------
# Model 2: Predict Usable Space
# ---------------------------
X_usable_space = df_train[["Total Power (MW)", "Colocation Flag", "Year", "Zipcode"]]
y_usable_space = df_train["Usable Space"]

grid_usable = GridSearchCV(rf, param_grid, cv=kf, scoring='r2')
grid_usable.fit(X_usable_space, y_usable_space)
best_rf_us = grid_usable.best_estimator_

# Cross-validation
cv_rf_us = cross_validate(best_rf_us, X_usable_space, y_usable_space, cv=kf, scoring=scoring)

# Feature importance
feature_importance_us = pd.Series(best_rf_us.feature_importances_, index=X_usable_space.columns)

# Append results
results.append({
    "Target": "Usable Space",
    "Model": "Random Forest",
    "R2": np.mean(cv_rf_us['test_R2']),
    "MAE": np.mean(cv_rf_us['test_MAE']),
    "RMSE": np.mean(cv_rf_us['test_RMSE']),
    "MAPE": np.mean(cv_rf_us['test_MAPE'])
})

print("Best params for Usable Space:", grid_usable.best_params_)
print("Feature importance (Usable Space):")
print(feature_importance_us.sort_values(ascending=False))

# ---------------------------
# Convert results to DataFrame
# ---------------------------
results_df = pd.DataFrame(results)
print("\n=== Cross-Validation Results ===")
print(results_df)
Cell [18]
from xgboost import XGBRegressor

#Using Lat/Long instead of Zipcode that produced slightly higher results

# ---------------------------
# Model 1: Predict Total Power
# ---------------------------

X_total_power = df_train[["Usable Space", "Colocation Flag", "Year", "Latitude", "Longitude"]]
y_total_power = df_train["Total Power (MW)"]

param_grid = {
    'n_estimators': [50, 100, 200],
    'max_depth': [3, 5, 10],
    'learning_rate': [0.01, 0.1, 0.2],
    'subsample': [0.7, 1.0],
    'random_state': [42]
}

xgb = XGBRegressor(objective='reg:squarederror')
grid_total = GridSearchCV(xgb, param_grid, cv=kf, scoring='r2')
grid_total.fit(X_total_power, y_total_power)
best_xgb_total = grid_total.best_estimator_

# Cross-validation
cv_xgb_tp = cross_validate(best_xgb_total, X_total_power, y_total_power, cv=kf, scoring=scoring)

# Feature importance
feature_importance_total = pd.Series(best_xgb_total.feature_importances_, index=X_total_power.columns)

# Append results
results.append({
    "Target": "Total Power",
    "Model": "XGBoost",
    "R2": np.mean(cv_xgb_tp['test_R2']),
    "MAE": np.mean(cv_xgb_tp['test_MAE']),
    "RMSE": np.mean(cv_xgb_tp['test_RMSE']),
    "MAPE": np.mean(cv_xgb_tp['test_MAPE'])
})

print("Best params for Total Power:", grid_total.best_params_)
print("Feature importance (Total Power):")
print(feature_importance_total.sort_values(ascending=False))


# ---------------------------
# Model 2: Predict Usable Space
# ---------------------------
X_usable_space = df_train[["Total Power (MW)", "Colocation Flag", "Year", "Latitude", "Longitude"]]
y_usable_space = df_train["Usable Space"]

grid_usable = GridSearchCV(xgb, param_grid, cv=kf, scoring='r2')
grid_usable.fit(X_usable_space, y_usable_space)
best_xgb_us = grid_usable.best_estimator_

# Cross-validation
cv_xgb_us = cross_validate(best_xgb_us, X_usable_space, y_usable_space, cv=kf, scoring=scoring)

# Feature importance
feature_importance_us = pd.Series(best_xgb_us.feature_importances_, index=X_usable_space.columns)

# Append results
results.append({
    "Target": "Usable Space",
    "Model": "XGBoost",
    "R2": np.mean(cv_xgb_us['test_R2']),
    "MAE": np.mean(cv_xgb_us['test_MAE']),
    "RMSE": np.mean(cv_xgb_us['test_RMSE']),
    "MAPE": np.mean(cv_xgb_us['test_MAPE'])
})

print("Best params for Usable Space:", grid_usable.best_params_)
print("Feature importance (Usable Space):")
print(feature_importance_us.sort_values(ascending=False))

# ---------------------------
# Convert results to DataFrame
# ---------------------------
results_df = pd.DataFrame(results)
print(results_df)
Cell [20]
from tabulate import tabulate

print(tabulate(results_df, headers='keys', tablefmt='fancy_grid', showindex=False))
Cell [23]
# -----------------------------
# Step 0: Create a copy of the original df
# -----------------------------
df_fully_imputed = df.copy()  # keep original index


# -----------------------------
# Step 1: Create subsets for imputation (keep original indices)
# -----------------------------
# Missing Total Power but Usable Space available
df_impute_tp = df_fully_imputed.loc[
    df_fully_imputed["Total Power (MW)"].isna() & df_fully_imputed["Usable Space"].notna()
]

# Missing Usable Space but Total Power available
df_impute_us = df_fully_imputed.loc[
    df_fully_imputed["Total Power (MW)"].notna() & df_fully_imputed["Usable Space"].isna()
]

# Missing both Total Power and Usable Space
df_impute_all = df_fully_imputed.loc[
    df_fully_imputed["Total Power (MW)"].isna() & df_fully_imputed["Usable Space"].isna()
]

print(f"Rows from training: {len(df_train)}")
print(f"Rows to impute Total Power only: {len(df_impute_tp)}")
print(f"Rows to impute Usable Space only: {len(df_impute_us)}")
print(f"Rows to impute both: {len(df_impute_all)}")
Cell [24]
# Initialize Imputed column with 'Yes'
df_fully_imputed["Imputed"] = "UKN"

# For rows that originally had both Total Power and Usable Space present, mark as "No"
df_fully_imputed.loc[
    df_fully_imputed["Total Power (MW)"].notna() & df_fully_imputed["Usable Space"].notna(),
    "Imputed"
] = "No"

# Check result
print(df_fully_imputed[["Total Power (MW)", "Usable Space", "Imputed"]].head(10))
Cell [26]
# -----------------------------
# Step 2: Scale numeric features using training scaling
# -----------------------------
numeric_cols = ["Usable Space", "Year", "Latitude", "Longitude"]

for col in numeric_cols:
    mean = df_train[col + "_scaled"].mean()
    std = df_train[col + "_scaled"].std()
    df_impute_tp[col + "_scaled"] = (df_impute_tp[col] - mean) / std

# Keep feature order consistent with training
X_impute_knn = df_impute_tp[
    ["Usable Space_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]
]

# -----------------------------
# Step 3: Predict using trained KNN
# -----------------------------
y_pred_scaled = best_knn_total.predict(X_impute_knn)

# -----------------------------
# Step 4: Rescale predictions back to original units
# -----------------------------
tp_mean = df_train["Total Power (MW)_scaled"].mean()
tp_std = df_train["Total Power (MW)_scaled"].std()
y_pred_original = y_pred_scaled * tp_std + tp_mean

# -----------------------------
# Step 5: Update df_fully_imputed
# -----------------------------
df_fully_imputed.loc[df_impute_tp.index, "Total Power (MW)"] = y_pred_original
df_fully_imputed.loc[df_impute_tp.index, "Imputed"] = "KNN"

# -----------------------------
# Step 6: Verify imputation
# -----------------------------
null_counts = df_fully_imputed.loc[df_impute_tp.index].isna().sum()
print("Number of missing values per column in df_impute_tp after imputation:")
print(null_counts)
Cell [28]

# -----------------------------
# Step 2: Prepare predictors for imputation
# -----------------------------
predictor_cols = ["Usable Space", "Colocation Flag", "Year", "Zipcode"]
X_impute_tp = df_impute_tp[predictor_cols].copy()

# -----------------------------
# Step 3: Predict missing Total Power using trained Random Forest
# -----------------------------
y_pred_tp = best_rf_total.predict(X_impute_tp)

# -----------------------------
# Step 4: Update df_fully_imputed in place using the original indices
# -----------------------------
df_fully_imputed.loc[df_impute_tp.index, "Total Power (MW)"] = y_pred_tp
df_fully_imputed.loc[df_impute_tp.index, "Imputed"] = "Random Forest"

# -----------------------------
# Step 5: Check results
# -----------------------------
print(f"Imputed Usable Space for {len(df_impute_tp)} rows in df_fully_imputed:")
print(df_fully_imputed.loc[df_impute_tp.index, predictor_cols + ["Total Power (MW)", "Imputed"]])

# -----------------------------
# Step 6: Optional - count remaining nulls
# -----------------------------
null_counts = df_fully_imputed.loc[df_impute_tp.index].isna().sum()
print("\nNull counts in df_fully_imputed for these rows after imputation:")
print(null_counts)
Cell [29]
len(df_fully_imputed[df_fully_imputed["Imputed"]=="Random Forest"])
Cell [30]
df_fully_imputed[df_fully_imputed["Imputed"]=="Random Forest"]
Cell [32]

# -----------------------------
# Step 2: Prepare predictors for imputation
# -----------------------------
predictor_cols = ["Total Power (MW)", "Colocation Flag", "Year", "Latitude", "Longitude"]
X_impute_us = df_impute_us[predictor_cols].copy()

# -----------------------------
# Step 3: Predict missing Usable Space using trained XGBoost
# -----------------------------
y_pred_us = best_xgb_us.predict(X_impute_us)

# -----------------------------
# Step 4: Update df_fully_imputed in place using the original indices
# -----------------------------
df_fully_imputed.loc[df_impute_us.index, "Usable Space"] = y_pred_us
df_fully_imputed.loc[df_impute_us.index, "Imputed"] = "XGBoost"

# -----------------------------
# Step 5: Check results
# -----------------------------
print(f"Imputed Usable Space for {len(df_impute_us)} rows in df_fully_imputed:")
print(df_fully_imputed.loc[df_impute_us.index, predictor_cols + ["Usable Space", "Imputed"]])

# -----------------------------
# Step 6: Optional - count remaining nulls
# -----------------------------
null_counts = df_fully_imputed.loc[df_impute_us.index].isna().sum()
print("\nNull counts in df_fully_imputed for these rows after imputation:")
print(null_counts)
Cell [33]
len(df_fully_imputed[df_fully_imputed["Imputed"]=="XGBoost"])
Cell [35]
import pandas as pd
import numpy as np


# Identify rows to impute
mask_both_null = df_fully_imputed["Total Power (MW)"].isna() & df_fully_imputed["Usable Space"].isna()
print(f"Rows to impute both Total Power and Usable Space: {mask_both_null.sum()}")

# -----------------------------
# Step 1: Create decade column
# -----------------------------
df_fully_imputed["Decade"] = (df_fully_imputed["Year"] // 10).astype(int) * 10

# -----------------------------
# Step 2: Compute median values from *original df* only where not null
# -----------------------------
median_by_company_decade = (
    df[df["Total Power (MW)"].notna() & df["Usable Space"].notna()]
    .groupby(["Company", (df["Year"] // 10).astype(int) * 10])[["Total Power (MW)", "Usable Space"]]
    .median()
)

median_by_decade = (
    df[df["Total Power (MW)"].notna() & df["Usable Space"].notna()]
    .groupby((df["Year"] // 10).astype(int) * 10)[["Total Power (MW)", "Usable Space"]]
    .median()
)

overall_medians = (
    df[df["Total Power (MW)"].notna() & df["Usable Space"].notna()][["Total Power (MW)", "Usable Space"]]
    .median()
)

# -----------------------------
# Step 3: Define safe imputation function
# -----------------------------
def impute_row_safe(row):
    company = row["Company"]
    decade = row["Decade"]

    # Default: overall median
    tp_value = overall_medians["Total Power (MW)"]
    us_value = overall_medians["Usable Space"]
    imputed_label = "Overall"

    # Try company + decade median
    if (company, decade) in median_by_company_decade.index:
        tp_value = median_by_company_decade.loc[(company, decade), "Total Power (MW)"]
        us_value = median_by_company_decade.loc[(company, decade), "Usable Space"]
        imputed_label = "Company/Decade"

    # Else, try decade median
    elif decade in median_by_decade.index:
        tp_value = median_by_decade.loc[decade, "Total Power (MW)"]
        us_value = median_by_decade.loc[decade, "Usable Space"]
        imputed_label = "Decade"

    # Assign imputed values
    row["Total Power (MW)"] = tp_value
    row["Usable Space"] = us_value
    row["Imputed"] = imputed_label
    return row

# -----------------------------
# Step 4: Apply to rows needing imputation
# -----------------------------
df_fully_imputed.loc[mask_both_null] = (
    df_fully_imputed.loc[mask_both_null].apply(impute_row_safe, axis=1)
)

# -----------------------------
# Step 5: Verify results
# -----------------------------
print("Sample imputed rows for both Total Power and Usable Space:")
print(df_fully_imputed.loc[mask_both_null, ["Company", "Year", "Decade", "Total Power (MW)", "Usable Space", "Imputed"]])

# Check nulls in the imputed subset
null_counts = df_fully_imputed.loc[mask_both_null, ["Total Power (MW)", "Usable Space"]].isna().sum()
print("\nNull counts in df_fully_imputed for imputed rows:")
print(null_counts)
Cell [36]
print(df_fully_imputed.head())
Cell [37]
df_fully_imputed.isna().sum()
Cell [38]
len(df_fully_imputed[df_fully_imputed["Imputed"]=="No"])
Cell [39]
df_fully_imputed.to_excel("df_fully_imputed_11.15.25.xlsx", index=False)
Scraper — Size, Power, Colocation – dc_scrape_size_power_colocation.ipynb4 cells · supporting · large

Web-scrapes DataCenters.com for facility names, addresses, total power (MW), total and colocation space. The large second cell contains a scraped-data dump as a Python literal and has been trimmed here to the first ~2 KB for readability — the full literal is available by re-running the scraper (Cell 3 in this notebook).

Cell [1]
!apt-get update > /dev/null
!apt install chromium-chromedriver > /dev/null
!pip install selenium --quiet
Cell [2]
final_name_address = [{
"id": 6621,
"locationId": 5280,
"name": "Riverside 1 Data Center",
"fullAddress": "1550 Marlborough Avenue, Riverside, CA, USA",
"sponsoredGlobal": False,
"sponsoredCountry": False,
"sponsoredState": False,
"sponsoredCity": False,
"url": "/lumen-riverside-1",
"providerId": 200038,
"providerName": "Lumen",
"providerAgreement": True,
"mainImageUrl": "https://res.cloudinary.com/hjlz68xhm/image/upload/hgewxpznetxgbporzgmu",
"latitude": 33.9970355,
"longitude": -117.3456799,
"providerLogo": "https://res.cloudinary.com/hjlz68xhm/image/upload/v1707770321/utkipf0eps6go6nejfmy.png"
},
{
"id": 6910,
"locationId": 7730,
"name": "Los Angeles - Century Data Center",
"fullAddress": "6171 West Century Boulevard, Los Angeles, CA, USA",
"sponsoredGlobal": False,
"sponsoredCountry": False,
"sponsoredState": False,
"sponsoredCity": False,
"url": "/quadranet-los-angeles-century",
"providerId": 120,
"providerName": "QuadraNet",
"providerAgreement": False,
"mainImageUrl": "https://res.cloudinary.com/hjlz68xhm/image/upload/m0cq9zex5emyc1ytmkgg",
"latitude": 33.9459317,
"longitude": -118.3935351,
"providerLogo": "https://res.cloudinary.com/hjlz68xhm/image/upload/lm94nvmn6wib4xpf63ap"
},
{
"id": 6737,
"locationId": 5587,
"name": "Sunnyvale 1 Data Center",
"fullAddress": "1380 Kifer Road, Sunnyvale, CA, USA",
"sponsoredGlobal": False,
"sponsoredCountry": False,
"sponsoredState": False,
"sponsoredCity": False,
"url": "/lumen-sunnyvale-1",
"providerId": 200038,
"providerName": "Lumen",
"providerAgreement": True,
"mainImageUrl": "https://res.cloudinary.com/hjlz68xhm/image/upload/v1/locations/lumen-sunnyvale-1/images/ublca3eg4jlbtdgdzpmz",
"latitude": 37.3734237,
"longitude": -121.9876548,
"providerLogo": "https://res.cloudinary.com/hjlz68xhm/image/upload/v1707770321/utkipf0eps6go6nejfmy.png"
},
{
"id": 9398,
"locationId": 8650,
"name": "San Diego Data Center",
"fullAddress": "9725 Scranton Rd, San Diego, CA, USA",
"sponsoredGlobal": False,
"sponsoredCountry": False,
"spo

# ... [TRIMMED: original cell is 211,209 chars, showing first 2,000. This cell contains a large scraped-data dump embedded as a Python literal — the full list of 197 data-center records. Full notebook available on request.] ...
Cell [4]
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import pandas as pd
import time

# Setup Selenium Chrome driver for Colab
options = Options()
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev_shm_usage')
driver = webdriver.Chrome(options=options)

# Use the existing datacenter_name_address list from cell -VzhPOHaHK7J
# Ensure datacenter_name_address is a list of dictionaries as expected
if 'final_name_address' in locals() and isinstance(final_name_address, list):
    data_centers_combined = []
    for dc_info in final_name_address:
        if isinstance(dc_info, dict):
            # Extract basic info already available
            data_centers_combined.append({
                'Company': dc_info.get('providerName'), # Assuming providerName is the company
                'Name': dc_info.get('name'),
                'Full Address': dc_info.get('fullAddress'),
                'Latitude': dc_info.get('latitude'),
                'Longitude': dc_info.get('longitude'),
                'URL': dc_info.get('url'),
                'Total Space': None, # Initialize, will be scraped later
                'Total Power': None,  # Initialize, will be scraped later
                'Colocation Space': None # Initialize for colocation space
            })
        else:
            print(f"Skipping unexpected item in datacenter_name_address: {dc_info}")

    print(f"Starting to scrape detail pages for Total Space, Total Power, and Colocation Space for {len(data_centers_combined)} data centers...")

    # Now, iterate through the collected data centers and scrape details using Selenium
    for dc in data_centers_combined:
        url_path = dc['URL']
        # Construct the full URL if the URL in the list is just the path
        if url_path and not url_path.startswith('http'):
             url = "https://www.datacenters.com" + url_path
        else:
             url = url_path # Use as is if it's already a full URL

        name = dc['Name']

        if not url:
            print(f"Skipping scraping for {name} due to missing URL.")
            continue

        print(f"Scraping Total Space, Power, and Colocation Space for: {name} ({url})")

        try:
            driver.get(url)
            # Wait for an element that indicates the page has loaded
            WebDriverWait(driver, 10).until(
                EC.presence_of_element_located((By.CSS_SELECTOR, 'h1')) # Wait for the main heading as a general indicator
            )
            dc_soup = BeautifulSoup(driver.page_source, 'html.parser')

            # Total Space
            total_space = None
            total_space_div = dc_soup.find('div', id='totalSpace')
            if total_space_div:
                stat_info_div = total_space_div.find('div', id='statInfo')
                if stat_info_div:
                    strong_tag = stat_info_div.find('strong')
                    if strong_tag:
                        total_space = strong_tag.get_text(strip=True)
                        dc['Total Space'] = total_space # Update the list

            # Total Power
            total_power = None
            total_power_div = dc_soup.find('div', id='power')
            if total_power_div:
                stat_info_div = total_power_div.find('div', id='statInfo')
                if stat_info_div:
                    strong_tag = stat_info_div.find('strong')
                    if strong_tag:
                        total_power = strong_tag.get_text(strip=True)
                        dc['Total Power'] = total_power # Update the list

            # Colocation Space (Assuming the ID is 'colocationSpace' based on typical patterns)
            colocation_space = None
            colocation_space_div = dc_soup.find('div', id='colocationSpace')
            if colocation_space_div:
                stat_info_div = colocation_space_div.find('div', id='statInfo')
                if stat_info_div:
                    strong_tag = stat_info_div.find('strong')
                    if strong_tag:
                        colocation_space = strong_tag.get_text(strip=True)
                        dc['Colocation Space'] = colocation_space # Update the list


        except Exception as e:
            print(f"Failed to scrape data for {url}: {e}")
            # Total Space, Total Power, and Colocation Space will remain None as initialized

        time.sleep(1) # polite pause between detail pages

    driver.quit()

    df_final_combined = pd.DataFrame(data_centers_combined)
    print(df_final_combined)

else:
    print("Error: datacenter_name_address variable is not defined or is not a list.")
Cell [5]
df_final_combined.to_csv('datacenters_space_power_colocation.csv', index=False)
from google.colab import files
files.download('datacenters_space_power_colocation.csv')
Scraper — Founding Dates – dc_scrape_founding_dates.ipynb5 cells · supporting

Fills in year-of-construction for each facility — with parent-company founding year as a last-resort fallback.

Cell [5]
company_names = ["Amazon AWS",
"AT&T",
"Atlantic Metro Communications",
"Atlantic.net",
"BreezeHost.io",
"Cato Digital, Inc",
"CBRE",
"CBTS LLC",
"Claranet",
"Colo Locker - Evocative",
"Cologix",
"CoreSite",
"Data Canopy",
"Datacate, Inc",
"EdgeConnex",
"Enzu",
"Fortress Data Centers",
"Hivelocity",
"IBM Cloud",
"Krypt",
"LeaseWeb",
"Megaport",
"Microsoft Azure",
"NOVVA",
"One Data Center America",
"Psychz Networks",
"Quest Technology Management",
"Summit",
"Synoptek",
"unWired Broadband LLC",
"Vantage Data Centers",
"Vultr"]




base_url = "https://www.datacenters.com/providers/"
Cell [8]
urls = []
for name in company_names:
  formatted_name = name.lower().replace(" ", "-")
  full_url = base_url + formatted_name
  urls.append(full_url)

print(urls)
Cell [11]
import requests
from bs4 import BeautifulSoup

results = []
for url in urls:
    try:
        response = requests.get(url)
        response.raise_for_status()  # Raise an exception for bad status codes
        soup = BeautifulSoup(response.content, 'html.parser')
        year_founded_div = soup.find('div', id='yearFounded')
        if year_founded_div:
            year_founded = year_founded_div.get_text(strip=True)
            # Extract company name from URL
            company_name_formatted = url.split('/')[-1]
            company_name = company_name_formatted.replace('-', ' ').title()
            results.append({'company_name': company_name, 'year_founded': year_founded})
        else:
            print(f"Could not find year founded for {url}")
    except requests.exceptions.RequestException as e:
        print(f"Error fetching {url}: {e}")

display(results)
Cell [13]
cleaned_results = []
for result in results:
    year_founded = result['year_founded'].replace('Calendar', '').replace('year founded', '')
    cleaned_results.append({'company_name': result['company_name'], 'year_founded': year_founded.strip()})

display(cleaned_results)
Cell [16]
import pandas as pd

df_results = pd.DataFrame(cleaned_results)
display(df_results)
Broadband instrument (FCC Form 477, 2014) – dc_broadband_fiber_instrument.ipynb7 cells · supporting

Builds the fiber-broadband instrument (provider count per county, mean max upload / download speeds) from FCC Form 477 2014 vintage — the exogenous variation used in the 2SLS.

Cell [5]
import pandas as pd

df = pd.read_csv('/content/bdc_06_FibertothePremises_fixed_broadband_D24_24dec2025.csv')

print("First 5 rows of the DataFrame:")
print(df.head())

print("\nDataFrame Information (columns and data types):")
print(df.info())
Cell [8]
df['Geoid'] = df['block_geoid'].astype(str).str[:4].astype(int)

print("First 5 rows with new 'Geoid' column:")
print(df.head())

print("\nDataFrame Information (columns and data types) with new 'Geoid' column:")
print(df.info())
Cell [11]
df = df[df['business_residential_code'] != 'R']

print("First 5 rows after filtering out 'R' from 'business_residential_code':")
print(df.head())

print("\nDataFrame Information after filtering:")
print(df.info())
Cell [14]
df['year_month'] = '2024-12'

print("First 5 rows with new 'year_month' column:")
print(df.head())

print("\nDataFrame Information with new 'year_month' column:")
print(df.info())
Cell [17]
aggregated_df = df.groupby('Geoid').agg(
    avg_max_advertised_download_speed=('max_advertised_download_speed', 'mean'),
    avg_max_advertised_upload_speed=('max_advertised_upload_speed', 'mean'),
    unique_provider_count=('provider_id', 'nunique')
).reset_index()

df = pd.merge(df, aggregated_df, on='Geoid', how='left')

print("First 5 rows with new aggregated columns:")
print(df.head())

print("\nDataFrame Information with new aggregated columns:")
print(df.info())
Cell [19]
final_df = aggregated_df.copy()
final_df['year_month'] = '2024-12'

print("First 5 rows of the new DataFrame:")
print(final_df.head())

print("\nDataFrame Information for the new DataFrame:")
print(final_df.info())
Cell [20]
final_df.to_excel('/content/broadband_2024-12.xlsx', index=False)
print("DataFrame exported successfully to '/content/broadband_2024-12.xlsx'")

Upstream: raw data via CAISO OASIS, CEC, CalEnviroScreen 4.0, NOAA GHCN, EIA-860, FCC Form 477 (2014), ACS 5-year, CPUC NEM, and a scraped data center directory (197 facilities). Data warehouse: BigQuery. Transformations: dbt (staging → intermediate → domain). Analysis notebooks: Google Colab / Deepnote. IV estimation: linearmodels.iv.IV2SLS with cluster-robust SE; two-way FE via explicit dummies or iterative Gauss-Seidel demeaning per Frisch-Waugh-Lovell.