New York University · Stern School of Business
MS in Business Analytics & AI · Capstone
April 2026
Electric vehicles and data centers are reshaping the grid. The question nobody is asking: who pays when the grid can't predict itself?
Arguelles · Graff · Harrison · Jessup · Lindley
↓ Scroll to explore · Charts are interactive · Methodology at the bottom
Key Findings
Sign Reversal. EV adoption causally increases electricity price volatility — the opposite of what naïve regressions suggest.
Income Gradient. Lower-income counties bear a disproportionate share of that burden. The gap widens as adoption scales.
Data Centers: Null. Stable baseload demand does not destabilize grids. The concern is cost allocation, not volatility.
The Convergence
Three forces are converging: AI data centers demanding continuous power, a state mandate electrifying millions of vehicles, and climate extremes straining generation capacity. The grid is bending. But the critical question isn't whether it can handle the volume — it's whether the unpredictability that follows falls equally on everyone. It doesn't.
0
EVs on CA Roads
18.7 GW
DC Capacity Requested
36 mo
Price Data Analyzed
Background
California's wholesale market uses locational marginal pricing — the price at a specific grid node reflects the real-time cost of delivering one more megawatt there. When transmission lines congest, prices at different nodes diverge. A node behind a bottleneck becomes cut off from cheaper generation. These differentials reflect decades of infrastructure investment decisions.
Key mechanism: Wholesale volatility doesn't hit bills directly. It flows through utilities' procurement costs, hedging, and infrastructure planning — passed to ratepayers over annual or multi-year cycles. The burden is real, but delayed and opaque.
The Trend
California's EV fleet has more than doubled since April 2022, but charging infrastructure hasn't kept pace — reaching only 1.65× while adoption hit 2.3×. The stress isn't about total energy; it's about timing. Home charging clusters in the evening, precisely when solar ramps down.
EV Adoption vs. Tesla Store Expansion (Indexed, Base = 100)
Data Landscape
Multiple data sources mapped across 52 counties. Toggle the base layer to change county coloring. Check point layers to overlay infrastructure.
County Color
⏱ VARIES OVER TIME
▦ STATIC (COUNTY-LEVEL)
Point Layers
Time Period
Price Volatility (MAD) by County
The Key Insight
Average prices peaked in late 2022 (heat waves, fuel spikes) then declined. That story is over. But volatility — price unpredictability — is episodic, appearing in clusters throughout. The forces that drove the 2022 peak are not the forces creating recurring instability.
Monthly Avg. Wholesale LMP ($/MWh)
Monthly Volatility (MAD)
The forces that drove prices to their 2022 peak are not the same forces creating recurring instability. It's the unpredictability that affects utilities' planning — and flows through to household costs.
The Causal Finding
OLS says more EVs = less volatility. But that's confounded — wealthier counties have both more EVs and better grids. Using Tesla store openings as an instrument (methodology ↓) (corporate decisions exogenous to local electricity prices), the sign flips:
OLS (Naïve)
−0.000347OLS coefficient — biased by confounders (wealth, grid infrastructure, urban density) that correlate with both EV adoption and lower volatility.
p = 0.016
2SLS (IV)
+0.0005932SLS coefficient from IV regression using Tesla store IDW exposure as instrument. County + month FE, cluster-robust SE. N=1,976.
p = 0.001
Each additional EV per 1,000 residents causally increases monthly price volatility. For +50 EVs/1,000, that's 0.7 standard deviations — roughly half the interquartile range (~48%) of observed volatility. A substantial shift in grid unpredictability.
The Equity Question
Each $10,000 increase in county median income reduces the EV-induced volatility effect by ~$0.0003/MWh per EV/1,000 (p < 0.01). Lower-income counties absorb meaningfully more instability for the same adoption increase.
Explore: County Median Household Income
Income: $70,000
+0.029
$/MWh volatility increase per 50 EVs/1,000
Model specification (interaction 2SLS)
β1 = +0.003554 (main EV effect, p < 0.01) · β2 = −0.000303 (income interaction, p < 0.01). Income is scaled in $10,000 units. The slider above evaluates the marginal effect (β1 + β2·income10k) × 50 EVs. Controls X: max temperature, cumulative data center power, generator capacity, lagged solar capacity. μc = county FE; λt = month FE. SEs clustered by county.
County Clusters: Income vs. Volatility
Key: DAC designation (CalEnviroScreen) does not capture where this burden concentrates. Income is the more precise policy lever.
The Other Driver
No statistically significant causal effect on volatility or spikes. A 1% increase in DC power associates with a $0.088/MWh decrease in average prices — predictable baseload demand smoothing operations. The equity concern isn't volatility; it's cost allocation — see the full results table ↓.
IV Coefficient Estimates: Data Center Power by Outcome
Whiskers show 95% confidence intervals. Values of zero (dashed line) indicate no effect.
Case Study
Two California policies shifted inside our study window from opposite sides of the grid. Exploiting their timing gives us a second identification strategy and a sharper picture of how short-run shocks propagate to volatility.
The Clean Vehicle Rebate Project paid income-scaled rebates for EV purchases. California announced its closure in August 2023 and the program terminated in November. That three-month window produced a "run on the bank" — a burst of subsidy-anchored purchases followed by a sustained decline in adoption as the price signal disappeared.
We use the closure timing interacted with each county's pre-period CVRP reliance weight (share of applications that received low-income boosted rebates) as a shift-share instrument. The second-stage coefficient on EV adoption is −0.0007 (p = 0.11) — marginally negative but below conventional significance. We report this transparently: the CVRP-identified effect does not clear the 5% bar, and we interpret it as identifying a different complier population than the Tesla IV (rebate-sensitive rather than access-driven).
Monthly EV Adoption (New EVs per 1,000, All Counties)
California's Net Energy Metering framework transitioned from NEM 2.0 (retail-rate compensation) to NEM 3.0 (avoided-cost compensation) in April 2023. A large backlog of NEM 2.0 applications — systems grandfathered into the legacy payment structure — physically connected to the grid in the months that followed, peaking in early 2024. These were mostly battery-less solar-only installations that rushed to secure favorable rates before the rule changed.
We include a post-January-2024 structural break control (solar_x_nem3) in the final 2SLS specification to isolate this solar shock from the concurrent EV dynamics. The coefficient is +0.0177 (p < 0.06): the rapid wave of solar-only interconnection temporarily raised volatility, consistent with the duck-curve intuition — new capacity deepens midday troughs and steepens evening ramps without storage to buffer it.
A supply-side mirror of the EV problem. EVs add load when there isn't enough generation capacity; NEM 2.0 solar added generation when there isn't enough storage. Both are physical grid-timing problems rather than aggregate-capacity problems, and both point at the same policy priority: flexible infrastructure that buffers short-run ramps.
NEM 2.0 Rush: Solar Capacity (Cumulative MW per 1,000)
+0.0177**
The wave of battery-less NEM 2.0 systems connecting to the grid increased volatility
Summary
| Analysis | Outcome | OLS | 2SLS (IV) | Significance |
|---|---|---|---|---|
| EV → Volatility | MAD rate | −0.000347 ** | +0.000593 *** | Sign reversal ✓ |
| EV × Income | MAD rate | −0.000080 * | −0.000303 *** | Consistent ✓ |
| EV × DAC | MAD rate | Positive n.s. | Positive n.s. | Null result |
| EV → Price level | Avg LMP | Small, unstable | Small, unstable | No robust effect |
| DC → Volatility | MAD rate | Not significant | Not significant | Robust null ✓ |
| DC → Avg price | Avg LMP | Not significant | −8.8¢/MWh *** | Significant ✓ |
| CVRP (EV alt. IV) | MAD rate | — | −0.0007 (p=0.11) | Not Significant |
| NEM 2.0 backlog | MAD rate | — | +0.0177 ** | Policy timing |
*** p<0.01 | ** p<0.05 | * p<0.10 | n.s. = not significant
What This Means
Transportation electrification is a state mandate. But its grid impacts are not equity-neutral.
1. Target infrastructure investment by income, not just DAC status. CalEnviroScreen alone doesn't capture where volatility burden concentrates.
2. Prioritize distribution grid upgrades in lower-income counties as EV adoption scales.
3. Data center policy should focus on cost allocation, not volatility. Stable loads don't destabilize grids.
Under the Hood
All public data. BigQuery + dbt + Python. 1,976 county-month observations, 52 counties, Apr 2022 – May 2025.
Full regression ladder · first-stage diagnostics · interaction spec · robustness grid
Looking Ahead
Data center power is proxied. Operational load is estimated from web-scraped nameplate capacity using the CEC YoY ramp and JLARC seasonality profile, not observed consumption. DC price results are directional.
Two-month causal window. The Tesla IV coefficient retains ≈60% of its magnitude through a two-month lag before dropping sharply. EVs are a stock variable, so a purely cumulative-load mechanism would be expected to persist longer. This ambiguity between a genuine short-lived effect and a transient correlated shock is reported transparently.
Tract-level re-estimation does not replicate. At census-tract granularity the coefficient drops to near zero. We attribute this to a scale mismatch between LMP node geography and tracts, not to evidence against the county-level result, but it is a constraint on generalization.
CVRP IV is marginally significant. The policy-discontinuity IV produces a coefficient of −0.0007 (p=0.11) after full controls, below conventional thresholds. We interpret the Tesla and CVRP instruments as identifying different complier populations rather than as contradictory evidence — see the statistical appendix.
LLM-powered interactive dashboard. AI-driven application allowing stakeholders to query grid equity data in natural language.
Utility-level consumption data. Partner with CAISO for metered load profiles to validate data center estimates.
Sub-monthly pricing + charging. Hourly LMP data matched with EV charging event data to isolate the temporal mechanism.
For the Defense
Questions we expect from reviewers, with concise answers. Click any question to expand.
Standard-deviation-based volatility measures are dominated by extreme outliers. California's wholesale prices exhibit heavy tails with occasional extreme spikes during grid emergencies — in our sample, monthly mean LMP reached ≈$270/MWh during the late-2022 heat wave. Including those spikes in a variance calculation would make volatility look like a story about rare crisis moments rather than the chronic planning-uncertainty problem utilities actually face.
MAD captures the typical dispersion households and utilities experience while limiting outlier influence. It also aligns more closely with the kind of volatility that translates into utility hedging costs and procurement reserves — the channel through which wholesale dispersion eventually reaches retail rates.
Three reasons. First, coverage: LMP nodes don't map cleanly to utility territories, and many tracts contain no pricing nodes at all. County aggregation achieves near-complete statewide coverage. Second, consistency: every other variable in our panel (EV registrations, CalEnviroScreen tracts, census demographics, weather) is available or clean to aggregate at the county level. Third, policy relevance: California's environmental justice framework, disadvantaged community designations, and many infrastructure funding programs operate at the county level.
We tested tract-level re-estimation and report the null result transparently (Section H of the statistical appendix) — we attribute it to scale mismatch, not to evidence against the county result.
Both instruments are reported. The Tesla IV is the preferred main specification because (a) its first stage is strong and clean, (b) it passes the permutation test at p = 0.014, and (c) store-opening timing is plausibly exogenous on multi-year corporate real-estate horizons.
The CVRP shift-share IV is reported as a complementary specification. The two instruments identify different compliers — the Tesla instrument captures access-driven adoption (higher-income, urban), the CVRP instrument captures rebate-sensitive adoption (lower-income, higher-reliance). That the two produce different second-stage signs is informative about heterogeneity, not contradictory evidence.
This is the single most-asked question about the paper and is fully addressed in Section B of the statistical appendix. Briefly: without fixed effects, cross-county variation dominates and the sign is positive — counties with more EVs have higher volatility on average. Adding county FE strips the cross-sectional signal and reveals the within-county pattern: months where a given county has higher-than-usual EV adoption see lower volatility. Adding month FE on its own does not flip the sign because it only absorbs statewide time shocks.
The TWFE-OLS gives −0.000347 (p=0.015), but this is still biased — it confounds EV adoption with time-varying county investments in grid infrastructure (battery rollouts, TOU rates, distribution upgrades) that are negatively correlated with volatility. Once the Tesla IV removes that confounded variation, the coefficient returns to positive: +0.000593 (p=0.001). The sign flips are each driven by a specific source of bias being removed.
Intentionally. The dominant source of uncertainty in the forecast is structural, not statistical — it depends on whether the causal coefficients β₁ and β₂ stay approximately constant over the next decade, which they may not. Adding parametric 95% CIs from the 2SLS would imply a false precision that our honest uncertainty doesn't support.
The parametric CI on the $40k / 2030 point would be roughly ±40% of the central estimate. If the real mechanism attenuates as utilities deploy storage, V2G, or dynamic rates, the actual 2030 burden could be substantially lower. We frame the forecast as a "no-intervention projection" that motivates action, not as a point forecast.
It is a fair concern and we report it transparently. EVs are a stock variable; a mechanism operating purely through cumulative charging load would be expected to persist longer than two months. The permutation test (p = 0.014) confirms the effect at contemporaneous timing is statistically real, but cannot distinguish between a genuine short-lived causal effect and a transient correlated shock.
Two non-exclusive interpretations: (a) utilities and distribution networks respond to new EV load within a quarter (demand response, rate adjustments, feeder upgrades), transiently reducing the marginal volatility contribution of any new cohort; (b) the signal is real at contemporaneous timing but attenuates as charging habits stabilize. Distinguishing these requires sub-monthly charging event data, which is in the paper's "Next Steps."
DAC designation at the tract level, when aggregated to a county share, is a noisy and coarse measure of economic vulnerability compared to continuous county median household income. CalEnviroScreen weights environmental burdens, pollution exposure, health outcomes, and socioeconomic factors — income is only one of many inputs. Two counties can have similar DAC shares but very different income distributions.
Our result — significant income gradient, null DAC gradient — is consistent with income being the precise policy lever for EV-induced volatility burden, not the broader cumulative-burden framework DAC captures. This is actually a policy-relevant finding in its own right: DAC designation is useful for many things but does not predict volatility exposure here.
We treat it as directional, not precise. The collapsed IV coefficient (−0.174, p=0.001, translating to −$0.088/MWh per 1% DC power) is statistically significant and directionally consistent with PSM. The PSM robustness check produces −0.087 on arcsinh-transformed prices, in the same direction.
The caveat: our DC operational load is proxied from imputed nameplate capacity using the CEC YoY ramp methodology, not observed consumption. Variation is primarily cross-sectional across a limited set of host counties, constraining precision. Our strongest DC claim is the robust null on volatility; the negative price effect is reported as suggestive.
Sources
California Air Resources Board. (2022). 2022 Scoping Plan for Achieving Carbon Neutrality.
California Public Utilities Commission. (2025). How will data center growth impact California ratepayers?
Elmallah, S., Brockway, A. M., & Callaway, D. (2022). Can distribution grid infrastructure accommodate residential electrification and EV adoption in Northern California? Environmental Research: Infrastructure and Sustainability, 2, 045005.
Jenn, A., & Highwayman, J. (2021). Distribution grid impacts of electric vehicles: A California case study. iScience, 25(1), 103686.
Li, Y., & Jenn, A. (2024). Impact of EV charging demand on power distribution grid congestion. Proceedings of the National Academy of Sciences, 121(18).
Padilla, S. (2025). Senate Bill 57: Ratepayer and Technological Innovation Protection Act. California Legislature.
Ren, S. (2025). An assessment of California data centers' environmental and public health impacts. Next 10.
Schweppe, F. C., Caramanis, M. C., Tabors, R. D., & Bohn, R. E. (1988). Spot pricing of electricity. Kluwer Academic Publishers.
Shehabi, A., et al. (2024). 2024 United States data center energy usage report. Lawrence Berkeley National Laboratory.
Talkington, S., West, A., & Haider, R. (2024). Locational marginal burden: Quantifying the equity of optimal power flow solutions. E-Energy '24.
U.S. Department of Energy, EERE. (n.d.). Energy accessibility.
How to Cite
Arguelles, S., Graff, V., Harrison, E., Jessup, R., & Lindley, S. (2026). Electricity Price Volatility and Grid Pressure in California, 2022–2025: A Data Infrastructure for Analyzing Household Impacts. NYU Stern School of Business, MS in Business Analytics & AI Capstone Project.
BibTeX available on request · Data pipeline: BigQuery + dbt + Python · All sources publicly available
Dashboard-only · Not in the Capstone Report
The content in this section sits outside the scope of the formal capstone report. It is included here because the dashboard serves a different purpose than the report, and a few things make sense in one venue but not the other.
Why these live in the dashboard but not the report. The report is the empirical deliverable — it defends claims about 2022–2025, the window the data actually covers, and every coefficient reported there is bounded by what the panel can identify. Extrapolating the two interaction coefficients β1 and β2 forward to 2035 requires assumptions the data cannot test: that utilities do not deploy managed charging at scale, that V2G and dynamic rates remain limited, that the EV fleet's charging behavior stays similar to today's, and that the causal mechanism itself stays stable over a decade. Those are assumptions a stakeholder-facing decision aid can legitimately make (“if nothing changes, here is where we land”), but that a peer-reviewable empirical paper should not. The forecast therefore belongs in a policy-monitoring tool, not in the paper's results section. The same logic applies to the raw source code: it's essential for reproducibility and anyone who wants to audit the pipeline, but it would consume half the report page count without adding to the argument. Keeping both here keeps the report focused and the dashboard useful.
California targets 5 million EVs by 2030 and 100% ZEV sales by 2035. Under the no-intervention assumption — β1 and β2 held constant at their Apr 2022 – May 2025 values — the slider below projects how the volatility burden would shift across income levels over the next decade. Treat as a directional scenario, not a point forecast.
Projected EV Adoption & Volatility Impact by County Income
Projection Year
2025
Est. EVs Statewide
1.5M
Volatility Burden ($40k County)
0.03%
of electricity spend
Volatility Burden ($120k County)
~0%
of electricity spend
Projection formula
Uses the same β1 = +0.003554 and β2 = −0.000303 from the interaction 2SLS. EV trajectory anchors: 1.5M (2025), 5.0M (2030, state target), 12.2M (2035, 100% ZEV sales). Household consumption fixed at 7 MWh/yr (CA residential average). Spend benchmarks by income tier: $2,200 ($40k county), $2,700 ($70k), $3,800 ($120k) — from the energy-burden literature. Floored at 0. No parametric CIs are shaded — the dominant uncertainty here is structural (coefficient stability over a decade), not sampling variance, and narrow parametric bands would imply a precision this extrapolation does not have.
Every cell from the seven production notebooks is included below, grouped by workstream. Collapsed by default — click to expand. These are the complete notebooks as run; nothing has been edited for brevity aside from trimming one very long scraped-data literal for readability (flagged inline). Long; use Ctrl/Cmd+F once a notebook is expanded.
EV Ecosystem · causal identification
Two complementary identification strategies for EV adoption: Tesla-store IDW exposure as the preferred IV, CVRP closure as a shift-share falsification check.
ev_iv_tesla_2sls.ipynb46 cellsimport pandas as _hex_pandas
import datetime as _hex_datetime
import json as _hex_jsonpip install linearmodelsfrom google.colab import drive
drive.mount('/content/drive')from pathlib import Path
## YOU WILL NEED TO UPDATE MYDRIVE TO SHAREDWITHME IF YOU ARE TRYING TO RUN THIS AND AREN'T RACHEL
DATA_DIR = Path(
"/content/drive/MyDrive/Capstone: Watts the Problem/Datasets - Colab"
)
DATA_DIR.exists()import pandas as pd
panel_df = pd.read_csv(DATA_DIR / "panel_df.csv")
county_shape_df = pd.read_csv(DATA_DIR / "county_shapefile_df.csv")
tesla_loc_df = pd.read_csv(DATA_DIR / "tesla_loc_df.csv")
census_df = pd.read_csv(DATA_DIR / "census_df.csv")
solar_df = pd.read_csv(DATA_DIR / "solar_control_variable.csv")
tract_shape_df = pd.read_csv(DATA_DIR / "tract_shapefile_df.csv")
tract_panel_df = pd.read_csv(DATA_DIR / "tract_panel_df.csv")
tract_census_df = pd.read_csv(DATA_DIR / "tract_census_df.csv")import numpy as np
import pandas as pd
# -------------------------
# Helpers
# -------------------------
def haversine_km(lat1, lon1, lat2, lon2):
R = 6371.0
lat1, lon1, lat2, lon2 = map(np.radians, [lat1, lon1, lat2, lon2])
dlat = lat2 - lat1
dlon = lon2 - lon1
a = np.sin(dlat / 2) ** 2 + np.cos(lat1) * np.cos(lat2) * np.sin(dlon / 2) ** 2
return 2 * R * np.arcsin(np.sqrt(a))
def to_month_start(x):
return pd.to_datetime(x, errors="coerce").dt.to_period("M").dt.to_timestamp()
# -------------------------
# Prepare inputs
# -------------------------
df = panel_df.copy()
df["month"] = to_month_start(df["month"])
county_points_df = (
county_shape_df[["county_geoid", "internal_point_latitude", "internal_point_longitude"]]
.drop_duplicates("county_geoid")
.rename(columns={"internal_point_latitude": "county_lat", "internal_point_longitude": "county_lon"})
.copy()
)
county_points_df["county_lat"] = pd.to_numeric(county_points_df["county_lat"], errors="coerce")
county_points_df["county_lon"] = pd.to_numeric(county_points_df["county_lon"], errors="coerce")
county_points_df = county_points_df.dropna(subset=["county_lat", "county_lon"])
tesla_stores_df = tesla_loc_df[["latitude", "longitude", "open_month"]].copy()
tesla_stores_df["latitude"] = pd.to_numeric(tesla_stores_df["latitude"], errors="coerce")
tesla_stores_df["longitude"] = pd.to_numeric(tesla_stores_df["longitude"], errors="coerce")
tesla_stores_df["open_month"] = to_month_start(tesla_stores_df["open_month"])
tesla_stores_df = tesla_stores_df.dropna(subset=["latitude", "longitude", "open_month"])
# -------------------------
# Compute tesla_store_exposure_log across all months
# This measures inverse-distance-weighted exposure to open Tesla stores
# -------------------------
months = pd.DatetimeIndex(sorted(df["month"].dropna().unique()))
counties = county_points_df[["county_geoid", "county_lat", "county_lon"]].copy()
eps_km = 1.0 # minimum distance cap to avoid division by zero
iv_rows = []
for m in months:
open_stores_m = tesla_stores_df.loc[tesla_stores_df["open_month"] <= m]
if open_stores_m.empty:
continue
tmp_m = counties.assign(_k=1).merge(open_stores_m.assign(_k=1), on="_k").drop(columns="_k")
tmp_m["dist_km"] = haversine_km(tmp_m["county_lat"], tmp_m["county_lon"], tmp_m["latitude"], tmp_m["longitude"])
tmp_m["dist_km_capped"] = np.maximum(tmp_m["dist_km"], eps_km)
iv_m = tmp_m.groupby("county_geoid", as_index=False).agg(
idw_exposure=("dist_km_capped", lambda s: float(np.sum(1.0 / s))),
)
iv_m["month"] = m
iv_rows.append(iv_m)
iv_table = pd.concat(iv_rows, ignore_index=True)
iv_table["tesla_store_exposure"] = iv_table["idw_exposure"]
iv_table["iv_tesla_store_exposure_log"] = np.log1p(iv_table["idw_exposure"])
iv_table = iv_table[["county_geoid", "month", "tesla_store_exposure", "iv_tesla_store_exposure_log"]]
# Merge baseline IV onto the panel
df_iv = df.merge(iv_table, on=["county_geoid", "month"], how="left", validate="m:1")
df_ivimport pandas as pd
import numpy as np
import geopandas as gpd
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
from matplotlib.lines import Line2D
# Get the maximum month in the sample
max_month = pd.to_datetime(df_iv['month'].max())
print(f"Mapping Instrument for: {max_month.strftime('%B %Y')}")
# Filter data for max month
iv_max_month = df_iv[df_iv['month'] == df_iv['month'].max()].copy()
# Get stores open by max month
tesla_open_dates = pd.to_datetime(tesla_loc_df['open_month']).dt.tz_localize(None)
stores_open = tesla_loc_df[tesla_open_dates <= max_month].copy()
print(f"Tesla stores open: {len(stores_open)}")
# Create GeoDataFrame with correct geometry column
county_gdf = gpd.GeoDataFrame(
county_shape_df,
geometry=gpd.GeoSeries.from_wkt(county_shape_df['county_geometry']),
crs='EPSG:4326'
)
# Merge with IV values
county_map_data = county_gdf.merge(
iv_max_month[['county_geoid', 'iv_tesla_store_exposure_log', 'tesla_store_exposure']],
on='county_geoid',
how='left'
)
# Get county centroids
county_centroids = county_points_df.copy()
# Create figure
fig, ax = plt.subplots(1, 1, figsize=(16, 14))
# Plot county boundaries colored by log exposure
county_map_data.plot(
column='iv_tesla_store_exposure_log',
ax=ax,
cmap='YlOrRd',
edgecolor='black',
linewidth=0.3,
legend=True,
legend_kwds={'label': 'Tesla Store Exposure (log)\nlog(1 + Σ(1/dist))', 'shrink': 0.6},
missing_kwds={'color': 'lightgrey', 'edgecolor': 'black', 'linewidth': 0.3}
)
# Plot county centroids
ax.scatter(
county_centroids['county_lon'],
county_centroids['county_lat'],
c='blue',
s=20,
alpha=0.7,
marker='o',
edgecolors='darkblue',
linewidths=0.5,
label='County Centroids',
zorder=3
)
# Plot Tesla stores
ax.scatter(
stores_open['longitude'],
stores_open['latitude'],
c='red',
s=180,
alpha=0.95,
marker='*',
edgecolors='darkred',
linewidths=1.5,
label=f'Tesla Stores (n={len(stores_open)})',
zorder=4
)
# Add title
ax.set_title(
f'Instrument Construction: Inverse-Distance Weighted Tesla Store Exposure\n{max_month.strftime("%B %Y")}',
fontsize=16, fontweight='bold', pad=20
)
ax.set_xlabel('Longitude', fontsize=12, fontweight='bold')
ax.set_ylabel('Latitude', fontsize=12, fontweight='bold')
# Customize legend
ax.legend(loc='lower left', fontsize=11, framealpha=0.95, edgecolor='black')
# Clean axes
ax.set_xticks([])
ax.set_yticks([])
for spine in ax.spines.values():
spine.set_visible(False)
plt.tight_layout()
plt.show()
# Print summary
print("\n" + "=" * 80)
print(f"TESLA STORE EXPOSURE SUMMARY ({max_month.strftime('%B %Y')})")
print("=" * 80)
print(f"\nRaw Exposure (Σ 1/distance_km):")
print(iv_max_month['tesla_store_exposure'].describe().round(3))
print(f"\nLog Exposure (used in regressions):")
print(iv_max_month['iv_tesla_store_exposure_log'].describe().round(3))
# Top counties
print("\n" + "=" * 80)
print("TOP 10 COUNTIES BY TESLA STORE EXPOSURE")
print("=" * 80)
top_counties = iv_max_month.nlargest(10, 'tesla_store_exposure')[
['county_name', 'tesla_store_exposure', 'iv_tesla_store_exposure_log', 'evs_per_1000']
].round(3)
display(top_counties)
print("\n" + "=" * 80)
print("MAP INTERPRETATION")
print("=" * 80)
print("• DARK RED counties = High exposure (near multiple Tesla stores)")
print("• YELLOW counties = Low exposure (rural, far from stores)")
print("• BLUE DOTS = County centroids (distances calculated from this point)")
print("• RED STARS = Tesla stores (70 locations statewide)")
print("\nGEOGRAPHIC PATTERNS:")
print("• Bay Area: Highest exposure (San Francisco, Santa Clara, Alameda)")
print("• Southern CA: Concentrated around Los Angeles, Orange, San Diego")
print("• Central Valley & Rural: Minimal exposure (far from all stores)")
print("\nCALCULATION: For each county, exposure = log(1 + Σ(1/distance_km to each store))")
print("• Closer stores contribute MORE (1/10km = 0.10 vs 1/100km = 0.01)")
print("• Multiple nearby stores compound the exposure")
print("• This creates continuous spatial variation exploited by the instrument")
print("=" * 80)import pandas as pd
import numpy as np
# Convert both to string for merge
df_iv_temp = df_iv.copy()
df_iv_temp['county_geoid'] = df_iv_temp['county_geoid'].astype(str)
# Drop old census columns if they exist (from previous merge attempts)
census_cols = ['percent_dac_tracts', 'median_household_income',
'percent_bachelor_or_higher', 'percent_no_vehicle_available']
for col in census_cols:
if col in df_iv_temp.columns:
df_iv_temp = df_iv_temp.drop(columns=[col])
census_temp = census_df.copy()
census_temp['county_geoid'] = census_temp['county_geoid'].astype(str)
# Merge census demographics onto IV panel
df_with_census = df_iv_temp.merge(
census_temp[['county_geoid'] + census_cols],
on='county_geoid',
how='left',
validate='m:1'
)
print("=" * 80)
print("CENSUS DATA MERGE")
print("=" * 80)
print(f"\nOriginal df_iv: {len(df_iv)} rows, {df_iv['county_geoid'].nunique()} counties")
print(f"After merge: {len(df_with_census)} rows, {df_with_census['county_geoid'].nunique()} counties")
print(f"\nMissing census data by variable:")
for var in census_cols:
missing = df_with_census[var].isna().sum()
pct_missing = missing/len(df_with_census)*100
print(f" {var}: {missing} missing ({pct_missing:.1f}%)")
# -------------------------
# Create Interaction Terms for Heterogeneity Analysis
# -------------------------
print("\n" + "=" * 80)
print("CREATING INTERACTION TERMS")
print("=" * 80)
# Standardize income to avoid numerical issues (convert to $10k units)
df_with_census["median_household_income_10k"] = df_with_census["median_household_income"] / 10000
# Create EV × Demographics interactions (for 2nd stage)
df_with_census["ev_x_dac"] = df_with_census["evs_per_1000"] * df_with_census["percent_dac_tracts"]
df_with_census["ev_x_income"] = df_with_census["evs_per_1000"] * df_with_census["median_household_income_10k"]
# Create IV × Demographics interactions (for 1st stage)
df_with_census["iv_x_dac"] = df_with_census["iv_tesla_store_exposure_log"] * df_with_census["percent_dac_tracts"]
df_with_census["iv_x_income"] = df_with_census["iv_tesla_store_exposure_log"] * df_with_census["median_household_income_10k"]
print("\nInteraction terms created:")
print(" • ev_x_dac (evs_per_1000 × percent_dac_tracts)")
print(" • ev_x_income (evs_per_1000 × median_household_income_10k)")
print(" • iv_x_dac (iv_tesla_store_exposure_log × percent_dac_tracts)")
print(" • iv_x_income (iv_tesla_store_exposure_log × median_household_income_10k)")
print("\n" + "=" * 80)
print("Census data and interaction terms ready for heterogeneity analysis")
print("=" * 80)
df_with_censusimport pandas as pd
import numpy as np
# -------------------------
# Create Interaction Terms for Heterogeneity Analysis
# -------------------------
df_interactions = df_with_census.copy()
print("=" * 80)
print("CREATING INTERACTION TERMS FOR HETEROGENEITY ANALYSIS")
print("=" * 80)
# Standardize income to avoid numerical issues (convert to $10k units)
df_interactions["median_household_income_10k"] = df_interactions["median_household_income"] / 10000
print("\n1. Income standardized to $10k units")
print(f" Original range: ${df_interactions['median_household_income'].min():.0f} - ${df_interactions['median_household_income'].max():.0f}")
print(f" Standardized range: {df_interactions['median_household_income_10k'].min():.1f} - {df_interactions['median_household_income_10k'].max():.1f}")
# Create EV × Demographics interactions (for 2nd stage)
df_interactions["ev_x_dac"] = df_interactions["evs_per_1000"] * df_interactions["percent_dac_tracts"]
df_interactions["ev_x_income"] = df_interactions["evs_per_1000"] * df_interactions["median_household_income_10k"]
# Create IV × Demographics interactions (for 1st stage instruments)
df_interactions["iv_x_dac"] = df_interactions["iv_tesla_store_exposure_log"] * df_interactions["percent_dac_tracts"]
df_interactions["iv_x_income"] = df_interactions["iv_tesla_store_exposure_log"] * df_interactions["median_household_income_10k"]
print("\n2. Interaction terms created:")
print(" For 2nd stage (endogenous interactions):")
print(" • ev_x_dac = evs_per_1000 × percent_dac_tracts")
print(" • ev_x_income = evs_per_1000 × median_household_income_10k")
print("\n For 1st stage (instruments for interactions):")
print(" • iv_x_dac = iv_tesla_store_exposure_log × percent_dac_tracts")
print(" • iv_x_income = iv_tesla_store_exposure_log × median_household_income_10k")
# Check for missing values in interaction terms
print("\n3. Missing values check:")
interaction_cols = ["ev_x_dac", "ev_x_income", "iv_x_dac", "iv_x_income"]
for col in interaction_cols:
missing = df_interactions[col].isna().sum()
pct = missing / len(df_interactions) * 100
print(f" {col:<20} {missing:>6} missing ({pct:>5.1f}%)")
print("\n" + "=" * 80)
print("Interaction terms ready for 2SLS heterogeneity analysis")
print("=" * 80)
print("\nNote: Time-invariant moderators (DAC %, income) will be absorbed by")
print(" county fixed effects. Only the interaction terms are identified,")
print(" allowing us to test how EV effects vary across the demographic gradient.")
print("=" * 80)
df_interactionsimport pandas as pd
# -------------------------
# Load & merge solar control variable
# -------------------------
#solar_df = pd.read_csv("solar_control_variable.csv")
solar_df["month"] = pd.to_datetime(solar_df["month"])
# -------------------------
# Validate solar lookup coverage
# -------------------------
solar_keys = set(zip(solar_df["county_name"], solar_df["month"]))
panel_keys = set(zip(df_interactions["county_name"], pd.to_datetime(df_interactions["month"])))
solar_only = solar_keys - panel_keys
panel_only_missing = panel_keys - solar_keys
print("SOLAR MERGE VALIDATION")
print("=" * 80)
print(f" Solar rows: {len(solar_df):,}")
print(f" Panel rows: {len(df_interactions):,}")
print(f" Solar keys matched: {len(solar_keys & panel_keys):,} / {len(solar_keys):,}")
if solar_only:
print(f"\n ⚠ {len(solar_only)} solar rows have NO match in df_interactions:")
for cn, m in sorted(solar_only)[:20]:
print(f" {cn} | {m.strftime('%Y-%m-%d')}")
if len(solar_only) > 20:
print(f" ... and {len(solar_only) - 20} more")
else:
print("\n ✓ Every solar row matched a panel row")
if panel_only_missing:
print(f"\n ⚠ {len(panel_only_missing)} panel rows will have NULL solar values (no solar data):")
for cn, m in sorted(panel_only_missing)[:20]:
print(f" {cn} | {m.strftime('%Y-%m-%d')}")
if len(panel_only_missing) > 20:
print(f" ... and {len(panel_only_missing) - 20} more")
else:
print("\n ✓ Every panel row has a solar match")
print("=" * 80 + "\n")
# -------------------------
# Create Final Analysis Dataset
# -------------------------
print("=" * 80)
print("CREATING FINAL ANALYSIS DATASET")
print("=" * 80)
# Select only the columns needed for 2SLS analysis
analysis_cols = [
# Identifiers
"county_geoid",
"county_name",
"month",
# Outcomes (dependent variables)
"avg_lmp_weighted",
"mad_rate",
"median_of_medians",
# Endogenous variable
"evs_per_1000",
# Instrument
"iv_tesla_store_exposure_log",
# Controls
"tmax_c",
"cumulative_data_center_power_mw",
"total_generator_active_capacity_mw",
"solar_mw_per_1000_lag1",
# Demographics (moderators)
"percent_dac_tracts",
"median_household_income_10k",
# Interaction terms (2nd stage)
"ev_x_dac",
"ev_x_income",
# Interaction instruments (1st stage)
"iv_x_dac",
"iv_x_income"
]
df_final = df_interactions.merge(
solar_df, on=["county_name", "month"], how="left"
)[analysis_cols].copy()
print("\n1. DATASET STRUCTURE")
print("=" * 80)
print(f"Rows: {len(df_final):,}")
print(f"Columns: {len(df_final.columns)}")
print(f"Counties: {df_final['county_geoid'].nunique()}")
print(f"Time periods: {df_final['month'].nunique()}")
print(f"Date range: {df_final['month'].min()} to {df_final['month'].max()}")
print("\n2. COLUMN GROUPS")
print("=" * 80)
print("\n IDENTIFIERS (3):")
print(" • county_geoid, county_name, month")
print("\n OUTCOMES (3):")
print(" • avg_lmp_weighted - Average electricity price")
print(" • mad_rate - Price volatility (MAD)")
print(" • median_of_medians - Median electricity price")
print("\n TREATMENT & INSTRUMENT (2):")
print(" • evs_per_1000 - EV adoption (endogenous)")
print(" • iv_tesla_store_exposure_log - Tesla store exposure (instrument)")
print("\n CONTROLS (4):")
print(" • tmax_c - Max temperature")
print(" • cumulative_data_center_power_mw - Data center demand")
print(" • total_generator_active_capacity_mw - Generation capacity")
print(" • solar_mw_per_1000_lag1 - Solar capacity (lagged)")
print("\n DEMOGRAPHICS (2):")
print(" • percent_dac_tracts - % disadvantaged tracts")
print(" • median_household_income_10k - Median income ($10k units)")
print("\n INTERACTIONS (4):")
print(" • ev_x_dac - EV × DAC (2nd stage)")
print(" • ev_x_income - EV × Income (2nd stage)")
print(" • iv_x_dac - IV × DAC (1st stage instrument)")
print(" • iv_x_income - IV × Income (1st stage instrument)")
print("\n3. MISSING DATA")
print("=" * 80)
total_missing = df_final.isna().sum()
if total_missing.sum() > 0:
print("\nColumns with missing values:")
for col, missing in total_missing[total_missing > 0].items():
pct = missing / len(df_final) * 100
print(f" {col:<40} {missing:>6} ({pct:>5.1f}%)")
else:
print("\n ✓ No missing values in any column")
print("\n" + "=" * 80)
print("Final analysis dataset ready for 2SLS regressions")
print("=" * 80)
df_finaldf_final.columns.tolist()import scipy.stats as stats
# County-level: mean solar capacity and income
county_level = df_final.groupby('county_name').agg(
mean_solar=('solar_mw_per_1000_lag1', 'mean'),
income=('median_household_income_10k', 'first')
).reset_index()
# Correlation
r, p = stats.pearsonr(county_level['mean_solar'], county_level['income'])
print(f"Pearson r: {r:.3f}, p-value: {p:.4f}")
# Scatter plot
import matplotlib.pyplot as plt
plt.scatter(county_level['income'], county_level['mean_solar'])
for _, row in county_level.iterrows():
plt.annotate(row['county_name'], (row['income'], row['mean_solar']), fontsize=6)
plt.xlabel('Median Household Income ($10k)')
plt.ylabel('Mean Solar Capacity per 1,000')
plt.title('Solar Capacity vs Income by County')
plt.tight_layout()
plt.show()import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
# -------------------------------------------------------------------
# First-stage diagnostics: testing instrument relevance
# Including interaction instruments for heterogeneity analysis
# with cluster-robust standard errors (clustered by county_geoid)
# -------------------------------------------------------------------
df = df_final.copy()
# Ensure proper types
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)
num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce").astype(float)
# Controls (matching second-stage specification)
controls = [
"tmax_c",
"cumulative_data_center_power_mw",
"total_generator_active_capacity_mw",
"solar_mw_per_1000_lag1"
]
print("=" * 80)
print("FIRST-STAGE DIAGNOSTICS")
print("=" * 80)
print(f"\nControls: {', '.join(controls)}")
print(f"Fixed Effects: County FE + Month FE")
print(f"Standard Errors: Clustered by county")
print("\n" + "=" * 80)
first_stage_results = []
# ==========================================
# 1. Main Effect: evs_per_1000 ~ iv_tesla_store_exposure_log
# ==========================================
print("\n1. MAIN EFFECT INSTRUMENT")
print("-" * 80)
needed_cols = ["county_geoid", "month", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls
data_main = df[needed_cols].dropna().copy()
formula_main = (
"evs_per_1000 ~ iv_tesla_store_exposure_log"
+ " + " + " + ".join(controls)
+ " + C(county_geoid) + C(month)"
)
model_main_classic = smf.ols(formula_main, data=data_main).fit()
model_main_cluster = smf.ols(formula_main, data=data_main).fit(
cov_type="cluster",
cov_kwds={"groups": data_main["county_geoid"]},
)
classic_test_main = model_main_classic.f_test("iv_tesla_store_exposure_log = 0")
classic_f_main = float(np.asarray(classic_test_main.fvalue).squeeze())
robust_test_main = model_main_cluster.wald_test("iv_tesla_store_exposure_log = 0", scalar=True)
robust_stat_main = float(robust_test_main.statistic)
first_stage_results.append({
"Endogenous Variable": "evs_per_1000",
"Instrument": "iv_tesla_store_exposure_log",
"Coefficient": float(model_main_cluster.params["iv_tesla_store_exposure_log"]),
"SE (cluster)": float(model_main_cluster.bse["iv_tesla_store_exposure_log"]),
"P-value": float(model_main_cluster.pvalues["iv_tesla_store_exposure_log"]),
"F-stat (classic)": classic_f_main,
"Wald (cluster)": robust_stat_main,
"N": int(model_main_cluster.nobs),
})
print(f"Dependent: evs_per_1000")
print(f"Instrument: iv_tesla_store_exposure_log")
print(f"F-stat (classic): {classic_f_main:.2f} | Wald (cluster): {robust_stat_main:.2f}")
status_main = '✓ STRONG' if classic_f_main > 10 else '✗ WEAK'
print(f"Status: {status_main} (F > 10)")
# ==========================================
# 2. DAC Interaction: ev_x_dac ~ iv_x_dac
# ==========================================
print("\n2. DAC INTERACTION INSTRUMENT")
print("-" * 80)
needed_cols_dac = ["county_geoid", "month", "ev_x_dac", "iv_x_dac"] + controls
data_dac = df[needed_cols_dac].dropna().copy()
formula_dac = (
"ev_x_dac ~ iv_x_dac"
+ " + " + " + ".join(controls)
+ " + C(county_geoid) + C(month)"
)
model_dac_classic = smf.ols(formula_dac, data=data_dac).fit()
model_dac_cluster = smf.ols(formula_dac, data=data_dac).fit(
cov_type="cluster",
cov_kwds={"groups": data_dac["county_geoid"]},
)
classic_test_dac = model_dac_classic.f_test("iv_x_dac = 0")
classic_f_dac = float(np.asarray(classic_test_dac.fvalue).squeeze())
robust_test_dac = model_dac_cluster.wald_test("iv_x_dac = 0", scalar=True)
robust_stat_dac = float(robust_test_dac.statistic)
first_stage_results.append({
"Endogenous Variable": "ev_x_dac",
"Instrument": "iv_x_dac",
"Coefficient": float(model_dac_cluster.params["iv_x_dac"]),
"SE (cluster)": float(model_dac_cluster.bse["iv_x_dac"]),
"P-value": float(model_dac_cluster.pvalues["iv_x_dac"]),
"F-stat (classic)": classic_f_dac,
"Wald (cluster)": robust_stat_dac,
"N": int(model_dac_cluster.nobs),
})
print(f"Dependent: ev_x_dac (EV × percent_dac_tracts)")
print(f"Instrument: iv_x_dac (IV × percent_dac_tracts)")
print(f"F-stat (classic): {classic_f_dac:.2f} | Wald (cluster): {robust_stat_dac:.2f}")
status_dac = '✓ STRONG' if classic_f_dac > 10 else '✗ WEAK'
print(f"Status: {status_dac} (F > 10)")
# ==========================================
# 3. Income Interaction: ev_x_income ~ iv_x_income
# ==========================================
print("\n3. INCOME INTERACTION INSTRUMENT")
print("-" * 80)
needed_cols_income = ["county_geoid", "month", "ev_x_income", "iv_x_income"] + controls
data_income = df[needed_cols_income].dropna().copy()
formula_income = (
"ev_x_income ~ iv_x_income"
+ " + " + " + ".join(controls)
+ " + C(county_geoid) + C(month)"
)
model_income_classic = smf.ols(formula_income, data=data_income).fit()
model_income_cluster = smf.ols(formula_income, data=data_income).fit(
cov_type="cluster",
cov_kwds={"groups": data_income["county_geoid"]},
)
classic_test_income = model_income_classic.f_test("iv_x_income = 0")
classic_f_income = float(np.asarray(classic_test_income.fvalue).squeeze())
robust_test_income = model_income_cluster.wald_test("iv_x_income = 0", scalar=True)
robust_stat_income = float(robust_test_income.statistic)
first_stage_results.append({
"Endogenous Variable": "ev_x_income",
"Instrument": "iv_x_income",
"Coefficient": float(model_income_cluster.params["iv_x_income"]),
"SE (cluster)": float(model_income_cluster.bse["iv_x_income"]),
"P-value": float(model_income_cluster.pvalues["iv_x_income"]),
"F-stat (classic)": classic_f_income,
"Wald (cluster)": robust_stat_income,
"N": int(model_income_cluster.nobs),
})
print(f"Dependent: ev_x_income (EV × median_household_income_10k)")
print(f"Instrument: iv_x_income (IV × median_household_income_10k)")
print(f"F-stat (classic): {classic_f_income:.2f} | Wald (cluster): {robust_stat_income:.2f}")
status_income = '✓ STRONG' if classic_f_income > 10 else '✗ WEAK'
print(f"Status: {status_income} (F > 10)")
# ==========================================
# Summary Table
# ==========================================
print("\n" + "=" * 80)
print("SUMMARY: ALL FIRST-STAGE INSTRUMENTS")
print("=" * 80)
summary_df = pd.DataFrame(first_stage_results)
display(summary_df)
print("\n" + "=" * 80)
print("INTERPRETATION")
print("=" * 80)
all_strong = all(r["F-stat (classic)"] > 10 for r in first_stage_results)
if all_strong:
print("\n✓ ALL instruments are STRONG (F > 10)")
print("✓ Main effect AND interaction effects are well-identified")
print("✓ Can proceed with heterogeneity analysis with confidence")
else:
print("\n⚠️ Some instruments are WEAK (F < 10)")
print("⚠️ Weak instrument bias may affect coefficient estimates")
print("=" * 80)
# ==========================================
# NAIVE OLS: mad_rate ~ evs_per_1000 (no FEs, no IV)
# ==========================================
data_ols_nofe = df[
["county_geoid", "mad_rate", "evs_per_1000"] + controls
].dropna().copy()
formula_ols_nofe = "mad_rate ~ evs_per_1000 + " + " + ".join(controls)
model_ols_nofe = smf.ols(formula_ols_nofe, data=data_ols_nofe).fit(
cov_type="cluster",
cov_kwds={"groups": data_ols_nofe["county_geoid"]},
)
print(model_ols_nofe.summary())import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
# ==========================================
# FULL FIRST-STAGE REGRESSION TABLES
# ==========================================
# This cell shows detailed regression output for all three first-stage regressions
# (Main effect, DAC interaction, Income interaction)
df = df_final.copy()
# Ensure proper types
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)
num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce").astype(float)
controls = [
"tmax_c",
"cumulative_data_center_power_mw",
"total_generator_active_capacity_mw"
]
print("=" * 80)
print("DETAILED FIRST-STAGE REGRESSION OUTPUTS")
print("=" * 80)
print("\nShowing full regression tables with cluster-robust standard errors")
print("=" * 80)
# ==========================================
# 1. Main Effect
# ==========================================
print("\n\n1. MAIN EFFECT: evs_per_1000 ~ iv_tesla_store_exposure_log")
print("=" * 80)
needed_cols = ["county_geoid", "month", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls
data_main = df[needed_cols].dropna().copy()
formula_main = (
"evs_per_1000 ~ iv_tesla_store_exposure_log"
+ " + " + " + ".join(controls)
+ " + C(county_geoid) + C(month)"
)
model_main = smf.ols(formula_main, data=data_main).fit(
cov_type="cluster",
cov_kwds={"groups": data_main["county_geoid"]},
)
print(model_main.summary())
# ==========================================
# 2. DAC Interaction
# ==========================================
print("\n\n2. DAC INTERACTION: ev_x_dac ~ iv_x_dac")
print("=" * 80)
needed_cols_dac = ["county_geoid", "month", "ev_x_dac", "iv_x_dac"] + controls
data_dac = df[needed_cols_dac].dropna().copy()
formula_dac = (
"ev_x_dac ~ iv_x_dac"
+ " + " + " + ".join(controls)
+ " + C(county_geoid) + C(month)"
)
model_dac = smf.ols(formula_dac, data=data_dac).fit(
cov_type="cluster",
cov_kwds={"groups": data_dac["county_geoid"]},
)
print(model_dac.summary())
# ==========================================
# 3. Income Interaction
# ==========================================
print("\n\n3. INCOME INTERACTION: ev_x_income ~ iv_x_income")
print("=" * 80)
needed_cols_income = ["county_geoid", "month", "ev_x_income", "iv_x_income"] + controls
data_income = df[needed_cols_income].dropna().copy()
formula_income = (
"ev_x_income ~ iv_x_income"
+ " + " + " + ".join(controls)
+ " + C(county_geoid) + C(month)"
)
model_income = smf.ols(formula_income, data=data_income).fit(
cov_type="cluster",
cov_kwds={"groups": data_income["county_geoid"]},
)
print(model_income.summary())
print("\n" + "=" * 80)
print("END OF FIRST-STAGE REGRESSION TABLES")
print("=" * 80)import numpy as np
import pandas as pd
from linearmodels.iv import IV2SLS
df = df_final.copy()
# Type prep
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)
# Convert numeric columns (excluding identifiers)
num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")
# EXPANDED CONTROLS (matching first-stage)
# Adding generator capacity to control for supply-side factors
controls = [
"tmax_c",
"cumulative_data_center_power_mw",
"total_generator_active_capacity_mw",
"solar_mw_per_1000_lag1"
]
control_str = " + ".join(controls)
results = {}
print("=" * 80)
print("TWO-STAGE LEAST SQUARES REGRESSION")
print("=" * 80)
print(f"\nExpanded controls ({len(controls)}):")
for ctrl in controls:
print(f" • {ctrl}")
print(f"\nFixed Effects: County FE + Month FE")
print(f"Standard Errors: Clustered by county")
print("\n" + "=" * 80)
# ==========================================
# Spec 1: avg_lmp_weighted
# ==========================================
iv_cols = ["avg_lmp_weighted", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]
iv_data = df[iv_cols].dropna().copy()
for c in iv_data.columns:
if pd.api.types.is_numeric_dtype(iv_data[c]):
iv_data[c] = iv_data[c].astype(float)
iv_formula = f"""
avg_lmp_weighted ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""
iv_model = IV2SLS.from_formula(iv_formula, data=iv_data).fit(
cov_type="clustered",
clusters=iv_data["county_geoid"],
)
results["avg_lmp_weighted"] = {
"outcome": "avg_lmp_weighted",
"coef": iv_model.params["evs_per_1000"],
"se": iv_model.std_errors["evs_per_1000"],
"pval": iv_model.pvalues["evs_per_1000"],
"n": int(iv_model.nobs),
}
print("\nSpec 1: EV Adoption → Average Price (avg_lmp_weighted)")
print("=" * 80)
print(iv_model.summary.tables[1])
# ==========================================
# Spec 2: mad_rate
# ==========================================
iv_cols_mad = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]
iv_data_mad = df[iv_cols_mad].dropna().copy()
for c in iv_data_mad.columns:
if pd.api.types.is_numeric_dtype(iv_data_mad[c]):
iv_data_mad[c] = iv_data_mad[c].astype(float)
iv_formula_mad = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""
iv_model_mad = IV2SLS.from_formula(iv_formula_mad, data=iv_data_mad).fit(
cov_type="clustered",
clusters=iv_data_mad["county_geoid"],
)
results["mad_rate"] = {
"outcome": "mad_rate",
"coef": iv_model_mad.params["evs_per_1000"],
"se": iv_model_mad.std_errors["evs_per_1000"],
"pval": iv_model_mad.pvalues["evs_per_1000"],
"n": int(iv_model_mad.nobs),
}
print("\n" + "=" * 80)
print("Spec 2: EV Adoption → Price Volatility (mad_rate)")
print("=" * 80)
print(iv_model_mad.summary.tables[1])
# ==========================================
# Spec 3: median_of_medians
# ==========================================
iv_cols_median = ["median_of_medians", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]
iv_data_median = df[iv_cols_median].dropna().copy()
for c in iv_data_median.columns:
if pd.api.types.is_numeric_dtype(iv_data_median[c]):
iv_data_median[c] = iv_data_median[c].astype(float)
iv_formula_median = f"""
median_of_medians ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""
iv_model_median = IV2SLS.from_formula(iv_formula_median, data=iv_data_median).fit(
cov_type="clustered",
clusters=iv_data_median["county_geoid"],
)
results["median_of_medians"] = {
"outcome": "median_of_medians",
"coef": iv_model_median.params["evs_per_1000"],
"se": iv_model_median.std_errors["evs_per_1000"],
"pval": iv_model_median.pvalues["evs_per_1000"],
"n": int(iv_model_median.nobs),
}
print("\n" + "=" * 80)
print("Spec 3: EV Adoption → Median Price (median_of_medians)")
print("=" * 80)
print(iv_model_median.summary.tables[1])
# ==========================================
# Spec 4: mad_rate with DAC Interaction
# ==========================================
print("\n" + "=" * 80)
print("Spec 4: EV × DAC Interaction → Price Volatility (mad_rate)")
print("=" * 80)
iv_cols_mad_dac = ["mad_rate", "evs_per_1000", "ev_x_dac",
"iv_tesla_store_exposure_log", "iv_x_dac"] + controls + ["county_geoid", "month"]
iv_data_mad_dac = df[iv_cols_mad_dac].dropna().copy()
for c in iv_data_mad_dac.columns:
if pd.api.types.is_numeric_dtype(iv_data_mad_dac[c]):
iv_data_mad_dac[c] = iv_data_mad_dac[c].astype(float)
iv_formula_mad_dac = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 + ev_x_dac ~ iv_tesla_store_exposure_log + iv_x_dac]
"""
iv_model_mad_dac = IV2SLS.from_formula(iv_formula_mad_dac, data=iv_data_mad_dac).fit(
cov_type="clustered",
clusters=iv_data_mad_dac["county_geoid"],
)
results["mad_rate_dac_interaction"] = {
"outcome": "mad_rate (DAC interaction)",
"main_coef": iv_model_mad_dac.params["evs_per_1000"],
"main_se": iv_model_mad_dac.std_errors["evs_per_1000"],
"main_pval": iv_model_mad_dac.pvalues["evs_per_1000"],
"interaction_coef": iv_model_mad_dac.params["ev_x_dac"],
"interaction_se": iv_model_mad_dac.std_errors["ev_x_dac"],
"interaction_pval": iv_model_mad_dac.pvalues["ev_x_dac"],
"n": int(iv_model_mad_dac.nobs),
}
print(iv_model_mad_dac.summary.tables[1])
# ==========================================
# Spec 5: mad_rate with Income Interaction
# ==========================================
print("\n" + "=" * 80)
print("Spec 5: EV × Income Interaction → Price Volatility (mad_rate)")
print("=" * 80)
iv_cols_mad_income = ["mad_rate", "evs_per_1000", "ev_x_income",
"iv_tesla_store_exposure_log", "iv_x_income"] + controls + ["county_geoid", "month"]
iv_data_mad_income = df[iv_cols_mad_income].dropna().copy()
for c in iv_data_mad_income.columns:
if pd.api.types.is_numeric_dtype(iv_data_mad_income[c]):
iv_data_mad_income[c] = iv_data_mad_income[c].astype(float)
iv_formula_mad_income = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 + ev_x_income ~ iv_tesla_store_exposure_log + iv_x_income]
"""
iv_model_mad_income = IV2SLS.from_formula(iv_formula_mad_income, data=iv_data_mad_income).fit(
cov_type="clustered",
clusters=iv_data_mad_income["county_geoid"],
)
results["mad_rate_income_interaction"] = {
"outcome": "mad_rate (Income interaction)",
"main_coef": iv_model_mad_income.params["evs_per_1000"],
"main_se": iv_model_mad_income.std_errors["evs_per_1000"],
"main_pval": iv_model_mad_income.pvalues["evs_per_1000"],
"interaction_coef": iv_model_mad_income.params["ev_x_income"],
"interaction_se": iv_model_mad_income.std_errors["ev_x_income"],
"interaction_pval": iv_model_mad_income.pvalues["ev_x_income"],
"n": int(iv_model_mad_income.nobs),
}
print(iv_model_mad_income.summary.tables[1])
results
import pandas as pd
models = {
"mad_rate": iv_model_mad,
"mad_rate (DAC int.)": iv_model_mad_dac,
"mad_rate (Income int.)": iv_model_mad_income,
}
controls_to_report = [
"tmax_c",
"cumulative_data_center_power_mw",
"total_generator_active_capacity_mw",
"solar_mw_per_1000_lag1",
]
rows = []
for spec_name, mod in models.items():
for var in controls_to_report:
if var in mod.params.index:
pval = mod.pvalues[var]
sig = "***" if pval < 0.01 else "**" if pval < 0.05 else "*" if pval < 0.10 else ""
rows.append({
"Specification": spec_name,
"Variable": var,
"Coefficient": mod.params[var],
"P-value": pval,
"Sig.": sig,
})
control_summary = pd.DataFrame(rows)
control_summary["Coefficient"] = control_summary["Coefficient"].map("{:.6f}".format)
control_summary["P-value"] = control_summary["P-value"].map("{:.4f}".format)
#print("=" * 70)
#print("SECOND-STAGE 2SLS: CONTROL VARIABLE COEFFICIENTS")
#print("=" * 70)
control_summary
import pandas as pd
import statsmodels.formula.api as smf
print("=" * 80)
print("OLS vs 2SLS COMPARISON: mad_rate (Price Volatility) ")
print("=" * 80)
print("\nAll specifications include:")
print(" • Controls: Temperature, Data Centers, Generation Capacity")
print(" • Fixed Effects: County FE + Month FE")
print(" • Standard Errors: Clustered by county")
print("\n" + "=" * 80)
print("IMPORTANT NOTES ON INTERPRETATION:")
print("=" * 80)
print("\n1. SIGN CHANGES:")
print(" 'YES' means OLS and 2SLS have opposite signs → severe endogeneity bias")
print("\n2. INTERACTION MODELS (Sections 2 & 3):")
print(" • Main effects are conditional on moderator = 0 (theoretical/extrapolated)")
print(" • DAC model: evs_per_1000 coef = effect when percent_dac_tracts = 0")
print(" • Income model: evs_per_1000 coef = effect when income = $0")
print(" • These are absorbed by county FE and not directly interpretable")
print("\n3. WHAT MATTERS:")
print(" ✓ Baseline spec: evs_per_1000 coefficient = average causal effect")
print(" ✓ Interaction specs: INTERACTION coefficient = how effect varies by demographic")
print(" ✗ DO NOT compare main effects across specifications with/without interactions")
print("\n4. INTERPRETING INTERACTIONS:")
print(" Interaction coef = change in EV effect per unit change in moderator")
print(" • If negative: EV effect decreases as moderator increases")
print(" • If positive: EV effect increases as moderator increases")
print("\n" + "=" * 80)
# ==========================================
# RUN OLS REGRESSIONS
# ==========================================
df = df_final.copy()
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)
num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")
controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
control_str = " + ".join(controls)
# Baseline OLS
ols_cols_baseline = ["mad_rate", "evs_per_1000"] + controls + ["county_geoid", "month"]
ols_data_baseline = df[ols_cols_baseline].dropna().copy()
ols_formula_baseline = f"mad_rate ~ evs_per_1000 + {control_str} + C(county_geoid) + C(month)"
ols_baseline = smf.ols(ols_formula_baseline, data=ols_data_baseline).fit(
cov_type="cluster", cov_kwds={"groups": ols_data_baseline["county_geoid"]})
# DAC Interaction OLS
ols_cols_dac = ["mad_rate", "evs_per_1000", "ev_x_dac"] + controls + ["county_geoid", "month"]
ols_data_dac = df[ols_cols_dac].dropna().copy()
ols_formula_dac = f"mad_rate ~ evs_per_1000 + ev_x_dac + {control_str} + C(county_geoid) + C(month)"
ols_dac = smf.ols(ols_formula_dac, data=ols_data_dac).fit(
cov_type="cluster", cov_kwds={"groups": ols_data_dac["county_geoid"]})
# Income Interaction OLS
ols_cols_income = ["mad_rate", "evs_per_1000", "ev_x_income"] + controls + ["county_geoid", "month"]
ols_data_income = df[ols_cols_income].dropna().copy()
ols_formula_income = f"mad_rate ~ evs_per_1000 + ev_x_income + {control_str} + C(county_geoid) + C(month)"
ols_income = smf.ols(ols_formula_income, data=ols_data_income).fit(
cov_type="cluster", cov_kwds={"groups": ols_data_income["county_geoid"]})
# ==========================================
# SECTION 1: BASELINE SPECIFICATION
# ==========================================
print("\n\n" + "=" * 80)
print("SECTION 1: BASELINE SPECIFICATION")
print("=" * 80)
print("\nModel: mad_rate ~ evs_per_1000 + controls + FE")
tsls_base = results["mad_rate"]
baseline_table = pd.DataFrame([{
"Method": "OLS",
"Coefficient": f"{ols_baseline.params['evs_per_1000']:.6f}",
"Std. Error": f"{ols_baseline.bse['evs_per_1000']:.6f}",
"P-value": f"{ols_baseline.pvalues['evs_per_1000']:.4f}",
"Sig.": "**" if ols_baseline.pvalues['evs_per_1000'] < 0.05 else "*" if ols_baseline.pvalues['evs_per_1000'] < 0.10 else "",
"N": int(ols_baseline.nobs)
}, {
"Method": "2SLS (IV)",
"Coefficient": f"{tsls_base['coef']:.6f}",
"Std. Error": f"{tsls_base['se']:.6f}",
"P-value": f"{tsls_base['pval']:.4f}",
"Sig.": "***" if tsls_base['pval'] < 0.01 else "**" if tsls_base['pval'] < 0.05 else "*" if tsls_base['pval'] < 0.10 else "",
"N": tsls_base['n']
}])
print("\n")
display(baseline_table)
sign_change_baseline = (ols_baseline.params['evs_per_1000'] * tsls_base['coef']) < 0
print(f"\nSign Change: {'YES' if sign_change_baseline else 'NO'}")
# ==========================================
# SECTION 2: DAC INTERACTION
# ==========================================
print("\n\n" + "=" * 80)
print("SECTION 2: DAC INTERACTION SPECIFICATION")
print("=" * 80)
print("\nModel: mad_rate ~ evs_per_1000 + (evs_per_1000 × percent_dac_tracts) + controls + FE")
tsls_dac = results["mad_rate_dac_interaction"]
dac_table = pd.DataFrame([{
"Variable": "evs_per_1000 (main)",
"OLS Coef": f"{ols_dac.params['evs_per_1000']:.6f}",
"OLS P-val": f"{ols_dac.pvalues['evs_per_1000']:.4f}",
"2SLS Coef": f"{tsls_dac['main_coef']:.6f}",
"2SLS P-val": f"{tsls_dac['main_pval']:.4f}",
}, {
"Variable": "ev_x_dac (interaction)",
"OLS Coef": f"{ols_dac.params['ev_x_dac']:.6f}",
"OLS P-val": f"{ols_dac.pvalues['ev_x_dac']:.4f}",
"2SLS Coef": f"{tsls_dac['interaction_coef']:.6f}",
"2SLS P-val": f"{tsls_dac['interaction_pval']:.4f}",
}])
print("\n")
display(dac_table)
sign_change_dac_main = (ols_dac.params['evs_per_1000'] * tsls_dac['main_coef']) < 0
sign_change_dac_int = (ols_dac.params['ev_x_dac'] * tsls_dac['interaction_coef']) < 0
print(f"\nSign Changes: Main = {'YES' if sign_change_dac_main else 'NO'}, Interaction = {'YES' if sign_change_dac_int else 'NO'}")
# ==========================================
# SECTION 3: INCOME INTERACTION
# ==========================================
print("\n\n" + "=" * 80)
print("SECTION 3: INCOME INTERACTION SPECIFICATION")
print("=" * 80)
print("\nModel: mad_rate ~ evs_per_1000 + (evs_per_1000 × median_income_10k) + controls + FE")
tsls_income = results["mad_rate_income_interaction"]
income_table = pd.DataFrame([{
"Variable": "evs_per_1000 (main)",
"OLS Coef": f"{ols_income.params['evs_per_1000']:.6f}",
"OLS P-val": f"{ols_income.pvalues['evs_per_1000']:.4f}",
"2SLS Coef": f"{tsls_income['main_coef']:.6f}",
"2SLS P-val": f"{tsls_income['main_pval']:.4f}",
}, {
"Variable": "ev_x_income (interaction)",
"OLS Coef": f"{ols_income.params['ev_x_income']:.6f}",
"OLS P-val": f"{ols_income.pvalues['ev_x_income']:.4f}",
"2SLS Coef": f"{tsls_income['interaction_coef']:.6f}",
"2SLS P-val": f"{tsls_income['interaction_pval']:.4f}",
}])
print("\n")
display(income_table)
sign_change_income_main = (ols_income.params['evs_per_1000'] * tsls_income['main_coef']) < 0
sign_change_income_int = (ols_income.params['ev_x_income'] * tsls_income['interaction_coef']) < 0
print(f"\nSign Changes: Main = {'YES' if sign_change_income_main else 'NO'}, Interaction = {'YES' if sign_change_income_int else 'NO'}")
print("\n" + "=" * 80)
print("Significance: *** p<0.01, ** p<0.05, * p<0.10")
print("=" * 80)
# Store results for markdown interpretation
ols_vs_2sls_results = {
"baseline": {
"ols_coef": ols_baseline.params['evs_per_1000'],
"ols_pval": ols_baseline.pvalues['evs_per_1000'],
"tsls_coef": tsls_base['coef'],
"tsls_pval": tsls_base['pval'],
},
"dac_main": {
"ols_coef": ols_dac.params['evs_per_1000'],
"ols_pval": ols_dac.pvalues['evs_per_1000'],
"tsls_coef": tsls_dac['main_coef'],
"tsls_pval": tsls_dac['main_pval'],
},
"dac_interaction": {
"ols_coef": ols_dac.params['ev_x_dac'],
"ols_pval": ols_dac.pvalues['ev_x_dac'],
"tsls_coef": tsls_dac['interaction_coef'],
"tsls_pval": tsls_dac['interaction_pval'],
},
"income_main": {
"ols_coef": ols_income.params['evs_per_1000'],
"ols_pval": ols_income.pvalues['evs_per_1000'],
"tsls_coef": tsls_income['main_coef'],
"tsls_pval": tsls_income['main_pval'],
},
"income_interaction": {
"ols_coef": ols_income.params['ev_x_income'],
"ols_pval": ols_income.pvalues['ev_x_income'],
"tsls_coef": tsls_income['interaction_coef'],
"tsls_pval": tsls_income['interaction_pval'],
}
}
ols_vs_2sls_resultsimport pandas as pd
import matplotlib.pyplot as plt
import re
def parse_and_filter_summary(model, spec_name):
"""Extract parameter table, drop FE rows, return as DataFrame."""
table_str = str(model.summary.tables[1])
lines = table_str.split('\n')
rows = []
for line in lines:
# Skip county and month FE rows
if 'C(county_geoid)' in line or 'C(month)' in line:
continue
# Skip separator lines
if set(line.strip()) <= set('=-'):
continue
# Skip empty lines
if not line.strip():
continue
# Skip header line
if 'Parameter' in line and 'Std. Err' in line:
continue
rows.append(line)
# Parse into columns
data = []
for row in rows:
parts = row.split()
if len(parts) >= 6:
# Handle multi-word row names
try:
# Last 6 values are numeric
nums = parts[-6:]
name = ' '.join(parts[:-6])
data.append([name] + nums)
except:
pass
df = pd.DataFrame(data, columns=['Variable', 'Coef', 'Std. Err.', 'T-stat', 'P-value', 'Lower CI', 'Upper CI'])
return df
def save_summary_png(model, spec_name, filename):
# Header info
h = model.summary.tables[0]
header_str = str(h)
# Parameter table
df = parse_and_filter_summary(model, spec_name)
fig = plt.figure(figsize=(10, len(df) * 0.35 + 3))
# Header text
ax_header = fig.add_axes([0, 0.85, 1, 0.15])
ax_header.axis('off')
ax_header.text(0.01, 0.8, f"Appendix: Full Regression Output — {spec_name}",
fontsize=10, fontweight='bold', va='top')
ax_header.text(0.01, 0.4,
f"N={int(model.nobs):,} R²={model.rsquared:.4f} "
f"Clustered SEs by county County & Month FE included (coefficients omitted)",
fontsize=8, color='#555555', va='top')
# Parameter table
ax_table = fig.add_axes([0, 0, 1, 0.85])
ax_table.axis('off')
tbl = ax_table.table(
cellText=df.values,
colLabels=df.columns,
cellLoc='center',
loc='center',
)
tbl.auto_set_font_size(False)
tbl.set_fontsize(8.5)
tbl.auto_set_column_width(col=list(range(len(df.columns))))
for (row, col), cell in tbl.get_celld().items():
cell.set_edgecolor('#dddddd')
cell.set_linewidth(0.5)
if row == 0:
cell.set_facecolor('#f0f0f0')
cell.set_text_props(fontweight='bold')
else:
cell.set_facecolor('white')
if col == 0:
cell.set_text_props(ha='left')
cell.set_height(0.06)
plt.savefig(filename, dpi=150, bbox_inches='tight')
plt.close()
print(f"Saved: {filename}")
# ── Save all three ─────────────────────────────────────────────────────────────
save_summary_png(iv_model_mad, "Baseline", "appendix_baseline.png")
save_summary_png(iv_model_mad_dac, "DAC Interaction", "appendix_dac.png")
save_summary_png(iv_model_mad_income, "Income Interaction","appendix_income.png")import pandas as pd
import matplotlib.pyplot as plt
def parse_and_filter_summary(model, spec_name):
"""Extract parameter table, drop FE rows, return as DataFrame."""
table_str = str(model.summary.tables[1])
lines = table_str.split('\n')
rows = []
for line in lines:
if 'C(county_geoid)' in line or 'C(month)' in line:
continue
if set(line.strip()) <= set('=-'):
continue
if not line.strip():
continue
if 'Parameter' in line and 'Std. Err' in line:
continue
rows.append(line)
data = []
for row in rows:
parts = row.split()
if len(parts) >= 6:
try:
nums = parts[-6:]
name = ' '.join(parts[:-6])
data.append([name] + nums)
except:
pass
df = pd.DataFrame(data, columns=['Variable', 'Coef', 'Std. Err.', 'T-stat', 'P-value', 'Lower CI', 'Upper CI'])
return df
def save_summary_png(model, spec_name, filename):
df = parse_and_filter_summary(model, spec_name)
fig = plt.figure(figsize=(10, len(df) * 0.35 + 3))
ax_header = fig.add_axes([0, 0.85, 1, 0.15])
ax_header.axis('off')
ax_header.text(0.01, 0.8, f"Appendix: Full Regression Output — {spec_name}",
fontsize=10, fontweight='bold', va='top')
ax_header.text(0.01, 0.4,
f"N={int(model.nobs):,} R²={model.rsquared:.4f} "
f"Clustered SEs by county County & Month FE included (coefficients omitted)",
fontsize=8, color='#555555', va='top')
ax_table = fig.add_axes([0, 0, 1, 0.85])
ax_table.axis('off')
tbl = ax_table.table(
cellText=df.values,
colLabels=df.columns,
cellLoc='center',
loc='center',
)
tbl.auto_set_font_size(False)
tbl.set_fontsize(8.5)
tbl.auto_set_column_width(col=list(range(len(df.columns))))
for (row, col), cell in tbl.get_celld().items():
cell.set_edgecolor('#dddddd')
cell.set_linewidth(0.5)
if row == 0:
cell.set_facecolor('#f0f0f0')
cell.set_text_props(fontweight='bold')
else:
cell.set_facecolor('white')
if col == 0:
cell.set_text_props(ha='left')
cell.set_height(0.06)
plt.savefig(filename, dpi=150, bbox_inches='tight')
plt.close()
print(f"Saved: {filename}")
# ── Save all five specs ────────────────────────────────────────────────────────
save_summary_png(iv_model_mad, "Baseline (mad_rate)", "appendix_baseline_mad.png")
save_summary_png(iv_model_mad_dac, "DAC Interaction (mad_rate)", "appendix_dac_mad.png")
save_summary_png(iv_model_mad_income, "Income Interaction (mad_rate)", "appendix_income_mad.png")
save_summary_png(iv_model, "Baseline (avg_lmp_weighted)", "appendix_baseline_lmp.png")
save_summary_png(iv_model_median, "Baseline (median_of_medians)", "appendix_baseline_median.png")import matplotlib.pyplot as plt
from collections import OrderedDict
# ── Extract F-stats ───────────────────────────────────────────────────────────
wald_main = float(iv_model_mad.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
wald_dac = float(iv_model_mad_dac.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
wald_income = float(iv_model_mad_income.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
classical_main = float(model_main.fvalue)
classical_dac = float(model_dac.fvalue)
classical_income = float(model_income.fvalue)
import matplotlib.pyplot as plt
def get_stats(model):
is_iv = hasattr(model, 'std_errors')
if is_iv:
params = model.params
bse = model.std_errors
pvals = model.pvalues
nobs = int(model.nobs)
rsq = float(model.rsquared)
else:
params = model.params
bse = model.bse
pvals = model.pvalues
nobs = int(model.nobs)
rsq = float(model.rsquared)
mask = ~(
params.index.str.startswith('C(county_geoid)') |
params.index.str.startswith('C(month)')
)
return params[mask], bse[mask], pvals[mask], nobs, rsq
def sig_stars(p):
if p < 0.01: return '***'
if p < 0.05: return '**'
return ''
def make_reg_table(
models,
col_labels,
var_labels,
county_fe,
month_fe,
filename,
dep_var_label='Dependent Variable: Price Volatility (mad_rate)',
note='Clustered standard errors in parentheses (county level). *** p<0.01, ** p<0.05',
fstats=None,
row_h=0.30,
top_h=0.90,
bot_h=0.55,
var_col_w=2.8,
data_col_w=1.55,
):
all_stats = [get_stats(m) for m in models]
seen, var_order = set(), []
for params, *_ in all_stats:
for v in params.index:
if v not in seen:
var_order.append(v)
seen.add(v)
data_rows = []
for var in var_order:
display = var_labels.get(var, var)
coef_cells, se_cells = [], []
for params, bse, pvals, *_ in all_stats:
if var in params.index:
c, s, p = params[var], bse[var], pvals[var]
coef_cells.append(f"{c:.4f}{sig_stars(p)}")
se_cells.append(f"({s:.4f})")
else:
coef_cells.append('')
se_cells.append('')
data_rows.append((True, display, coef_cells))
data_rows.append((False, display, se_cells))
footer = []
if fstats:
wald_row, classical_row = [], []
for col in col_labels:
if col in fstats:
wald, classical = fstats[col]
wald_row.append(f"{wald:.2f}")
classical_row.append(f"{classical:.2f}")
else:
wald_row.append('—')
classical_row.append('—')
footer.append(('Cluster-Robust Wald F', wald_row))
footer.append(('Classical F-stat', classical_row))
footer += [
('County FEs', [('Yes' if county_fe[i] else 'No') for i in range(len(models))]),
('Month FEs', [('Yes' if month_fe[i] else 'No') for i in range(len(models))]),
('Observations', [f"{s[3]:,}" for s in all_stats]),
('R-squared', [f"{s[4]:.4f}" for s in all_stats]),
]
n_cols = len(models)
n_data = len(data_rows)
n_foot = len(footer)
n_total = n_data + n_foot
fig_w = var_col_w + n_cols * data_col_w
fig_h = top_h + n_total * row_h + bot_h
fig, ax = plt.subplots(figsize=(fig_w, fig_h))
ax.set_xlim(0, fig_w)
ax.set_ylim(0, fig_h)
ax.axis('off')
BLACK = '#000000'
GRAY = '#444444'
FONT = 'DejaVu Sans'
col_x = [var_col_w + (i + 0.5) * data_col_w for i in range(n_cols)]
def row_y(i):
return fig_h - top_h - (i + 0.5) * row_h
# dep var label
ax.text(fig_w / 2, fig_h - 0.12, dep_var_label,
ha='center', va='center', fontsize=8, fontfamily=FONT, color=GRAY)
# column headers
for cx, label in zip(col_x, col_labels):
ax.text(cx, fig_h - 0.52, label,
ha='center', va='center', fontsize=8.5,
fontfamily=FONT, color=BLACK)
# thick rule above column headers
ax.plot([0.05, fig_w - 0.05], [fig_h - 0.28, fig_h - 0.28],
color=BLACK, lw=1.4)
# thick rule below column headers
ax.plot([0.05, fig_w - 0.05], [fig_h - top_h, fig_h - top_h],
color=BLACK, lw=1.4)
# data rows
for ri, (is_coef, var_display, cells) in enumerate(data_rows):
y = row_y(ri)
if is_coef:
ax.text(0.10, y, var_display,
ha='left', va='center', fontsize=8,
fontfamily=FONT, color=BLACK, fontweight='bold')
for cx, val in zip(col_x, cells):
ax.text(cx, y, val,
ha='center', va='center', fontsize=8,
fontfamily=FONT, color=BLACK)
# footer rows — no rule, flows directly from data rows
for fi, (label, vals) in enumerate(footer):
y = row_y(n_data + fi)
ax.text(0.10, y, label,
ha='left', va='center', fontsize=8,
fontfamily=FONT, color=BLACK)
for cx, val in zip(col_x, vals):
ax.text(cx, y, val,
ha='center', va='center', fontsize=8,
fontfamily=FONT, color=BLACK)
# thick bottom rule
bot_rule_y = fig_h - top_h - n_total * row_h
ax.plot([0.05, fig_w - 0.05], [bot_rule_y, bot_rule_y],
color=BLACK, lw=1.4)
# notes
ax.text(0.10, bot_rule_y - 0.10, note,
ha='left', va='top', fontsize=7,
fontfamily=FONT, color=GRAY, style='italic')
plt.savefig(filename, dpi=200, bbox_inches='tight', facecolor='white')
plt.close()
print(f"Saved: {filename}")
VAR_LABELS = {
'evs_per_1000': 'EVs per 1,000 Residents',
'ev_x_dac': 'EVs × DAC Share',
'ev_x_income': 'EVs × Median Income ($10k)',
'percent_dac_tracts': 'DAC Share',
'median_household_income_10k': 'Median Income ($10k)',
'tmax_c': 'Max Temperature (°C)',
'cumulative_data_center_power_mw': 'Cumulative Data Center Power (MW)',
'total_generator_active_capacity_mw': 'Total Generator Capacity (MW)',
'solar_mw_per_1000_lag1': 'Solar Capacity per 1,000 (lagged)',
'Intercept': 'Constant',
}
print("Functions defined.")
wald_main = float(iv_model_mad.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
wald_dac = float(iv_model_mad_dac.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
wald_income = float(iv_model_mad_income.first_stage.diagnostics.loc['evs_per_1000', 'f.stat'])
classical_main = float(model_main.fvalue)
classical_dac = float(model_dac.fvalue)
classical_income = float(model_income.fvalue)
make_reg_table(
models = [model_ols_nofe, ols_baseline, iv_model_mad],
col_labels = ['(1) Naive OLS', '(2) OLS', '(3) 2SLS'],
var_labels = VAR_LABELS,
county_fe = [False, True, True],
month_fe = [False, True, True],
dep_var_label = 'Dependent Variable: Price Volatility (mad_rate)',
filename = 'table_A_baseline.png',
fstats = {'(3) 2SLS': (wald_main, classical_main)},
)
make_reg_table(
models = [ols_dac, iv_model_mad_dac],
col_labels = ['(1) OLS', '(2) 2SLS'],
var_labels = VAR_LABELS,
county_fe = [True, True],
month_fe = [True, True],
dep_var_label = 'Dependent Variable: Price Volatility (mad_rate)',
filename = 'table_B_dac.png',
fstats = {'(2) 2SLS': (wald_dac, classical_dac)},
)
make_reg_table(
models = [ols_income, iv_model_mad_income],
col_labels = ['(1) OLS', '(2) 2SLS'],
var_labels = VAR_LABELS,
county_fe = [True, True],
month_fe = [True, True],
dep_var_label = 'Dependent Variable: Price Volatility (mad_rate)',
filename = 'table_C_income.png',
fstats = {'(2) 2SLS': (wald_income, classical_income)},
)import pandas as pd
import statsmodels.formula.api as smf
from linearmodels.iv import IV2SLS
# ==========================================
# MANUAL 2SLS: Showing the Two Stages Explicitly
# ==========================================
print("=" * 80)
print("MANUAL 2SLS VERIFICATION")
print("=" * 80)
print("\nThis cell manually runs the two stages to verify the IV2SLS package")
print("is using first-stage predictions correctly in the second stage.")
print("\n" + "=" * 80)
df = df_final.copy()
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)
num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")
controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw"]
control_str = " + ".join(controls)
# Prepare data
needed_cols = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]
data = df[needed_cols].dropna().copy()
# ==========================================
# STAGE 1: Predict evs_per_1000 using IV
# ==========================================
print("\nSTAGE 1: First-Stage Regression")
print("-" * 80)
print("Regressing: evs_per_1000 ~ iv_tesla_store_exposure_log + controls + FE")
first_stage_formula = (
f"evs_per_1000 ~ iv_tesla_store_exposure_log + {control_str} + C(county_geoid) + C(month)"
)
first_stage_model = smf.ols(first_stage_formula, data=data).fit()
# Get predicted values from first stage
data["evs_per_1000_predicted"] = first_stage_model.fittedvalues
print(f"\nFirst-stage predictions created:")
print(f" Mean predicted EV adoption: {data['evs_per_1000_predicted'].mean():.4f}")
print(f" Std dev: {data['evs_per_1000_predicted'].std():.4f}")
print(f" Min: {data['evs_per_1000_predicted'].min():.4f}")
print(f" Max: {data['evs_per_1000_predicted'].max():.4f}")
# ==========================================
# STAGE 2: Use predictions in second stage
# ==========================================
print("\n\nSTAGE 2: Second-Stage Regression")
print("-" * 80)
print("Regressing: mad_rate ~ evs_per_1000_predicted + controls + FE")
print("\nNOTE: Standard errors from naive OLS would be WRONG here.")
print(" We're just showing the coefficient for comparison.")
second_stage_formula = (
f"mad_rate ~ evs_per_1000_predicted + {control_str} + C(county_geoid) + C(month)"
)
second_stage_model = smf.ols(second_stage_formula, data=data).fit()
manual_2sls_coef = second_stage_model.params["evs_per_1000_predicted"]
manual_2sls_se_naive = second_stage_model.bse["evs_per_1000_predicted"] # WRONG SE!
print(f"\nManual 2SLS coefficient: {manual_2sls_coef:.6f}")
print(f"Naive SE (INCORRECT): {manual_2sls_se_naive:.6f}")
# ==========================================
# COMPARE TO IV2SLS PACKAGE
# ==========================================
print("\n\n" + "=" * 80)
print("COMPARISON: Manual vs IV2SLS Package")
print("=" * 80)
# Run IV2SLS package version
iv_formula = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""
iv_model = IV2SLS.from_formula(iv_formula, data=data).fit(
cov_type="clustered",
clusters=data["county_geoid"],
)
package_coef = iv_model.params["evs_per_1000"]
package_se = iv_model.std_errors["evs_per_1000"]
package_pval = iv_model.pvalues["evs_per_1000"]
print("\nCoefficient Comparison:")
print(f" Manual 2SLS: {manual_2sls_coef:.6f}")
print(f" IV2SLS Package: {package_coef:.6f}")
print(f" Difference: {abs(manual_2sls_coef - package_coef):.10f}")
if abs(manual_2sls_coef - package_coef) < 0.000001:
print("\n ✓ COEFFICIENTS MATCH (within rounding error)")
print(" ✓ The package IS using first-stage predictions correctly")
else:
print("\n ✗ Coefficients differ - something is wrong")
print("\nStandard Error Comparison:")
print(f" Manual 2SLS (naive OLS SE): {manual_2sls_se_naive:.6f} ← WRONG!")
print(f" IV2SLS Package (correct SE): {package_se:.6f} ← CORRECT")
print(f"\n The package computes proper 2SLS standard errors that account for")
print(f" the two-stage estimation process. Naive OLS SE underestimates uncertainty.")
print("\nFinal Result:")
print(f" Coefficient: {package_coef:.6f}")
print(f" Std. Error: {package_se:.6f}")
print(f" P-value: {package_pval:.4f}")
print("\n" + "=" * 80)
print("CONCLUSION: We are doing 2SLS correctly!")
print("=" * 80)
print("\nThe IV2SLS package:")
print(" ✓ Runs first stage internally")
print(" ✓ Uses predicted values in second stage")
print(" ✓ Computes correct standard errors (accounting for two-stage uncertainty)")
print(" ✓ Handles fixed effects and clustering properly")
print("\nYou can trust the previous 2SLS results - they're methodologically sound.")
print("=" * 80)
# -------------------------
# Falsification IVs: Lag and Lead variants recomputed at source
# -------------------------
variants = {
"lag2": lambda m: m - pd.DateOffset(months=2),
"lag3": lambda m: m - pd.DateOffset(months=3),
"lag6": lambda m: m - pd.DateOffset(months=6),
"lead3": lambda m: m + pd.DateOffset(months=3),
"lead6": lambda m: m + pd.DateOffset(months=6),
"lead12": lambda m: m + pd.DateOffset(months=12),
}
iv_falsification_tables = {}
for label, shift_fn in variants.items():
rows = []
is_lead = label.startswith("lead")
for m in months:
ref_month = shift_fn(m)
if is_lead:
open_stores_m = tesla_stores_df.loc[
(tesla_stores_df["open_month"] > m) & (tesla_stores_df["open_month"] <= ref_month)
]
else:
open_stores_m = tesla_stores_df.loc[tesla_stores_df["open_month"] <= ref_month]
if open_stores_m.empty:
iv_m = counties[["county_geoid"]].copy()
iv_m["idw_exposure"] = 0.0
iv_m["month"] = m
rows.append(iv_m)
continue
tmp_m = counties.assign(_k=1).merge(open_stores_m.assign(_k=1), on="_k").drop(columns="_k")
tmp_m["dist_km"] = haversine_km(
tmp_m["county_lat"], tmp_m["county_lon"],
tmp_m["latitude"], tmp_m["longitude"]
)
tmp_m["dist_km_capped"] = np.maximum(tmp_m["dist_km"], eps_km)
iv_m = tmp_m.groupby("county_geoid", as_index=False).agg(
idw_exposure=("dist_km_capped", lambda s: float(np.sum(1.0 / s))),
)
iv_m["month"] = m
rows.append(iv_m)
if rows:
tbl = pd.concat(rows, ignore_index=True)
col_name = f"iv_tesla_store_exposure_log_{label}"
tbl[col_name] = np.log1p(tbl["idw_exposure"])
iv_falsification_tables[label] = tbl[["county_geoid", "month", col_name]]
# Merge all falsification IVs onto df_iv
df_iv_falsification = df_iv.copy()
for label, tbl in iv_falsification_tables.items():
df_iv_falsification = df_iv_falsification.merge(
tbl, on=["county_geoid", "month"], how="left", validate="m:1"
)
falsification_cols = [c for c in df_iv_falsification.columns if "lag" in c or "lead" in c]
print(f"Falsification IVs merged onto df_iv_falsification ({len(df_iv_falsification):,} rows)")
for c in falsification_cols:
nn = df_iv_falsification[c].notna().sum()
print(f" {c:<45} {nn:>6} non-null")
df_iv_falsification
import pandas as pd
# -------------------------
# Robustness Dataset: lag/lead IV variants recomputed at source
# -------------------------
robustness_iv_cols = [
"iv_tesla_store_exposure_log_lag2",
"iv_tesla_store_exposure_log_lag3",
"iv_tesla_store_exposure_log_lag6",
"iv_tesla_store_exposure_log_lead3",
"iv_tesla_store_exposure_log_lead6",
"iv_tesla_store_exposure_log_lead12",
]
# Start from df_final columns + lag/lead IVs from df_iv_falsification
base_cols = list(df_final.columns)
extra_cols = [c for c in robustness_iv_cols if c in df_iv_falsification.columns]
# Merge the falsification IVs onto df_final via county_geoid + month
falsification_merge = df_iv_falsification[["county_geoid", "month"] + extra_cols].copy()
falsification_merge["county_geoid"] = falsification_merge["county_geoid"].astype(str)
df_final_falsification = df_final.merge(
falsification_merge,
on=["county_geoid", "month"],
how="left",
validate="m:1",
)
# Create interaction instruments for each lag/lead variant
for iv_col in extra_cols:
suffix = iv_col.replace("iv_tesla_store_exposure_log", "") # e.g. "_lag6"
df_final_falsification[f"iv_x_dac{suffix}"] = (
df_final_falsification[iv_col] * df_final_falsification["percent_dac_tracts"]
)
df_final_falsification[f"iv_x_income{suffix}"] = (
df_final_falsification[iv_col] * df_final_falsification["median_household_income_10k"]
)
print(f"df_final_falsification: {len(df_final_falsification):,} rows × {len(df_final_falsification.columns)} cols")
print(f"\nLag/lead IV columns carried through:")
for c in extra_cols:
non_null = df_final_falsification[c].notna().sum()
print(f" {c:<45} {non_null:>6} non-null")
print(f"\nNew interaction instruments:")
for c in df_final_falsification.columns:
if c.startswith("iv_x_") and c not in ["iv_x_dac", "iv_x_income"]:
non_null = df_final_falsification[c].notna().sum()
print(f" {c:<45} {non_null:>6} non-null")
df_final_falsification
ts = tesla_loc_df.copy()
ts["open_month"] = pd.to_datetime(ts["open_month"], errors="coerce").dt.to_period("M").dt.to_timestamp()
ts = ts.dropna(subset=["open_month"])
openings = ts.groupby("open_month").size().rename("new_stores")
cumulative = openings.cumsum().rename("cumulative_stores")
summary = pd.concat([openings, cumulative], axis=1).reset_index()
summary.rename(columns={"open_month": "month"}, inplace=True)
# Panel time range for reference
panel_min = pd.to_datetime(panel_df["month"], errors="coerce").min()
panel_max = pd.to_datetime(panel_df["month"], errors="coerce").max()
print(f"Panel range: {panel_min.strftime('%Y-%m')} to {panel_max.strftime('%Y-%m')}")
print(f"Total store openings: {openings.sum()}")
print(f"Openings within panel range: {openings.loc[(openings.index >= panel_min) & (openings.index <= panel_max)].sum()}")
print()
print(summary.to_string(index=False))from linearmodels.iv import IV2SLS
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
# -------------------------------------------------------------------
# LAG DECAY: Run 2SLS for lag-0 through lag-6 in 1-month increments
# Plot coefficient path to see exactly where the signal breaks down
# -------------------------------------------------------------------
controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
control_str = " + ".join(controls)
# Prep base panel (same as main spec)
df_base = df_final.copy()
df_base["month"] = pd.to_datetime(df_base["month"], errors="coerce")
df_base["month_str"] = df_base["month"].dt.to_period("M").astype(str)
df_base["county_geoid"] = df_base["county_geoid"].astype(str)
for c in ["mad_rate", "evs_per_1000"] + controls:
df_base[c] = pd.to_numeric(df_base[c], errors="coerce")
# Precompute inverse distances (reuse county-level setup)
county_arr = counties[["county_geoid", "county_lat", "county_lon"]].reset_index(drop=True)
store_arr = tesla_stores_df[["latitude", "longitude", "open_month"]].reset_index(drop=True)
clat = county_arr["county_lat"].values[:, None]
clon = county_arr["county_lon"].values[:, None]
slat = store_arr["latitude"].values[None, :]
slon = store_arr["longitude"].values[None, :]
R = 6371.0
la1, lo1, la2, lo2 = map(np.radians, [clat, clon, slat, slon])
a = np.sin((la2 - la1) / 2)**2 + np.cos(la1) * np.cos(la2) * np.sin((lo2 - lo1) / 2)**2
inv_dist = 1.0 / np.maximum(2 * R * np.arcsin(np.sqrt(a)), 1.0)
inv_dist_county = inv_dist.copy() # ← save county-level before anything overwrites inv_dist
county_ids = county_arr["county_geoid"].values
open_dates = pd.to_datetime(store_arr["open_month"]).values
panel_months = pd.DatetimeIndex(sorted(df_base["month"].dropna().unique()))
# Run lag-0 through lag-6
lag_results = []
for lag_k in range(7):
# Compute IV with lag
shifted_months = panel_months - pd.DateOffset(months=lag_k)
is_open = (open_dates[:, None] <= shifted_months.values[None, :]).astype(float)
exposure = inv_dist_county @ is_open # ← use county-level matrix
iv_log = np.log1p(exposure)
# Build IV table
cids = np.repeat(county_ids, len(panel_months))
mths = np.tile(panel_months.values, len(county_ids))
iv_df = pd.DataFrame({
"county_geoid": cids.astype(str),
"month": mths,
"iv_lag": iv_log.ravel(),
})
# Merge and run 2SLS
d = df_base[["county_geoid", "month", "month_str", "mad_rate", "evs_per_1000"] + controls].dropna().copy()
d = d.merge(iv_df, on=["county_geoid", "month"], how="inner").dropna()
for c in d.columns:
if pd.api.types.is_numeric_dtype(d[c]):
d[c] = d[c].astype(float)
formula = f"mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month_str) [evs_per_1000 ~ iv_lag]"
model = IV2SLS.from_formula(formula, data=d).fit(
cov_type="clustered", clusters=d["county_geoid"]
)
coef = model.params["evs_per_1000"]
se = model.std_errors["evs_per_1000"]
pval = model.pvalues["evs_per_1000"]
fstat = float(model.first_stage.diagnostics.loc["evs_per_1000", "f.stat"])
lag_results.append({"Lag": lag_k, "Coef": coef, "SE": se, "P-value": pval, "F-stat": fstat, "N": int(model.nobs)})
print(f"Lag-{lag_k}: coef={coef:.6f} SE={se:.6f} p={pval:.4f} F={fstat:.1f}")
results_df = pd.DataFrame(lag_results)
fig, ax = plt.subplots(figsize=(8, 5))
ax.errorbar(results_df["Lag"], results_df["Coef"],
yerr=1.96 * results_df["SE"], fmt="o-", color="steelblue",
capsize=4, linewidth=2, markersize=6)
ax.axhline(0, color="gray", linewidth=0.8, linestyle="-", alpha=0.4)
ax.set_xlabel("Lag (months)", fontsize=11)
ax.set_ylabel("2SLS Coefficient (evs_per_1000)", fontsize=11)
ax.set_title("Lag Decay: 2SLS Coefficient by IV Timing",
fontsize=12, fontweight="bold")
ax.set_xticks(range(7))
ax.set_xticklabels([f"Lag-{k}" for k in range(7)])
ax.spines["top"].set_visible(False)
ax.spines["right"].set_visible(False)
ax.spines["left"].set_color("#666666")
ax.spines["bottom"].set_color("#666666")
ax.tick_params(colors="#666666")
plt.tight_layout()
plt.savefig("lag_decay.png", dpi=150, bbox_inches="tight")
plt.show()
display(results_df)from linearmodels.iv import IV2SLS
import numpy as np
import pandas as pd
df_rob = df_final_falsification.copy()
df_rob["month"] = pd.to_datetime(df_rob["month"], errors="coerce").dt.to_period("M").astype(str)
df_rob["county_geoid"] = df_rob["county_geoid"].astype(str)
num_cols = [c for c in df_rob.columns if c not in ["county_geoid", "county_name", "month"]]
df_rob[num_cols] = df_rob[num_cols].apply(pd.to_numeric, errors="coerce").astype(float)
controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
control_str = " + ".join(controls)
# --- Run each spec ---
specs = {
"Main (contemporaneous IV)": "iv_tesla_store_exposure_log",
"Lag-2 IV (exposure 2mo prior)": "iv_tesla_store_exposure_log_lag2",
"Lag-3 IV (exposure 3mo prior)": "iv_tesla_store_exposure_log_lag3",
"Lag-6 IV (exposure 6mo prior)": "iv_tesla_store_exposure_log_lag6",
}
rows = []
for label, iv_col in specs.items():
cols = ["mad_rate", "evs_per_1000", iv_col] + controls + ["county_geoid", "month"]
d = df_rob[cols].dropna().copy()
for c in d.columns:
if pd.api.types.is_numeric_dtype(d[c]):
d[c] = d[c].astype(float)
formula = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ {iv_col}]
"""
model = IV2SLS.from_formula(formula, data=d).fit(
cov_type="clustered", clusters=d["county_geoid"]
)
rows.append({
"Specification": label,
"2SLS Coef": round(model.params["evs_per_1000"], 4),
"SE (clustered)": round(model.std_errors["evs_per_1000"], 4),
"P-value": round(model.pvalues["evs_per_1000"], 4),
"First-stage F": round(float(model.first_stage.diagnostics.loc["evs_per_1000", "f.stat"]), 2),
"N": int(model.nobs),
})
summary = pd.DataFrame(rows)
print("=" * 80)
print("LAG ROBUSTNESS: mad_rate ~ evs_per_1000 (instrumented)")
print("=" * 80)
print(f"Controls: {', '.join(controls)}")
print(f"FEs: county + month | SEs: clustered by county")
print()
display(summary)import statsmodels.formula.api as smf
# -------------------------------------------------------------------
# LEAD FALSIFICATION: Reduced-form OLS (NOT 2SLS)
#
# Purpose: Test whether proximity to stores that don't exist yet
# predicts current electricity market outcomes.
#
# Construction: Each lead variable measures IDW exposure ONLY to
# stores opening strictly after month m but within k months.
# These stores are not yet open — they cannot causally affect
# current EV adoption or electricity prices.
#
# Why OLS, not 2SLS: We are not instrumenting anything. We are
# directly testing whether future store openings (which shouldn't
# matter yet) correlate with current outcomes. This is a pure
# reduced-form check of the exclusion restriction.
#
# Interpretation:
# - Insignificant lead coefficients → future stores don't predict
# current outcomes → supports instrument validity
# - Significant lead coefficients → counties that will later get
# stores were already trending differently → confounding concern
#
# Caveat: With only 29 store openings in the panel, these tests
# are underpowered. A null result is necessary but not sufficient
# evidence of instrument validity.
# -------------------------------------------------------------------
lead_ivs = [
"iv_tesla_store_exposure_log_lead3", # stores opening in next 3 months
"iv_tesla_store_exposure_log_lead6", # stores opening in next 6 months
"iv_tesla_store_exposure_log_lead12", # stores opening in next 12 months
]
print("=" * 80)
print("LEAD IVs: REDUCED-FORM FALSIFICATION (should be insignificant)")
print("=" * 80)
falsification_results = []
outcomes = ["mad_rate"]
print(f"\ndf_rob shape: {df_rob.shape}")
print(f"Controls: {controls}")
print(f"control_str: {control_str}")
for lc in lead_ivs:
present = lc in df_rob.columns
nn = df_rob[lc].notna().sum() if present else 0
print(f" {lc}: present={present}, non-null={nn}")
for outcome in outcomes:
for lead_col in lead_ivs:
cols = [outcome, lead_col] + controls + ["county_geoid", "month"]
missing_cols = [c for c in cols if c not in df_rob.columns]
if missing_cols:
print(f"\n⚠ {lead_col}: missing columns in df_rob: {missing_cols}")
continue
d = df_rob[cols].dropna().copy()
print(f"\n{lead_col}: {len(d)} rows after dropna (from {len(df_rob)})")
if len(d) == 0:
print(f" ⚠ No rows remaining — skipping")
continue
for c in d.columns:
if pd.api.types.is_numeric_dtype(d[c]):
d[c] = d[c].astype(float)
formula = f"{outcome} ~ {lead_col} + {control_str} + C(county_geoid) + C(month)"
model = smf.ols(formula, data=d).fit(
cov_type="cluster", cov_kwds={"groups": d["county_geoid"]}
)
coef = model.params[lead_col]
se = model.bse[lead_col]
pval = model.pvalues[lead_col]
sig = "*" if pval < 0.1 else ""
falsification_results.append({
"Outcome": outcome,
"Lead IV": lead_col.replace("iv_tesla_store_exposure_log_", ""),
"Coef": round(coef, 4),
"SE": round(se, 4),
"P-value": round(pval, 4),
"Sig": sig,
"N": int(model.nobs),
})
falsification_df = pd.DataFrame(falsification_results)
print("\n")
print(falsification_df.to_string(index=False))
print("\n* = p < 0.10. Significant leads suggest confounding / pre-trends.")import numpy as np
import pandas as pd
np.random.seed(42)
# Original store opening dates
store_arr = tesla_stores_df[["latitude", "longitude", "open_month"]].copy()
store_arr["open_month"] = pd.to_datetime(store_arr["open_month"], errors="coerce")
store_arr = store_arr.dropna(subset=["open_month"]).reset_index(drop=True)
real_open = store_arr["open_month"].values
print(f"Stores: {len(real_open)}")
print(f"Date range: {pd.Timestamp(real_open.min()).strftime('%Y-%m')} to {pd.Timestamp(real_open.max()).strftime('%Y-%m')}")
# One example shuffle — dates reassigned across locations
shuffled_open = np.random.permutation(real_open)
# Show a few rows to confirm
check = pd.DataFrame({
"lat": store_arr["latitude"].values[:10],
"lon": store_arr["longitude"].values[:10],
"real_open": pd.to_datetime(real_open[:10]),
"shuffled_open": pd.to_datetime(shuffled_open[:10]),
})
print("\nSample (first 10 stores):")
print(check.to_string(index=False))# Precompute pairwise inverse distances (fixed — locations don't change)
county_arr = counties[["county_geoid", "county_lat", "county_lon"]].reset_index(drop=True)
clat = county_arr["county_lat"].values[:, None]
clon = county_arr["county_lon"].values[:, None]
slat = store_arr["latitude"].values[None, :]
slon = store_arr["longitude"].values[None, :]
R = 6371.0
la1, lo1, la2, lo2 = map(np.radians, [clat, clon, slat, slon])
a = np.sin((la2 - la1) / 2)**2 + np.cos(la1) * np.cos(la2) * np.sin((lo2 - lo1) / 2)**2
inv_dist = 1.0 / np.maximum(2 * R * np.arcsin(np.sqrt(a)), 1.0)
inv_dist_county = inv_dist.copy()
print(f"inv_dist shape: {inv_dist.shape} (counties × stores)")
# Compute IV for one set of opening dates
panel_months = pd.DatetimeIndex(sorted(df_final["month"].dropna().unique()))
def compute_iv_table(open_dates):
"""opening dates → DataFrame of (county_geoid, month, iv_perm_log)"""
is_open = (open_dates[:, None] <= panel_months.values[None, :]).astype(float)
exposure = inv_dist_county @ is_open # was inv_dist
iv_log = np.log1p(exposure)
# Reshape to long DataFrame
cids = np.repeat(county_arr["county_geoid"].values, len(panel_months))
mths = np.tile(panel_months.values, len(county_arr))
vals = iv_log.ravel()
return pd.DataFrame({"county_geoid": cids, "month": mths, "iv_perm_log": vals})
# Test with real dates
iv_real = compute_iv_table(real_open)
print(f"\nIV table: {len(iv_real)} rows")
print(iv_real.head())import statsmodels.formula.api as smf
# -------------------------------------------------------------------
# REDUCED-FORM REGRESSION: REAL STORE OPENING DATES
#
# This is the baseline for the permutation test.
# We regress mad_rate directly on the IV (Tesla store exposure)
# with controls and two-way FEs. The coefficient tells us whether
# the IV predicts price volatility in the actual data.
#
# The permutation test will then ask: does this coefficient remain
# significant when we randomly reassign opening dates to stores?
# -------------------------------------------------------------------
df_reg = df_final[["county_geoid", "month", "mad_rate",
"tmax_c", "cumulative_data_center_power_mw",
"total_generator_active_capacity_mw",
"solar_mw_per_1000_lag1"]].copy()
df_reg["county_geoid"] = df_reg["county_geoid"].astype(str)
df_reg["month"] = pd.to_datetime(df_reg["month"], errors="coerce")
iv_real["county_geoid"] = iv_real["county_geoid"].astype(str)
df_reg = df_reg.merge(iv_real, on=["county_geoid", "month"], how="inner")
df_reg = df_reg.dropna().reset_index(drop=True)
df_reg["month_str"] = df_reg["month"].dt.to_period("M").astype(str)
controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
formula = ("mad_rate ~ iv_perm_log + tmax_c + cumulative_data_center_power_mw"
" + total_generator_active_capacity_mw + solar_mw_per_1000_lag1 + C(county_geoid) + C(month_str)")
model = smf.ols(formula, data=df_reg).fit(
cov_type="cluster", cov_kwds={"groups": df_reg["county_geoid"]}
)
real_coef = model.params["iv_perm_log"]
real_se = model.bse["iv_perm_log"]
real_pval = model.pvalues["iv_perm_log"]
print("=" * 60)
print("BASELINE: REDUCED-FORM WITH REAL OPENING DATES")
print("=" * 60)
print(f"Outcome: mad_rate (price volatility)")
print(f"Key regressor: iv_perm_log (log IDW Tesla store exposure)")
print(f"Controls: {', '.join(controls)}")
print(f"Fixed effects: county + month")
print(f"Standard errors: clustered by county")
print(f"N: {int(model.nobs)}")
print(f"")
print(f"Coefficient: {real_coef:.6f}")
print(f"SE (clustered): {real_se:.6f}")
print(f"P-value: {real_pval:.4f}")
print(f"Significant: {'Yes (p < 0.05)' if real_pval < 0.05 else 'No'}")
print(f"")
print(f"Interpretation: A one-unit increase in log Tesla store")
print(f" exposure is associated with a {real_coef:.4f}")
print(f" increase in mad_rate, after absorbing")
print(f" county and month fixed effects.")
print("=" * 60)N_PERMS = 500
perm_coefs = []
print(f"Running {N_PERMS} permutations...")
for i in range(N_PERMS):
shuffled = np.random.permutation(real_open)
iv_perm = compute_iv_table(shuffled)
iv_perm["county_geoid"] = iv_perm["county_geoid"].astype(str)
df_p = df_final[["county_geoid", "month", "mad_rate",
"tmax_c", "cumulative_data_center_power_mw",
"total_generator_active_capacity_mw",
"solar_mw_per_1000_lag1"]].copy()
df_p["county_geoid"] = df_p["county_geoid"].astype(str)
df_p["month"] = pd.to_datetime(df_p["month"], errors="coerce")
df_p = df_p.merge(iv_perm, on=["county_geoid", "month"], how="inner").dropna()
df_p["month_str"] = df_p["month"].dt.to_period("M").astype(str)
try:
m = smf.ols(formula, data=df_p).fit(
cov_type="cluster", cov_kwds={"groups": df_p["county_geoid"]}
)
perm_coefs.append(m.params["iv_perm_log"])
except Exception as e:
if i == 0: # only print error on first failure to avoid spam
print(f"Model error (iteration {i}): {e}")
if (i + 1) % 100 == 0:
print(f" {i + 1}/{N_PERMS} done ({len(perm_coefs)} valid)")
perm_coefs = np.array(perm_coefs)
p_two_sided = np.mean(np.abs(perm_coefs) >= np.abs(real_coef))
print(f"\nReal coef: {real_coef:.6f}")
print(f"Valid permutations: {len(perm_coefs)}/{N_PERMS}")
print(f"Permutation p-value (two-sided): {p_two_sided:.4f}")
fig, ax = plt.subplots(figsize=(8, 4))
ax.hist(perm_coefs, bins=40, color="steelblue", alpha=0.7, edgecolor="white")
ax.axvline(real_coef, color="red", linewidth=2, label=f"Real = {real_coef:.3f}")
ax.set_xlabel("Reduced-form coefficient (mad_rate ~ IV)")
ax.set_ylabel("Count")
ax.set_title(f"Permutation Test (n={len(perm_coefs)}, p={p_two_sided:.3f})")
ax.legend()
ax.spines["top"].set_visible(False)
ax.spines["right"].set_visible(False)
ax.spines["left"].set_color("#666666")
ax.spines["bottom"].set_color("#666666")
ax.tick_params(colors="#666666")
plt.tight_layout()
plt.savefig("permutation_test.png", dpi=150, bbox_inches="tight")
plt.show()from linearmodels.iv import IV2SLS
import numpy as np
import pandas as pd
# -------------------------------------------------------------------
# COMPARISON: Dummy Variable 2SLS vs Demeaned 2SLS
#
# Both estimate:
# mad_rate ~ evs_per_1000 (instrumented) + controls + county FE + month FE
#
# Dummy: includes FEs as explicit regressors
# Demeaned: iteratively absorbs county + month means from all
# variables, then runs IV2SLS on residualized data
# -------------------------------------------------------------------
df = df_final.copy()
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)
num_cols = [c for c in df.columns if c not in ["county_geoid", "county_name", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")
controls = ["tmax_c", "cumulative_data_center_power_mw", "total_generator_active_capacity_mw", "solar_mw_per_1000_lag1"]
control_str = " + ".join(controls)
iv_cols = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["county_geoid", "month"]
data = df[iv_cols].dropna().copy()
for c in data.columns:
if pd.api.types.is_numeric_dtype(data[c]):
data[c] = data[c].astype(float)
# ==========================================
# Method 1: Dummy Variable 2SLS (current)
# ==========================================
formula_dummy = f"""
mad_rate ~ 1 + {control_str} + C(county_geoid) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""
model_dummy = IV2SLS.from_formula(formula_dummy, data=data).fit(
cov_type="clustered", clusters=data["county_geoid"]
)
# ==========================================
# Method 2: Demeaned 2SLS
# ==========================================
# Iterative two-way demeaning to absorb county + month FEs
def demean_twoway(vals, entity, time, max_iter=100, tol=1e-10):
s = vals.copy().astype(float)
ent = pd.Series(entity)
tm = pd.Series(time)
for _ in range(max_iter):
old = s.copy()
s -= pd.Series(s).groupby(ent).transform("mean").values
s -= pd.Series(s).groupby(tm).transform("mean").values
if np.max(np.abs(s - old)) < tol:
break
return s
entity = data["county_geoid"].values
time = data["month"].values
# Demean all variables
vars_to_demean = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls
data_dm = pd.DataFrame()
for v in vars_to_demean:
data_dm[v] = demean_twoway(data[v].values, entity, time)
# Run IV2SLS on demeaned data (no FEs needed)
formula_dm = f"""
mad_rate ~ {control_str}
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""
# Clustered SEs still need the original county IDs
model_dm = IV2SLS.from_formula(formula_dm, data=data_dm).fit(
cov_type="clustered", clusters=data["county_geoid"]
)
# ==========================================
# Comparison Table
# ==========================================
n_fe = data["county_geoid"].nunique() + data["month"].nunique()
summary = pd.DataFrame([
{
"Method": "Dummy Variable 2SLS",
"Coef": round(model_dummy.params["evs_per_1000"], 6),
"SE (clustered)": round(model_dummy.std_errors["evs_per_1000"], 6),
"P-value": round(model_dummy.pvalues["evs_per_1000"], 4),
"First-stage F": round(float(model_dummy.first_stage.diagnostics.loc["evs_per_1000", "f.stat"]), 2),
"N": int(model_dummy.nobs),
},
{
"Method": "Demeaned 2SLS",
"Coef": round(model_dm.params["evs_per_1000"], 6),
"SE (clustered)": round(model_dm.std_errors["evs_per_1000"], 6),
"P-value": round(model_dm.pvalues["evs_per_1000"], 4),
"First-stage F": round(float(model_dm.first_stage.diagnostics.loc["evs_per_1000", "f.stat"]), 2),
"N": int(model_dm.nobs),
},
])
print("=" * 70)
print("COMPARISON: DUMMY VARIABLE vs DEMEANED 2SLS")
print("=" * 70)
print(f"Outcome: mad_rate")
print(f"Endogenous: evs_per_1000")
print(f"Instrument: iv_tesla_store_exposure_log")
print(f"Controls: {', '.join(controls)}")
print(f"FEs: county + month (~{n_fe} parameters)")
print(f"SEs: clustered by county")
print()
display(summary)
se_dummy = model_dummy.std_errors["evs_per_1000"]
se_dm = model_dm.std_errors["evs_per_1000"]
pct_diff = 100 * (se_dm - se_dummy) / se_dummy
print(f"\nCoefficient difference: {abs(model_dummy.params['evs_per_1000'] - model_dm.params['evs_per_1000']):.8f}")
print(f"SE difference: {pct_diff:+.1f}%")
if abs(pct_diff) < 5:
print("Negligible — both methods yield effectively the same inference.")
else:
print(f"Demeaned SEs are {'tighter' if pct_diff < 0 else 'wider'} — "
f"degrees-of-freedom correction matters with {len(data)} obs "
f"and ~{n_fe} FE parameters.")import geopandas as gpd
from shapely import wkt
tract_geo = gpd.GeoDataFrame(
tract_shape_df,
geometry=tract_shape_df["geometry"].apply(wkt.loads),
crs="EPSG:4326",
)
tract_geo = tract_geo.to_crs("EPSG:3310")
centroids = tract_geo.geometry.centroid.to_crs("EPSG:4326")
tract_geo["tract_lat"] = centroids.y
tract_geo["tract_lon"] = centroids.x
tract_points = tract_geo[["census_tract", "tract_lat", "tract_lon"]].copy()
print(f"Tract centroids: {len(tract_points)}")
print(tract_points.head())import numpy as np
# Prep Tesla stores
tesla_stores = tesla_loc_df[["latitude", "longitude", "open_month"]].copy()
tesla_stores["latitude"] = pd.to_numeric(tesla_stores["latitude"], errors="coerce")
tesla_stores["longitude"] = pd.to_numeric(tesla_stores["longitude"], errors="coerce")
tesla_stores["open_month"] = pd.to_datetime(tesla_stores["open_month"], errors="coerce").dt.to_period("M").dt.to_timestamp()
tesla_stores = tesla_stores.dropna().reset_index(drop=True)
# Pairwise inverse distances: (n_tracts, n_stores)
tlat = tract_points["tract_lat"].values[:, None]
tlon = tract_points["tract_lon"].values[:, None]
slat = tesla_stores["latitude"].values[None, :]
slon = tesla_stores["longitude"].values[None, :]
R = 6371.0
la1, lo1, la2, lo2 = map(np.radians, [tlat, tlon, slat, slon])
a = np.sin((la2 - la1) / 2)**2 + np.cos(la1) * np.cos(la2) * np.sin((lo2 - lo1) / 2)**2
dist = 2 * R * np.arcsin(np.sqrt(a))
inv_dist = 1.0 / np.maximum(dist, 1.0)
print(f"inv_dist shape: {inv_dist.shape} (tracts × stores)")
# Compute IV for each tract-month
panel_months = pd.DatetimeIndex(sorted(panel_df["month"].dropna().unique()))
open_dates = tesla_stores["open_month"].values
is_open = (open_dates[:, None] <= panel_months.values[None, :]).astype(float) # (stores, months)
exposure = inv_dist @ is_open # (tracts, months)
# Reshape to long DataFrame
tract_ids = np.repeat(tract_points["census_tract"].values, len(panel_months))
months_tile = np.tile(panel_months.values, len(tract_points))
iv_tract = pd.DataFrame({
"census_tract": tract_ids,
"month": months_tile,
"tesla_store_exposure": exposure.ravel(),
"iv_tesla_store_exposure_log": np.log1p(exposure.ravel()),
})
print(f"IV table: {len(iv_tract):,} rows")
print(iv_tract.head())print(tract_points.columns.tolist())
print(tract_points.head(3))from linearmodels.iv import IV2SLS
df = tract_panel_df.copy()
df["month"] = pd.to_datetime(df["month"], errors="coerce").dt.to_period("M").astype(str)
df["census_tract"] = df["census_tract"].astype(str)
df["county_geoid"] = df["county_geoid"].astype(str)
num_cols = [c for c in df.columns if c not in ["census_tract", "county_geoid", "month"]]
df[num_cols] = df[num_cols].apply(pd.to_numeric, errors="coerce")
# Controls (dropping tmax_c and data center — not available at tract level)
controls = ["total_active_capacity_mw"]
control_str = " + ".join(controls)
cols = ["mad_rate", "evs_per_1000", "iv_tesla_store_exposure_log"] + controls + ["census_tract", "county_geoid", "month"]
data = df[cols].dropna().copy()
for c in data.columns:
if pd.api.types.is_numeric_dtype(data[c]):
data[c] = data[c].astype(float)
# --- Tract FE + Month FE, clustered by county ---
formula = f"""
mad_rate ~ 1 + {control_str} + C(census_tract) + C(month)
[evs_per_1000 ~ iv_tesla_store_exposure_log]
"""
model = IV2SLS.from_formula(formula, data=data).fit(
cov_type="clustered", clusters=data["county_geoid"]
)
print("=" * 70)
print("TRACT-LEVEL 2SLS: mad_rate ~ evs_per_1000 (instrumented)")
print("=" * 70)
print(f"FEs: tract + month")
print(f"Controls: {', '.join(controls)}")
print(f"SEs: clustered by county ({data['county_geoid'].nunique()} clusters)")
print(f"N: {int(model.nobs)}")
print(f"Tracts: {data['census_tract'].nunique()}")
print(f"")
print(f"Coef: {model.params['evs_per_1000']:.6f}")
print(f"SE: {model.std_errors['evs_per_1000']:.6f}")
print(f"P-value: {model.pvalues['evs_per_1000']:.4f}")
print(f"First-stg F: {float(model.first_stage.diagnostics.loc['evs_per_1000', 'f.stat']):.2f}")
print("=" * 70)
#print(model.summary.tables[1])print([c for c in tract_panel_df.columns if "iv" in c.lower() or "tesla" in c.lower() or "exposure" in c.lower()])
# ── mad_rate distribution stats ───────────────────────────────────────────────
mad_stats = df_final['mad_rate'].describe(percentiles=[0.25, 0.5, 0.75])
mad_std = df_final['mad_rate'].std()
mad_iqr = df_final['mad_rate'].quantile(0.75) - df_final['mad_rate'].quantile(0.25)
mad_median = df_final['mad_rate'].median()
# ── Income percentiles (county-level, one obs per county) ─────────────────────
income_by_county = df_final.groupby('county_geoid')['median_household_income_10k'].first()
p10_income = income_by_county.quantile(0.10)
p25_income = income_by_county.quantile(0.25)
p75_income = income_by_county.quantile(0.75)
p90_income = income_by_county.quantile(0.90)
# ── EV adoption distribution ──────────────────────────────────────────────────
ev_stats = df_final['evs_per_1000'].describe(percentiles=[0.25, 0.5, 0.75, 0.90, 0.95])
# ── Pull coefficients directly from model object ───────────────────────────────
coef_ev = iv_model_mad_income.params['evs_per_1000']
coef_int = iv_model_mad_income.params['ev_x_income']
breakeven = -coef_ev / coef_int
print("=== Income-interaction 2SLS coefficients ===")
print(f"Main EV coef: {coef_ev:.6f}")
print(f"Interaction coef: {coef_int:.6f}")
print(f"Break-even income: ${breakeven * 10:.0f}k")
# ── Marginal effect function ───────────────────────────────────────────────────
def marginal_effect(income_10k):
return coef_ev + coef_int * income_10k
# ── Marginal effects at income percentiles ────────────────────────────────────
print("\n=== Marginal EV effect on mad_rate by income percentile ===")
income_anchors = [
("P10", p10_income),
("P25", p25_income),
("P75", p75_income),
("P90", p90_income),
]
for label, inc in income_anchors:
me = marginal_effect(inc)
scaled = me * 50
print(f" {label} (${inc*10:.0f}k): marginal = {me:.6f}, scaled (50 EVs) = {scaled:.4f} $/MWh ({scaled/mad_std:.2f} std devs)")
# ── P25 vs P75 ratio ──────────────────────────────────────────────────────────
me_p25 = marginal_effect(p25_income)
me_p75 = marginal_effect(p75_income)
me_p10 = marginal_effect(p10_income)
me_p90 = marginal_effect(p90_income)
print(f"\nP25/P75 effect ratio: {me_p25/me_p75:.2f}x")
print(f"P10/P90 effect ratio: {me_p10/me_p90:.2f}x")
# ── Scaled effect context (50 EV scenario at median income) ───────────────────
me_median = marginal_effect(income_by_county.median())
scaled_median = me_median * 50
print(f"\n=== 50-EV scenario at median county income (${income_by_county.median()*10:.0f}k) ===")
print(f"Implied mad_rate increase: {scaled_median:.4f} $/MWh")
print(f"As % of std dev: {scaled_median/mad_std:.2%}")
print(f"As % of IQR: {scaled_median/mad_iqr:.2%}")
print(f"As % of median: {scaled_median/mad_median:.2%}")
# ── EV adoption context (justifying the 50-EV scenario) ───────────────────────
print("\n=== evs_per_1000 distribution (justifying 50-EV scenario) ===")
print(ev_stats)coef_baseline = iv_model_mad.params['evs_per_1000']
delta_ev = 50
scaled = coef_baseline * delta_ev
print(f"Baseline coef: {coef_baseline:.6f}")
print(f"Scaled (50 EVs): {scaled:.4f} $/MWh")
print(f"As % of std dev: {scaled/mad_std:.2%}")
print(f"As % of IQR: {scaled/mad_iqr:.2%}")
print(f"As % of median: {scaled/mad_median:.2%}")
print(f"\n90th pctile adoption: {ev_stats['90%']:.1f} EVs/1,000")
print(f"75th pctile adoption: {ev_stats['75%']:.1f} EVs/1,000")import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import numpy as np
# ── Pull coefficients from model ───────────────────────────────────────────────
coef_ev = iv_model_mad_income.params['evs_per_1000']
coef_int = iv_model_mad_income.params['ev_x_income']
se_ev = iv_model_mad_income.std_errors['evs_per_1000']
se_int = iv_model_mad_income.std_errors['ev_x_income']
cov_ev_int = iv_model_mad_income.cov.loc['evs_per_1000', 'ev_x_income']
# ── Income range for plot ──────────────────────────────────────────────────────
income_min = income_by_county.min()
income_max = income_by_county.max()
income_grid = np.linspace(income_min, income_max, 200)
# ── Marginal effect and 95% CI ─────────────────────────────────────────────────
me = coef_ev + coef_int * income_grid
me_var = se_ev**2 + income_grid**2 * se_int**2 + 2 * income_grid * cov_ev_int
me_se = np.sqrt(me_var)
me_upper = me + 1.96 * me_se
me_lower = me - 1.96 * me_se
breakeven_k = (-coef_ev / coef_int) * 10
# ── Plot ───────────────────────────────────────────────────────────────────────
fig, ax = plt.subplots(figsize=(8, 4.5))
# CI band
ax.fill_between(income_grid * 10, me_lower, me_upper,
alpha=0.15, color='steelblue', label='95% Confidence Interval')
# Main line
ax.plot(income_grid * 10, me, color='steelblue', linewidth=2,
label='Marginal Effect of EV Adoption')
# Zero line (solid gray)
ax.axhline(0, color='gray', linewidth=0.8, linestyle='-', alpha=0.4)
# Break-even line (red dotted)
ax.axvline(breakeven_k, color='firebrick', linewidth=1.2, linestyle=':',
label=f'Break-even income (${breakeven_k:.0f}k)')
# Income percentile markers (solid gray, low alpha)
markers = {
'P25 ($57k)': p25_income,
'P75 ($85k)': p75_income,
}
for label, val in markers.items():
ax.axvline(val * 10, color='gray', linewidth=0.8, linestyle='-', alpha=0.4)
ax.text(val * 10, me_upper.max() * 1.02, label,
ha='center', va='bottom', fontsize=7.5, color='gray')
# Rug plot
county_incomes_k = income_by_county.values * 10
ax.plot(county_incomes_k,
np.full_like(county_incomes_k, me_lower.min() * 0.85),
'|', color='steelblue', alpha=0.5, markersize=8, markeredgewidth=1.2,
label='California counties')
# Notable county labels
county_income_df = df_final.groupby('county_name')['median_household_income_10k'].first()
label_counties = ['Los Angeles', 'San Francisco', 'Fresno', 'Santa Clara', 'Tulare']
for county in label_counties:
if county in county_income_df.index:
inc_k = county_income_df[county] * 10
me_at_inc = coef_ev + coef_int * (inc_k / 10)
ax.annotate(county, xy=(inc_k, me_at_inc),
xytext=(inc_k, me_at_inc + 0.0004),
fontsize=7, color='dimgray', ha='center',
arrowprops=dict(arrowstyle='-', color='lightgray', lw=0.8))
# Labels and formatting
ax.set_xlabel('Median Household Income ($k)', fontsize=11)
ax.set_ylabel('Marginal Effect on mad_rate\n($/MWh per EV per 1,000)', fontsize=10)
ax.set_title('Marginal Effect of EV Adoption on Price Volatility\nby County Income Level (2SLS Estimates)',
fontsize=12, fontweight='bold')
ax.legend(fontsize=9, framealpha=0.9)
xticks = np.arange(np.floor(income_min * 10 / 10) * 10,
np.ceil(income_max * 10 / 10) * 10 + 1, 10)
ax.set_xticks(xticks)
ax.set_xticklabels([f'${int(x)}k' for x in xticks])
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
plt.tight_layout()
plt.savefig('ev_income_marginal_effect.png', dpi=150, bbox_inches='tight')
plt.show()# Filter out zero or near-zero counties for cleaner plot
dac_plot = (df_final.groupby('county_name')['percent_dac_tracts']
.first()
.sort_values(ascending=True))
# Drop counties with 0% DAC share, add a note instead
dac_plot_filtered = dac_plot[dac_plot > 0]
n_zero = (dac_plot == 0).sum()
fig, ax = plt.subplots(figsize=(8, 7)) # shorter now
colors = ['firebrick' if v >= 50 else 'steelblue' if v >= 25 else 'lightsteelblue'
for v in dac_plot_filtered.values]
bars = ax.barh(dac_plot_filtered.index, dac_plot_filtered.values,
color=colors, edgecolor='none', height=0.7)
# Reference lines
ax.axvline(25, color='gray', linewidth=0.8, linestyle='--', alpha=0.5)
ax.axvline(50, color='gray', linewidth=0.8, linestyle='--', alpha=0.5)
ax.text(25.5, -1.2, '25%', fontsize=7.5, color='gray', va='top')
ax.text(50.5, -1.2, '50%', fontsize=7.5, color='gray', va='top')
# Value labels
for bar, val in zip(bars, dac_plot_filtered.values):
ax.text(val + 0.5, bar.get_y() + bar.get_height()/2,
f'{val:.0f}%', va='center', ha='left', fontsize=7.5, color='dimgray')
# Note about zero counties
ax.text(0, -2.5, f'Note: {n_zero} counties have 0% DAC-designated tracts and are excluded.',
fontsize=7.5, color='gray', style='italic')
ax.set_xlabel('Share of Census Tracts Designated as Disadvantaged (%)', fontsize=10)
ax.set_title('DAC Intensity by California County\n(Share of Disadvantaged Community Tracts)',
fontsize=12, fontweight='bold')
ax.set_xlim(0, dac_plot_filtered.max() * 1.15)
import matplotlib.patches as mpatches
legend_elements = [
mpatches.Patch(color='firebrick', label='≥50% DAC tracts'),
mpatches.Patch(color='steelblue', label='25–50% DAC tracts'),
mpatches.Patch(color='lightsteelblue', label='<25% DAC tracts'),
]
ax.legend(handles=legend_elements, fontsize=8, loc='lower right')
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
ax.spines['left'].set_visible(False)
ax.tick_params(left=False)
plt.tight_layout()
plt.savefig('dac_share_by_county.png', dpi=150, bbox_inches='tight')
plt.show()county_summary = (df_final.groupby('county_name')
.agg(
dac_share=('percent_dac_tracts', 'first'),
median_income_10k=('median_household_income_10k', 'first'),
mean_evs_per_1000=('evs_per_1000', 'mean'),
mean_mad_rate=('mad_rate', 'mean')
)
.reset_index()
.sort_values('dac_share', ascending=False))
# High DAC + low income counties
print("=== High DAC share + low income ===")
print(county_summary[county_summary['dac_share'] > 20]
[['county_name', 'dac_share', 'median_income_10k', 'mean_evs_per_1000']]
.sort_values('median_income_10k').to_string())ev_iv_cvrp_shift_share.ipynb46 cellsfrom google.colab import drive
drive.mount('/content/drive')
!pip install linearmodels
!pip install pandas numpy
# Setup & Imports
import os
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
import statsmodels.api as sm
from linearmodels.iv import IV2SLS
import warnings
warnings.filterwarnings('ignore', category=pd.errors.SettingWithCopyWarning)
# Define DATA_DIR - Modify this path as needed for your Colab environment
from pathlib import Path
DATA_DIR = Path(
"/content/drive/MyDrive/Capstone_Datasets"
)
DATA_DIR.exists()
# Initialize a list to store all regression results for the Step 5 summary table
all_model_results = []
def print_sanity_check(df, df_name="Dataset"):
if 'county_name' in df.columns:
print(f"[{df_name}] Sanity Check - Current N: {len(df)}, Unique Counties: {df['county_name'].nunique()}")
else:
print(f"[{df_name}] Sanity Check - Current N: {len(df)}")# STEP 0A — COMPUTE RELIANCE WEIGHT
print("--- STEP 0A: COMPUTE RELIANCE WEIGHT ---")
try:
cvrp_path = os.path.join(DATA_DIR, "CVRPStats_data.csv")
cvrp_df = pd.read_csv(cvrp_path)
# Clean Air District (actually county)
cvrp_df['county_name'] = cvrp_df['Air District'].astype(str).str.replace(' County', '', case=False).str.title().str.strip()
# Parse Application Date
cvrp_df['Application Date'] = pd.to_datetime(cvrp_df['Application Date'], format='%m/%d/%Y', errors='coerce')
# Filter to pre-treatment baseline (2019-2021)
cvrp_baseline = cvrp_df[(cvrp_df['Application Date'].dt.year >= 2019) &
(cvrp_df['Application Date'].dt.year <= 2021)]
# Identify low-income boosted applications (Assuming column name 'Low-/Moderate-Income Increased Rebate' or similar; adjust if needed)
li_col = 'Low-/Moderate-Income Increased Rebate'
if li_col not in cvrp_baseline.columns:
# Fallback if exact column name differs slightly in your local file
li_col = [c for c in cvrp_baseline.columns if 'Low' in c and 'Income' in c][0]
cvrp_baseline['is_low_income'] = (cvrp_baseline[li_col] > 0).astype(int)
# Aggregate to county level
county_reliance = cvrp_baseline.groupby('county_name').agg(
total_applications=('ID', 'count'),
li_applications=('is_low_income', 'sum')
).reset_index()
county_reliance['reliance_weight'] = county_reliance['li_applications'] / county_reliance['total_applications']
# Apply minimum threshold
threshold_mask = county_reliance['total_applications'] < 10
excluded_counties = county_reliance[threshold_mask]
county_reliance.loc[threshold_mask, 'reliance_weight'] = np.nan
print(f"Counties below the 10-application threshold (assigned NaN reliance):")
for _, row in excluded_counties.iterrows():
print(f" - {row['county_name']}: {row['total_applications']} applications")
print("\nCounty reliance dataframe created.")
except Exception as e:
print(f"ERROR loading or processing CVRP data: {e}")# STEP 0B — LOCK THE ANALYSIS SAMPLE
print("--- STEP 0B: LOCK THE ANALYSIS SAMPLE ---")
try:
# Load Main Panel
panel_path = os.path.join(DATA_DIR, "panel_df.csv")
panel_df = pd.read_csv(panel_path)
panel_df['month'] = pd.to_datetime(panel_df['month'])
# Load Solar Control
solar_path = os.path.join(DATA_DIR, "solar_control_variable.csv")
solar_df = pd.read_csv(solar_path)
solar_df['month'] = pd.to_datetime(solar_df['month'])
# Load Voter Registration (Political Data)
voter_path = os.path.join(DATA_DIR, "voter_registration.csv")
voter_df = pd.read_csv(voter_path)
# Date Parsing
voter_df['report_date'] = pd.to_datetime(voter_df[['year', 'month', 'day']])
# CRITICAL: Filter for Pre-April 2022 Data (Strict Exogeneity)
study_start_date = '2022-04-01'
history_mask = voter_df['report_date'] < study_start_date
baseline_voter = voter_df[history_mask].copy()
# Calculate Dem Share
baseline_voter['dem_share'] = baseline_voter['democratic'] / baseline_voter['registered']
# Clean County Names
baseline_voter['county_clean'] = baseline_voter['county'].astype(str).str.strip().str.title()
# Collapse to Static County Level
# Result has columns: ['county_clean', 'dem_share']
county_politics = baseline_voter.groupby('county_clean')['dem_share'].mean().reset_index()
# --- ERROR FIX IS HERE ---
county_politics = county_politics.rename(columns={
'county_clean': 'county_name'
})
voter_df = county_politics.copy()
# Merge datasets
analysis_df = panel_df.merge(county_reliance[['county_name', 'reliance_weight']], on='county_name', how='left')
analysis_df = analysis_df.merge(solar_df[['county_name', 'month', 'solar_mw_per_1000_lag1']], on=['county_name', 'month'], how='left')
analysis_df = analysis_df.merge(voter_df[['county_name', 'dem_share']], on='county_name', how='left')
# Check completeness per county
county_completeness = analysis_df.groupby('county_name').agg(
has_reliance=('reliance_weight', lambda x: x.notnull().any()),
has_solar=('solar_mw_per_1000_lag1', lambda x: x.notnull().any()),
has_voter=('dem_share', lambda x: x.notnull().any())
).reset_index()
# Lock Sample
valid_counties = county_completeness[
county_completeness['has_reliance'] &
county_completeness['has_solar'] &
county_completeness['has_voter']
]['county_name'].tolist()
excluded_counties = county_completeness[~county_completeness['county_name'].isin(valid_counties)]
analysis_df = analysis_df[analysis_df['county_name'].isin(valid_counties)].copy()
print(f"Included Counties ({len(valid_counties)}): {', '.join(valid_counties)}\n")
print(f"Excluded Counties ({len(excluded_counties)}):")
for _, row in excluded_counties.iterrows():
reasons = []
if not row['has_reliance']: reasons.append("Missing Reliance Weight")
if not row['has_solar']: reasons.append("Missing Solar Data")
if not row['has_voter']: reasons.append("Missing Voter Data")
print(f" - {row['county_name']}: {', '.join(reasons)}")
# Ensure Categorical Data for Clustering & Demeaning
analysis_df['county_cat'] = pd.Categorical(analysis_df['county_name'])
analysis_df['month_cat'] = pd.Categorical(analysis_df['month'])
print(f"\nLOCKED SAMPLE SANITY CHECK: N = {len(analysis_df)}, Unique Counties = {analysis_df['county_cat'].nunique()}")
except Exception as e:
print(f"ERROR merging datasets: {e}")print(panel_df.columns)# STEP 1 — DESCRIPTIVE STATISTICS & BALANCE TABLE
print("--- STEP 1: DESCRIPTIVE STATISTICS & BALANCE TABLE ---")
# Sanity Check
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")
# Plot over time
time_agg = analysis_df.groupby('month')[['evs_per_1000', 'monthly_new_evs']].mean().reset_index()
fig, ax1 = plt.subplots(figsize=(10, 5))
ax2 = ax1.twinx()
ax1.plot(time_agg['month'], time_agg['evs_per_1000'], color='blue', label='EVs per 1000 (Cumulative)')
ax2.plot(time_agg['month'], time_agg['monthly_new_evs'], color='red', label='Monthly New EVs', linestyle='--')
ax1.set_xlabel('Month')
ax1.set_ylabel('EVs per 1000', color='blue')
ax2.set_ylabel('Monthly New EVs', color='red')
plt.title('Average EV Adoption Over Time (Locked Sample)')
fig.legend(loc="upper left", bbox_to_anchor=(0.1,0.9))
plt.show()
# Balance Table (Pre-treatment implies prior to Sept 2023)
# We reconstruct panel_df to get excluded counties for comparison
pre_treat = panel_df[panel_df['month'] < '2023-09-01'].copy()
pre_treat['included'] = pre_treat['county_name'].isin(valid_counties)
# We also merge CVRP total_applications for the balance table
pre_treat = pre_treat.merge(county_reliance[['county_name', 'total_applications']], on='county_name', how='left')
balance_cols = ['evs_per_1000', 'mad_rate', 'tmax_c', 'total_applications']
balance_table = pre_treat.groupby('included')[balance_cols].mean().T
balance_table.columns = ['Excluded Counties', 'Included Counties']
print("\nPRE-TREATMENT BALANCE TABLE (Means):")
print(balance_table.round(3))# STEP 2 — BUILD THE HIGH/LOW RELIANCE DUMMIES
print("--- STEP 2: BUILD HIGH/LOW RELIANCE DUMMIES ---")
# Sanity Check
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")
# Median Split
median_reliance = analysis_df[['county_name', 'reliance_weight']].drop_duplicates()['reliance_weight'].median()
analysis_df['is_high_reliance'] = (analysis_df['reliance_weight'] > median_reliance).astype(int)
# Tercile Split
terciles = pd.qcut(analysis_df[['county_name', 'reliance_weight']].drop_duplicates()['reliance_weight'], 3, labels=[0, 1, 2])
county_tercile_map = dict(zip(analysis_df[['county_name']].drop_duplicates()['county_name'], terciles))
# Map to analysis_df: 0 = Bottom, 2 = Top (we rename to 1), 1 = Middle (NaN)
def map_tercile(county):
t = county_tercile_map[county]
if t == 2: return 1
elif t == 0: return 0
else: return np.nan
analysis_df['reliance_tercile'] = analysis_df['county_name'].apply(map_tercile)
# Print lists
high_rel_counties = analysis_df[analysis_df['is_high_reliance']==1]['county_name'].unique()
low_rel_counties = analysis_df[analysis_df['is_high_reliance']==0]['county_name'].unique()
top_tercile = analysis_df[analysis_df['reliance_tercile']==1]['county_name'].unique()
bottom_tercile = analysis_df[analysis_df['reliance_tercile']==0]['county_name'].unique()
print(f"Median Threshold: {median_reliance:.4f}")
print(f"High Reliance Counties (Median Split): {', '.join(high_rel_counties)}")
print(f"Low Reliance Counties (Median Split): {', '.join(low_rel_counties)}\n")
print(f"Top Tercile Counties: {', '.join(top_tercile)}")
print(f"Bottom Tercile Counties: {', '.join(bottom_tercile)}\n")
# Plot Histogram
rel_unique = analysis_df[['county_name', 'reliance_weight']].drop_duplicates()
plt.figure(figsize=(8,4))
plt.hist(rel_unique['reliance_weight'], bins=15, color='gray', edgecolor='black')
plt.axvline(median_reliance, color='red', linestyle='dashed', linewidth=2, label='Median')
tercile_bounds = np.percentile(rel_unique['reliance_weight'].dropna(), [33.33, 66.67])
plt.axvline(tercile_bounds[0], color='blue', linestyle='dotted', linewidth=2, label='Tercile 1/2')
plt.axvline(tercile_bounds[1], color='blue', linestyle='dotted', linewidth=2, label='Tercile 2/3')
plt.title('Distribution of CVRP Reliance Weight')
plt.xlabel('Share of Low-Income Boosted Applications')
plt.ylabel('Count of Counties')
plt.legend()
plt.show()# STEP 3 — BUILD INSTRUMENTS AND POLICY VARIABLES
print("--- STEP 3: BUILD INSTRUMENTS & POLICY VARIABLES ---")
# Sanity Check
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")
# Create temporal logic
analysis_df['post_cvrp'] = (analysis_df['month'] >= '2023-09-01').astype(int)
# Trend break: months elapsed since Nov 2023
analysis_df['iv_trend_break'] = ((analysis_df['month'].dt.year - 2023) * 12 + (analysis_df['month'].dt.month - 11)).clip(lower=0)
# Anticipation stock shift (Sep 2023 - Nov 2023)
analysis_df['iv_anticipation_stock_shift'] = ((analysis_df['month'] > '2023-08-01') &
(analysis_df['month'] <= '2023-11-30')).astype(int)
# NEM 3.0 control
analysis_df['post_nem3'] = (analysis_df['month'] >= '2024-01-01').astype(int)
analysis_df['solar_x_nem3'] = analysis_df['solar_mw_per_1000_lag1'] * analysis_df['post_nem3']
# Interactions
analysis_df['ev_x_high_reliance'] = analysis_df['evs_per_1000'] * analysis_df['is_high_reliance']
analysis_df['iv_trend_x_high_reliance'] = analysis_df['iv_trend_break'] * analysis_df['is_high_reliance']
analysis_df['iv_stock_x_high_reliance'] = analysis_df['iv_anticipation_stock_shift'] * analysis_df['is_high_reliance']
# Add continuous reliance interactions for later
analysis_df['ev_x_reliance_weight'] = analysis_df['evs_per_1000'] * analysis_df['reliance_weight']
analysis_df['iv_trend_x_reliance_weight'] = analysis_df['iv_trend_break'] * analysis_df['reliance_weight']
analysis_df['iv_stock_x_reliance_weight'] = analysis_df['iv_anticipation_stock_shift'] * analysis_df['reliance_weight']
# Print Phase Counts
phase_counts = pd.Series({
'Pre-Announcement (Before Sept 2023)': len(analysis_df[analysis_df['month'] < '2023-09-01']),
'Anticipation (Sept - Nov 2023)': len(analysis_df[(analysis_df['month'] >= '2023-09-01') & (analysis_df['month'] <= '2023-11-30')]),
'Post-Program (Dec 2023 Onward)': len(analysis_df[analysis_df['month'] > '2023-11-30'])
})
print(phase_counts)# STEP 4 — DEMEANING FUNCTIONS
print("--- STEP 4: DEMEANING FUNCTIONS ---")
def iterative_demean(df, value_cols, cat1='county_cat', cat2='month_cat',
tol=1e-10, max_iter=1000, verbose=True):
"""
Iterative (Gauss-Seidel) two-way demeaning implementing Frisch-Waugh-Lovell.
Returns df with new columns `{col}_dm` for each value_col.
IMPORTANT CAVEAT ON DEGREES OF FREEDOM
--------------------------------------
When you run OLS/IV on the demeaned variables, the software does NOT know
that `N_cat1 + N_cat2 - 1` parameters have been absorbed. Residual degrees
of freedom and (non-cluster) standard errors will be too small. Under
cluster-robust SEs with enough clusters (~50+), the distortion is modest
because the cluster count dominates the adjustment — but for exact
inference, use linearmodels.panel.PanelOLS with entity_effects=True /
time_effects=True, or the dof-corrected IV helper defined below.
"""
df_out = df.copy()
subset = [cat1, cat2] + value_cols
valid_idx = df_out[subset].dropna().index
n_valid = len(valid_idx)
n_total = len(df_out)
if verbose:
print(f" Demeaning {len(value_cols)} columns on {n_valid} rows "
f"({n_total - n_valid} dropped due to NaN)")
max_iter_used = 0
convergence_issues = []
for col in value_cols:
demeaned = df_out.loc[valid_idx, col].values.astype(float).copy()
c1 = df_out.loc[valid_idx, cat1].values
c2 = df_out.loc[valid_idx, cat2].values
diff = np.inf
iteration = 0
while diff > tol and iteration < max_iter:
old_demeaned = demeaned.copy()
# Sweep cat1
means1 = pd.Series(demeaned).groupby(c1, observed=True).transform('mean').values
demeaned = demeaned - means1
# Sweep cat2
means2 = pd.Series(demeaned).groupby(c2, observed=True).transform('mean').values
demeaned = demeaned - means2
diff = np.max(np.abs(demeaned - old_demeaned))
iteration += 1
max_iter_used = max(max_iter_used, iteration)
if iteration >= max_iter:
convergence_issues.append(col)
df_out.loc[valid_idx, f"{col}_dm"] = demeaned
if verbose:
print(f" Max iterations across columns: {max_iter_used} "
f"(tol={tol:.0e}, max_iter={max_iter})")
if convergence_issues:
print(f" WARNING: these columns hit max_iter without converging: {convergence_issues}")
return df_out
def county_demean(df, cols, verbose=False):
"""One-way demeaning by county. Single pass (exact)."""
df_out = df.copy()
valid_idx = df_out[['county_cat'] + cols].dropna().index
for c in cols:
means = df_out.loc[valid_idx, c].groupby(
df_out.loc[valid_idx, 'county_cat'], observed=True
).transform('mean')
df_out.loc[valid_idx, f"{c}_cdm"] = df_out.loc[valid_idx, c] - means
if verbose:
print(f" County-demeaned {len(cols)} columns on {len(valid_idx)} rows")
return df_out
def k_absorbed_twfe(df, entity='county_cat', time='month_cat'):
"""
Number of linearly independent FE parameters absorbed by TWFE demeaning:
N_entity + N_time - 1 (the -1 because one common constant is shared).
Used for dof correction when running IV2SLS on demeaned data.
"""
n_e = df[entity].nunique()
n_t = df[time].nunique()
return n_e + n_t - 1
print("Demeaning functions defined (iterative_demean, county_demean, k_absorbed_twfe).")
# MODEL 1: Naive OLS (no FE, no IV, no controls)
print("--- MODEL 1: NAIVE OLS ---")
# Sanity Check
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
# Formula String
print("Formula: mad_rate ~ evs_per_1000\n")
# Model
df_clean = analysis_df[['mad_rate', 'evs_per_1000', 'county_cat']].dropna()
exog = sm.add_constant(df_clean[['evs_per_1000']])
mod1 = IV2SLS(dependent=df_clean['mad_rate'],
exog=exog,
endog=None,
instruments=None)
res1 = mod1.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res1.summary)# MODEL 2: OLS with Controls
print("--- MODEL 2: OLS WITH CONTROLS ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate ~ evs_per_1000 + tmax_c + cumulative_data_center_power_mw + solar_mw_per_1000_lag1 + total_generator_active_capacity_mw\n")
cols = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw', 'county_cat']
df_clean = analysis_df[cols].dropna()
exog = sm.add_constant(df_clean[['evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw']])
mod2 = IV2SLS(dependent=df_clean['mad_rate'],
exog=exog,
endog=None,
instruments=None)
res2 = mod2.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res2.summary)# MODEL 2 FE VARIANTS (2a: county FE, 2b: month FE, 2c: TWFE)
from linearmodels.panel import PanelOLS
print("--- MODEL 2 FE VARIANTS ---")
cols = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw']
df_m2fe = analysis_df[cols + ['county_name', 'month']].dropna().copy()
# PanelOLS requires a MultiIndex: (entity, time)
df_m2fe = df_m2fe.set_index(['county_name', 'month'])
y = df_m2fe['mad_rate']
X = df_m2fe[['evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw']]
X_const = sm.add_constant(X) # constant absorbed when FE included, but PanelOLS handles it
print(f"N = {len(df_m2fe)}, Unique counties = {df_m2fe.index.get_level_values(0).nunique()}, "
f"Unique months = {df_m2fe.index.get_level_values(1).nunique()}\n")
# -------- MODEL 2a: County FE only --------
print("=" * 70)
print("MODEL 2a: OLS + Controls + County FE")
print("Formula: mad_rate ~ evs_per_1000 + tmax_c + data_center + solar_lag1 + total_generator_active_capacity_mw + county_FE")
print("=" * 70)
mod2a = PanelOLS(y, X_const, entity_effects=True, drop_absorbed=True)
res2a = mod2a.fit(cov_type='clustered', cluster_entity=True)
print(res2a.summary)
print()
# -------- MODEL 2b: Month FE only --------
print("=" * 70)
print("MODEL 2b: OLS + Controls + Month FE")
print("Formula: mad_rate ~ evs_per_1000 + tmax_c + data_center + solar_lag1 + total_generator_active_capacity_mw + month_FE")
print("=" * 70)
mod2b = PanelOLS(y, X_const, time_effects=True, drop_absorbed=True)
res2b = mod2b.fit(cov_type='clustered', cluster_entity=True)
print(res2b.summary)
print()
# -------- MODEL 2c: County + Month FE (TWFE) --------
print("=" * 70)
print("MODEL 2c: OLS + Controls + County FE + Month FE (TWFE)")
print("Formula: mad_rate ~ evs_per_1000 + tmax_c + data_center + solar_lag1 + total_generator_active_capacity_mw + county_FE + month_FE")
print("=" * 70)
mod2c = PanelOLS(y, X_const, entity_effects=True, time_effects=True, drop_absorbed=True)
res2c = mod2c.fit(cov_type='clustered', cluster_entity=True)
print(res2c.summary)
print()
# -------- Side-by-side comparison of evs_per_1000 coefficient --------
print("=" * 70)
print("COMPARISON: evs_per_1000 coefficient across Model 2 variants")
print("=" * 70)
comparison = pd.DataFrame({
'Model': ['2 (no FE)', '2a (county FE)', '2b (month FE)', '2c (TWFE)'],
'evs_per_1000 coef': [
res2e.params['evs_per_1000'],
res2a.params['evs_per_1000'],
res2b.params['evs_per_1000'],
res2c.params['evs_per_1000'],
],
'Std. Err': [
res2e.std_errors['evs_per_1000'],
res2a.std_errors['evs_per_1000'],
res2b.std_errors['evs_per_1000'],
res2c.std_errors['evs_per_1000'],
],
'p-value': [
res2e.pvalues['evs_per_1000'],
res2a.pvalues['evs_per_1000'],
res2b.pvalues['evs_per_1000'],
res2c.pvalues['evs_per_1000'],
],
'Within R2': [
res2e.rsquared,
res2a.rsquared_within,
res2b.rsquared_within,
res2c.rsquared_within,
],
})
print(comparison.to_string(index=False))
# MODEL 3: TWFE OLS
print("--- MODEL 3: TWFE OLS ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ evs_per_1000_dm + tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + total_generator_active_capacity_mw_dm \n")
cols_to_demean = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw']
df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c+"_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['evs_per_1000_dm', 'tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'solar_mw_per_1000_lag1_dm','total_generator_active_capacity_mw_dm']]
mod3 = IV2SLS(dependent=df_clean['mad_rate_dm'],
exog=exog,
endog=None,
instruments=None)
res3 = mod3.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res3.summary)# EVENT STUDY PLOT — PRE/POST PARALLEL TRENDS
print("--- EVENT STUDY PLOT ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")
event_df = analysis_df.dropna(subset=['mad_rate', 'is_high_reliance'])
agg_trends = event_df.groupby(['month', 'is_high_reliance'])['mad_rate'].mean().unstack()
plt.figure(figsize=(12, 6))
plt.plot(agg_trends.index, agg_trends[1], label='High Reliance Counties', color='darkred', linewidth=2)
plt.plot(agg_trends.index, agg_trends[0], label='Low Reliance Counties', color='darkblue', linewidth=2)
# Policy timeline
plt.axvline(pd.to_datetime('2023-09-01'), color='black', linestyle='--', label='CVRP Closure Announced (Sep 2023)')
plt.axvline(pd.to_datetime('2023-11-01'), color='black', linestyle='--', label='CVRP Effectively Ended (Nov 2023)')
# Shade Anticipation Window
plt.axvspan(pd.to_datetime('2023-09-01'), pd.to_datetime('2023-11-01'), color='gray', alpha=0.3, label='Anticipation Window')
plt.title('Parallel Trends: Mean Absolute Deviation (MAD) of LMP')
plt.ylabel('MAD Rate')
plt.xlabel('Month')
plt.legend()
plt.tight_layout()
plt.show()# MODEL 4a: Weak IV (post_cvrp as Standalone Instrument)
print("--- MODEL 4a: WEAK IV (STRAW MAN) ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + total_generator_active_capacity_mw_dm + [evs_per_1000_dm ~ post_cvrp_dm]\n")
cols_to_demean = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'post_cvrp','total_generator_active_capacity_mw']
df_dm = iterative_demean(analysis_df, cols_to_demean)
dm_cols = [c+"_dm" for c in cols_to_demean]
df_clean = df_dm[dm_cols + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'solar_mw_per_1000_lag1_dm','total_generator_active_capacity_mw_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['post_cvrp_dm']]
try:
mod4a = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res4a = mod4a.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res4a.summary)
except ValueError as e:
print(f"ECONOMETRIC RANK FAILURE CAUGHT:\nValueError: {e}\n")
print("""
DIAGNOSTIC — WHY THIS INSTRUMENT FAILS (AND THROWS AN ERROR):
1. Perfect Collinearity: 'post_cvrp' is a universal time indicator. After two-way demeaning (subtracting month fixed effects), it has strictly zero cross-sectional variation. The demeaned column is perfectly zeros, causing the design matrix to drop rank, hence the ValueError.
2. Identification Failure: This proves you cannot use a universal macro-shock as an instrument in a TWFE model. The instrument is completely absorbed by the time fixed effects.
3. This specification is presented deliberately as a failed straw man to motivate the preferred shift-share instruments (iv_trend_break × reliance), which introduce necessary cross-sectional variation.
""")# MODEL 4b: Simple DiD (county FE only)
print("--- MODEL 4b: SIMPLE DiD ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate ~ evs_per_1000 + post_cvrp + post_cvrp × is_high_reliance + tmax_c + cumulative_data_center_power_mw + solar_mw_per_1000_lag1 + total_generator_active_capacity_mw + county_FE\n")
analysis_df['post_x_reliance'] = analysis_df['post_cvrp'] * analysis_df['is_high_reliance']
cols = ['mad_rate', 'evs_per_1000', 'post_cvrp', 'post_x_reliance', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'total_generator_active_capacity_mw']
df_cdm = county_demean(analysis_df, cols)
df_clean = df_cdm[[c+"_cdm" for c in cols] + ['county_cat']].dropna()
exog = df_clean[[c+"_cdm" for c in cols if c != 'mad_rate']]
mod4b = IV2SLS(dependent=df_clean['mad_rate_cdm'], exog=exog, endog=None, instruments=None)
res4b = mod4b.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res4b.summary)
print("\nNOTE: Time fixed effects are intentionally excluded from this model. With full TWFE, the post_cvrp dummy would be collinear with the month fixed effects since the CVRP closure is a universal California event. This model is a deliberate straw man demonstrating the identification problem that motivates the reliance-based IV design.")# MODEL 4c: Simple DiD (county + month FE )
print("--- MODEL 4C: SIMPLE DiD County + Month FE ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ evs_per_1000_dm + (post_cvrp × is_high_reliance)_dm + tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + total_generator_active_capacity_mw_dm")
analysis_df['post_x_reliance'] = analysis_df['post_cvrp'] * analysis_df['is_high_reliance']
cols = ['mad_rate', 'evs_per_1000', 'post_x_reliance', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'total_generator_active_capacity_mw']
df_dm = iterative_demean(analysis_df, cols)
df_clean = df_dm[[c+"_dm" for c in cols] + ['county_cat']].dropna()
exog = df_clean[[c+"_dm" for c in cols if c != 'mad_rate']]
mod4c = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=None, instruments=None)
res4c = mod4c.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res4c.summary)
# MODEL 4d: Simple DiD (county + month FE)
print("--- MODEL 4D: SIMPLE DiD County + Month FE ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ evs_per_1000_dm + (post_cvrp × reliance_weight)_dm + tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + total_generator_active_capacity_mw_dm")
analysis_df['post_x_reliance_weight'] = analysis_df['post_cvrp'] * analysis_df['reliance_weight']
cols = ['mad_rate', 'evs_per_1000', 'post_x_reliance_weight', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'total_generator_active_capacity_mw']
df_dm = iterative_demean(analysis_df, cols)
df_clean = df_dm[[c+"_dm" for c in cols] + ['county_cat']].dropna()
exog = df_clean[[c+"_dm" for c in cols if c != 'mad_rate']]
mod4d = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=None, instruments=None)
res4d = mod4d.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res4d.summary)# === MODIFICATION 1 (REVISED): MODEL 5 REDUCED FORM ===
# MODEL 5: Reduced Form (TWFE)
# The base instruments iv_trend_break and iv_anticipation_stock_shift are purely
# time-varying (identical across counties within each month). After subtracting
# month means in TWFE they collapse to zero and cause rank failure. The correct
# instruments for a TWFE design are the shift-share versions that interact the
# time shock with the pre-treatment county reliance weight, preserving
# cross-sectional variation after month demeaning.
print("--- MODEL 5: REDUCED FORM ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm\n")
cols_to_demean = [
'mad_rate', 'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
'tmax_c', 'cumulative_data_center_power_mw','total_generator_active_capacity_mw','solar_mw_per_1000_lag1', 'solar_x_nem3'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
dm_cols = [c + "_dm" for c in cols_to_demean]
df_clean = df_dm[dm_cols + ['county_cat']].dropna()
exog = df_clean[[c for c in dm_cols if c != 'mad_rate_dm']]
mod5 = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=None, instruments=None)
res5 = mod5.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res5.summary)print("--- MODEL 5A: REDUCED FORM ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ evs_per_1000_dm + iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm\n")
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
'total_generator_active_capacity_mw'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
dm_cols = [c + "_dm" for c in cols_to_demean]
df_clean = df_dm[dm_cols + ['county_cat']].dropna()
exog = df_clean[[c for c in dm_cols if c != 'mad_rate_dm']]
mod5 = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=None, instruments=None)
res5 = mod5.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res5.summary)# === MODIFICATION 2 (REVISED): MODEL 6 IV-2SLS BASELINE ===
# MODEL 6: IV-2SLS Baseline (TWFE)
print("--- MODEL 6: IV-2SLS BASELINE ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1 + [evs_per_1000_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm]\n")
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight', 'total_generator_active_capacity_mw'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]
mod6 = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res6 = mod6.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res6.summary)
f_stat = res6.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat:.4f}")
if f_stat < 10:
print("WARNING: First-stage F-statistic is below 10. Instrument may be weak.")
# SARGAN-HANSEN OVERIDENTIFICATION TEST — MODEL 6 (manual computation)
# S = n * R² from regressing 2SLS residuals on the full instrument matrix
# Distributed as chi2(L - K) where L = excluded instruments, K = endogenous vars
# Model 6: L=2, K=1 -> df=1
from scipy import stats as scipy_stats
try:
e = res6.resids.values
Z_full = np.column_stack([exog.values, instr.values]) # exog + excluded instruments
n = len(e)
# Project residuals onto instrument space
ZtZ_inv = np.linalg.pinv(Z_full.T @ Z_full)
e_fitted = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
sargan_stat = float(n * (e_fitted @ e_fitted) / (e @ e))
df_sargan = instr.shape[1] - endog.shape[1] # 2 - 1 = 1
sargan_pval = 1 - scipy_stats.chi2.cdf(sargan_stat, df=df_sargan)
print("\n--- SARGAN-HANSEN OVERIDENTIFICATION TEST ---")
print(f"Statistic: {sargan_stat:.4f}")
print(f"P-value: {sargan_pval:.4f}")
print(f"Distribution: chi2({df_sargan}) "
f"[degrees of freedom = {instr.shape[1]} instruments - {endog.shape[1]} endogenous var]")
if sargan_pval > 0.10:
print("RESULT: Fail to reject H0 (p > 0.10). Instruments pass the overidentification test.")
print(" Both instruments are jointly consistent with instrument validity.")
elif sargan_pval > 0.05:
print("RESULT: Marginal rejection at 10% level. Interpret instrument validity with caution.")
else:
print("RESULT: Reject H0 (p < 0.05). At least one instrument may be invalid.")
print(" Consider whether exclusion restriction holds for both instruments separately.")
except Exception as e_err:
print(f"Sargan test could not be computed: {e_err}")df_clean.columns# === MODIFICATION 3 (REVISED): MODEL 7 IV-2SLS FINAL ===
# MODEL 7: IV-2SLS Final (TWFE)
print("--- MODEL 7: IV-2SLS FINAL ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + total_generator_active_capacity_mw_dm + [evs_per_1000_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm]\n")
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
'solar_mw_per_1000_lag1', 'solar_x_nem3', 'total_generator_active_capacity_mw' ,
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm', 'total_generator_active_capacity_mw_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]
mod7 = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res7 = mod7.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res7.summary)
f_stat = res7.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat:.4f}")
if f_stat < 10:
print("WARNING: First-stage F-statistic is below 10. Instrument may be weak.")
# SARGAN-HANSEN OVERIDENTIFICATION TEST — MODEL 7 (manual computation)
# S = n * R² from regressing 2SLS residuals on the full instrument matrix
# Distributed as chi2(L - K) where L = excluded instruments, K = endogenous vars
# Model 7: L=2, K=1 -> df=1
from scipy import stats as scipy_stats
try:
e = res7.resids.values
Z_full = np.column_stack([exog.values, instr.values]) # exog + excluded instruments
n = len(e)
# Project residuals onto instrument space
ZtZ_inv = np.linalg.pinv(Z_full.T @ Z_full)
e_fitted = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
sargan_stat = float(n * (e_fitted @ e_fitted) / (e @ e))
df_sargan = instr.shape[1] - endog.shape[1] # 2 - 1 = 1
sargan_pval = 1 - scipy_stats.chi2.cdf(sargan_stat, df=df_sargan)
print("\n--- SARGAN-HANSEN OVERIDENTIFICATION TEST ---")
print(f"Statistic: {sargan_stat:.4f}")
print(f"P-value: {sargan_pval:.4f}")
print(f"Distribution: chi2({df_sargan}) "
f"[degrees of freedom = {instr.shape[1]} instruments - {endog.shape[1]} endogenous var]")
if sargan_pval > 0.10:
print("RESULT: Fail to reject H0 (p > 0.10). Instruments pass the overidentification test.")
print(" Both instruments are jointly consistent with instrument validity.")
elif sargan_pval > 0.05:
print("RESULT: Marginal rejection at 10% level. Interpret instrument validity with caution.")
else:
print("RESULT: Reject H0 (p < 0.05). At least one instrument may be invalid.")
print(" Consider whether exclusion restriction holds for both instruments separately.")
except Exception as e_err:
print(f"Sargan test could not be computed: {e_err}")# === MODEL 7R IV-2SLS FINAL ===
# MODEL 7R: IV-2SLS Final - No FE
print("--- MODEL 7R: IV-2SLS NO FE ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate ~ tmax_c + cumulative_data_center_power_mw + solar_mw_per_1000_lag1 + solar_x_nem3 + total_generator_active_capacity_mw + [evs_per_1000 ~ iv_trend_x_reliance_weight + iv_stock_x_reliance_weight]\n")
exog = sm.add_constant(analysis_df[['tmax_c', 'cumulative_data_center_power_mw',
'solar_mw_per_1000_lag1', 'solar_x_nem3',
'total_generator_active_capacity_mw']])
endog = analysis_df[['evs_per_1000']]
instr = analysis_df[['iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']]
mod7r = IV2SLS(dependent=analysis_df['mad_rate'], exog=exog, endog=endog, instruments=instr)
res7r = mod7r.fit(cov_type='clustered', clusters=analysis_df['county_cat'])
print(res7r.summary)
f_stat = res7r.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat:.4f}")
if f_stat < 10:
print("WARNING: First-stage F-statistic is below 10. Instrument may be weak.")
# SARGAN-HANSEN OVERIDENTIFICATION TEST — MODEL 7 (manual computation)
# S = n * R² from regressing 2SLS residuals on the full instrument matrix
# Distributed as chi2(L - K) where L = excluded instruments, K = endogenous vars
# Model 7: L=2, K=1 -> df=1
from scipy import stats as scipy_stats
try:
e = res7r.resids.values
Z_full = np.column_stack([exog.values, instr.values]) # exog + excluded instruments
n = len(e)
# Project residuals onto instrument space
ZtZ_inv = np.linalg.pinv(Z_full.T @ Z_full)
e_fitted = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
sargan_stat = float(n * (e_fitted @ e_fitted) / (e @ e))
df_sargan = instr.shape[1] - endog.shape[1] # 2 - 1 = 1
sargan_pval = 1 - scipy_stats.chi2.cdf(sargan_stat, df=df_sargan)
print("\n--- SARGAN-HANSEN OVERIDENTIFICATION TEST ---")
print(f"Statistic: {sargan_stat:.4f}")
print(f"P-value: {sargan_pval:.4f}")
print(f"Distribution: chi2({df_sargan}) "
f"[degrees of freedom = {instr.shape[1]} instruments - {endog.shape[1]} endogenous var]")
if sargan_pval > 0.10:
print("RESULT: Fail to reject H0 (p > 0.10). Instruments pass the overidentification test.")
print(" Both instruments are jointly consistent with instrument validity.")
elif sargan_pval > 0.05:
print("RESULT: Marginal rejection at 10% level. Interpret instrument validity with caution.")
else:
print("RESULT: Reject H0 (p < 0.05). At least one instrument may be invalid.")
print(" Consider whether exclusion restriction holds for both instruments separately.")
except Exception as e_err:
print(f"Sargan test could not be computed: {e_err}")# === MODEL 7S IV-2SLS FINAL ===
# MODEL 7S: IV-2SLS Final - No FE
print("--- MODEL 7S: IV-2SLS NO FE ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate ~ tmax_c + cumulative_data_center_power_mw + solar_mw_per_1000_lag1 + total_generator_active_capacity_mw + [evs_per_1000 ~ iv_trend_x_reliance_weight + iv_stock_x_reliance_weight]\n")
exog = sm.add_constant(analysis_df[['tmax_c', 'cumulative_data_center_power_mw',
'solar_mw_per_1000_lag1',
'total_generator_active_capacity_mw']])
endog = analysis_df[['evs_per_1000']]
instr = analysis_df[['iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']]
mod7s = IV2SLS(dependent=analysis_df['mad_rate'], exog=exog, endog=endog, instruments=instr)
res7s = mod7s.fit(cov_type='clustered', clusters=analysis_df['county_cat'])
print(res7s.summary)
f_stat = res7s.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat:.4f}")
if f_stat < 10:
print("WARNING: First-stage F-statistic is below 10. Instrument may be weak.")
# SARGAN-HANSEN OVERIDENTIFICATION TEST — MODEL 7 (manual computation)
# S = n * R² from regressing 2SLS residuals on the full instrument matrix
# Distributed as chi2(L - K) where L = excluded instruments, K = endogenous vars
# Model 7: L=2, K=1 -> df=1
from scipy import stats as scipy_stats
try:
e = res7s.resids.values
Z_full = np.column_stack([exog.values, instr.values]) # exog + excluded instruments
n = len(e)
# Project residuals onto instrument space
ZtZ_inv = np.linalg.pinv(Z_full.T @ Z_full)
e_fitted = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
sargan_stat = float(n * (e_fitted @ e_fitted) / (e @ e))
df_sargan = instr.shape[1] - endog.shape[1] # 2 - 1 = 1
sargan_pval = 1 - scipy_stats.chi2.cdf(sargan_stat, df=df_sargan)
print("\n--- SARGAN-HANSEN OVERIDENTIFICATION TEST ---")
print(f"Statistic: {sargan_stat:.4f}")
print(f"P-value: {sargan_pval:.4f}")
print(f"Distribution: chi2({df_sargan}) "
f"[degrees of freedom = {instr.shape[1]} instruments - {endog.shape[1]} endogenous var]")
if sargan_pval > 0.10:
print("RESULT: Fail to reject H0 (p > 0.10). Instruments pass the overidentification test.")
print(" Both instruments are jointly consistent with instrument validity.")
elif sargan_pval > 0.05:
print("RESULT: Marginal rejection at 10% level. Interpret instrument validity with caution.")
else:
print("RESULT: Reject H0 (p < 0.05). At least one instrument may be invalid.")
print(" Consider whether exclusion restriction holds for both instruments separately.")
except Exception as e_err:
print(f"Sargan test could not be computed: {e_err}")# MODEL 7a: IV-2SLS Trend Only (TWFE)
print("--- MODEL 7a: IV-2SLS TREND INSTRUMENT ONLY ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm ~ iv_trend_x_reliance_weight_dm]\n")
print("NOTE: Exactly identified (1 instrument, 1 endogenous variable). Sargan test not available.\n")
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'total_generator_active_capacity_mw',
'solar_mw_per_1000_lag1', 'solar_x_nem3',
'iv_trend_x_reliance_weight'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm',
'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm']]
mod7a = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res7a = mod7a.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res7a.summary)
f_stat_7a = res7a.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat_7a:.4f}")
if f_stat_7a < 10:
print("WARNING: First-stage F-statistic is below 10. Weak instrument bias is a concern.")
else:
print("First-stage F-statistic is above 10. Instrument passes relevance threshold.")
print(f"\nEV Coefficient: {res7a.params['evs_per_1000_dm']:.6f}")
print(f"P-value: {res7a.pvalues['evs_per_1000_dm']:.4f}")# MODEL 7b: IV-2SLS Anticipation Only (TWFE)
print("--- MODEL 7b: IV-2SLS ANTICIPATION INSTRUMENT ONLY ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm ~ iv_stock_x_reliance_weight_dm]\n")
print("NOTE: Exactly identified (1 instrument, 1 endogenous variable). Sargan test not available.\n")
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'total_generator_active_capacity_mw',
'solar_mw_per_1000_lag1', 'solar_x_nem3',
'iv_stock_x_reliance_weight'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm',
'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_stock_x_reliance_weight_dm']]
mod7b = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res7b = mod7b.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res7b.summary)
f_stat_7b = res7b.first_stage.diagnostics['f.stat'].iloc[0]
print(f"\nFirst-Stage F-Statistic: {f_stat_7b:.4f}")
if f_stat_7b < 10:
print("WARNING: First-stage F-statistic is below 10. Weak instrument bias is a concern.")
else:
print("First-stage F-statistic is above 10. Instrument passes relevance threshold.")
print(f"\nEV Coefficient: {res7b.params['evs_per_1000_dm']:.6f}")
print(f"P-value: {res7b.pvalues['evs_per_1000_dm']:.4f}")# ROBUSTNESS: Native-FE re-estimation of Models 3, 4b/4c/4d (OLS) and 5, 6, 7 (IV)
from linearmodels.panel import PanelOLS
from linearmodels.iv import IV2SLS
print("=" * 78)
print("ROBUSTNESS — Native FE Re-estimation")
print("=" * 78)
# Build a panel-indexed dataframe once
panel_df = analysis_df.copy()
if 'post_x_reliance' not in panel_df.columns:
panel_df['post_x_reliance'] = panel_df['post_cvrp'] * panel_df['is_high_reliance']
if 'post_x_reliance_weight' not in panel_df.columns:
panel_df['post_x_reliance_weight'] = panel_df['post_cvrp'] * panel_df['reliance_weight']
panel_df = panel_df.set_index(['county_name', 'month']).sort_index()
def run_panel_ols(y_col, x_cols, entity=True, time=True, label=''):
"""Fit PanelOLS with requested FE and clustered SEs by entity."""
sub = panel_df[[y_col] + x_cols].dropna()
y = sub[y_col]
X = sub[x_cols]
mod = PanelOLS(y, X, entity_effects=entity, time_effects=time, drop_absorbed=True)
res = mod.fit(cov_type='clustered', cluster_entity=True)
print(f"\n{'=' * 78}\n{label}\n{'=' * 78}")
print(f"N={int(res.nobs)}, entities={sub.index.get_level_values(0).nunique()}, "
f"FE: entity={entity}, time={time}")
print(res.summary)
return res
# ---- Model 3 native (TWFE) ----
res3_native = run_panel_ols(
'mad_rate',
['evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw'],
entity=True, time=True,
label='MODEL 3 (native TWFE via PanelOLS)'
)
# ---- Model 4b native (county FE only, deliberate straw man) ----
res4b_native = run_panel_ols(
'mad_rate',
['evs_per_1000', 'post_cvrp', 'post_x_reliance', 'tmax_c',
'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw'],
entity=True, time=False,
label='MODEL 4b (native county FE only — straw man)'
)
# ---- Model 4c native (TWFE + binary interaction) ----
res4c_native = run_panel_ols(
'mad_rate',
['evs_per_1000', 'post_x_reliance', 'tmax_c',
'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw'],
entity=True, time=True,
label='MODEL 4c (native TWFE, binary post×reliance)'
)
# ---- Model 4d native (TWFE + continuous interaction) ----
res4d_native = run_panel_ols(
'mad_rate',
['evs_per_1000', 'post_x_reliance_weight', 'tmax_c',
'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1','total_generator_active_capacity_mw'],
entity=True, time=True,
label='MODEL 4d (native TWFE, continuous post×reliance_weight)'
)
# ROBUSTNESS (cont'd) — dof-corrected IV2SLS for Models 5, 6, 7
# ------------------------------------------------------------------
# Strategy: demean via FWL (which we've already validated gives exact
# coefficients), then fit IV2SLS on the demeaned data, but rescale the
# cluster-robust variance by (N-1)/(N - k - k_fe) vs. linearmodels' default
# (N-1)/(N - k), where k_fe = N_county + N_month - 1.
#
# This mirrors what Stata's ivreghdfe and R's fixest do internally.
# ------------------------------------------------------------------
from scipy import stats as scipy_stats
def fit_iv_dof_corrected(df_dm_clean, dep_col, exog_cols, endog_cols, instr_cols,
k_fe, cluster_col='county_cat', label=''):
"""
Fit IV2SLS on pre-demeaned data and apply FE dof correction to the
cluster-robust standard errors.
linearmodels reports SEs using a small-sample factor that assumes
residual dof = N - k (where k = # non-FE regressors). The true residual
dof after absorbing FEs is N - k - k_fe. We rescale the variance matrix
by (N - k) / (N - k - k_fe) and recompute t-stats, p-values, CIs.
"""
dep = df_dm_clean[dep_col]
exog = df_dm_clean[exog_cols] if exog_cols else None
endog = df_dm_clean[endog_cols] if endog_cols else None
instr = df_dm_clean[instr_cols] if instr_cols else None
mod = IV2SLS(dependent=dep, exog=exog, endog=endog, instruments=instr)
res = mod.fit(cov_type='clustered', clusters=df_dm_clean[cluster_col], debiased=True)
# Rescale
N = int(res.nobs)
k_total = len(res.params) # non-FE regressors including constant if present
naive_dof = N - k_total
corrected_dof = N - k_total - k_fe
scale = naive_dof / corrected_dof # variance inflation factor
V_corrected = res.cov * scale
se_corrected = np.sqrt(np.diag(V_corrected))
tstats = res.params.values / se_corrected
pvals = 2 * (1 - scipy_stats.t.cdf(np.abs(tstats), df=corrected_dof))
print(f"\n{'=' * 78}\n{label}\n{'=' * 78}")
print(f"N={N}, non-FE params k={k_total}, absorbed FE params k_fe={k_fe}")
print(f"Variance inflation factor (dof correction): {scale:.4f}")
print(f"\n{'Parameter':<45s} {'Coef':>12s} {'Naive SE':>12s} {'Corr. SE':>12s} {'Corr. p':>10s}")
print('-' * 93)
for i, name in enumerate(res.params.index):
print(f"{name:<45s} {res.params.values[i]:>12.4e} "
f"{res.std_errors.values[i]:>12.4e} "
f"{se_corrected[i]:>12.4e} {pvals[i]:>10.4f}")
return {
'res': res,
'N': N,
'k_fe': k_fe,
'scale': scale,
'params': res.params,
'se_naive': res.std_errors,
'se_corrected': pd.Series(se_corrected, index=res.params.index),
'pvals_corrected': pd.Series(pvals, index=res.params.index),
}
k_fe = k_absorbed_twfe(analysis_df)
print(f"\nAbsorbed TWFE parameters (N_county + N_month - 1): {k_fe}")
# ---- Model 5 (Reduced Form) ----
cols5 = ['mad_rate', 'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3','total_generator_active_capacity_mw']
df5 = iterative_demean(analysis_df, cols5, verbose=False)
df5_clean = df5[[c + '_dm' for c in cols5] + ['county_cat']].dropna()
r5_corr = fit_iv_dof_corrected(
df5_clean, 'mad_rate_dm',
exog_cols=[c + '_dm' for c in cols5[1:]],
endog_cols=None, instr_cols=None,
k_fe=k_fe, label='MODEL 5 (Reduced Form) — dof-corrected'
)
# ---- Model 6 (IV Baseline) ----
cols6 = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight','total_generator_active_capacity_mw']
df6 = iterative_demean(analysis_df, cols6, verbose=False)
df6_clean = df6[[c + '_dm' for c in cols6] + ['county_cat']].dropna()
r6_corr = fit_iv_dof_corrected(
df6_clean, 'mad_rate_dm',
exog_cols=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm'],
endog_cols=['evs_per_1000_dm'],
instr_cols=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
k_fe=k_fe, label='MODEL 6 (IV Baseline) — dof-corrected'
)
# ---- Model 7 (IV Final — preferred specification) ----
cols7 = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
'solar_mw_per_1000_lag1', 'solar_x_nem3', 'total_generator_active_capacity_mw',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']
df7 = iterative_demean(analysis_df, cols7, verbose=False)
df7_clean = df7[[c + '_dm' for c in cols7] + ['county_cat']].dropna()
r7_corr = fit_iv_dof_corrected(
df7_clean, 'mad_rate_dm',
exog_cols=['tmax_c_dm', 'cumulative_data_center_power_mw_dm',
'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm','total_generator_active_capacity_mw_dm'],
endog_cols=['evs_per_1000_dm'],
instr_cols=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
k_fe=k_fe, label='MODEL 7 (IV Final — PREFERRED) — dof-corrected'
)
# ---- Summary comparison of evs_per_1000 across Models 3, 6, 7 (original vs corrected) ----
print(f"\n{'=' * 78}\nSUMMARY: evs_per_1000 across demeaned vs native/corrected\n{'=' * 78}")
summary_rows = []
summary_rows.append({
'Model': '3 (TWFE OLS)',
'Source': 'Demeaned (original)',
'Coef': res3.params['evs_per_1000_dm'],
'SE': res3.std_errors['evs_per_1000_dm'],
'p': res3.pvalues['evs_per_1000_dm'],
})
summary_rows.append({
'Model': '3 (TWFE OLS)',
'Source': 'PanelOLS native',
'Coef': res3_native.params['evs_per_1000'],
'SE': res3_native.std_errors['evs_per_1000'],
'p': res3_native.pvalues['evs_per_1000'],
})
summary_rows.append({
'Model': '6 (IV baseline)',
'Source': 'Demeaned (original)',
'Coef': res6.params['evs_per_1000_dm'],
'SE': res6.std_errors['evs_per_1000_dm'],
'p': res6.pvalues['evs_per_1000_dm'],
})
summary_rows.append({
'Model': '6 (IV baseline)',
'Source': 'dof-corrected',
'Coef': r6_corr['params']['evs_per_1000_dm'],
'SE': r6_corr['se_corrected']['evs_per_1000_dm'],
'p': r6_corr['pvals_corrected']['evs_per_1000_dm'],
})
summary_rows.append({
'Model': '7 (IV final)',
'Source': 'Demeaned (original)',
'Coef': res7.params['evs_per_1000_dm'],
'SE': res7.std_errors['evs_per_1000_dm'],
'p': res7.pvalues['evs_per_1000_dm'],
})
summary_rows.append({
'Model': '7 (IV final)',
'Source': 'dof-corrected',
'Coef': r7_corr['params']['evs_per_1000_dm'],
'SE': r7_corr['se_corrected']['evs_per_1000_dm'],
'p': r7_corr['pvals_corrected']['evs_per_1000_dm'],
})
print(pd.DataFrame(summary_rows).to_string(index=False))
print()
print("Interpretation: if Coef matches across Source rows within a Model, the")
print("demeaning is correct. SE differences reveal the magnitude of the dof bias.")
fs6 = res6.first_stage.individual['evs_per_1000_dm']
instruments = ['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']
first_stage_table = pd.DataFrame({
'Coef': [fs6.params[i] for i in instruments],
'SE': [fs6.std_errors[i] for i in instruments],
't-stat': [fs6.tstats[i] for i in instruments],
'p-value': [fs6.pvalues[i] for i in instruments],
}, index=['Trend × Reliance', 'Anticipation × Reliance'])
print("\nModel 6 First Stage (dep var: evs_per_1000_dm):")
print(first_stage_table.round(4))
print(f"\nFirst-stage F: {res6.first_stage.diagnostics['f.stat'].iloc[0]:.2f}")
print(f"Partial R²: {res6.first_stage.diagnostics['partial.rsquared'].iloc[0]:.4f}")
fs7 = res7.first_stage.individual['evs_per_1000_dm']
instruments = ['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']
first_stage_table = pd.DataFrame({
'Coef': [fs7.params[i] for i in instruments],
'SE': [fs7.std_errors[i] for i in instruments],
't-stat': [fs7.tstats[i] for i in instruments],
'p-value': [fs7.pvalues[i] for i in instruments],
}, index=['Trend × Reliance', 'Anticipation × Reliance'])
print("\nModel 7 First Stage (dep var: evs_per_1000_dm):")
print(first_stage_table.round(4))
print(f"\nFirst-stage F: {res7.first_stage.diagnostics['f.stat'].iloc[0]:.2f}")
print(f"Partial R²: {res7.first_stage.diagnostics['partial.rsquared'].iloc[0]:.4f}")# === MODIFICATION 4 (REVISED): MODEL 8a IV HETEROGENEITY MEDIAN SPLIT ===
# MODEL 8a: IV Heterogeneity, Median Split (TWFE)
print("--- MODEL 8a: IV HETEROGENEITY (MEDIAN SPLIT) ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + total_generator_active_capacity_mw + [evs_per_1000_dm + ev_x_high_reliance_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + iv_trend_x_high_reliance_dm + iv_stock_x_high_reliance_dm]\n")
# iv_trend_break and iv_anticipation_stock_shift are purely time-varying and become
# zero columns after month demeaning — they cannot serve as instruments in TWFE.
# iv_trend_x_reliance_weight and iv_stock_x_reliance_weight vary cross-sectionally
# (by county reliance weight) and survive two-way demeaning; they instrument evs_per_1000_dm.
# iv_trend_x_high_reliance and iv_stock_x_high_reliance instrument ev_x_high_reliance_dm.
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'ev_x_high_reliance', 'tmax_c',
'cumulative_data_center_power_mw','total_generator_active_capacity_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
'iv_trend_x_high_reliance', 'iv_stock_x_high_reliance'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm', 'ev_x_high_reliance_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm',
'iv_trend_x_high_reliance_dm', 'iv_stock_x_high_reliance_dm']]
mod8a = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res8a = mod8a.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res8a.summary)
print("\n--- FIRST STAGE DIAGNOSTICS ---")
print(res8a.first_stage.diagnostics['f.stat'])
for var in endog.columns:
f_stat = res8a.first_stage.diagnostics.loc[var, 'f.stat']
if f_stat < 10:
print(f"WARNING: First-stage F-statistic for {var} is {f_stat:.4f} — below 10. Instrument may be weak.")
base_coef = res8a.params['evs_per_1000_dm']
int_coef = res8a.params['ev_x_high_reliance_dm']
p_val = res8a.pvalues['ev_x_high_reliance_dm']
tot_effect = base_coef + int_coef
conclusion = "significant at 10%" if p_val < 0.10 else "not significant"
print(f"""
Base Effect (Low Reliance Counties): {base_coef:.6f}
Interaction Effect (Differential): {int_coef:.6f}
Interaction P-Value: {p_val:.4f}
Total Effect (High Reliance Counties): {tot_effect:.6f}
Conclusion: {conclusion}
""")# === MODIFICATION 5 (REVISED): MODEL 8b IV HETEROGENEITY CONTINUOUS RELIANCE ===
# MODEL 8b: IV Heterogeneity, Continuous Reliance Weight (TWFE)
print("--- MODEL 8b: IV HETEROGENEITY (CONTINUOUS) ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm + ev_x_reliance_weight_dm ~ iv_trend_x_high_reliance_dm + iv_stock_x_high_reliance_dm + iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm]\n")
# For the continuous model: iv_trend_x_high_reliance and iv_stock_x_high_reliance
# instrument evs_per_1000_dm (binary cross-sectional variation survives TWFE);
# iv_trend_x_reliance_weight and iv_stock_x_reliance_weight instrument
# ev_x_reliance_weight_dm (continuous cross-sectional variation survives TWFE).
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'ev_x_reliance_weight', 'tmax_c',
'cumulative_data_center_power_mw','total_generator_active_capacity_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
'iv_trend_x_high_reliance', 'iv_stock_x_high_reliance',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm','solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm', 'ev_x_reliance_weight_dm']]
instr = df_clean[['iv_trend_x_high_reliance_dm', 'iv_stock_x_high_reliance_dm',
'iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]
mod8b = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res8b = mod8b.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res8b.summary)
print("\n--- FIRST STAGE DIAGNOSTICS ---")
print(res8b.first_stage.diagnostics['f.stat'])
for var in endog.columns:
f_stat = res8b.first_stage.diagnostics.loc[var, 'f.stat']
if f_stat < 10:
print(f"WARNING: First-stage F-statistic for {var} is {f_stat:.4f} — below 10. Instrument may be weak.")# === MODIFICATION 6 (REVISED): MODEL 8c IV HETEROGENEITY TERCILE SPLIT ===
# MODEL 8c: IV Heterogeneity, Tercile Split (TWFE)
print("--- MODEL 8c: IV HETEROGENEITY (TERCILE) ---")
tercile_df = analysis_df.dropna(subset=['reliance_tercile']).copy()
print(f"Current N (Excluding Middle Tercile): {len(tercile_df)}, Unique counties: {tercile_df['county_cat'].nunique()}")
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + total_generator_active_capacity_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm + ev_x_tercile_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + iv_trend_x_tercile_dm + iv_stock_x_tercile_dm]\n")
tercile_df['ev_x_tercile'] = tercile_df['evs_per_1000'] * tercile_df['reliance_tercile']
tercile_df['iv_trend_x_tercile'] = tercile_df['iv_trend_break'] * tercile_df['reliance_tercile']
tercile_df['iv_stock_x_tercile'] = tercile_df['iv_anticipation_stock_shift'] * tercile_df['reliance_tercile']
# iv_trend_x_reliance_weight and iv_stock_x_reliance_weight instrument evs_per_1000_dm;
# iv_trend_x_tercile and iv_stock_x_tercile instrument ev_x_tercile_dm.
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'ev_x_tercile', 'tmax_c',
'cumulative_data_center_power_mw','total_generator_active_capacity_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
'iv_trend_x_tercile', 'iv_stock_x_tercile'
]
df_dm = iterative_demean(tercile_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm', 'ev_x_tercile_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm',
'iv_trend_x_tercile_dm', 'iv_stock_x_tercile_dm']]
mod8c = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res8c = mod8c.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res8c.summary)
print("\n--- FIRST STAGE DIAGNOSTICS ---")
print(res8c.first_stage.diagnostics['f.stat'])
for var in endog.columns:
f_stat = res8c.first_stage.diagnostics.loc[var, 'f.stat']
if f_stat < 10:
print(f"WARNING: First-stage F-statistic for {var} is {f_stat:.4f} — below 10. Instrument may be weak.")# UPDATED COMBINED SUMMARY TABLE (includes Models 7a, 7b, and Sargan test)
def get_stars(pval):
if pval < 0.01: return '***'
elif pval < 0.05: return '**'
elif pval < 0.10: return '*'
else: return ''
def get_sargan(res, exog, instr, endog):
"""
Manual Sargan-Hansen test: n * R² from regressing 2SLS residuals
on the full instrument matrix (exog + excluded instruments).
Only valid for overidentified models (n_instruments > n_endogenous).
Returns (np.nan, np.nan) for exactly identified models.
"""
from scipy import stats as scipy_stats
try:
df_sargan = instr.shape[1] - endog.shape[1]
if df_sargan < 1:
return np.nan, np.nan # exactly identified — test undefined
e = res.resids.values
Z_full = np.column_stack([exog.values, instr.values])
n = len(e)
ZtZ_inv = np.linalg.pinv(Z_full.T @ Z_full)
e_fitted = Z_full @ (ZtZ_inv @ (Z_full.T @ e))
stat = float(n * (e_fitted @ e_fitted) / (e @ e))
pval = 1 - scipy_stats.chi2.cdf(stat, df=df_sargan)
return stat, pval
except:
return np.nan, np.nan
# Pre-compute Sargan results for overidentified models before the summary loop
sargan_results = {}
# Model 6 — re-run demeaning to recover exog/instr in scope
cols_m6 = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
'solar_mw_per_1000_lag1','total_generator_active_capacity_mw',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']
df_dm_m6 = iterative_demean(analysis_df, cols_m6)
df_c_m6 = df_dm_m6[[c+"_dm" for c in cols_m6] + ['county_cat']].dropna()
exog_m6 = df_c_m6[['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm']]
endog_m6 = df_c_m6[['evs_per_1000_dm']]
instr_m6 = df_c_m6[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]
sargan_results['Model 6: IV-2SLS Base'] = get_sargan(res6, exog_m6, instr_m6, endog_m6)
# Model 7 — re-run demeaning to recover exog/instr in scope
cols_m7 = ['mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'total_generator_active_capacity_mw',
'solar_mw_per_1000_lag1', 'solar_x_nem3',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight']
df_dm_m7 = iterative_demean(analysis_df, cols_m7)
df_c_m7 = df_dm_m7[[c+"_dm" for c in cols_m7] + ['county_cat']].dropna()
exog_m7 = df_c_m7[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'total_generator_active_capacity_mw_dm',
'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog_m7 = df_c_m7[['evs_per_1000_dm']]
instr_m7 = df_c_m7[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]
sargan_results['Model 7: IV-2SLS Final (Both)'] = get_sargan(res7, exog_m7, instr_m7, endog_m7)
models = [
('Model 1: Naive OLS', res1, 'evs_per_1000', 'None', False),
('Model 2: OLS w/ Controls', res2, 'evs_per_1000', 'None', False),
('Model 3: TWFE OLS', res3, 'evs_per_1000_dm', 'None', False),
('Model 4b: Simple DiD + County FE',res4b, 'evs_per_1000_cdm', 'None', False),
('Model 6: IV-2SLS Base', res6, 'evs_per_1000_dm', 'Trend + Anticip.', True),
('Model 7: IV-2SLS Final (Both)', res7, 'evs_per_1000_dm', 'Trend + Anticip.', True),
('Model 7a: IV-2SLS Trend Only', res7a, 'evs_per_1000_dm', 'Trend only', False),
('Model 7b: IV-2SLS Anticip. Only', res7b, 'evs_per_1000_dm', 'Anticip. only', False),
('Model 8a: IV Het (Base)', res8a, 'evs_per_1000_dm', 'Trend + Anticip. (×reliance)', True),
('Model 8a: IV Het (Interaction)', res8a, 'ev_x_high_reliance_dm','Trend + Anticip. (×reliance)', True),
]
# Note: Models 6 and 7 are overidentified (2 instruments, 1 endogenous var) -> Sargan available
# Models 7a and 7b are exactly identified (1 instrument, 1 endogenous var) -> Sargan undefined
# Model 8a has 4 instruments and 2 endogenous vars -> overidentified by 2 -> Sargan available
# OLS models have no instruments -> Sargan undefined
summary_data = []
for name, res, var, iv_label, is_iv in models:
# Second-stage coefficient, SE, p-value
coef = res.params.get(var, np.nan)
se = res.std_errors.get(var, np.nan)
pval = res.pvalues.get(var, np.nan)
stars = get_stars(pval) if not pd.isna(pval) else ''
# First-stage F-statistic (IV models only)
fstat = np.nan
if is_iv:
try:
idx = var if var in res.first_stage.diagnostics.index \
else res.first_stage.diagnostics.index[0]
fstat = res.first_stage.diagnostics.loc[idx, 'f.stat']
except:
pass
# Sargan-Hansen overidentification test
# Only computed for overidentified models; exactly identified -> n_instr == n_endog
# We check overidentification by catching the test result
sargan_stat, sargan_pval = np.nan, np.nan
if is_iv:
sargan_stat, sargan_pval = sargan_results.get(name, (np.nan, np.nan))
summary_data.append({
'Model': name,
'N': int(res.nobs),
'EV Coef': f"{coef:.5f}{stars}" if not pd.isna(coef) else '-',
'Std. Err.': f"({se:.5f})" if not pd.isna(se) else '-',
'P-value': f"{pval:.3f}" if not pd.isna(pval) else '-',
'F-stat (1st)': f"{fstat:.2f}" if not pd.isna(fstat) else '-',
'Sargan stat': f"{sargan_stat:.4f}" if not pd.isna(sargan_stat) else 'n/a',
'Sargan p': f"{sargan_pval:.4f}" if not pd.isna(sargan_pval) else 'n/a',
'Instruments': iv_label,
})
sum_df = pd.DataFrame(summary_data)
print("=" * 130)
print("COMBINED SUMMARY TABLE — DEPENDENT VARIABLE: mad_rate")
print("=" * 130)
print(sum_df.to_string(index=False))
print("\nSignificance: * p<0.10 ** p<0.05 *** p<0.01")
print("\nNotes:")
print(" - All shift-share instruments defined as: time_shock × reliance_weight")
print(" Trend IV = months_since_CVRP_end × reliance_weight")
print(" Anticip IV = anticipation_phase_dummy × reliance_weight")
print(" - Sargan-Hansen test H0: instruments are valid (model not overidentified).")
print(" A p-value > 0.10 fails to reject H0 — instruments pass the validity test.")
print(" 'n/a' indicates exactly identified model (test undefined) or OLS (no instruments).")
print(" - Model 4b uses county FE only by design — post_cvrp would be absorbed by time FE.")# === MODIFICATION 9 (REVISED): STEP 6 NEM 3.0 SENSITIVITY ANALYSIS ===
print("--- STEP 6: NEM 3.0 SENSITIVITY ANALYSIS ---")
candidate_dates = pd.date_range(start='2023-04-01', end='2024-04-01', freq='MS')
results = []
for d in candidate_dates:
temp_df = analysis_df.copy()
temp_df['post_nem3_sens'] = (temp_df['month'] >= d).astype(int)
temp_df['solar_x_nem3_sens'] = temp_df['solar_mw_per_1000_lag1'] * temp_df['post_nem3_sens']
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw',
'solar_mw_per_1000_lag1', 'solar_x_nem3_sens',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight'
]
df_dm = iterative_demean(temp_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm',
'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_sens_dm']]
endog = df_clean[['evs_per_1000_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm']]
mod = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res = mod.fit(cov_type='clustered', clusters=df_clean['county_cat'])
results.append({
'Date': d,
'Coef': res.params['evs_per_1000_dm'],
'SE': res.std_errors['evs_per_1000_dm'],
'P-val': res.pvalues['evs_per_1000_dm']
})
sens_df = pd.DataFrame(results)
plt.figure(figsize=(10, 5))
plt.plot(sens_df['Date'], sens_df['Coef'], marker='o', color='blue', label='EV Coefficient')
plt.fill_between(sens_df['Date'],
sens_df['Coef'] - 1.645 * sens_df['SE'],
sens_df['Coef'] + 1.645 * sens_df['SE'],
color='blue', alpha=0.2, label='90% CI')
plt.axhline(0, color='black', linestyle='--')
plt.title('NEM 3.0 Structural Break Sensitivity Analysis')
plt.xlabel('Candidate Break Month')
plt.ylabel('EV Coefficient on MAD Rate')
plt.legend()
plt.show()
best_fit = sens_df.loc[sens_df['SE'].idxmin()]
print(f"Month with best model fit (lowest standard error): {best_fit['Date'].strftime('%Y-%m')}")
print("\nNOTE: Confirm visually from the plot whether the EV coefficient remains stable (confidence band excludes zero) across the tested break dates.")# STEP 7A: POLITICAL HETEROGENEITY DATA PREP
print("--- STEP 7A: POLITICAL DATA PROCESSING ---")
try:
med_dem = analysis_df[['county_name', 'dem_share']].drop_duplicates()['dem_share'].median()
analysis_df['is_high_dem'] = (analysis_df['dem_share'] > med_dem).astype(int)
analysis_df['ev_x_high_dem'] = analysis_df['evs_per_1000'] * analysis_df['is_high_dem']
analysis_df['iv_trend_x_high_dem'] = analysis_df['iv_trend_break'] * analysis_df['is_high_dem']
analysis_df['iv_stock_x_high_dem'] = analysis_df['iv_anticipation_stock_shift'] * analysis_df['is_high_dem']
print(f"Median Democratic Share: {med_dem:.4f}")
print(f"Counties coded as High Dem: {len(analysis_df[analysis_df['is_high_dem']==1]['county_name'].unique())}")
except Exception as e:
print(f"ERROR processing political data: {e}")# === MODIFICATION 7 (REVISED): STEP 7B POLITICAL HETEROGENEITY MODEL ===
# STEP 7B: POLITICAL HETEROGENEITY MODEL (TWFE)
print("--- STEP 7B: POLITICAL HETEROGENEITY MODEL ---")
try:
print("Formula: mad_rate_dm ~ tmax_c_dm + cumulative_data_center_power_mw_dm + solar_mw_per_1000_lag1_dm + solar_x_nem3_dm + [evs_per_1000_dm + ev_x_high_dem_dm ~ iv_trend_x_reliance_weight_dm + iv_stock_x_reliance_weight_dm + iv_trend_x_high_dem_dm + iv_stock_x_high_dem_dm]\n")
# iv_trend_x_reliance_weight and iv_stock_x_reliance_weight instrument evs_per_1000_dm;
# iv_trend_x_high_dem and iv_stock_x_high_dem instrument ev_x_high_dem_dm.
cols_to_demean = [
'mad_rate', 'evs_per_1000', 'ev_x_high_dem', 'tmax_c',
'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1', 'solar_x_nem3',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
'iv_trend_x_high_dem', 'iv_stock_x_high_dem'
]
df_dm = iterative_demean(analysis_df, cols_to_demean)
df_clean = df_dm[[c + "_dm" for c in cols_to_demean] + ['county_cat']].dropna()
exog = df_clean[['tmax_c_dm', 'cumulative_data_center_power_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm']]
endog = df_clean[['evs_per_1000_dm', 'ev_x_high_dem_dm']]
instr = df_clean[['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm',
'iv_trend_x_high_dem_dm', 'iv_stock_x_high_dem_dm']]
mod_pol = IV2SLS(dependent=df_clean['mad_rate_dm'], exog=exog, endog=endog, instruments=instr)
res_pol = mod_pol.fit(cov_type='clustered', clusters=df_clean['county_cat'])
print(res_pol.summary)
print("\n--- FIRST STAGE DIAGNOSTICS ---")
print(res_pol.first_stage.diagnostics['f.stat'])
for var in endog.columns:
f_stat = res_pol.first_stage.diagnostics.loc[var, 'f.stat']
if f_stat < 10:
print(f"WARNING: First-stage F-statistic for {var} is {f_stat:.4f} — below 10. Instrument may be weak.")
base_coef = res_pol.params['evs_per_1000_dm']
int_coef = res_pol.params['ev_x_high_dem_dm']
p_val = res_pol.pvalues['ev_x_high_dem_dm']
tot_effect = base_coef + int_coef
sig_label = "significant at 10%" if p_val < 0.10 else "not significant"
if p_val < 0.10:
pol_conclusion = f"{sig_label} — EV grid impacts are concentrated in politically distinct county groups, suggesting the burden or benefit is not distributed uniformly across the political geography of California."
else:
pol_conclusion = f"{sig_label} — EV grid impacts do not appear to differ systematically by county-level political affiliation; the evidence does not support a politically concentrated distribution of grid volatility effects."
print(f"""
Base Effect (Low Dem / Conservative Counties): {base_coef:.6f}
Interaction Effect (Differential): {int_coef:.6f}
Interaction P-Value: {p_val:.4f}
Total Effect (High Dem / Liberal Counties): {tot_effect:.6f}
Conclusion: {pol_conclusion}
""")
except Exception as e:
print(f"ERROR running political heterogeneity model: {e}")# ============================================================
# CVRP DAILY APPLICATIONS — EVENT CHART
# Visualization style: Sosulski (clean, annotated, purposeful)
# ============================================================
# ---- EASILY EDITABLE PARAMETERS ----
PLOT_START = '2022-12-31' # Start of date range to display
PLOT_END = '2023-12-31' # End of date range to display
ROLLING_WINDOW = 1 # Days for rolling average smoothing (set to 1 for raw daily)
# Key event dates
DATE_ANNOUNCE = '2023-08-21' # CVRP closure announced
DATE_STANDBY = '2023-09-06' # Applications placed on standby
DATE_CLOSURE = '2023-11-08' # Effective closure
# Colors — edit here to restyle the entire chart
COLOR_RAW = '#C8D8E8' # Raw daily bars (muted blue-grey)
COLOR_ROLLING = '#1A5276' # Rolling average line (dark navy)
COLOR_SHADE = '#F5CBA7' # Anticipation window shading (warm amber)
COLOR_ANNOUNCE = '#E74C3C' # Announcement line (red)
COLOR_STANDBY = '#E74C3C' # Standby line (purple)
COLOR_CLOSURE = '#E74C3C' # Closure line (green)
# ---- END EDITABLE PARAMETERS ----
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import matplotlib.ticker as mticker
import warnings
warnings.filterwarnings('ignore')
# ---- BUILD DAILY APPLICATION COUNTS ----
# Ensure app_date column exists; re-parse if needed
cvrp_plot = cvrp_df.copy()
cvrp_plot['app_date'] = pd.to_datetime(cvrp_plot['Application Date'], errors='coerce')
cvrp_plot = cvrp_plot.dropna(subset=['app_date'])
# Aggregate to daily counts
daily = (
cvrp_plot
.groupby('app_date')
.size()
.reset_index(name='applications')
.rename(columns={'app_date': 'date'})
)
# Filter to plot range
daily = daily[
(daily['date'] >= PLOT_START) &
(daily['date'] <= PLOT_END)
].copy()
# Fill missing dates with zero (no applications = 0, not missing)
full_range = pd.DataFrame({'date': pd.date_range(PLOT_START, PLOT_END)})
daily = full_range.merge(daily, on='date', how='left').fillna(0)
# Rolling average
daily['rolling'] = daily['applications'].rolling(window=ROLLING_WINDOW, center=True).mean()
# Convert key dates
d_announce = pd.Timestamp(DATE_ANNOUNCE)
d_standby = pd.Timestamp(DATE_STANDBY)
d_closure = pd.Timestamp(DATE_CLOSURE)
# ---- CHART ----
fig, ax = plt.subplots(figsize=(14, 6))
fig.patch.set_facecolor('white')
ax.set_facecolor('white')
# Shade the anticipation window (announcement → standby)
ax.axvspan(d_announce, d_closure, color=COLOR_SHADE, alpha=0.35, zorder=1,
label='Anticipation Window')
# Raw daily bars — thin, muted, background layer
ax.bar(daily['date'], daily['applications'],
color=COLOR_RAW, width=1.0, alpha=0.6, zorder=2, label='Daily Applications')
# Rolling average — foreground signal
ax.plot(daily['date'], daily['rolling'],
color=COLOR_ROLLING, linewidth=2.2, zorder=3,
label=f'{ROLLING_WINDOW}-Day Rolling Average')
# ---- VERTICAL EVENT LINES ----
line_ymax = daily['applications'].max() * 1.05
for date, color, ls in [
(d_announce, COLOR_ANNOUNCE, '--'),
(d_standby, COLOR_STANDBY, '-'),
(d_closure, COLOR_CLOSURE, ':'),
]:
ax.axvline(date, color=color, linewidth=1.6, linestyle=ls, zorder=4, alpha=0.9)
# ---- ANNOTATIONS (Sosulski style: horizontal callout with leader line) ----
annotation_cfg = [
dict(
date = d_announce,
label = "Closure\nAnnounced\nAug 21, 2023",
color = COLOR_ANNOUNCE,
x_off = -2, # days offset for text box (negative = left)
y_pos = 0.7, # fraction of y-axis height
ha = 'right',
),
dict(
date = d_standby,
label = "Application\nStandby\nSep 6, 2023",
color = COLOR_STANDBY,
x_off = 2,
y_pos = 0.7,
ha = 'left',
),
dict(
date = d_closure,
label = "Effective\nClosure\nNov 8, 2023",
color = COLOR_CLOSURE,
x_off = 2,
y_pos = 0.7,
ha = 'left',
),
]
y_top = daily['applications'].max()
for cfg in annotation_cfg:
text_x = cfg['date'] + pd.Timedelta(days=cfg['x_off'])
text_y = y_top * cfg['y_pos']
ax.annotate(
cfg['label'],
xy = (cfg['date'], text_y * 0.80), # arrow tip on the line
xytext = (text_x, text_y), # text box position
color = cfg['color'],
fontsize = 12,
fontweight= 'semibold',
ha = cfg['ha'],
va = 'top',
arrowprops= dict(
arrowstyle = '-',
color = cfg['color'],
lw = 1.2,
linestyle = 'dashed',
),
bbox = dict(
boxstyle = 'round,pad=0.25',
facecolor = 'white',
edgecolor = cfg['color'],
linewidth = 1.0,
alpha = 0.9,
),
zorder = 5,
)
# ---- SPIKE CALLOUT ----
# Find the peak day and annotate it directly
#peak_row = daily.loc[daily['applications'].idxmax()]
#peak_date = peak_row['date']
#peak_val = peak_row['applications']
#ax.annotate(
# f"Peak: {int(peak_val):,} applications\n{peak_date.strftime('%b %d, %Y')}",
# xy = (peak_date, peak_val),
# xytext = (peak_date - pd.Timedelta(days=18), peak_val * 0.88),
# fontsize = 12,
# color = '#1A1A1A',
# ha = 'right',
# va = 'top',
# arrowprops= dict(
# arrowstyle = '-|>',
# color = '#555555',
# lw = 1.2,
# ),
# bbox = dict(
# boxstyle = 'round,pad=0.3',
# facecolor = '#FDFEFE',
# edgecolor = '#AAAAAA',
# linewidth = 0.8,
# alpha = 0.9,
# ),
# zorder = 6,
#)
# ---- AXES STYLING (Sosulski: remove top/right spines, minimal grid) ----
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
ax.spines['left'].set_color('#CCCCCC')
ax.spines['bottom'].set_color('#CCCCCC')
ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, _: f'{int(x):,}'))
import matplotlib.dates as mdates
ax.xaxis.set_major_locator(mdates.MonthLocator())
ax.xaxis.set_major_formatter(mdates.DateFormatter('%b-%y'))
ax.tick_params(axis='both', labelsize=9, color='#AAAAAA')
ax.yaxis.grid(True, color='#EEEEEE', linewidth=0.8, zorder=0)
ax.set_axisbelow(True)
# ---- LABELS ----
# ax.set_xlabel('Application Date', fontsize=10, color='#444444', labelpad=8)
# ax.set_ylabel('Daily Applications', fontsize=10, color='#444444', labelpad=8)
#ax.set_title(
# 'CVRP Daily Applications — Behavioral Surge Around Program Closure',
# fontsize=13, fontweight='bold', color='#1A1A1A', pad=14, loc='left'
#)
# Subtitle
#fig.text(
# 0.013, 0.91,
# f'Shaded region = anticipation window (Aug 21 – Sep 6, 2023) | '
# f'{ROLLING_WINDOW}-day rolling average overlaid on raw daily counts',
# fontsize=8.5, color='#777777'
#)
# ---- LEGEND ----
#handles = [
# mpatches.Patch(color=COLOR_RAW, alpha=0.6, label='Daily Applications'),
# plt.Line2D([0], [0], color=COLOR_ROLLING, lw=2.2, label=f'{ROLLING_WINDOW}-Day Rolling Avg'),
# mpatches.Patch(color=COLOR_SHADE, alpha=0.35, label='Anticipation Window'),
# plt.Line2D([0], [0], color=COLOR_ANNOUNCE, lw=1.6, ls='--', label='Closure Announced'),
# plt.Line2D([0], [0], color=COLOR_STANDBY, lw=1.6, ls='-', label='Applications on Standby'),
# plt.Line2D([0], [0], color=COLOR_CLOSURE, lw=1.6, ls=':', label='Effective Closure'),
#]
#ax.legend(
# handles = handles,
# loc = 'upper left',
# fontsize = 8.5,
# frameon = True,
# framealpha= 0.9,
# edgecolor = '#DDDDDD',
# ncol = 2,
#)
ax.set_xlim(pd.Timestamp(PLOT_START), pd.Timestamp(PLOT_END))
plt.show()
print(f"\nDate range plotted: {PLOT_START} to {PLOT_END}")
# print(f"Peak day: {peak_date.strftime('%Y-%m-%d')} with {int(peak_val):,} applications")
print(f"Rolling window: {ROLLING_WINDOW} days")# ============================================================
# THE POLICY CLIFF — MONTHLY EV ADOPTION
# Visualization style: Sosulski (clean, annotated, purposeful)
# ============================================================
# ---- EASILY EDITABLE PARAMETERS ----
PLOT_START = '2022-05-01' # Start of date range
PLOT_END = '2025-05-01' # End of date range
# Key event date
DATE_ANNOUNCE = '2023-08-01' # CVRP closure announced (month-level)
# Colors
COLOR_LINE = '#1A3A4A' # Main line color (dark navy)
COLOR_FILL = '#E8F4F8' # Area fill under line (light blue)
COLOR_ANNOUNCE = '#E74C3C' # Announcement line (red)
COLOR_ANNOT_BG = 'white' # Annotation box background
# ---- END EDITABLE PARAMETERS ----
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
import matplotlib.ticker as mticker
import warnings
warnings.filterwarnings('ignore')
# ---- BUILD MONTHLY EV TOTALS ----
monthly_ev = (
analysis_df
.groupby('month')
.agg(
total_new_evs = ('monthly_new_evs', 'sum')
)
.reset_index()
)
# Filter to plot range
monthly_ev = monthly_ev[
(monthly_ev['month'] >= PLOT_START) &
(monthly_ev['month'] <= PLOT_END)
].copy()
# Key date
d_announce = pd.Timestamp(DATE_ANNOUNCE)
# ---- CHART ----
fig, ax = plt.subplots(figsize=(10, 6))
fig.patch.set_facecolor('white')
ax.set_facecolor('white')
# Area fill under the line
ax.fill_between(
monthly_ev['month'],
monthly_ev['total_new_evs'],
color=COLOR_FILL,
alpha=1.0,
zorder=1
)
# Main line
ax.plot(
monthly_ev['month'],
monthly_ev['total_new_evs'],
color=COLOR_LINE,
linewidth=2.0,
zorder=2
)
# ---- VERTICAL EVENT LINE ----
ax.axvline(
d_announce,
color = COLOR_ANNOUNCE,
linewidth = 1.6,
linestyle = '--',
zorder = 3,
alpha = 0.9
)
# ---- ANNOTATION ----
y_top = monthly_ev['total_new_evs'].max()
text_x = d_announce - pd.Timedelta(days=10)
text_y = y_top * 0.88
ax.annotate(
"Closure\nAnnounced\nAug, 2023",
xy = (d_announce, text_y * 0.78),
xytext = (text_x, text_y),
color = COLOR_ANNOUNCE,
fontsize = 11,
fontweight= 'semibold',
ha = 'right',
va = 'top',
arrowprops= dict(
arrowstyle = '-',
color = COLOR_ANNOUNCE,
lw = 1.2,
linestyle = 'dashed',
),
bbox = dict(
boxstyle = 'round,pad=0.25',
facecolor = COLOR_ANNOT_BG,
edgecolor = COLOR_ANNOUNCE,
linewidth = 1.0,
alpha = 0.9,
),
zorder = 4,
)
# ---- AXES STYLING ----
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
ax.spines['left'].set_color('#CCCCCC')
ax.spines['bottom'].set_color('#CCCCCC')
ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, _: f'{int(x):,}'))
ax.xaxis.set_major_locator(mdates.MonthLocator(interval=4))
ax.xaxis.set_major_formatter(mdates.DateFormatter('%b %Y'))
ax.tick_params(axis='both', labelsize=9, color='#AAAAAA')
ax.yaxis.grid(True, color='#EEEEEE', linewidth=0.8, zorder=0)
ax.set_axisbelow(True)
# Force x-axis to match plot range exactly
ax.set_xlim(pd.Timestamp(PLOT_START), pd.Timestamp(PLOT_END))
# Floor y-axis at zero
ax.set_ylim(bottom=0)
# ---- LABELS ----
ax.set_xlabel('', labelpad=8)
ax.set_ylabel(
'New EVs Registered (Monthly, All Counties)',
fontsize = 10,
color = '#444444',
labelpad = 8
)
#ax.set_title(
# 'The Policy Cliff: Monthly EV Adoption',
# fontsize = 13,
# fontweight= 'bold',
# color = '#1A1A1A',
# pad = 14,
# loc = 'left'
#)
plt.tight_layout()
plt.show()
print(f"\nDate range plotted: {PLOT_START} to {PLOT_END}")
print(f"Peak month: {monthly_ev.loc[monthly_ev['total_new_evs'].idxmax(), 'month'].strftime('%Y-%m')} "
f"with {monthly_ev['total_new_evs'].max():,.0f} new EVs")
print(f"Announcement month ({DATE_ANNOUNCE}): "
f"{monthly_ev.loc[monthly_ev['month'] == DATE_ANNOUNCE, 'total_new_evs'].values[0]:,.0f} new EVs"
if DATE_ANNOUNCE in monthly_ev['month'].astype(str).values
else f"Announcement date {DATE_ANNOUNCE} not in plotted range")# ============================================================
# THE NEM 2.0 RUSH — MONTHLY SOLAR CAPACITY ADDITIONS
# Visualization style: Sosulski (clean, annotated, purposeful)
# ============================================================
# ---- EASILY EDITABLE PARAMETERS ----
PLOT_START = '2022-05-01' # Start of date range
PLOT_END = '2025-05-01' # End of date range
# Key NEM 3.0 event dates
DATE_NEM3_APPROVE = '2022-12-01' # CPUC approves NEM 3.0
DATE_NEM3_DEADLINE= '2023-04-01' # NEM 2.0 grandfathering deadline (Apr 14)
DATE_BACKLOG = '2024-01-01' # NEM 2.0 backlog physically connects to grid
# Colors
COLOR_LINE = '#1A3A4A' # Main line (dark navy)
COLOR_FILL = '#E8F4F8' # Area fill (light blue)
COLOR_NEM_APPROVE = '#E67E22' # Approval line (orange)
COLOR_NEM_DEAD = '#E74C3C' # Deadline line (red)
COLOR_BACKLOG = '#1E8449' # Backlog line (green)
COLOR_ANNOT_BG = 'white'
# ---- END EDITABLE PARAMETERS ----
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
import matplotlib.ticker as mticker
import warnings
warnings.filterwarnings('ignore')
# ---- BUILD MONTHLY NEW SOLAR CAPACITY (TOTAL MW, SUM ACROSS COUNTIES) ----
# Step 1: recover total MW from per-1000 variable by multiplying back by population
# solar_mw_per_1000_lag1 × total_population / 1000 = total MW (lagged)
df_solar = analysis_df[['county_name', 'month',
'solar_mw_per_1000_lag1',
'total_population']].copy()
df_solar['solar_mw_total'] = (
df_solar['solar_mw_per_1000_lag1'] * df_solar['total_population'] / 1000
)
# Step 2: sum across all counties to get state-level cumulative (still lagged)
solar_monthly = (
df_solar
.groupby('month')
.agg(solar_cumulative_lagged = ('solar_mw_total', 'sum'))
.reset_index()
.sort_values('month')
)
# Step 3: undo the 1-month lag to recover the current-month cumulative
solar_monthly['solar_cumulative'] = solar_monthly['solar_cumulative_lagged'].shift(-1)
# Step 4: first difference → new MW installed each month
solar_monthly['new_solar_mw'] = solar_monthly['solar_cumulative'].diff()
# Step 5: clip negatives (data artifacts) and filter to plot range
solar_monthly['new_solar_mw'] = solar_monthly['new_solar_mw'].clip(lower=0)
solar_monthly = solar_monthly[
(solar_monthly['month'] >= PLOT_START) &
(solar_monthly['month'] <= PLOT_END)
].dropna(subset=['new_solar_mw']).copy()
# Sanity check
print(f"Months in series: {len(solar_monthly)}")
print(f"Peak month: {solar_monthly.loc[solar_monthly['new_solar_mw'].idxmax(), 'month'].strftime('%Y-%m')} "
f"— {solar_monthly['new_solar_mw'].max():,.1f} MW")
print(f"Mean new MW/month: {solar_monthly['new_solar_mw'].mean():,.1f} MW")
# Key dates
d_approve = pd.Timestamp(DATE_NEM3_APPROVE)
d_deadline = pd.Timestamp(DATE_NEM3_DEADLINE)
d_backlog = pd.Timestamp(DATE_BACKLOG)
# ---- CHART ----
fig, ax = plt.subplots(figsize=(10, 6))
fig.patch.set_facecolor('white')
ax.set_facecolor('white')
# Shade the NEM 2.0 rush window (approval → deadline)
ax.axvspan(
d_deadline, d_backlog,
color='#FDEBD0', alpha=0.5, zorder=1
)
# Area fill
ax.fill_between(
solar_monthly['month'],
solar_monthly['new_solar_mw'],
color=COLOR_FILL,
alpha=1.0,
zorder=2
)
# Main line
ax.plot(
solar_monthly['month'],
solar_monthly['new_solar_mw'],
color=COLOR_LINE,
linewidth=2.0,
zorder=3
)
# ---- VERTICAL EVENT LINES ----
for date, color, ls in [
# (d_approve, COLOR_NEM_APPROVE, '--'),
(d_deadline, COLOR_NEM_DEAD, '-'),
# (d_backlog, COLOR_BACKLOG, ':'),
]:
ax.axvline(date, color=color, linewidth=1.6, linestyle=ls, zorder=4, alpha=0.9)
# ---- ANNOTATIONS ----
y_top = solar_monthly['new_solar_mw'].max()
annotation_cfg = [
# dict(
# date = d_approve,
# label = "NEM 3.0\nApproved\nDec 2022",
# color = COLOR_NEM_APPROVE,
# x_off = -15,
# y_pos = 0.98,
# ha = 'right',
# ),
dict(
date = d_deadline,
label = "Application\nDeadline\nApr 14, 2023",
color = COLOR_NEM_DEAD,
x_off = -10,
y_pos = 0.92,
ha = 'right',
),
# dict(
# date = d_backlog,
# label = "Backlog Connects\nto Grid\nJan 2024",
# color = COLOR_BACKLOG,
# x_off = 2,
# y_pos = 0.98,
# ha = 'left',
# ),
]
for cfg in annotation_cfg:
text_x = cfg['date'] + pd.Timedelta(days=cfg['x_off'])
text_y = y_top * cfg['y_pos']
ax.annotate(
cfg['label'],
xy = (cfg['date'], text_y * 0.80),
xytext = (text_x, text_y),
color = cfg['color'],
fontsize = 11,
fontweight= 'semibold',
ha = cfg['ha'],
va = 'top',
arrowprops= dict(
arrowstyle = '-',
color = cfg['color'],
lw = 1.2,
linestyle = 'dashed',
),
bbox = dict(
boxstyle = 'round,pad=0.25',
facecolor = COLOR_ANNOT_BG,
edgecolor = cfg['color'],
linewidth = 1.0,
alpha = 0.9,
),
zorder = 5,
)
# ---- PEAK CALLOUT ----
peak_row = solar_monthly.loc[solar_monthly['new_solar_mw'].idxmax()]
peak_date = peak_row['month']
peak_val = peak_row['new_solar_mw']
ax.annotate(
f"Peak: {peak_val:,.0f} MW\n{peak_date.strftime('%b %Y')}",
xy = (peak_date, peak_val),
xytext = (peak_date + pd.Timedelta(days=45), peak_val * 0.88),
fontsize = 12,
color = '#1A1A1A',
ha = 'left',
va = 'top',
arrowprops= dict(
arrowstyle = '-|>',
color = '#555555',
lw = 1.2,
),
bbox = dict(
boxstyle = 'round,pad=0.3',
facecolor = '#FDFEFE',
edgecolor = '#AAAAAA',
linewidth = 0.8,
alpha = 0.9,
),
zorder = 6,
)
# ---- AXES STYLING ----
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
ax.spines['left'].set_color('#CCCCCC')
ax.spines['bottom'].set_color('#CCCCCC')
ax.yaxis.set_major_formatter(mticker.FuncFormatter(lambda x, _: f'{x:,.0f}'))
ax.xaxis.set_major_locator(mdates.MonthLocator(interval=4))
ax.xaxis.set_major_formatter(mdates.DateFormatter('%b %Y'))
ax.tick_params(axis='both', labelsize=9, color='#AAAAAA')
ax.yaxis.grid(True, color='#EEEEEE', linewidth=0.8, zorder=0)
ax.set_axisbelow(True)
ax.set_xlim(pd.Timestamp(PLOT_START), pd.Timestamp(PLOT_END))
ax.set_ylim(bottom=0)
# ---- LABELS ----
ax.set_xlabel('', labelpad=8)
ax.set_ylabel(
'New Solar Capacity Installed — Total MW',
fontsize = 12,
color = '#444444',
labelpad = 8
)
#ax.set_title(
# 'The NEM 2.0 Rush: Monthly Solar Capacity Additions',
# fontsize = 13,
# fontweight= 'bold',
# color = '#1A1A1A',
# pad = 14,
# loc = 'left'
#)
plt.tight_layout()
plt.show()# ============================================================
# CALIFORNIA COUNTY MAP — CVRP LEGACY RELIANCE RATE
# Visualization style: Sosulski (clean, purposeful, annotated)
# ============================================================
# ---- EASILY EDITABLE PARAMETERS ----
COLORMAP = 'Blues' # Matplotlib colormap — try 'Blues', 'RdYlGn_r', 'Oranges', 'YlOrRd'
FIGSIZE = (10, 12) # Figure dimensions
LABEL_COUNTIES = True # Show county name labels on map
LABEL_MIN_SIZE = 6 # Minimum font size for labels
SHOW_EXCLUDED = True # Show excluded counties (NaN reliance) in gray
COLOR_EXCLUDED = '#D5D8DC' # Color for excluded/NaN counties
COLOR_BORDER = 'white' # County border color
BORDER_WIDTH = 0.5 # County border line width
# ---- END EDITABLE PARAMETERS ----
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import matplotlib.patches as mpatches
import matplotlib.ticker as mticker
from matplotlib.colors import Normalize
from matplotlib.cm import ScalarMappable
import warnings
warnings.filterwarnings('ignore')
# ---- INSTALL AND IMPORT GEOPANDAS ----
try:
import geopandas as gpd
except ImportError:
import subprocess
subprocess.run(['pip', 'install', 'geopandas', '-q'], check=True)
import geopandas as gpd
# ---- DOWNLOAD CALIFORNIA COUNTY SHAPEFILE FROM CENSUS BUREAU ----
import urllib.request, zipfile, os
shp_url = "https://www2.census.gov/geo/tiger/GENZ2022/shp/cb_2022_us_county_5m.zip"
shp_zip = "/tmp/us_counties.zip"
shp_dir = "/tmp/us_counties"
if not os.path.exists(shp_dir):
print("Downloading US county shapefile from Census Bureau...")
urllib.request.urlretrieve(shp_url, shp_zip)
with zipfile.ZipFile(shp_zip, 'r') as z:
z.extractall(shp_dir)
print("Download complete.")
shp_file = [f for f in os.listdir(shp_dir) if f.endswith('.shp')][0]
counties = gpd.read_file(os.path.join(shp_dir, shp_file))
# Filter to California (FIPS = 06)
ca_counties = counties[counties['STATEFP'] == '06'].copy()
ca_counties['county_name'] = ca_counties['NAME'].str.strip().str.title()
print(f"California counties in shapefile: {len(ca_counties)}")
# ---- BUILD RELIANCE DATA ----
# Get one row per county from analysis_df
reliance_data = (
analysis_df[['county_name', 'reliance_weight', 'is_high_reliance']]
.drop_duplicates(subset='county_name')
.copy()
)
# Also pull from county_reliance to get excluded counties
all_reliance = county_reliance[['county_name', 'reliance_weight']].copy()
print(f"Counties in reliance data: {len(reliance_data)}")
print(f"Counties in locked sample: {reliance_data['reliance_weight'].notna().sum()}")
# ---- MERGE SHAPEFILE WITH RELIANCE DATA ----
# Merge locked sample first
ca_map = ca_counties.merge(
reliance_data[['county_name', 'reliance_weight', 'is_high_reliance']],
on='county_name',
how='left'
)
# For counties excluded from locked sample but present in county_reliance,
# fill in their reliance weight so they still show data
ca_map = ca_map.merge(
all_reliance.rename(columns={'reliance_weight': 'reliance_all'}),
on='county_name',
how='left'
)
# Use locked sample weight where available, otherwise use all_reliance
ca_map['plot_reliance'] = ca_map['reliance_weight'].fillna(ca_map['reliance_all'])
# Flag: in locked sample, excluded by threshold, or not in CVRP at all
ca_map['status'] = 'no_data'
ca_map.loc[ca_map['reliance_weight'].notna(), 'status'] = 'included'
ca_map.loc[
ca_map['reliance_weight'].isna() & ca_map['reliance_all'].notna(),
'status'
] = 'excluded'
print(f"\nMap county status breakdown:")
print(ca_map['status'].value_counts().to_string())
# ---- COMPUTE CENTROID FOR LABELS ----
ca_map = ca_map.copy()
ca_map['centroid_x'] = ca_map.geometry.centroid.x
ca_map['centroid_y'] = ca_map.geometry.centroid.y
# ---- CHART ----
fig, ax = plt.subplots(figsize=FIGSIZE)
fig.patch.set_facecolor('white')
ax.set_facecolor('white')
# Normalize colormap to reliance range (included counties only)
vmin = ca_map.loc[ca_map['status'] == 'included', 'plot_reliance'].min()
vmax = ca_map.loc[ca_map['status'] == 'included', 'plot_reliance'].max()
norm = Normalize(vmin=vmin, vmax=vmax)
cmap = plt.get_cmap(COLORMAP)
# ---- PLOT INCLUDED COUNTIES (colored by reliance) ----
included = ca_map[ca_map['status'] == 'included'].copy()
included.plot(
column = 'plot_reliance',
ax = ax,
cmap = COLORMAP,
norm = norm,
linewidth = BORDER_WIDTH,
edgecolor = COLOR_BORDER,
zorder = 2
)
# ---- PLOT EXCLUDED COUNTIES (gray, below threshold) ----
if SHOW_EXCLUDED:
excluded = ca_map[ca_map['status'] == 'excluded'].copy()
if len(excluded) > 0:
excluded.plot(
ax = ax,
color = COLOR_EXCLUDED,
linewidth = BORDER_WIDTH,
edgecolor = COLOR_BORDER,
zorder = 2
)
# ---- PLOT NO-DATA COUNTIES (light gray) ----
no_data = ca_map[ca_map['status'] == 'no_data'].copy()
if len(no_data) > 0:
no_data.plot(
ax = ax,
color = '#F2F3F4',
linewidth = BORDER_WIDTH,
edgecolor = COLOR_BORDER,
zorder = 2
)
'''
# ---- COUNTY LABELS ----
if LABEL_COUNTIES:
for _, row in ca_map.iterrows():
if pd.isna(row['plot_reliance']) and row['status'] == 'no_data':
continue
# Shorten long county names for readability
name = row['county_name']
short = (name
.replace(' County', '')
.replace('San ', 'S. ')
.replace('Santa ', 'Sta. ')
.replace('Los Angeles', 'L.A.')
.replace('San Francisco', 'S.F.'))
# Scale font size by county area (larger counties get bigger labels)
area = row.geometry.area
fsize = np.clip(6 + np.log10(area + 1) * 0.4, LABEL_MIN_SIZE, 8)
# Bold label for high-reliance counties
weight = 'bold' if row.get('is_high_reliance', 0) == 1 else 'normal'
color = 'white' if (
pd.notna(row['plot_reliance']) and
row['plot_reliance'] > vmin + 0.65 * (vmax - vmin)
) else '#1A1A1A'
ax.text(
row['centroid_x'], row['centroid_y'],
short,
fontsize = fsize,
fontweight= weight,
ha = 'center',
va = 'center',
color = color,
zorder = 3,
)
'''
# ---- COLORBAR ----
sm = ScalarMappable(cmap=cmap, norm=norm)
sm.set_array([])
cbar = fig.colorbar(sm, ax=ax, fraction=0.03, pad=0.02, aspect=30)
'''
cbar.set_label(
'Share of Low-Income Boosted CVRP Applications (2019–2021)',
fontsize=20, color='#444444', labelpad=10
)
'''
cbar.ax.tick_params(labelsize=8)
cbar.outline.set_edgecolor('#CCCCCC')
'''
# ---- LEGEND FOR EXCLUDED / NO DATA ----
legend_handles = []
if SHOW_EXCLUDED:
legend_handles.append(
mpatches.Patch(color=COLOR_EXCLUDED, label='Excluded (<10 applications, NaN reliance)')
)
legend_handles.append(
mpatches.Patch(color='#F2F3F4', label='Not in analysis sample')
)
legend_handles.append(
mpatches.Patch(
facecolor='none', edgecolor='#1A1A1A',
linewidth=1.5, label='Bold label = High Reliance (above median)'
)
)
ax.legend(
handles = legend_handles,
loc = 'lower left',
fontsize = 8,
frameon = True,
framealpha= 0.9,
edgecolor = '#DDDDDD',
)
# ---- MEDIAN THRESHOLD ANNOTATION ----
median_rel = reliance_data['reliance_weight'].median()
ax.text(
0.02, 0.12,
f"Median reliance threshold: {median_rel:.3f}\n"
f"High reliance: {(reliance_data['is_high_reliance']==1).sum()} counties\n"
f"Low reliance: {(reliance_data['is_high_reliance']==0).sum()} counties",
transform = ax.transAxes,
fontsize = 8,
color = '#555555',
va = 'bottom',
bbox = dict(
boxstyle = 'round,pad=0.4',
facecolor = 'white',
edgecolor = '#CCCCCC',
alpha = 0.9
)
)
'''
# ---- AXES STYLING ----
ax.set_axis_off()
'''
ax.set_title(
'CVRP Legacy Reliance Rate by California County',
fontsize = 20,
fontweight= 'bold',
color = '#1A1A1A',
pad = 16,
loc = 'left'
)
#fig.text(
# 0.02, 0.96,
# 'Share of CVRP applications receiving low-income boost, 2019–2021 baseline | '
# 'Bold county names = above-median reliance (high treatment intensity)',
# fontsize = 8,
# color = '#777777'
#)
'''
plt.tight_layout()
plt.show()
# print(f"\nMedian reliance threshold: {median_rel:.4f}")
print(f"Min reliance (included): {vmin:.4f}")
print(f"Max reliance (included): {vmax:.4f}")import statsmodels.api as sm
# DIAGNOSTIC 1A: County-level means + Cook's distance / DFBETA on Model 2
print("--- DIAGNOSTIC 1A: COUNTY MEANS & INFLUENCE STATISTICS ---")
print(f"Current N: {len(analysis_df)}, Unique counties: {analysis_df['county_cat'].nunique()}\n")
controls_pc = ['evs_per_1000', 'tmax_c', 'cumulative_data_center_power_mw', 'solar_mw_per_1000_lag1']
# County-level means (for scatter and reference)
county_means = analysis_df.groupby('county_name').agg(
mean_evs=('evs_per_1000', 'mean'),
mean_mad=('mad_rate', 'mean'),
mean_tmax=('tmax_c', 'mean'),
mean_dc=('cumulative_data_center_power_mw', 'mean'),
mean_solar=('solar_mw_per_1000_lag1', 'mean'),
pop=('total_population', 'mean')
).reset_index()
print("Top 10 counties by EVs/1000:")
print(county_means.sort_values('mean_evs', ascending=False).head(10).to_string(index=False))
print("\nTop 10 counties by mad_rate:")
print(county_means.sort_values('mean_mad', ascending=False).head(10).to_string(index=False))
# Compute influence stats via statsmodels OLS (Model 2 specification)
df2 = analysis_df[['mad_rate'] + controls_pc + ['county_cat']].dropna()
ols_sm = sm.OLS(df2['mad_rate'], sm.add_constant(df2[controls_pc])).fit()
infl = ols_sm.get_influence()
df2 = df2.copy()
df2['cooks_d'] = infl.cooks_distance[0]
df2['dfbeta_evs'] = infl.dfbetas[:, 1] # column 1 = evs_per_1000 (after const)
df2['leverage'] = infl.hat_matrix_diag
county_infl = df2.groupby('county_cat', observed=True).agg(
max_cooks=('cooks_d', 'max'),
sum_abs_dfbeta_evs=('dfbeta_evs', lambda s: s.abs().sum()),
max_leverage=('leverage', 'max')
).reset_index().sort_values('sum_abs_dfbeta_evs', ascending=False)
n = len(df2)
print(f"\nInfluence thresholds: Cook's D > {4/n:.5f}, |DFBETA| > {2/np.sqrt(n):.5f}")
print("\nCounties ranked by total influence on evs_per_1000 coefficient:")
print(county_infl.head(15).to_string(index=False))# DIAGNOSTIC 1B: Sequential leave-out tests
# Re-run Model 2 dropping the top counties by (i) EV penetration, (ii) mad_rate, (iii) DFBETA influence
print("--- DIAGNOSTIC 1B: SEQUENTIAL LEAVE-OUT TESTS ---\n")
def fit_model2(sub_df):
d = sub_df[['mad_rate'] + controls_pc + ['county_cat']].dropna()
e = sm.add_constant(d[controls_pc])
r = IV2SLS(d['mad_rate'], e, None, None).fit(cov_type='clustered', clusters=d['county_cat'])
return {
'N': len(d),
'counties': d['county_cat'].nunique(),
'evs_coef': r.params['evs_per_1000'],
'evs_se': r.std_errors['evs_per_1000'],
'evs_pval': r.pvalues['evs_per_1000'],
}
def sequential_drop(rank_list, label):
rows = []
for k in [0, 1, 2, 3, 4, 5]:
if k == 0:
sub = analysis_df
dropped = "none"
else:
sub = analysis_df[~analysis_df['county_name'].isin(rank_list[:k])]
dropped = ", ".join(rank_list[:k])
row = {'k': k, 'dropped': dropped, **fit_model2(sub)}
rows.append(row)
print(f"### Drop top counties by {label} ###")
print(pd.DataFrame(rows).to_string(index=False))
print()
drops_by_ev = county_means.sort_values('mean_evs', ascending=False)['county_name'].tolist()
drops_by_mad = county_means.sort_values('mean_mad', ascending=False)['county_name'].tolist()
drops_by_infl = county_infl['county_cat'].astype(str).tolist()
sequential_drop(drops_by_ev, "EV penetration")
sequential_drop(drops_by_mad, "mad_rate")
sequential_drop(drops_by_infl, "DFBETA influence on evs_per_1000")
# Winsorization check
print("### Winsorize mad_rate at 1st/99th percentile ###")
lo, hi = analysis_df['mad_rate'].quantile([0.01, 0.99])
df_w = analysis_df.copy()
df_w['mad_rate'] = df_w['mad_rate'].clip(lo, hi)
res = fit_model2(df_w)
print(f"Trimmed to [{lo:.4f}, {hi:.4f}]")
print(f"evs_coef={res['evs_coef']:.6e}, SE={res['evs_se']:.6e}, p={res['evs_pval']:.4f}")# DIAGNOSTIC 1C: Scatter plot of county means with Bay Area & high-volatility groups labeled
try:
from adjustText import adjust_text
except ImportError:
!pip install adjustText -q
from adjustText import adjust_text
print("--- DIAGNOSTIC 1C: COUNTY MEANS SCATTER ---\n")
bay_area = ['Santa Clara', 'San Mateo', 'Marin', 'Alameda', 'Contra Costa',
'San Francisco', 'Napa', 'Sonoma', 'Solano']
high_mad = ['Lake', 'Kings', 'Mendocino', 'Nevada', 'Sutter', 'Madera', 'Yuba', 'Colusa']
cm = county_means.copy()
cm['group'] = 'Other'
cm.loc[cm['county_name'].isin(bay_area), 'group'] = 'Bay Area'
cm.loc[cm['county_name'].isin(high_mad), 'group'] = 'High-volatility rural'
fig, ax = plt.subplots(figsize=(11, 7.5))
colors = {'Other': '#bdbdbd', 'Bay Area': '#2166ac', 'High-volatility rural': '#b2182b'}
sizes = {'Other': 60, 'Bay Area': 110, 'High-volatility rural': 110}
for grp, sub in cm.groupby('group'):
ax.scatter(sub['mean_evs'], sub['mean_mad'],
s=sizes[grp], c=colors[grp], alpha=0.85,
edgecolors='black', linewidths=0.6, label=grp, zorder=3)
# OLS through county means (illustrative — between variation only)
slope, intercept = np.polyfit(cm['mean_evs'], cm['mean_mad'], 1)
xx = np.linspace(cm['mean_evs'].min(), cm['mean_evs'].max(), 100)
ax.plot(xx, intercept + slope * xx, color='#444', linestyle='--', linewidth=1.5,
label=f'OLS through county means (slope={slope:.2e})', zorder=2)
labels_to_show = bay_area + high_mad + ['Los Angeles', 'Orange', 'San Diego']
texts = []
for _, row in cm.iterrows():
if row['county_name'] in labels_to_show:
texts.append(ax.text(row['mean_evs'], row['mean_mad'], row['county_name'], fontsize=8.5, zorder=4))
adjust_text(texts, ax=ax, arrowprops=dict(arrowstyle='-', color='#666', lw=0.5))
ax.set_xlabel('Mean EVs per 1,000 residents (county average)', fontsize=11)
ax.set_ylabel('Mean MAD rate (county average)', fontsize=11)
ax.set_title('Between-county relationship: EV penetration vs LMP volatility\n'
f'(48 counties, {analysis_df["month"].min().strftime("%Y-%m")} to {analysis_df["month"].max().strftime("%Y-%m")})',
fontsize=12.5, fontweight='bold')
ax.legend(loc='upper right', frameon=True, framealpha=0.95)
ax.grid(True, alpha=0.3)
ax.spines['top'].set_visible(False)
ax.spines['right'].set_visible(False)
plt.tight_layout()
plt.show()
print(f"\nSlope through county means: {slope:.4e}")
print("For reference, Model 2 panel coef on evs_per_1000 was ~+8.86e-05.")
print("The cross-sectional pattern is essentially flat — Bay Area counties sit on")
print("the line, and high-volatility rural counties sit *above* it (pulling slope DOWN).")# DIAGNOSTIC 2A: Re-run Models 2, 2a, 2c using cumulative_evs instead of evs_per_1000
print("--- DIAGNOSTIC 2A: OLS / FE MODELS WITH LEVELS ---\n")
controls_lvl = ['cumulative_evs', 'tmax_c', 'cumulative_data_center_power_mw','total_generator_active_capacity_mw', 'solar_mw_per_1000_lag1']
# Model 2 (no FE) — levels
df_l = analysis_df[['mad_rate'] + controls_lvl + ['county_cat']].dropna()
exog = sm.add_constant(df_l[controls_lvl])
r2_lvl = IV2SLS(df_l['mad_rate'], exog, None, None).fit(cov_type='clustered', clusters=df_l['county_cat'])
print("### Model 2 (no FE) — LEVELS ###")
#print(f" cumulative_evs: coef={r2_lvl.params['cumulative_evs']:.6e}, SE={r2_lvl.std_errors['cumulative_evs']:.6e}, p={r2_lvl.pvalues['cumulative_evs']:.4f}")
#print(f" R2={r2_lvl.rsquared:.4f}\n")
print(r2_lvl.summary)
# Models 2a and 2c via PanelOLS — levels
df_panel = analysis_df[['mad_rate'] + controls_lvl + ['county_name', 'month']].dropna().set_index(['county_name', 'month'])
y = df_panel['mad_rate']
X = sm.add_constant(df_panel[controls_lvl])
print("### Model 2a (county FE) — LEVELS ###")
m2a_lvl = PanelOLS(y, X, entity_effects=True, drop_absorbed=True).fit(cov_type='clustered', cluster_entity=True)
#print(f" cumulative_evs: coef={m2a_lvl.params['cumulative_evs']:.6e}, SE={m2a_lvl.std_errors['cumulative_evs']:.6e}, p={m2a_lvl.pvalues['cumulative_evs']:.4f}")
#print(f" Within R2={m2a_lvl.rsquared_within:.4f}\n")
print(m2a_lvl.summary)
print("### Model 2c (TWFE) — LEVELS ###")
m2c_lvl = PanelOLS(y, X, entity_effects=True, time_effects=True, drop_absorbed=True).fit(cov_type='clustered', cluster_entity=True)
#print(f" cumulative_evs: coef={m2c_lvl.params['cumulative_evs']:.6e}, SE={m2c_lvl.std_errors['cumulative_evs']:.6e}, p={m2c_lvl.pvalues['cumulative_evs']:.4f}")
##print(f" Within R2={m2c_lvl.rsquared_within:.4f}\n")
print(m2c_lvl.summary)
# Side-by-side: per-capita vs levels (TWFE)
df_pc = analysis_df[['mad_rate'] + controls_pc + ['county_name', 'month']].dropna().set_index(['county_name', 'month'])
y_pc = df_pc['mad_rate']
X_pc = sm.add_constant(df_pc[controls_pc])
m2c_pc = PanelOLS(y_pc, X_pc, entity_effects=True, time_effects=True, drop_absorbed=True).fit(cov_type='clustered', cluster_entity=True)
sd_pc = analysis_df['evs_per_1000'].std()
sd_lvl = analysis_df['cumulative_evs'].std()
comp = pd.DataFrame({
'spec': ['per-capita (evs_per_1000)', 'levels (cumulative_evs)'],
'coef': [m2c_pc.params['evs_per_1000'], m2c_lvl.params['cumulative_evs']],
'SE': [m2c_pc.std_errors['evs_per_1000'], m2c_lvl.std_errors['cumulative_evs']],
'pval': [m2c_pc.pvalues['evs_per_1000'], m2c_lvl.pvalues['cumulative_evs']],
'sd_x': [sd_pc, sd_lvl],
'std_effect_per_1sd': [m2c_pc.params['evs_per_1000'] * sd_pc, m2c_lvl.params['cumulative_evs'] * sd_lvl],
})
print("### TWFE side-by-side: per-capita vs levels ###")
print(comp.to_string(index=False))
print("\n(std_effect_per_1sd = change in mad_rate from a 1 SD increase in the EV variable)")# DIAGNOSTIC 2B: IV-2SLS Models 6 and 7 with cumulative_evs (LEVELS)
# This is the critical test — does instrument relevance survive when we drop per-capita?
print("--- DIAGNOSTIC 2B: IV-2SLS WITH LEVELS ---\n")
# >>> CHANGE 1: Demean ALL variables we'll need, in one call, into a new dataframe.
# This avoids relying on _dm columns that may or may not exist on analysis_df
# from earlier cells, and avoids mutating the global.
cols_needed = [
'mad_rate', 'evs_per_1000', 'cumulative_evs',
'iv_trend_x_reliance_weight', 'iv_stock_x_reliance_weight',
'tmax_c', 'cumulative_data_center_power_mw', 'total_generator_active_capacity_mw',
'solar_mw_per_1000_lag1', 'solar_x_nem3',
]
df_diag = iterative_demean(analysis_df, cols_needed)
# >>> CHANGE 2: Function now takes `data` as an argument instead of
# reaching out to the global analysis_df. Also validates columns up front.
def run_iv_diagnostic(data, dep, endog, instruments, exog_controls, label):
required = [dep, endog] + list(instruments) + list(exog_controls) + ['county_cat']
missing = [c for c in required if c not in data.columns]
if missing:
available_dm = [c for c in data.columns if c.endswith('_dm')]
raise ValueError(
f"Missing columns: {missing}\nAvailable _dm columns: {available_dm}"
)
d = data[required].dropna()
mod = IV2SLS(
dependent=d[dep],
exog=d[exog_controls],
endog=d[[endog]],
instruments=d[instruments]
)
res = mod.fit(cov_type='clustered', clusters=d['county_cat'])
print(f"### {label} ###")
print(f" N = {len(d)}")
print(f" {endog}: coef={res.params[endog]:.6e}, SE={res.std_errors[endog]:.6e}, p={res.pvalues[endog]:.4f}")
ci = res.conf_int().loc[endog]
print(f" 95% CI: [{ci['lower']:.6e}, {ci['upper']:.6e}]")
diag = res.first_stage.diagnostics
f_col = [c for c in diag.columns if 'f.stat' in c.lower()][0]
pr2_col = [c for c in diag.columns if 'partial' in c.lower()]
print(f" First-stage F: {diag.loc[endog, f_col]:.3f}")
if pr2_col:
print(f" Partial R2: {diag.loc[endog, pr2_col[0]]:.4f}")
fs = res.first_stage.individual[endog]
for inst in instruments:
print(f" {inst}: coef={fs.params[inst]:.4f}, t={fs.tstats[inst]:.3f}, p={fs.pvalues[inst]:.4f}")
try:
sh = res.sargan
print(f" Sargan-Hansen: stat={sh.stat:.4f}, p={sh.pval:.4f}")
except Exception:
pass
print()
return res
# Per-capita replication (sanity check — should match Models 6 and 7)
print("=" * 70)
print("PER-CAPITA REPLICATION (should match Models 6 and 7)")
print("=" * 70)
m6_pc = run_iv_diagnostic(
data=df_diag, # >>> CHANGE 3: pass df_diag explicitly
dep='mad_rate_dm', endog='evs_per_1000_dm',
instruments=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
exog_controls=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm'],
label="Model 6 (per-capita)"
)
m7_pc = run_iv_diagnostic(
data=df_diag,
dep='mad_rate_dm', endog='evs_per_1000_dm',
instruments=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
exog_controls=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm'],
label="Model 7 (per-capita)"
)
# Levels versions
print("=" * 70)
print("LEVELS VERSIONS (cumulative_evs)")
print("=" * 70)
m6_lvl = run_iv_diagnostic(
data=df_diag,
dep='mad_rate_dm', endog='cumulative_evs_dm',
instruments=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
exog_controls=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm'],
label="Model 6 (levels)"
)
m7_lvl = run_iv_diagnostic(
data=df_diag,
dep='mad_rate_dm', endog='cumulative_evs_dm',
instruments=['iv_trend_x_reliance_weight_dm', 'iv_stock_x_reliance_weight_dm'],
exog_controls=['tmax_c_dm', 'cumulative_data_center_power_mw_dm','total_generator_active_capacity_mw_dm', 'solar_mw_per_1000_lag1_dm', 'solar_x_nem3_dm'],
label="Model 7 (levels)"
)
# >>> CHANGE 4: Compute sd_pc and sd_lvl here rather than assuming they
# already exist from an earlier cell (which is the next KeyError waiting to happen).
sd_pc = df_diag['evs_per_1000'].std()
sd_lvl = df_diag['cumulative_evs'].std()
# Standardized comparison
print("=" * 70)
print("STANDARDIZED COMPARISON (effect of 1 SD change in EV variable)")
print("=" * 70)
f_col = [c for c in m6_pc.first_stage.diagnostics.columns if 'f.stat' in c.lower()][0]
summary = pd.DataFrame({
'spec': ['M6 per-capita', 'M6 levels', 'M7 per-capita', 'M7 levels'],
'coef': [m6_pc.params['evs_per_1000_dm'], m6_lvl.params['cumulative_evs_dm'],
m7_pc.params['evs_per_1000_dm'], m7_lvl.params['cumulative_evs_dm']],
'pval': [m6_pc.pvalues['evs_per_1000_dm'], m6_lvl.pvalues['cumulative_evs_dm'],
m7_pc.pvalues['evs_per_1000_dm'], m7_lvl.pvalues['cumulative_evs_dm']],
'first_stage_F': [
m6_pc.first_stage.diagnostics.iloc[0][f_col],
m6_lvl.first_stage.diagnostics.iloc[0][f_col],
m7_pc.first_stage.diagnostics.iloc[0][f_col],
m7_lvl.first_stage.diagnostics.iloc[0][f_col],
],
'sd_x': [sd_pc, sd_lvl, sd_pc, sd_lvl],
})
summary['effect_per_1sd'] = summary['coef'] * summary['sd_x']
print(summary.to_string(index=False))import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from scipy import stats
# --- 1. DEFINING TIME PERIODS FROM LOCKED SAMPLE ---
# Using your locked panel dataframe: 'analysis_df'
# Define first 6 months (April 2022 - Sept 2022)
start_period_mask = (analysis_df['month'] >= '2022-04-01') & (analysis_df['month'] <= '2022-09-30')
df_start = analysis_df[start_period_mask].groupby('county_name')[['evs_per_1000', 'mad_rate']].mean().reset_index()
# Define last 6 months (Dec 2024 - May 2025)
end_period_mask = (analysis_df['month'] >= '2024-12-01') & (analysis_df['month'] <= '2025-05-31')
df_end = analysis_df[end_period_mask].groupby('county_name')[['evs_per_1000', 'mad_rate']].mean().reset_index()
# --- 2. FIGURE 1: First 6 Months Cross-Section ---
plt.figure(figsize=(10, 6))
sns.scatterplot(data=df_start, x='evs_per_1000', y='mad_rate', s=80, alpha=0.7)
# Add OLS trendline
slope1, intercept1, r_value1, p_value1, std_err1 = stats.linregress(df_start['evs_per_1000'], df_start['mad_rate'])
plt.plot(df_start['evs_per_1000'], intercept1 + slope1 * df_start['evs_per_1000'], 'k--', label=f'OLS (slope={slope1:.2e})')
# Formatting
plt.title('Figure 1: EV Penetration vs LMP Volatility (First 6 Months: Apr-Sep 2022)', fontsize=14)
plt.xlabel('Mean EVs per 1,000 residents', fontsize=12)
plt.ylabel('Mean MAD rate', fontsize=12)
plt.legend()
#plt.grid(True, alpha=0.3)
plt.tight_layout()
plt.savefig('figure1_start_period.png', dpi=300)
plt.show()
# --- 3. FIGURE 2: Changes Over Time (End Period - Start Period) ---
# Merge start and end periods to calculate changes
df_changes = pd.merge(df_start, df_end, on='county_name', suffixes=('_start', '_end'))
df_changes['delta_evs'] = df_changes['evs_per_1000_end'] - df_changes['evs_per_1000_start']
df_changes['delta_mad'] = df_changes['mad_rate_end'] - df_changes['mad_rate_start']
plt.figure(figsize=(10, 6))
sns.scatterplot(data=df_changes, x='delta_evs', y='delta_mad', s=80, alpha=0.7, color='teal')
# Add OLS trendline for changes
slope2, intercept2, r_value2, p_value2, std_err2 = stats.linregress(df_changes['delta_evs'], df_changes['delta_mad'])
plt.plot(df_changes['delta_evs'], intercept2 + slope2 * df_changes['delta_evs'], 'k--', label=f'OLS (slope={slope2:.2e})')
# Formatting
plt.title('Figure 2: Change in LMP Volatility vs Change in EV Penetration\n(Apr-Sep 2022 to Dec 2024-May 2025)', fontsize=14)
plt.xlabel('Change in Mean EVs per 1,000 residents', fontsize=12)
plt.ylabel('Change in Mean MAD rate', fontsize=12)
plt.legend()
#plt.grid(True, alpha=0.3)
# Add zero lines to easily distinguish positive/negative changes
plt.axhline(0, color='gray', linewidth=1)
plt.axvline(0, color='gray', linewidth=1)
plt.tight_layout()
plt.savefig('figure2_changes.png', dpi=300)
plt.show()Data Centers · Main causal analysis
The substantive IV regression producing the headline DC findings.
dc_iv_fiber_2sls.ipynb10 cells · primaryFiber-instrumented 2SLS producing the headline DC results reported above: the −0.174 log-DC-power coefficient on arcsinh-transformed LMP (collapsed IV, first-stage F ≈ 60), the consistent null on volatility and spikes, and the PSM robustness check on DAC-heavy counties.
#Importing the dataset
import pandas as pd
df = pd.read_excel('/content/dmn_county_monthly_summary_final_with_DC.xlsx')
df.head()#Converting date to date format
df["date"] = pd.to_datetime(df["Year"].astype(str) + "-" + df["Month"].astype(str) + "-01")
#Rename Variables
df.columns = (
df.columns.str.strip()
.str.replace(r"[^\w]+", "_", regex=True)
.str.replace(r"__+", "_", regex=True)
)# Import census demographics info
df_census = pd.read_csv('/content/dmn_county_census_demographics_info.csv')
# Rename 'County GEOID' to 'county_geoid' for merging, if different
# And ensure 'median household income' is clean
df_census.rename(columns={'County GEOID': 'county_geoid', 'median household income': 'median_household_income', 'percent_dac_tracts': 'percent_dac_tracts'}, inplace=True)
# Select only unique county_geoid and their median household income and percent_dac_tracts
df_census_unique = df_census[['county_geoid', 'median_household_income', 'percent_dac_tracts', 'percent_below_poverty']].drop_duplicates(subset=['county_geoid'])
# Merge into the main dataframe
df = pd.merge(df, df_census_unique, on='county_geoid', how='left')
print("Median Household Income and Percent DAC Tracts added to df. First 5 rows:")
print(df[['county_geoid', 'county_name', 'median_household_income', 'percent_dac_tracts']].head())#STEP #1 - OLS BASIC REGRESSION
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
# ---------------------------
# 1) Create variables
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]
# Prices: asinh handles negatives
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])
# Create median_income_10k_units and then log-transform it
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])
# ---------------------------
# 2) Helper: clean coefficient table (hide FE rows)
# ---------------------------
def tidy_no_fe(res, drop_prefixes=("C(county_geoid)", "C(date)")):
tab = res.summary2().tables[1].copy() # includes Coef., Std.Err., P>|t|/P>|z|, [0.025, 0.975]
# Identify the p-value column name (can be P>|t| or P>|z|)
pcol = "P>|t|" if "P>|t|" in tab.columns else ("P>|z|" if "P>|z|" in tab.columns else None)
if pcol is None:
raise ValueError("Could not find p-value column in summary2 output.")
# Drop fixed effects rows
mask_fe = np.zeros(len(tab), dtype=bool)
for pref in drop_prefixes:
mask_fe |= tab.index.to_series().str.startswith(pref)
tab = tab[~mask_fe]
# Keep and rename columns nicely
out = tab.rename(columns={
"Coef.": "Coef",
"Std.Err.": "Std_err",
pcol: "P_value",
"[0.025": "C.I.95_lo",
"0.975]": "C.I.95_hi",
})[["Coef", "Std_err", "P_value", "C.I.95_lo", "C.I.95_hi"]]
return out
# ---------------------------
# 3) Run regressions (county FE + month FE, clustered by county)
# ---------------------------
# Spikes regression dataset
cols_spikes = ["log_spikes", "log_dc_power", "log_pop", "temp_f", "log_median_income_10k_units", "county_geoid", "date"]
d_spikes = df[cols_spikes].dropna().copy()
m_spikes = smf.ols(
"log_spikes ~ log_dc_power + log_pop + temp_f + log_median_income_10k_units + C(county_geoid) + C(date)",
data=d_spikes
).fit(cov_type="cluster", cov_kwds={"groups": d_spikes["county_geoid"]})
# Prices regression dataset
cols_prices = ["asinh_lmp", "log_dc_power", "log_pop", "temp_f", "log_median_income_10k_units", "county_geoid", "date"]
d_prices = df[cols_prices].dropna().copy()
m_prices = smf.ols(
"asinh_lmp ~ log_dc_power + log_pop + temp_f + log_median_income_10k_units + C(county_geoid) + C(date)",
data=d_prices
).fit(cov_type="cluster", cov_kwds={"groups": d_prices["county_geoid"]})
# ---------------------------
# 4) Print clean tables (no FE rows)
# ---------------------------
print("\n=== Spikes model (log1p) — FE included, FE rows hidden ===")
print(tidy_no_fe(m_spikes))
print("\n=== Prices model (asinh) — FE included, FE rows hidden ===")
print(tidy_no_fe(m_prices))# If needed in Colab (run once):
!pip -q install linearmodels
from linearmodels.iv import IV2SLS# STEP #2 — IV REGRESSION (2SLS) in the PANEL (county × month)
# Spec: Month FE only, clustered SE by county
# Outcomes: log_spikes and asinh_lmp
# Endogenous regressor: log_dc_power
# Instrument: fiber provider counts, average max upstream speed, average max downstream speed
# ============================================================
import numpy as np
import pandas as pd
# If needed in Colab (run once):
!pip -q install linearmodels
from linearmodels.iv import IV2SLS
# ---------------------------
# 1) Create variables
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]
# Prices: asinh handles negatives
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])
# Instrument (fiber providers)
df["fiber"] = df["Fiber_Provider_Counts"]
df["log_upstream"] = np.log1p(df["Average_Max_Upstream"])
df["log_downstream"] = np.log1p(df["Average_Max_Downstream"])
# Create median_income_10k_units and then log-transform it
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])
# ---------------------------
# 2) Helpers: tidy IV output + drop FE rows from display
# ---------------------------
def tidy_iv(res):
"""Return coef, std err, p-values, and 95% CI from linearmodels IV results."""
return pd.DataFrame({
"Coef": res.params,
"Std_err": res.std_errors,
"P_value": res.pvalues,
"C.I.95_lo": res.conf_int().iloc[:, 0],
"C.I.95_hi": res.conf_int().iloc[:, 1],
})
def drop_fe_rows(tab, prefixes=("C(date)",)):
"""Remove FE dummy rows from the printed table (FE stay in the model)."""
idx = tab.index.astype(str)
mask_fe = np.zeros(len(tab), dtype=bool)
for pref in prefixes:
mask_fe |= pd.Series(idx).str.startswith(pref).values
return tab[~mask_fe]
# ---------------------------
# 3) Build estimation datasets (drop NaNs consistently)
# ---------------------------
cols_spikes = ["county_geoid", "date", "log_spikes", "log_dc_power", "fiber", "log_pop", "temp_f", "log_median_income_10k_units","log_upstream", "log_downstream"]
d_spikes = df[cols_spikes].dropna().copy()
cols_prices = ["county_geoid", "date", "asinh_lmp", "log_dc_power", "fiber", "log_pop", "temp_f", "log_median_income_10k_units","log_upstream", "log_downstream"]
d_prices = df[cols_prices].dropna().copy()
# ---------------------------
# 4) Run IV regressions (2SLS) — Month FE only
# ---------------------------
# Outcome: spikes
iv_spikes_monthfe = IV2SLS.from_formula(
"log_spikes ~ 1 + log_pop + temp_f + log_median_income_10k_units + C(date) + [log_dc_power ~ fiber + log_upstream + log_downstream]",
data=d_spikes
).fit(cov_type="clustered", clusters=d_spikes["county_geoid"])
# Outcome: prices
iv_prices_monthfe = IV2SLS.from_formula(
"asinh_lmp ~ 1 + log_pop + temp_f + log_median_income_10k_units + C(date) + [log_dc_power ~ fiber + log_upstream + log_downstream]",
data=d_prices
).fit(cov_type="clustered", clusters=d_prices["county_geoid"])
# ---------------------------
# 5) Print clean coefficient tables (hide month FE dummies)
# ---------------------------
print("\n=== IV (2SLS) — Spikes | Month FE | Clustered by county ===")
t_spikes = drop_fe_rows(tidy_iv(iv_spikes_monthfe), prefixes=("C(date)",))
print(t_spikes)
print("\n=== IV (2SLS) — Prices | Month FE | Clustered by county ===")
t_prices = drop_fe_rows(tidy_iv(iv_prices_monthfe), prefixes=("C(date)",))
print(t_prices)
# ---------------------------
# 6) First-stage diagnostics (print for BOTH outcomes)
# ---------------------------
print("\n=== First-stage diagnostics (Spikes sample) ===")
print(iv_spikes_monthfe.first_stage)
print("\n=== First-stage diagnostics (Prices sample) ===")
print(iv_prices_monthfe.first_stage)
# STEP #3 — COLLAPSED REGRESSION (cross-county / between design)
# Collapse county × month -> county-level means
# Outcomes: log_spikes (mean) and asinh_lmp (mean)
# Regressor: log_dc_power (mean)
# Controls: log_pop (mean), temp_f (mean)
# Robust SE (HC3). No fixed effects needed after collapsing.
import numpy as np
import pandas as pd
import statsmodels.formula.api as smf
# ---------------------------
# 1) Create variables (same as before)
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])
# Create median_income_10k_units and then log-transform it
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])
# ---------------------------
# 2) Helpers: tidy OLS output (no FE rows to drop here, but keep consistent style)
# ---------------------------
def tidy_ols(res):
tab = res.summary2().tables[1].copy()
pcol = "P>|t|" if "P>|t|" in tab.columns else ("P>|z|" if "P>|z|" in tab.columns else None)
if pcol is None:
raise ValueError("Could not find p-value column in summary2 output.")
out = tab.rename(columns={
"Coef.": "Coef",
"Std.Err.": "Std_err",
pcol: "P_value",
"[0.025": "C.I.95_lo",
"0.975]": "C.I.95_hi",
})[["Coef", "Std_err", "P_value", "C.I.95_lo", "C.I.95_hi"]]
return out
# ---------------------------
# 3) Collapse to cross-county dataset (means by county)
# ---------------------------
cols_collapse = ["county_geoid", "log_spikes", "asinh_lmp", "log_dc_power", "log_pop", "temp_f", "log_median_income_10k_units"]
g = (df[cols_collapse]
.dropna()
.groupby("county_geoid", as_index=False)
.mean()
)
# ---------------------------
# 4) Run collapsed OLS regressions (cross-county)
# ---------------------------
# Outcome: spikes (mean)
collapsed_spikes = smf.ols(
"log_spikes ~ log_dc_power + log_pop + temp_f + log_median_income_10k_units",
data=g
).fit(cov_type="HC3")
# Outcome: prices (mean)
collapsed_prices = smf.ols(
"asinh_lmp ~ log_dc_power + log_pop + temp_f + log_median_income_10k_units",
data=g
).fit(cov_type="HC3")
# ---------------------------
# 5) Print clean coefficient tables
# ---------------------------
print("\n=== Collapsed OLS — Spikes (county means) | Robust SE (HC3) ===")
print(tidy_ols(collapsed_spikes))
print("\n=== Collapsed OLS — Prices (county means) | Robust SE (HC3) ===")
print(tidy_ols(collapsed_prices))# STEP #3 — COLLAPSED REGRESSION (cross-county / between design)
# Collapse county × month -> county-level means
# Outcomes: log_spikes (mean) and asinh_lmp (mean)
# Regressor: log_dc_power (mean)
# Instrument: fiber (mean), log upstream + downstream speeds
# Controls: log_pop (mean), temp_f (mean)
# Robust SE. No fixed effects needed after collapsing.
import numpy as np
import pandas as pd
from linearmodels.iv import IV2SLS
# ---------------------------
# 1) Create variables (same as before)
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])
df["fiber"] = df["Fiber_Provider_Counts"]
df["log_upstream"] = np.log1p(df["Average_Max_Upstream"])
df["log_downstream"] = np.log1p(df["Average_Max_Downstream"])
# Create median_income_10k_units and then log-transform it
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])
# ---------------------------
# 2) Helpers: tidy IV output
# ---------------------------
def tidy_iv(res):
"""Return coef, std err, p-values, and 95% CI from linearmodels IV results."""
return pd.DataFrame({
"Coef": res.params,
"Std_err": res.std_errors,
"P_value": res.pvalues,
"C.I.95_lo": res.conf_int().iloc[:, 0],
"C.I.95_hi": res.conf_int().iloc[:, 1],
})
# ---------------------------
# 3) Collapse to cross-county dataset (means by county)
# ---------------------------
cols_collapse = [
"county_geoid", "log_spikes", "asinh_lmp",
"log_dc_power", "log_pop", "temp_f", "fiber","log_median_income_10k_units",
"log_upstream", "log_downstream"
]
g = (
df[cols_collapse]
.dropna()
.groupby("county_geoid", as_index=False)
.mean()
)
# ---------------------------
# 4) Run collapsed IV regressions (2SLS) — cross-county
# ---------------------------
# Outcome: spikes (mean)
iv_collapsed_spikes = IV2SLS.from_formula(
"log_spikes ~ log_pop + temp_f + log_median_income_10k_units + [log_dc_power ~ fiber + log_upstream + log_downstream]",
data=g
).fit(cov_type="robust")
# Outcome: prices (mean)
iv_collapsed_prices = IV2SLS.from_formula(
"asinh_lmp ~ log_pop + temp_f + log_median_income_10k_units + [log_dc_power ~ fiber + log_upstream + log_downstream]",
data=g
).fit(cov_type="robust")
# ---------------------------
# 5) Print clean coefficient tables
# ---------------------------
print("\n=== Collapsed IV (2SLS) — Spikes (county means) | Robust SE ===")
print(tidy_iv(iv_collapsed_spikes))
print("\n=== Collapsed IV (2SLS) — Prices (county means) | Robust SE ===")
print(tidy_iv(iv_collapsed_prices))
# ---------------------------
# 6) First-stage diagnostics (print for BOTH outcomes)
# ---------------------------
print("\n=== First-stage diagnostics (Spikes sample - collapsed) ===")
print(iv_collapsed_spikes.first_stage)
print("\n=== First-stage diagnostics (Prices sample - collapsed) ===")
print(iv_collapsed_prices.first_stage)
import statsmodels.formula.api as smf
# Ensure all necessary variables are in the 'g' DataFrame
# The 'g' DataFrame was prepared in cell 'fo_kWVJVYhjR' with these columns.
print("\n=== OLS Robustness Check for Exclusion Restriction (Prices) ===")
print("\nOutcome: asinh_lmp | Explanatory: log_dc_power + Instruments + Controls")
m_prices_exclusion_check = smf.ols(
"asinh_lmp ~ log_dc_power + fiber + log_upstream + log_downstream + log_pop + temp_f + log_median_income_10k_units",
data=g
).fit(cov_type="HC3")
# Helper function to print clean tables (adapted from earlier helper)
def tidy_ols(res):
tab = res.summary2().tables[1].copy()
pcol = "P>|t|" if "P>|t|" in tab.columns else ("P>|z|" if "P>|z|" in tab.columns else None)
if pcol is None:
raise ValueError("Could not find p-value column in summary2 output.")
out = tab.rename(columns={
"Coef.": "Coef",
"Std.Err.": "Std_err",
pcol: "P_value",
"[0.025": "C.I.95_lo",
"0.975]": "C.I.95_hi",
})[["Coef", "Std_err", "P_value", "C.I.95_lo", "C.I.95_hi"]]
return out
print(tidy_ols(m_prices_exclusion_check))
print("\n--- Interpretation Hint ---")
print("If the instruments (fiber, log_upstream, log_downstream) were *valid*, their coefficients in this OLS regression should be statistically *insignificant* when log_dc_power is already in the model. A significant coefficient suggests a direct effect on prices, violating the exclusion restriction.")# ============================================================
# STEP #4 — PROPENSITY SCORE MATCHING (PSM) [Corrected v2] - CALIPER 0.15
# Treatment: above median percent_dac_tracts vs below median percent_dac_tracts
# Unit: County (collapsed to county-level means)
# Outcomes: log_spikes (mean) and asinh_lmp (mean)
# Matching: 1-to-1 nearest neighbor on propensity score + optional caliper
# ============================================================
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import NearestNeighbors
# --- Re-import and merge census data to ensure necessary columns are present ---
df_census = pd.read_csv('/content/dmn_county_census_demographics_info.csv')
df_census.rename(columns={'County GEOID': 'county_geoid', 'median household income': 'median_household_income', 'percent_dac_tracts': 'percent_dac_tracts'}, inplace=True)
df_census_unique = df_census[['county_geoid', 'median_household_income', 'percent_dac_tracts']].drop_duplicates(subset=['county_geoid'])
# Identify columns to be updated/added from df_census_unique
census_cols_to_merge = ['median_household_income', 'percent_dac_tracts']
# Drop existing versions of these columns from df before merging to avoid conflicts
for col in census_cols_to_merge:
if col in df.columns:
df = df.drop(columns=[col])
if col + '_orig' in df.columns: # Also drop any _orig versions that might clash if present
df = df.drop(columns=[col + '_orig'])
# Perform the merge. Since we dropped potential conflicts, no suffixes are needed.
# This ensures the columns from df_census_unique are added cleanly.
df = pd.merge(df, df_census_unique, on='county_geoid', how='left')
# ------------------------------------------------------------------------------------
# ---------------------------
# 1) Create variables
# ---------------------------
df["log_dc_power"] = np.log1p(df["Cumulative_Data_Center_Power"])
df["log_pop"] = np.log1p(df["County_Population"])
df["log_spikes"] = np.log1p(df["total_mad_spikes"])
df["temp_f"] = df["tavg_f"]
df["asinh_lmp"] = np.arcsinh(df["avg_lmp_weighted"])
df["fiber"] = df["Fiber_Provider_Counts"]
df["log_upstream"] = np.log1p(df["Average_Max_Upstream"])
df["log_downstream"] = np.log1p(df["Average_Max_Downstream"])
df["median_income_10k_units"] = df["median_household_income"] / 10000
df["log_median_income_10k_units"] = np.log1p(df["median_income_10k_units"])
# ---------------------------
# 2) Collapse to county-level means
# ---------------------------
cols_psm = ["county_geoid", "log_spikes", "asinh_lmp", "log_dc_power", "log_pop", "temp_f", "fiber", "log_upstream", "log_downstream", "log_median_income_10k_units", "percent_dac_tracts"]
g = (
df[cols_psm]
.dropna() # Drop NaNs after adding all columns needed for PSM
.groupby("county_geoid", as_index=False)
.mean()
)
# --- NEW: Filter to counties WITH Data Centers only ---
g_filtered_dc = g[g["log_dc_power"] > 0].copy()
# ---------------------------
# 3) Define treatment: above median percent_dac_tracts vs below median percent_dac_tracts
# ---------------------------
# Calculate median for the filtered group (counties with DC > 0)
median_dac_percent = g_filtered_dc["percent_dac_tracts"].median()
g_psm = g_filtered_dc.copy()
g_psm["treat"] = (g_psm["percent_dac_tracts"] >= median_dac_percent).astype(int)
print("=== Treatment counts (must include 0 and 1) ===")
print(g_psm["treat"].value_counts())
print("\nMedian Percent DAC Tracts (for counties with DC > 0):", median_dac_percent)
print("Counties in PSM sample:", g_psm.shape[0])
# ---------------------------
# 4) Estimate propensity scores
# ---------------------------
covariates = ["log_dc_power", "log_pop", "temp_f", "fiber", "log_upstream", "log_downstream", "log_median_income_10k_units"]
g_psm = g_psm.dropna(subset=covariates + ["treat", "log_spikes", "asinh_lmp", "percent_dac_tracts"]).copy() # Ensure all relevant columns are dropped consistently
X = g_psm[covariates].values
t = g_psm["treat"].values
scaler = StandardScaler()
Xz = scaler.fit_transform(X)
logit = LogisticRegression(max_iter=1000, solver="lbfgs")
logit.fit(Xz, t)
g_psm["pscore"] = logit.predict_proba(Xz)[:, 1]
# ---------------------------
# 5) Matching (reset indices to avoid index alignment errors)
# ---------------------------
treated = g_psm[g_psm["treat"] == 1].reset_index(drop=True).copy()
control = g_psm[g_psm["treat"] == 0].reset_index(drop=True).copy()
nn = NearestNeighbors(n_neighbors=1, metric="euclidean")
nn.fit(control[["pscore"]].values)
dist, idx = nn.kneighbors(treated[["pscore"]].values)
# idx is position in CONTROL (0..len(control)-1), aligned to treated rows (0..len(treated)-1)
matched_control = control.iloc[idx.flatten()].reset_index(drop=True).copy()
matched_treated = treated.reset_index(drop=True).copy()
matched_treated["match_dist"] = dist.flatten()
# Optional caliper
caliper = 0.3
keep = matched_treated["match_dist"] <= caliper
matched_treated = matched_treated.loc[keep].reset_index(drop=True)
matched_control = matched_control.loc[keep].reset_index(drop=True)
print("\nMatched pairs after caliper:", len(matched_treated))
# ---------------------------
# 6) ATT (Average Treatment effect on the Treated)
# ---------------------------
att_spikes = (matched_treated["log_spikes"] - matched_control["log_spikes"]).mean()
att_prices = (matched_treated["asinh_lmp"] - matched_control["asinh_lmp"]).mean()
# ---------------------------
# 7) Results tables
# ---------------------------
psm_results = pd.DataFrame({
"Outcome": ["log_spikes", "asinh_lmp"],
"ATT (Above Median % DACs - Below Median % DACs)": [att_spikes, att_prices],
"Matched pairs": [len(matched_treated), len(matched_treated)],
"Caliper": [caliper, caliper],
})
print("\n=== Propensity Score Matching Results (ATT) ===")
print(psm_results)
balance_post = pd.DataFrame({
"Mean Treated (matched)": matched_treated[covariates].mean(),
"Mean Control (matched)": matched_control[covariates].mean(),
"Diff (T - C)": matched_treated[covariates].mean() - matched_control[covariates].mean()
})
Data Centers · Supporting pipeline · data engineering
The four notebooks that build the DC dataset — imputation of missing fields, web-scraping facility attributes and founding dates, and constructing the fiber-broadband instrument.
dc_imputation_power_space.ipynb23 cells · supportingRandom Forest / XGBoost / KNN / Linear imputation for missing Total Power and Usable Space fields in the 197-facility dataset, with 5-fold cross-validation.
import pandas as pd
# Load the data
df = pd.read_excel('datacenter_list_11.6.25.xlsx', engine='openpyxl')
df.info()
# df_train: rows where both 'usable_space' and 'total_power' are not null
df_train = df[df['Usable Space'].notna() & df['Total Power (MW)'].notna()]
df_train.info()
from sklearn.model_selection import KFold, cross_val_score
# Define k-fold cross-validation
kf = KFold(n_splits=5, shuffle=True, random_state=42)from sklearn.model_selection import cross_validate
from sklearn.metrics import make_scorer, r2_score, mean_absolute_error, mean_squared_error
import numpy as np
# If sklearn >= 0.24, you can import this directly:
from sklearn.metrics import mean_absolute_percentage_error
def mape(y_true, y_pred):
return mean_absolute_percentage_error(y_true, y_pred) * 100 # percentage
# Define scoring dictionary
scoring = {
"R2": make_scorer(r2_score),
"MAE": make_scorer(mean_absolute_error),
"RMSE": make_scorer(lambda y_true, y_pred: np.sqrt(mean_squared_error(y_true, y_pred))),
"MAPE": make_scorer(mape)
}
# Initialize results list
results = []from sklearn.preprocessing import StandardScaler
# List of numeric features to scale
numeric_features = ["Usable Space", "Total Power (MW)", "Year", "Latitude", "Longitude"]
# Initialize scaler
scaler = StandardScaler()
# Fit and transform numeric features
scaled_values = scaler.fit_transform(df_train[numeric_features])
# Create a DataFrame with new column names for scaled features
scaled_df = pd.DataFrame(scaled_values,
columns=[f"{col}_scaled" for col in numeric_features],
index=df_train.index)
# Add the scaled features to df_train without replacing the originals
df_train = pd.concat([df_train, scaled_df], axis=1)
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import cross_validate
from sklearn.metrics import make_scorer, r2_score, mean_absolute_error, mean_squared_error
import statsmodels.api as sm
import numpy as np
# -------- Predicting Total Power --------
X_total_power = df_train[["Usable Space_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]] # exclude Total Power
y_total_power = df_train["Total Power (MW)"]
model_lr_tp = LinearRegression()
cv_lr_tp = cross_validate(model_lr_tp, X_total_power, y_total_power, cv=kf, scoring=scoring)
# Fit statsmodels OLS for summary
X_sm_tp = sm.add_constant(X_total_power)
ols_tp = sm.OLS(y_total_power, X_sm_tp).fit()
print("\n=== OLS Regression Summary: Predicting Total Power ===")
print(ols_tp.summary())
# Append metrics to results
results.append({
"Target": "Total Power",
"Model": "Linear Regression (OLS)",
"R2": np.mean(cv_lr_tp["test_R2"]),
"MAE": np.mean(cv_lr_tp["test_MAE"]),
"RMSE": np.mean(cv_lr_tp["test_RMSE"]),
"MAPE": np.mean(cv_lr_tp["test_MAPE"])
})
# -------- Predicting Usable Space --------
X_usable_space = df_train[["Total Power (MW)_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]]
y_usable_space = df_train["Usable Space"]
model_lr_us = LinearRegression()
cv_lr_us = cross_validate(model_lr_us, X_usable_space, y_usable_space, cv=kf, scoring=scoring)
# Fit statsmodels OLS for summary
X_sm_us = sm.add_constant(X_usable_space)
ols_us = sm.OLS(y_usable_space, X_sm_us).fit()
print("\n=== OLS Regression Summary: Predicting Usable Space ===")
print(ols_us.summary())
# Append metrics to results
results.append({
"Target": "Usable Space",
"Model": "Linear Regression (OLS)",
"R2": np.mean(cv_lr_us["test_R2"]),
"MAE": np.mean(cv_lr_us["test_MAE"]),
"RMSE": np.mean(cv_lr_us["test_RMSE"]),
"MAPE": np.mean(cv_lr_us["test_MAPE"])
})
# Convert results to DataFrame
results_df = pd.DataFrame(results)
print("\n=== Cross-Validation Results ===")
print(results_df)from sklearn.neighbors import KNeighborsRegressor
from sklearn.model_selection import GridSearchCV
from sklearn.preprocessing import StandardScaler
from sklearn.inspection import permutation_importance
# -------- Predicting Total Power --------
X_total_power = df_train[["Usable Space_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]] # exclude Total Power
y_total_power = df_train["Total Power (MW)"]
# Grid search for best k
param_grid = {'n_neighbors': list(range(1, 10))}
knn = KNeighborsRegressor()
grid_total = GridSearchCV(knn, param_grid, cv=kf, scoring='r2')
grid_total.fit(X_total_power, y_total_power)
best_knn_total = grid_total.best_estimator_
# Cross-validation
cv_knn_tp = cross_validate(best_knn_total, X_total_power, y_total_power, cv=kf, scoring=scoring)
# Permutation importance
perm_imp_total = permutation_importance(best_knn_total, X_total_power, y_total_power, n_repeats=10, random_state=42)
feature_importance_total = pd.Series(perm_imp_total.importances_mean, index=X_total_power.columns)
# Append results
results.append({
"Target": "Total Power",
"Model": "KNN",
"R2": np.mean(cv_knn_tp['test_R2']),
"MAE": np.mean(cv_knn_tp['test_MAE']),
"RMSE": np.mean(cv_knn_tp['test_RMSE']),
"MAPE": np.mean(cv_knn_tp['test_MAPE'])
})
print("Best k for Total Power:", grid_total.best_params_['n_neighbors'])
print("Feature importance (Total Power):")
print(feature_importance_total.sort_values(ascending=False))
# ---------------------------
# Model 2: Predict Usable Space
# ---------------------------
X_usable_space = df_train[["Total Power (MW)_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]]
y_usable_space = df_train["Usable Space"]
# Grid search for best k
grid_usable = GridSearchCV(knn, param_grid, cv=kf, scoring="r2")
grid_usable.fit(X_usable_space, y_usable_space)
best_knn_us = grid_usable.best_estimator_
# Cross-validation
cv_knn_us = cross_validate(best_knn_us, X_usable_space, y_usable_space, cv=kf, scoring=scoring)
# Permutation importance
perm_imp_us = permutation_importance(best_knn_us, X_usable_space, y_usable_space, n_repeats=10, random_state=42)
feature_importance_us = pd.Series(perm_imp_us.importances_mean, index=X_usable_space.columns)
# Append results
results.append({
"Target": "Usable Space",
"Model": "KNN",
"R2": np.mean(cv_knn_us['test_R2']),
"MAE": np.mean(cv_knn_us['test_MAE']),
"RMSE": np.mean(cv_knn_us['test_RMSE']),
"MAPE": np.mean(cv_knn_us['test_MAPE'])
})
print("Best k for Usable Space:", grid_usable.best_params_['n_neighbors'])
print("Feature importance (Usable Space):")
print(feature_importance_us.sort_values(ascending=False))
# Convert results to DataFrame
results_df = pd.DataFrame(results)
print(results_df)
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import GridSearchCV, cross_validate
import pandas as pd
import numpy as np
#Using zipcode instead of lat/long that had a slightly worse result
# -------- Predicting Total Power --------
X_total_power = df_train[["Usable Space", "Colocation Flag", "Year", "Zipcode"]]
y_total_power = df_train["Total Power (MW)"]
param_grid = {
'n_estimators': [50, 100, 200],
'max_depth': [None, 5, 10],
'random_state': [42]
}
rf = RandomForestRegressor()
grid_total = GridSearchCV(rf, param_grid, cv=kf, scoring='r2')
grid_total.fit(X_total_power, y_total_power)
best_rf_total = grid_total.best_estimator_
# Cross-validation
cv_rf_tp = cross_validate(best_rf_total, X_total_power, y_total_power, cv=kf, scoring=scoring)
# Feature importance (computed from the model trained on the same unscaled features)
feature_importance_total = pd.Series(best_rf_total.feature_importances_, index=X_total_power.columns)
# Append results
results.append({
"Target": "Total Power",
"Model": "Random Forest",
"R2": np.mean(cv_rf_tp['test_R2']),
"MAE": np.mean(cv_rf_tp['test_MAE']),
"RMSE": np.mean(cv_rf_tp['test_RMSE']),
"MAPE": np.mean(cv_rf_tp['test_MAPE'])
})
print("Best params for Total Power:", grid_total.best_params_)
print("Feature importance (Total Power):")
print(feature_importance_total.sort_values(ascending=False))
# ---------------------------
# Model 2: Predict Usable Space
# ---------------------------
X_usable_space = df_train[["Total Power (MW)", "Colocation Flag", "Year", "Zipcode"]]
y_usable_space = df_train["Usable Space"]
grid_usable = GridSearchCV(rf, param_grid, cv=kf, scoring='r2')
grid_usable.fit(X_usable_space, y_usable_space)
best_rf_us = grid_usable.best_estimator_
# Cross-validation
cv_rf_us = cross_validate(best_rf_us, X_usable_space, y_usable_space, cv=kf, scoring=scoring)
# Feature importance
feature_importance_us = pd.Series(best_rf_us.feature_importances_, index=X_usable_space.columns)
# Append results
results.append({
"Target": "Usable Space",
"Model": "Random Forest",
"R2": np.mean(cv_rf_us['test_R2']),
"MAE": np.mean(cv_rf_us['test_MAE']),
"RMSE": np.mean(cv_rf_us['test_RMSE']),
"MAPE": np.mean(cv_rf_us['test_MAPE'])
})
print("Best params for Usable Space:", grid_usable.best_params_)
print("Feature importance (Usable Space):")
print(feature_importance_us.sort_values(ascending=False))
# ---------------------------
# Convert results to DataFrame
# ---------------------------
results_df = pd.DataFrame(results)
print("\n=== Cross-Validation Results ===")
print(results_df)
from xgboost import XGBRegressor
#Using Lat/Long instead of Zipcode that produced slightly higher results
# ---------------------------
# Model 1: Predict Total Power
# ---------------------------
X_total_power = df_train[["Usable Space", "Colocation Flag", "Year", "Latitude", "Longitude"]]
y_total_power = df_train["Total Power (MW)"]
param_grid = {
'n_estimators': [50, 100, 200],
'max_depth': [3, 5, 10],
'learning_rate': [0.01, 0.1, 0.2],
'subsample': [0.7, 1.0],
'random_state': [42]
}
xgb = XGBRegressor(objective='reg:squarederror')
grid_total = GridSearchCV(xgb, param_grid, cv=kf, scoring='r2')
grid_total.fit(X_total_power, y_total_power)
best_xgb_total = grid_total.best_estimator_
# Cross-validation
cv_xgb_tp = cross_validate(best_xgb_total, X_total_power, y_total_power, cv=kf, scoring=scoring)
# Feature importance
feature_importance_total = pd.Series(best_xgb_total.feature_importances_, index=X_total_power.columns)
# Append results
results.append({
"Target": "Total Power",
"Model": "XGBoost",
"R2": np.mean(cv_xgb_tp['test_R2']),
"MAE": np.mean(cv_xgb_tp['test_MAE']),
"RMSE": np.mean(cv_xgb_tp['test_RMSE']),
"MAPE": np.mean(cv_xgb_tp['test_MAPE'])
})
print("Best params for Total Power:", grid_total.best_params_)
print("Feature importance (Total Power):")
print(feature_importance_total.sort_values(ascending=False))
# ---------------------------
# Model 2: Predict Usable Space
# ---------------------------
X_usable_space = df_train[["Total Power (MW)", "Colocation Flag", "Year", "Latitude", "Longitude"]]
y_usable_space = df_train["Usable Space"]
grid_usable = GridSearchCV(xgb, param_grid, cv=kf, scoring='r2')
grid_usable.fit(X_usable_space, y_usable_space)
best_xgb_us = grid_usable.best_estimator_
# Cross-validation
cv_xgb_us = cross_validate(best_xgb_us, X_usable_space, y_usable_space, cv=kf, scoring=scoring)
# Feature importance
feature_importance_us = pd.Series(best_xgb_us.feature_importances_, index=X_usable_space.columns)
# Append results
results.append({
"Target": "Usable Space",
"Model": "XGBoost",
"R2": np.mean(cv_xgb_us['test_R2']),
"MAE": np.mean(cv_xgb_us['test_MAE']),
"RMSE": np.mean(cv_xgb_us['test_RMSE']),
"MAPE": np.mean(cv_xgb_us['test_MAPE'])
})
print("Best params for Usable Space:", grid_usable.best_params_)
print("Feature importance (Usable Space):")
print(feature_importance_us.sort_values(ascending=False))
# ---------------------------
# Convert results to DataFrame
# ---------------------------
results_df = pd.DataFrame(results)
print(results_df)from tabulate import tabulate
print(tabulate(results_df, headers='keys', tablefmt='fancy_grid', showindex=False))
# -----------------------------
# Step 0: Create a copy of the original df
# -----------------------------
df_fully_imputed = df.copy() # keep original index
# -----------------------------
# Step 1: Create subsets for imputation (keep original indices)
# -----------------------------
# Missing Total Power but Usable Space available
df_impute_tp = df_fully_imputed.loc[
df_fully_imputed["Total Power (MW)"].isna() & df_fully_imputed["Usable Space"].notna()
]
# Missing Usable Space but Total Power available
df_impute_us = df_fully_imputed.loc[
df_fully_imputed["Total Power (MW)"].notna() & df_fully_imputed["Usable Space"].isna()
]
# Missing both Total Power and Usable Space
df_impute_all = df_fully_imputed.loc[
df_fully_imputed["Total Power (MW)"].isna() & df_fully_imputed["Usable Space"].isna()
]
print(f"Rows from training: {len(df_train)}")
print(f"Rows to impute Total Power only: {len(df_impute_tp)}")
print(f"Rows to impute Usable Space only: {len(df_impute_us)}")
print(f"Rows to impute both: {len(df_impute_all)}")
# Initialize Imputed column with 'Yes'
df_fully_imputed["Imputed"] = "UKN"
# For rows that originally had both Total Power and Usable Space present, mark as "No"
df_fully_imputed.loc[
df_fully_imputed["Total Power (MW)"].notna() & df_fully_imputed["Usable Space"].notna(),
"Imputed"
] = "No"
# Check result
print(df_fully_imputed[["Total Power (MW)", "Usable Space", "Imputed"]].head(10))
# -----------------------------
# Step 2: Scale numeric features using training scaling
# -----------------------------
numeric_cols = ["Usable Space", "Year", "Latitude", "Longitude"]
for col in numeric_cols:
mean = df_train[col + "_scaled"].mean()
std = df_train[col + "_scaled"].std()
df_impute_tp[col + "_scaled"] = (df_impute_tp[col] - mean) / std
# Keep feature order consistent with training
X_impute_knn = df_impute_tp[
["Usable Space_scaled", "Colocation Flag", "Year_scaled", "Latitude_scaled", "Longitude_scaled"]
]
# -----------------------------
# Step 3: Predict using trained KNN
# -----------------------------
y_pred_scaled = best_knn_total.predict(X_impute_knn)
# -----------------------------
# Step 4: Rescale predictions back to original units
# -----------------------------
tp_mean = df_train["Total Power (MW)_scaled"].mean()
tp_std = df_train["Total Power (MW)_scaled"].std()
y_pred_original = y_pred_scaled * tp_std + tp_mean
# -----------------------------
# Step 5: Update df_fully_imputed
# -----------------------------
df_fully_imputed.loc[df_impute_tp.index, "Total Power (MW)"] = y_pred_original
df_fully_imputed.loc[df_impute_tp.index, "Imputed"] = "KNN"
# -----------------------------
# Step 6: Verify imputation
# -----------------------------
null_counts = df_fully_imputed.loc[df_impute_tp.index].isna().sum()
print("Number of missing values per column in df_impute_tp after imputation:")
print(null_counts)
# -----------------------------
# Step 2: Prepare predictors for imputation
# -----------------------------
predictor_cols = ["Usable Space", "Colocation Flag", "Year", "Zipcode"]
X_impute_tp = df_impute_tp[predictor_cols].copy()
# -----------------------------
# Step 3: Predict missing Total Power using trained Random Forest
# -----------------------------
y_pred_tp = best_rf_total.predict(X_impute_tp)
# -----------------------------
# Step 4: Update df_fully_imputed in place using the original indices
# -----------------------------
df_fully_imputed.loc[df_impute_tp.index, "Total Power (MW)"] = y_pred_tp
df_fully_imputed.loc[df_impute_tp.index, "Imputed"] = "Random Forest"
# -----------------------------
# Step 5: Check results
# -----------------------------
print(f"Imputed Usable Space for {len(df_impute_tp)} rows in df_fully_imputed:")
print(df_fully_imputed.loc[df_impute_tp.index, predictor_cols + ["Total Power (MW)", "Imputed"]])
# -----------------------------
# Step 6: Optional - count remaining nulls
# -----------------------------
null_counts = df_fully_imputed.loc[df_impute_tp.index].isna().sum()
print("\nNull counts in df_fully_imputed for these rows after imputation:")
print(null_counts)len(df_fully_imputed[df_fully_imputed["Imputed"]=="Random Forest"])df_fully_imputed[df_fully_imputed["Imputed"]=="Random Forest"]
# -----------------------------
# Step 2: Prepare predictors for imputation
# -----------------------------
predictor_cols = ["Total Power (MW)", "Colocation Flag", "Year", "Latitude", "Longitude"]
X_impute_us = df_impute_us[predictor_cols].copy()
# -----------------------------
# Step 3: Predict missing Usable Space using trained XGBoost
# -----------------------------
y_pred_us = best_xgb_us.predict(X_impute_us)
# -----------------------------
# Step 4: Update df_fully_imputed in place using the original indices
# -----------------------------
df_fully_imputed.loc[df_impute_us.index, "Usable Space"] = y_pred_us
df_fully_imputed.loc[df_impute_us.index, "Imputed"] = "XGBoost"
# -----------------------------
# Step 5: Check results
# -----------------------------
print(f"Imputed Usable Space for {len(df_impute_us)} rows in df_fully_imputed:")
print(df_fully_imputed.loc[df_impute_us.index, predictor_cols + ["Usable Space", "Imputed"]])
# -----------------------------
# Step 6: Optional - count remaining nulls
# -----------------------------
null_counts = df_fully_imputed.loc[df_impute_us.index].isna().sum()
print("\nNull counts in df_fully_imputed for these rows after imputation:")
print(null_counts)len(df_fully_imputed[df_fully_imputed["Imputed"]=="XGBoost"])import pandas as pd
import numpy as np
# Identify rows to impute
mask_both_null = df_fully_imputed["Total Power (MW)"].isna() & df_fully_imputed["Usable Space"].isna()
print(f"Rows to impute both Total Power and Usable Space: {mask_both_null.sum()}")
# -----------------------------
# Step 1: Create decade column
# -----------------------------
df_fully_imputed["Decade"] = (df_fully_imputed["Year"] // 10).astype(int) * 10
# -----------------------------
# Step 2: Compute median values from *original df* only where not null
# -----------------------------
median_by_company_decade = (
df[df["Total Power (MW)"].notna() & df["Usable Space"].notna()]
.groupby(["Company", (df["Year"] // 10).astype(int) * 10])[["Total Power (MW)", "Usable Space"]]
.median()
)
median_by_decade = (
df[df["Total Power (MW)"].notna() & df["Usable Space"].notna()]
.groupby((df["Year"] // 10).astype(int) * 10)[["Total Power (MW)", "Usable Space"]]
.median()
)
overall_medians = (
df[df["Total Power (MW)"].notna() & df["Usable Space"].notna()][["Total Power (MW)", "Usable Space"]]
.median()
)
# -----------------------------
# Step 3: Define safe imputation function
# -----------------------------
def impute_row_safe(row):
company = row["Company"]
decade = row["Decade"]
# Default: overall median
tp_value = overall_medians["Total Power (MW)"]
us_value = overall_medians["Usable Space"]
imputed_label = "Overall"
# Try company + decade median
if (company, decade) in median_by_company_decade.index:
tp_value = median_by_company_decade.loc[(company, decade), "Total Power (MW)"]
us_value = median_by_company_decade.loc[(company, decade), "Usable Space"]
imputed_label = "Company/Decade"
# Else, try decade median
elif decade in median_by_decade.index:
tp_value = median_by_decade.loc[decade, "Total Power (MW)"]
us_value = median_by_decade.loc[decade, "Usable Space"]
imputed_label = "Decade"
# Assign imputed values
row["Total Power (MW)"] = tp_value
row["Usable Space"] = us_value
row["Imputed"] = imputed_label
return row
# -----------------------------
# Step 4: Apply to rows needing imputation
# -----------------------------
df_fully_imputed.loc[mask_both_null] = (
df_fully_imputed.loc[mask_both_null].apply(impute_row_safe, axis=1)
)
# -----------------------------
# Step 5: Verify results
# -----------------------------
print("Sample imputed rows for both Total Power and Usable Space:")
print(df_fully_imputed.loc[mask_both_null, ["Company", "Year", "Decade", "Total Power (MW)", "Usable Space", "Imputed"]])
# Check nulls in the imputed subset
null_counts = df_fully_imputed.loc[mask_both_null, ["Total Power (MW)", "Usable Space"]].isna().sum()
print("\nNull counts in df_fully_imputed for imputed rows:")
print(null_counts)
print(df_fully_imputed.head())df_fully_imputed.isna().sum()len(df_fully_imputed[df_fully_imputed["Imputed"]=="No"])df_fully_imputed.to_excel("df_fully_imputed_11.15.25.xlsx", index=False)
dc_scrape_size_power_colocation.ipynb4 cells · supporting · largeWeb-scrapes DataCenters.com for facility names, addresses, total power (MW), total and colocation space. The large second cell contains a scraped-data dump as a Python literal and has been trimmed here to the first ~2 KB for readability — the full literal is available by re-running the scraper (Cell 3 in this notebook).
!apt-get update > /dev/null
!apt install chromium-chromedriver > /dev/null
!pip install selenium --quiet
final_name_address = [{
"id": 6621,
"locationId": 5280,
"name": "Riverside 1 Data Center",
"fullAddress": "1550 Marlborough Avenue, Riverside, CA, USA",
"sponsoredGlobal": False,
"sponsoredCountry": False,
"sponsoredState": False,
"sponsoredCity": False,
"url": "/lumen-riverside-1",
"providerId": 200038,
"providerName": "Lumen",
"providerAgreement": True,
"mainImageUrl": "https://res.cloudinary.com/hjlz68xhm/image/upload/hgewxpznetxgbporzgmu",
"latitude": 33.9970355,
"longitude": -117.3456799,
"providerLogo": "https://res.cloudinary.com/hjlz68xhm/image/upload/v1707770321/utkipf0eps6go6nejfmy.png"
},
{
"id": 6910,
"locationId": 7730,
"name": "Los Angeles - Century Data Center",
"fullAddress": "6171 West Century Boulevard, Los Angeles, CA, USA",
"sponsoredGlobal": False,
"sponsoredCountry": False,
"sponsoredState": False,
"sponsoredCity": False,
"url": "/quadranet-los-angeles-century",
"providerId": 120,
"providerName": "QuadraNet",
"providerAgreement": False,
"mainImageUrl": "https://res.cloudinary.com/hjlz68xhm/image/upload/m0cq9zex5emyc1ytmkgg",
"latitude": 33.9459317,
"longitude": -118.3935351,
"providerLogo": "https://res.cloudinary.com/hjlz68xhm/image/upload/lm94nvmn6wib4xpf63ap"
},
{
"id": 6737,
"locationId": 5587,
"name": "Sunnyvale 1 Data Center",
"fullAddress": "1380 Kifer Road, Sunnyvale, CA, USA",
"sponsoredGlobal": False,
"sponsoredCountry": False,
"sponsoredState": False,
"sponsoredCity": False,
"url": "/lumen-sunnyvale-1",
"providerId": 200038,
"providerName": "Lumen",
"providerAgreement": True,
"mainImageUrl": "https://res.cloudinary.com/hjlz68xhm/image/upload/v1/locations/lumen-sunnyvale-1/images/ublca3eg4jlbtdgdzpmz",
"latitude": 37.3734237,
"longitude": -121.9876548,
"providerLogo": "https://res.cloudinary.com/hjlz68xhm/image/upload/v1707770321/utkipf0eps6go6nejfmy.png"
},
{
"id": 9398,
"locationId": 8650,
"name": "San Diego Data Center",
"fullAddress": "9725 Scranton Rd, San Diego, CA, USA",
"sponsoredGlobal": False,
"sponsoredCountry": False,
"spo
# ... [TRIMMED: original cell is 211,209 chars, showing first 2,000. This cell contains a large scraped-data dump embedded as a Python literal — the full list of 197 data-center records. Full notebook available on request.] ...
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import pandas as pd
import time
# Setup Selenium Chrome driver for Colab
options = Options()
options.add_argument('--headless')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev_shm_usage')
driver = webdriver.Chrome(options=options)
# Use the existing datacenter_name_address list from cell -VzhPOHaHK7J
# Ensure datacenter_name_address is a list of dictionaries as expected
if 'final_name_address' in locals() and isinstance(final_name_address, list):
data_centers_combined = []
for dc_info in final_name_address:
if isinstance(dc_info, dict):
# Extract basic info already available
data_centers_combined.append({
'Company': dc_info.get('providerName'), # Assuming providerName is the company
'Name': dc_info.get('name'),
'Full Address': dc_info.get('fullAddress'),
'Latitude': dc_info.get('latitude'),
'Longitude': dc_info.get('longitude'),
'URL': dc_info.get('url'),
'Total Space': None, # Initialize, will be scraped later
'Total Power': None, # Initialize, will be scraped later
'Colocation Space': None # Initialize for colocation space
})
else:
print(f"Skipping unexpected item in datacenter_name_address: {dc_info}")
print(f"Starting to scrape detail pages for Total Space, Total Power, and Colocation Space for {len(data_centers_combined)} data centers...")
# Now, iterate through the collected data centers and scrape details using Selenium
for dc in data_centers_combined:
url_path = dc['URL']
# Construct the full URL if the URL in the list is just the path
if url_path and not url_path.startswith('http'):
url = "https://www.datacenters.com" + url_path
else:
url = url_path # Use as is if it's already a full URL
name = dc['Name']
if not url:
print(f"Skipping scraping for {name} due to missing URL.")
continue
print(f"Scraping Total Space, Power, and Colocation Space for: {name} ({url})")
try:
driver.get(url)
# Wait for an element that indicates the page has loaded
WebDriverWait(driver, 10).until(
EC.presence_of_element_located((By.CSS_SELECTOR, 'h1')) # Wait for the main heading as a general indicator
)
dc_soup = BeautifulSoup(driver.page_source, 'html.parser')
# Total Space
total_space = None
total_space_div = dc_soup.find('div', id='totalSpace')
if total_space_div:
stat_info_div = total_space_div.find('div', id='statInfo')
if stat_info_div:
strong_tag = stat_info_div.find('strong')
if strong_tag:
total_space = strong_tag.get_text(strip=True)
dc['Total Space'] = total_space # Update the list
# Total Power
total_power = None
total_power_div = dc_soup.find('div', id='power')
if total_power_div:
stat_info_div = total_power_div.find('div', id='statInfo')
if stat_info_div:
strong_tag = stat_info_div.find('strong')
if strong_tag:
total_power = strong_tag.get_text(strip=True)
dc['Total Power'] = total_power # Update the list
# Colocation Space (Assuming the ID is 'colocationSpace' based on typical patterns)
colocation_space = None
colocation_space_div = dc_soup.find('div', id='colocationSpace')
if colocation_space_div:
stat_info_div = colocation_space_div.find('div', id='statInfo')
if stat_info_div:
strong_tag = stat_info_div.find('strong')
if strong_tag:
colocation_space = strong_tag.get_text(strip=True)
dc['Colocation Space'] = colocation_space # Update the list
except Exception as e:
print(f"Failed to scrape data for {url}: {e}")
# Total Space, Total Power, and Colocation Space will remain None as initialized
time.sleep(1) # polite pause between detail pages
driver.quit()
df_final_combined = pd.DataFrame(data_centers_combined)
print(df_final_combined)
else:
print("Error: datacenter_name_address variable is not defined or is not a list.")df_final_combined.to_csv('datacenters_space_power_colocation.csv', index=False)
from google.colab import files
files.download('datacenters_space_power_colocation.csv')dc_scrape_founding_dates.ipynb5 cells · supportingFills in year-of-construction for each facility — with parent-company founding year as a last-resort fallback.
company_names = ["Amazon AWS",
"AT&T",
"Atlantic Metro Communications",
"Atlantic.net",
"BreezeHost.io",
"Cato Digital, Inc",
"CBRE",
"CBTS LLC",
"Claranet",
"Colo Locker - Evocative",
"Cologix",
"CoreSite",
"Data Canopy",
"Datacate, Inc",
"EdgeConnex",
"Enzu",
"Fortress Data Centers",
"Hivelocity",
"IBM Cloud",
"Krypt",
"LeaseWeb",
"Megaport",
"Microsoft Azure",
"NOVVA",
"One Data Center America",
"Psychz Networks",
"Quest Technology Management",
"Summit",
"Synoptek",
"unWired Broadband LLC",
"Vantage Data Centers",
"Vultr"]
base_url = "https://www.datacenters.com/providers/"urls = []
for name in company_names:
formatted_name = name.lower().replace(" ", "-")
full_url = base_url + formatted_name
urls.append(full_url)
print(urls)import requests
from bs4 import BeautifulSoup
results = []
for url in urls:
try:
response = requests.get(url)
response.raise_for_status() # Raise an exception for bad status codes
soup = BeautifulSoup(response.content, 'html.parser')
year_founded_div = soup.find('div', id='yearFounded')
if year_founded_div:
year_founded = year_founded_div.get_text(strip=True)
# Extract company name from URL
company_name_formatted = url.split('/')[-1]
company_name = company_name_formatted.replace('-', ' ').title()
results.append({'company_name': company_name, 'year_founded': year_founded})
else:
print(f"Could not find year founded for {url}")
except requests.exceptions.RequestException as e:
print(f"Error fetching {url}: {e}")
display(results)cleaned_results = []
for result in results:
year_founded = result['year_founded'].replace('Calendar', '').replace('year founded', '')
cleaned_results.append({'company_name': result['company_name'], 'year_founded': year_founded.strip()})
display(cleaned_results)import pandas as pd
df_results = pd.DataFrame(cleaned_results)
display(df_results)dc_broadband_fiber_instrument.ipynb7 cells · supportingBuilds the fiber-broadband instrument (provider count per county, mean max upload / download speeds) from FCC Form 477 2014 vintage — the exogenous variation used in the 2SLS.
import pandas as pd
df = pd.read_csv('/content/bdc_06_FibertothePremises_fixed_broadband_D24_24dec2025.csv')
print("First 5 rows of the DataFrame:")
print(df.head())
print("\nDataFrame Information (columns and data types):")
print(df.info())df['Geoid'] = df['block_geoid'].astype(str).str[:4].astype(int)
print("First 5 rows with new 'Geoid' column:")
print(df.head())
print("\nDataFrame Information (columns and data types) with new 'Geoid' column:")
print(df.info())df = df[df['business_residential_code'] != 'R']
print("First 5 rows after filtering out 'R' from 'business_residential_code':")
print(df.head())
print("\nDataFrame Information after filtering:")
print(df.info())df['year_month'] = '2024-12'
print("First 5 rows with new 'year_month' column:")
print(df.head())
print("\nDataFrame Information with new 'year_month' column:")
print(df.info())aggregated_df = df.groupby('Geoid').agg(
avg_max_advertised_download_speed=('max_advertised_download_speed', 'mean'),
avg_max_advertised_upload_speed=('max_advertised_upload_speed', 'mean'),
unique_provider_count=('provider_id', 'nunique')
).reset_index()
df = pd.merge(df, aggregated_df, on='Geoid', how='left')
print("First 5 rows with new aggregated columns:")
print(df.head())
print("\nDataFrame Information with new aggregated columns:")
print(df.info())final_df = aggregated_df.copy()
final_df['year_month'] = '2024-12'
print("First 5 rows of the new DataFrame:")
print(final_df.head())
print("\nDataFrame Information for the new DataFrame:")
print(final_df.info())final_df.to_excel('/content/broadband_2024-12.xlsx', index=False)
print("DataFrame exported successfully to '/content/broadband_2024-12.xlsx'")Upstream: raw data via CAISO OASIS, CEC, CalEnviroScreen 4.0, NOAA GHCN, EIA-860, FCC Form 477 (2014), ACS 5-year, CPUC NEM, and a scraped data center directory (197 facilities). Data warehouse: BigQuery. Transformations: dbt (staging → intermediate → domain). Analysis notebooks: Google Colab / Deepnote. IV estimation: linearmodels.iv.IV2SLS with cluster-robust SE; two-way FE via explicit dummies or iterative Gauss-Seidel demeaning per Frisch-Waugh-Lovell.