Skip to content

Weight the ACS local release's calibration targets with the national release's loss formula, from one shared implementation - #1104

Merged
MaxGhenis merged 8 commits into
mainfrom
acs-local-target-loss-weights
Oct 5, 2026
Merged

MaxGhenis merged 8 commits into
mainfrom
acs-local-target-loss-weights

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

What this does

Until now the ACS local-area release calibrated with every target weighted equally. Max asked for this change on 2026-10-01, after Anthony raised it at Data weekly. The calibrate stage now weights each training target's loss with the national release's formula, sqrt_value_concept_budget_weighted_mape_50_50_amount_count_target_scale_cap_100pct, from one shared implementation, under an explicit ACS local row mapping.

Shared module. These national helpers move verbatim into microcosm.build.us_runtime.target_loss_weights:

  • _fiscal_target_loss_weights and its basis, concept-key and budget helpers;
  • the constants US_FISCAL_TARGET_VALUE_WEIGHT_POWER, US_FISCAL_TARGET_CONCEPT_METADATA_EXCLUSIONS and US_FISCAL_TARGET_LOSS_WEIGHTING;
  • the --target-family-loss-multiplier parser.

tools/build_us_fiscal_refresh_release.py imports them back under the names the scorers and the bundle contract read off it. For the national path this is a pure move for every input; see "National golden" below.

ACS local calibrate stage (tools/build_us_acs_local_release.py):

  • It computes the weights over the training targets, row-aligned with the calibration TargetSet, and passes them to calibrate() on every epoch batch.
  • It adds --target-family-loss-multiplier FAMILY=MULTIPLIER, parsed by the same function the national release uses.
  • solver_settings gains target_loss: the formula id, the row mapping, the multipliers and the weights' sha256. --resume and the completed-run shortcut therefore refuse weights calibrated under another weighting, or under none (every run before this change). Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier #1078's _stamped_settings does not fill this key.
  • The weights' stamp and digest are recorded in calibration_summary.json (target_loss, with the distribution by family × level × basis and the concept-group counts) and in consumer_export.json (via the settings).
  • The per-target weights are in calibration_diagnostics.json: per-target target_loss_weight/_share/_scale and the top-level target_loss_basis, as for the national release. Its build block carries the fields the calibration dashboard reads (target_loss_weighting, target_loss_family_multipliers, target_loss_cap), plus the row mapping, the digest and the distribution.
  • The package manifest's calibration block gains target_loss. The refresh recipe names each recorded multiplier, as Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier #1078's recipe names the penalty settings.
  • Finalize's ESS register keeps the "reviewed" concentration entry only for the equal-weight historical solve. A weighted solve records its own settings.

The ACS local row mapping

Reusing the national mapping as is would be wrong here in two ways. Both are measured in experiments/us-acs-local-target-loss-weights-20261004/.

  1. Ladder population rows (pop_state_*, pop_cd_*) carry only {"geography_level": …}. The national basis rule reads measure_mode and source_measure_id, so it silently files them as amounts. They then get 0.1% of the loss, and no district group shares a budget.
  2. District SOI rows. The ledger compiler stamps ledger_fact_label and ledger_layout_groupby_value_label on every row, and on district rows they name the district ("WV congressional district 1"). The state_cd rebase stamps state_cd_cd_file_value, which also differs by district. None of the three is in the national exclusion set. On the pinned feed, all 19,105 trained state_cd SOI district rows (and the 436 pop_cd_* rows) would be singleton groups.

US_ACS_LOCAL_TARGET_LOSS_ROW_MAPPING accepts two kinds of row and refuses everything else:

Rows Basis Concept budget
Ledger rows of any family (on the surface: SNAP, Medicaid enrollment, SOI) The national rule, which must agree with the row's ledger_measure_unit (count/usd). The row must carry measure_mode and source_measure_id. One reviewed exception, tax_filer_individual_count (people; national rule says amount), is mapped to count by ACS_LOCAL_LEDGER_BASIS_OVERRIDES. Any other disagreement is refused. National key, without the three per-district keys above
pop_state_SS (family census_population, geography_level=state) count, as the national Census population targets (measure_mode=indicator_sum) are Own group
pop_cd_SSDD count One group per state
A row that is neither Refused

validate_acs_local_concept_groups refuses three things in any SOI mode:

  • district rows of one concept in one state that land in more than one group (district_concept_identity);
  • a state_cd reconciliation block split across groups, or two blocks sharing one;
  • a district group spanning states.

On the pinned feed, every SOI mode (state, totals, full, state_cd) classifies. Injecting a per-district key into every district row is refused on each surface that has district ledger rows (results.json soi_modes).

Weight distribution, before and after

From experiments/us-acs-local-target-loss-weights-20261004/distribution.py (results.json, at f5f3806, clean tree). Shares are of total weight, which is the share of the loss each kind of target carries at equal scaled misses.

The 09-23 release's surface (--soi-mode state, the default): 4,459 targets, all trained.

Family Level Targets Before (equal) National mapping as is After
census_population district 436 9.8% 0.1% 1.7%
census_population state 51 1.1% 0.0% 4.4%
cms_medicaid state 51 1.1% 2.3% 2.0%
irs_soi state 3,819 85.6% 95.7% 90.2%
usda_snap state 102 2.3% 1.8% 1.7%

Counts and amounts carry 50% each after the change (55.5%/44.5% before). The 436 district population rows form 51 groups; California's 52 districts carry 1.51 of weight between them. The weights' digest is ba36f887….

--soi-mode state_cd on the pinned feed: 23,866 training targets after the default 10% district holdout.

Family Level Targets Before (equal) National mapping as is After
census_population district 436 1.8% 0.0% 1.4%
census_population state 51 0.2% 0.0% 3.6%
cms_medicaid state 51 0.2% 0.8% 1.6%
irs_soi district 19,105 80.1% 64.2% 16.7%
irs_soi state 4,121 17.3% 34.4% 75.4%
usda_snap state 102 0.4% 0.6% 1.3%

The 1,907 trained state_cd blocks each form exactly one group.

The formula balances counts against amounts, not families. SOI's share rises on the default surface (85.6% → 90.2%) and falls on state_cd (97.3% → 92.1%), where district rows stop multiplying their concept's weight. Route A doesn't balance families either: SOI is 74.0% of its targets and 74.1% of its loss weight. The family lever is --target-family-loss-multiplier; this PR sets no default.

Cost: district population fit without a multiplier

The full-scale weighted λ frontier in microcosm#1105 ("On the weighted loss" in experiments/us-acs-local-l2-basis-20260928/README.md) solved the 09-23 checkpoint on these weights.

  • Equal weights: every trained district population is within 0.7% at λ 0.
  • These weights, no multiplier: 11 of 436 miss by more than 10% at λ 0 (worst TX-14 −33%, CA-29 −32%). The penalty raises the count to 73 at λ 0.03 and 135 at λ 0.1.
  • Why: each state's pop_cd rows share one budget, as the national concept budget prescribes.
  • With census_population=8: district population is 9.5% of the loss (I recomputed this from these weights). No district misses by more than 10% at λ 0, and 3 do at λ 0.03 (worst 15%).

#1105 recommends that multiplier for the next build. The choice is decision d797, and nothing is rebuilt here.

National golden

The national weights are byte-identical before and after the move.

  • Route A, the certified release built from 4b57d15. tests/fixtures/us_route_a_target_loss_weights.json holds its 5,694 targets, with the fields the formula reads and the weights it recorded. The shared module and the tool's _fiscal_target_loss_weights reproduce all 5,694 weights bit for bit. With the library's default scales they also give the release's recorded loss basis hash, 206ada09584f6e84dd8f32e58d70fea8db257d530c6a1f4ff115630e500436e7. route_a_fixture.py pins the diagnostics' sha256 (64b55a02…), and it refuses to write unless the reduced fields reproduce the recorded weights. Route A has no district rows, so it pins steps 2, 4 and 5 of the formula.
  • Differential. test_support/microcosm_build/us_target_loss_weights_reference.py is a verbatim copy of the pre-move functions and constants from 48fa061. Three tests hold the shared module to it bit for bit:
    • a Hypothesis test on random registries whose district rows form multi-row concept groups (57-72% of draws across runs, most of them a group of 3+), with random multipliers;
    • a fixed registry with two three-district groups;
    • unchecked multipliers (0, −1, 1e308). The shared core adds no check the national helper lacked.
  • The national CLI's multiplier parse is tested.
  • The shared digest's canonical form equals the national _fiscal_target_loss_basis(...)["loss_vector_sha256"], tested on Route A.

Invariants (property-tested with Hypothesis)

For every input drawn:

  1. Pure move, as above.
  2. Positive, finite, mean 1 (multipliers in [0.05, 20]).
  3. Equal bases. Without multipliers, each basis present carries an equal share of the total weight.
  4. Concept budget. A group's rows sum to the group's largest value weight and keep their proportions. Every row of a basis is then scaled by one factor. Adding districts to a concept (1, 2, 7 or 52) leaves its share unchanged.
  5. Family multipliers. A multiplier scales its family relative to every other row by exactly that multiplier.
  6. Order-free. This holds for the formula and for the ACS local mapping.
  7. Monotone. Within a basis, a larger magnitude never gets a smaller singleton weight.
  8. ACS local mapping. Every surface row is classified explicitly. Each refusal raises:
    • a row that is neither a ledger row nor a ladder row;
    • an unknown unit, a unit/basis disagreement, or missing basis keys;
    • malformed population names (including non-digit FIPS) or levels;
    • a district row without state_fips;
    • a concept split across groups, a split or pooled block, or a cross-state group.

End-to-end. test_us_acs_local_target_loss_weights.py runs do_calibrate on a surface-shaped fixture: state SOI, SNAP, Medicaid, two state_cd blocks (one held out) and ladder population. It checks:

  • the weights reach every calibrate() batch, row-aligned and non-uniform;
  • the summary, weights_latest.npz and diagnostics record them (consumer_export.json is not written there, since the test builds no H5);
  • a multiplier reaches the solve and the recipe;
  • --resume refuses an unweighted or differently weighted stamp;
  • a misaligned TargetSet is refused;
  • the ESS register no longer certifies a weighted solve.

test_us_acs_local_release_tool.py checks that the package manifest records target_loss and the recipe flag.

A national finding, not changed here

experiments/…/national_cd_groups.py (national_cd_groups.json) compiles the national registry from the pinned feed and keeps its 24,340 district rows. --target-surface full (the parser default) calibrates them.

  • Under the national mapping they form 24,340 singleton groups and carry 47.0% of the loss.
  • With the two compiler labels excluded too, they form 3,125 groups and carry 11.6%.

The existing national budget test uses fixture rows without those labels. Route A calibrated --target-surface national_state, which has no district rows, so it is unaffected. This PR keeps the national path a pure move, so the fix is a separate follow-up task.

Scope and limits

  • This changes the ACS local calibration objective for future builds only. Nothing is rebuilt, republished or promoted. Rebuilding a local release, and which multiplier and λ it uses, is a separate call (d797).
  • The Modal stage plan (tools/modal_us_stage_plan.py) cannot yet pass --target-family-loss-multiplier, nor Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier #1078's --l2-basis and --mass-parametrization. That is a separate follow-up task. Default runs are unaffected: the weights always apply, and no multiplier is the default.
  • Route A's pinned worktree and commit are untouched.
  • The diff to the ACS tool stays inside calibrate, finalize's ESS entry, the CLI, the recipe and one manifest key.

axiom: n/a: calibration infrastructure, no policy logic.

Review

Independent review by in-session Opus 5.5 agents (Subfleet had no dispatchable lanes on 2026-10-04).

Round 1, at 4d87854. Five lenses: pure move, ACS mapping, executed mutation testing (45 mutants), claims audit, and CI/integration risk. An adversarial verifier then checked every medium-or-higher finding. The confirmed findings, all fixed in 2d56c32, 92ed562 and f5f3806:

  • CI was red because the us_runtime module was unclassified.
  • The Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier #1078 merge broke two of its tests.
  • --soi-mode full was refused on tax_filer_individual_count.
  • Finalize's ESS register claimed the reviewed solve.
  • The recipe did not pin multipliers.
  • The grouping guard covered state_cd only.
  • The pure-move differential never formed multi-row groups. Mutant M49 survived it then; it is killed now.
  • The national CLI multiplier parse had no test.
  • The shared core added a multiplier check.
  • Several claims were wrong: stale digests, the SOI share wording, "unknown family", the docstring's statement about national grouping, 19,105, 74.1%, and the provenance.

Round 2, at c4fd2e8: APPROVE. The reviewer re-checked all 11 findings:

  • reran the suites: 546 + 483 passed, plus 88 with the feed set;
  • ran 16 new mutants: 15 killed, and the survivor is inert on the real surfaces;
  • classified the real full, totals, state and state_cd surfaces: a one-to-one correspondence between district concepts and keys in every mode, with no false refusals;
  • reproduced the ba36f887… digest.

Round 2's non-blocking notes:

  • Two wording fixes, made in 7bde9d8, the only change since the approved head (3 lines): the override's reason string and the README now say only the full surface has tax_filer_individual_count rows. The draw-share figure above is now a range.
  • The guard's limits. A district-specific value inside one of the concept-identity keys would split the concept and its key together, and pass. Two concepts differing only outside the identity would be refused, which fails safe. Neither occurs on the pinned feed, and Bind state × AGI-band SOI targets and re-pin the US feed for TY2023 state bands (#940) #1102's re-pinned feed classifies cleanly (round 1, CI lens).

Reports: ~/PolicyEngine/_worktrees/.reviews/pr1104-*.md.

🤖 Generated with Claude Code

MaxGhenis and others added 5 commits October 4, 2026 14:37
The national fiscal release's loss-weight formula (sqrt value weight, concept
budgets for district groups, equal count/amount budgets, family multipliers)
moves verbatim into microcosm.build.us_runtime.target_loss_weights. The
release tool keeps importing the names other tools read off it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The calibrate stage computes the shared formula's weights over the training
targets under an explicit ACS local row mapping and passes them to
calibrate(). Population ladder rows are counts; each state's pop_cd rows and
each state_cd reconciliation block share one concept budget. The weights'
digest joins the solver settings, so a resume refuses weights from an
unweighted run, and the summary, diagnostics and package manifest record them.
Adds --target-family-loss-multiplier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ore/after

- Hypothesis properties of the formula and a differential against a frozen
  copy of the pre-move national helpers.
- A golden on Route A's 5,694 recorded weights and loss basis hash.
- The ACS local row mapping's classifications and refusals.
- do_calibrate end to end on a surface-shaped fixture.
- Evidence scripts and results for the 09-23 and state_cd surfaces, and the
  national district-label finding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… module

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolve the calibrate-stage conflicts: keep both the penalty settings and the
target_loss stamp in _solver_settings, keep _stamped_settings (it does not fill
target_loss, so a stamp from before the weights still never matches), name any
recorded --target-family-loss-multiplier in the refresh recipe, and give
#1078's two do_calibrate tests a surface the loss-weight mapping can classify.
Regenerate results.json: its weight digests predated the name@period row
names (every other field is byte-identical).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
Receipts in results/runs/w_*.json (kernel c028b6ad4, shared weights from
#1104's target_loss_weights.py 849fbede…, full-surface digest
ba36f887…). Every run's epoch-0 loss matches the recomputed weighted
loss within 1e-5 and its final loss exactly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 2 commits October 4, 2026 19:06
… guard, stricter pure-move tests

- Map tax_filer_individual_count (people; ledger unit count) to the count
  basis explicitly, so --soi-mode full calibrates; refuse any other
  unit/basis disagreement.
- Refuse district rows of one concept in one state that land in more than
  one concept group in any SOI mode, not only state_cd blocks.
- Drop the multiplier check the shared core added, so the move is pure for
  every input; refuse multipliers with no training targets in the ACS caller.
- Finalize records a weighted solve's ESS instead of claiming the reviewed
  equal-weight entry.
- Tests: the pure-move strategy now forms multi-row district groups (66% of
  draws), a fixed grouped differential, unchecked multipliers as before, the
  national CLI parse, the override, the guard, equal_share, the diagnostics
  distribution, the ESS entry; the package-manifest test moves beside its
  helpers.
- Evidence: an every-SOI-mode receipt, the clean-tree flag, the Route A
  diagnostics sha pin, the 09-23 registry receipt, and the district
  population cost from microcosm#1105; corrected wording and figures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… weights unchanged)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis MaxGhenis changed the title Weight the ACS local release's calibration targets like the national release, from one shared implementation Weight the ACS local release's calibration targets with the national release's loss formula, from one shared implementation Oct 4, 2026
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
…le hash

#1104's review fixes changed target_loss_weights.py (849fbede… to
0c2a6a9d…) without changing these weights. check_weights.py rebuilds every
w_ run's training and yardstick weights with the current module and
compares the digests with the receipts: 46 runs, 7 weightings, no mismatch
at #1104's head c4fd2e8. Rerun against main before merging.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis merged commit e34712c into main Oct 5, 2026
10 checks passed
@MaxGhenis
MaxGhenis deleted the acs-local-target-loss-weights branch October 5, 2026 00:28
MaxGhenis added a commit that referenced this pull request Oct 5, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant