Repository navigation
Weight the ACS local release's calibration targets with the national release's loss formula, from one shared implementation - #1104
Merged
Conversation
The national fiscal release's loss-weight formula (sqrt value weight, concept budgets for district groups, equal count/amount budgets, family multipliers) moves verbatim into microcosm.build.us_runtime.target_loss_weights. The release tool keeps importing the names other tools read off it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The calibrate stage computes the shared formula's weights over the training targets under an explicit ACS local row mapping and passes them to calibrate(). Population ladder rows are counts; each state's pop_cd rows and each state_cd reconciliation block share one concept budget. The weights' digest joins the solver settings, so a resume refuses weights from an unweighted run, and the summary, diagnostics and package manifest record them. Adds --target-family-loss-multiplier. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ore/after - Hypothesis properties of the formula and a differential against a frozen copy of the pre-move national helpers. - A golden on Route A's 5,694 recorded weights and loss basis hash. - The ACS local row mapping's classifications and refusals. - do_calibrate end to end on a surface-shaped fixture. - Evidence scripts and results for the 09-23 and state_cd surfaces, and the national district-label finding. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… module Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolve the calibrate-stage conflicts: keep both the penalty settings and the target_loss stamp in _solver_settings, keep _stamped_settings (it does not fill target_loss, so a stamp from before the weights still never matches), name any recorded --target-family-loss-multiplier in the refresh recipe, and give #1078's two do_calibrate tests a surface the loss-weight mapping can classify. Regenerate results.json: its weight digests predated the name@period row names (every other field is byte-identical). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Oct 4, 2026
Receipts in results/runs/w_*.json (kernel c028b6ad4, shared weights from #1104's target_loss_weights.py 849fbede…, full-surface digest ba36f887…). Every run's epoch-0 loss matches the recomputed weighted loss within 1e-5 and its final loss exactly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… guard, stricter pure-move tests - Map tax_filer_individual_count (people; ledger unit count) to the count basis explicitly, so --soi-mode full calibrates; refuse any other unit/basis disagreement. - Refuse district rows of one concept in one state that land in more than one concept group in any SOI mode, not only state_cd blocks. - Drop the multiplier check the shared core added, so the move is pure for every input; refuse multipliers with no training targets in the ACS caller. - Finalize records a weighted solve's ESS instead of claiming the reviewed equal-weight entry. - Tests: the pure-move strategy now forms multi-row district groups (66% of draws), a fixed grouped differential, unchecked multipliers as before, the national CLI parse, the override, the guard, equal_share, the diagnostics distribution, the ESS entry; the package-manifest test moves beside its helpers. - Evidence: an every-SOI-mode receipt, the clean-tree flag, the Route A diagnostics sha pin, the 09-23 registry receipt, and the district population cost from microcosm#1105; corrected wording and figures. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… weights unchanged) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Oct 4, 2026
…le hash #1104's review fixes changed target_loss_weights.py (849fbede… to 0c2a6a9d…) without changing these weights. check_weights.py rebuilds every w_ run's training and yardstick weights with the current module and compares the digests with the receipts: 46 runs, 7 weightings, no mismatch at #1104's head c4fd2e8. Rerun against main before merging. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Oct 5, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
Until now the ACS local-area release calibrated with every target weighted equally. Max asked for this change on 2026-10-01, after Anthony raised it at Data weekly. The calibrate stage now weights each training target's loss with the national release's formula,
sqrt_value_concept_budget_weighted_mape_50_50_amount_count_target_scale_cap_100pct, from one shared implementation, under an explicit ACS local row mapping.Shared module. These national helpers move verbatim into
microcosm.build.us_runtime.target_loss_weights:_fiscal_target_loss_weightsand its basis, concept-key and budget helpers;US_FISCAL_TARGET_VALUE_WEIGHT_POWER,US_FISCAL_TARGET_CONCEPT_METADATA_EXCLUSIONSandUS_FISCAL_TARGET_LOSS_WEIGHTING;--target-family-loss-multiplierparser.tools/build_us_fiscal_refresh_release.pyimports them back under the names the scorers and the bundle contract read off it. For the national path this is a pure move for every input; see "National golden" below.ACS local calibrate stage (
tools/build_us_acs_local_release.py):TargetSet, and passes them tocalibrate()on every epoch batch.--target-family-loss-multiplier FAMILY=MULTIPLIER, parsed by the same function the national release uses.solver_settingsgainstarget_loss: the formula id, the row mapping, the multipliers and the weights' sha256.--resumeand the completed-run shortcut therefore refuse weights calibrated under another weighting, or under none (every run before this change). Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier #1078's_stamped_settingsdoes not fill this key.calibration_summary.json(target_loss, with the distribution by family × level × basis and the concept-group counts) and inconsumer_export.json(via the settings).calibration_diagnostics.json: per-targettarget_loss_weight/_share/_scaleand the top-leveltarget_loss_basis, as for the national release. Itsbuildblock carries the fields the calibration dashboard reads (target_loss_weighting,target_loss_family_multipliers,target_loss_cap), plus the row mapping, the digest and the distribution.target_loss. The refresh recipe names each recorded multiplier, as Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier #1078's recipe names the penalty settings.The ACS local row mapping
Reusing the national mapping as is would be wrong here in two ways. Both are measured in
experiments/us-acs-local-target-loss-weights-20261004/.pop_state_*,pop_cd_*) carry only{"geography_level": …}. The national basis rule readsmeasure_modeandsource_measure_id, so it silently files them as amounts. They then get 0.1% of the loss, and no district group shares a budget.ledger_fact_labelandledger_layout_groupby_value_labelon every row, and on district rows they name the district ("WV congressional district 1"). Thestate_cdrebase stampsstate_cd_cd_file_value, which also differs by district. None of the three is in the national exclusion set. On the pinned feed, all 19,105 trainedstate_cdSOI district rows (and the 436pop_cd_*rows) would be singleton groups.US_ACS_LOCAL_TARGET_LOSS_ROW_MAPPINGaccepts two kinds of row and refuses everything else:ledger_measure_unit(count/usd). The row must carrymeasure_modeandsource_measure_id. One reviewed exception,tax_filer_individual_count(people; national rule says amount), is mapped to count byACS_LOCAL_LEDGER_BASIS_OVERRIDES. Any other disagreement is refused.pop_state_SS(familycensus_population,geography_level=state)measure_mode=indicator_sum) arepop_cd_SSDDvalidate_acs_local_concept_groupsrefuses three things in any SOI mode:district_concept_identity);state_cdreconciliation block split across groups, or two blocks sharing one;On the pinned feed, every SOI mode (
state,totals,full,state_cd) classifies. Injecting a per-district key into every district row is refused on each surface that has district ledger rows (results.jsonsoi_modes).Weight distribution, before and after
From
experiments/us-acs-local-target-loss-weights-20261004/distribution.py(results.json, at f5f3806, clean tree). Shares are of total weight, which is the share of the loss each kind of target carries at equal scaled misses.The 09-23 release's surface (
--soi-mode state, the default): 4,459 targets, all trained.Counts and amounts carry 50% each after the change (55.5%/44.5% before). The 436 district population rows form 51 groups; California's 52 districts carry 1.51 of weight between them. The weights' digest is
ba36f887….--soi-mode state_cdon the pinned feed: 23,866 training targets after the default 10% district holdout.The 1,907 trained
state_cdblocks each form exactly one group.The formula balances counts against amounts, not families. SOI's share rises on the default surface (85.6% → 90.2%) and falls on
state_cd(97.3% → 92.1%), where district rows stop multiplying their concept's weight. Route A doesn't balance families either: SOI is 74.0% of its targets and 74.1% of its loss weight. The family lever is--target-family-loss-multiplier; this PR sets no default.Cost: district population fit without a multiplier
The full-scale weighted λ frontier in microcosm#1105 ("On the weighted loss" in
experiments/us-acs-local-l2-basis-20260928/README.md) solved the 09-23 checkpoint on these weights.pop_cdrows share one budget, as the national concept budget prescribes.census_population=8: district population is 9.5% of the loss (I recomputed this from these weights). No district misses by more than 10% at λ 0, and 3 do at λ 0.03 (worst 15%).#1105 recommends that multiplier for the next build. The choice is decision d797, and nothing is rebuilt here.
National golden
The national weights are byte-identical before and after the move.
tests/fixtures/us_route_a_target_loss_weights.jsonholds its 5,694 targets, with the fields the formula reads and the weights it recorded. The shared module and the tool's_fiscal_target_loss_weightsreproduce all 5,694 weights bit for bit. With the library's default scales they also give the release's recorded loss basis hash,206ada09584f6e84dd8f32e58d70fea8db257d530c6a1f4ff115630e500436e7.route_a_fixture.pypins the diagnostics' sha256 (64b55a02…), and it refuses to write unless the reduced fields reproduce the recorded weights. Route A has no district rows, so it pins steps 2, 4 and 5 of the formula.test_support/microcosm_build/us_target_loss_weights_reference.pyis a verbatim copy of the pre-move functions and constants from 48fa061. Three tests hold the shared module to it bit for bit:_fiscal_target_loss_basis(...)["loss_vector_sha256"], tested on Route A.Invariants (property-tested with Hypothesis)
For every input drawn:
state_fips;End-to-end.
test_us_acs_local_target_loss_weights.pyrunsdo_calibrateon a surface-shaped fixture: state SOI, SNAP, Medicaid, twostate_cdblocks (one held out) and ladder population. It checks:calibrate()batch, row-aligned and non-uniform;weights_latest.npzand diagnostics record them (consumer_export.jsonis not written there, since the test builds no H5);--resumerefuses an unweighted or differently weighted stamp;test_us_acs_local_release_tool.pychecks that the package manifest recordstarget_lossand the recipe flag.A national finding, not changed here
experiments/…/national_cd_groups.py(national_cd_groups.json) compiles the national registry from the pinned feed and keeps its 24,340 district rows.--target-surface full(the parser default) calibrates them.The existing national budget test uses fixture rows without those labels. Route A calibrated
--target-surface national_state, which has no district rows, so it is unaffected. This PR keeps the national path a pure move, so the fix is a separate follow-up task.Scope and limits
tools/modal_us_stage_plan.py) cannot yet pass--target-family-loss-multiplier, nor Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier #1078's--l2-basisand--mass-parametrization. That is a separate follow-up task. Default runs are unaffected: the weights always apply, and no multiplier is the default.axiom: n/a: calibration infrastructure, no policy logic.
Review
Independent review by in-session Opus 5.5 agents (Subfleet had no dispatchable lanes on 2026-10-04).
Round 1, at 4d87854. Five lenses: pure move, ACS mapping, executed mutation testing (45 mutants), claims audit, and CI/integration risk. An adversarial verifier then checked every medium-or-higher finding. The confirmed findings, all fixed in 2d56c32, 92ed562 and f5f3806:
--soi-mode fullwas refused ontax_filer_individual_count.state_cdonly.Round 2, at c4fd2e8: APPROVE. The reviewer re-checked all 11 findings:
full,totals,stateandstate_cdsurfaces: a one-to-one correspondence between district concepts and keys in every mode, with no false refusals;ba36f887…digest.Round 2's non-blocking notes:
fullsurface hastax_filer_individual_countrows. The draw-share figure above is now a range.Reports:
~/PolicyEngine/_worktrees/.reviews/pr1104-*.md.🤖 Generated with Claude Code