Independent Review of the September 19 Report
Donāt trust the selling advice yet. Claude ignored cheaper comparable hay, misused freight costs, and compared different hay grades to justify selling earlier. Many source numbers are correct; the conclusions need fixing.
Reviewed commit 305edb9 against its parent, the saved primary-source data, current USDA/EIA/CPC pages, and executable checks on September 19, 2026. Verdict: request changes before relying on the price bracket or seller advice. Much of the data collection is useful and reproducible. The central valuation and timing conclusions are materially less secure than the report suggests.
This is a review, not a replacement forecast. Finding faults in $280 does not establish a different correct local price. Published reports, frozen predictions, archives, and production code were left unchanged.
Findings, ordered by consequence
1. High: the claimed regional floor is contradicted by a directly relevant trade omitted from parity
Locations: SeptemberHayReport.md:11, SeptemberHayReport.md:72, scripts/market_parity.py:44.
The report says Good hay has a $250ā260 FOB floor in every surrounding state and uses this to say āReject $180.ā But the September 18 Oklahoma report includes 380 tons of Good/Premium large-square 4x4 alfalfa at $170/ton FOB in Northwest Oklahoma. It also includes 720 tons of Premium/Supreme rounds at $240. Claude records the $170 square-bale trade in the research and parity row's note, but only the $240 rounds enter the calculation.
Using the report's own Northwest Oklahoma distance and freight assumptions, the square-bale comparison lands at approximately $170 + 265 Ć $3.60 / 23 = $211.48/ton, rather than the selected round-bale comparison's $281.48. That is a $70 difference within the same region. It does not prove this lot is still available or equivalent to the seller's stack; those questions apply to the higher-priced comparisons too. There is no documented exclusion rule that resolves the discrepancy.
Separately, Kansas has a Good large-square FOB new-crop trade at $230 (South Central, page 5), also below the asserted floor. Its lower-priced Fair/Good and old-crop transactions are additional evidence against assigning an untested carryover stack a universal minimum.
Correction: include the $170 trade, investigate its comparability, and show the sensitivity of the nowcast to its inclusion. Withdraw the categorical regional-floor claim. Obtain forage quality, condition, weights, and actual competing bids before recommending rejection of a particular offer. The report's recommendation to test and weigh is sensible; its declaration that untested carryover āis a Good-grade commercial lotā is not established by those descriptors.
Primary sources: Oklahoma AMS, page 3, Kansas AMS, page 5. Saved dated copies are in data/2026-09-19/.
2. High: $222 is not an observed export netback
Locations: scripts/market_parity.py:40, SeptemberHayReport.md:11, SeptemberHayReport.md:55.
The previous parity input was a delivered dairy transaction. The new input is $260 FOB-Farm/Ranch, but its calculation type remains export_netback: $260 minus $38 freight = $222. The Colorado report does not give the buyer's delivered price or destination. A regional origin price is not a delivered buying bid to which this subtraction can simply be applied. Greeley is also a regional proxy, not a disclosed destination of this transaction.
The subtraction itself is arithmetically correct. Presenting the result as what the McClave seller nets shipping to that buyer is unsupported. It could be retained as an explicitly assumed destination-price scenario, but it cannot substantiate the report's lower valuation bound.
Correction: use an actual delivered bid and its destination, or label this an unverified scenario. Preserve freight terms as structured input and reject unsupported netback calculations. Existing tests check arithmetic, not whether the source price has the right economic meaning.
Primary sources: Colorado AMS, page 2; USDA hay glossary, which defines FOB as origin price excluding transportation.
3. High: the earlier selling deadline relies on a changing mix of hay, not a measured same-product decline
Locations: SeptemberHayReport.md:13, SeptemberHayReport.md:53, SeptemberHayReport.md:70; research/17 sections 3.3ā3.4.
The $40 fall from November 2022 to February 2023 compares 1,000 tons of Good/Premium 4x4 new crop at $300 with 25 tons of Good 3x4 old crop at $260. December's $320 is a 50-ton Premium lot. Those differences confound seasonality with grade, lot size, crop age, and bale format. The ledger's same named Good-grade observations instead go $225 in September, $250 in October, and $260 in February; even these differ in crop age and are not a controlled index.
The stored +23% calculation is also more fragile than the prose: it compares AugāOct's 269 tons, weighted price about $244, with DecāFeb's 75 tons, weighted price $300. It excludes the November 1,000-ton transaction from the winter calculation, despite the report describing the lift as into NovemberāDecember. The February transaction is below the project's 100-ton grading threshold.
Correction: describe the historical rows as heterogeneous observations consistent with several explanations. The available evidence does not establish a comparable local market peak followed by a $40 decline. Keep a mid-December completion plan, if desired, explicitly as a risk-management judgment rather than a historically demonstrated optimal selling window. Do not interpret this finding as evidence that holding longer is better.
Evidence: data/2026-09-19/wayback_2905/build/analysis.txt, especially āLARGE-SQUARE ONLYā and āLS winter liftā; underlying dated report text and the historical CSV reproduce these rows.
4. Medium: archive coverage is materially incomplete, particularly for the latest winter
Locations: SeptemberHayReport.md:9, research/17-september-reckoning-local-cash-history-2026-09-19.md:47, research/predictions-2026-10.md:13.
I reproduced six qualifying reports out of the 242 collected reports. That sample arithmetic is correct. It is not a complete six-year publication record. The saved coverage table has no OctoberāNovember 2025 reports, no JanuaryāMarch 2026 reports, no May 2026 reports, and only eight reports total in 2026 through September 10. The ESMIS index stops in September 2025; later observations come from Wayback and local snapshots. The public statement that the 242 reports came from USDA's own file server is therefore also too broad.
Most notably, the ā2025ā26 winter ā11%ā result has just one December transaction and no January or February reports. It cannot measure a complete DecāFeb window. ā2 of 8 2026 reportsā is a retrieval-sample frequency, not an established current-year reporting probability. The research discloses missing months, but the public conclusions and categorical prediction anchors omit their significance. āNo primary copy exists anywhereā also exceeds what failed retrieval can establish.
Correction: disclose missing coverage alongside base rates; distinguish missing reports from reports with no qualifying trade; label incomplete seasonal windows as incomplete. Keep the useful finding that this target is sparse, with appropriately limited probability claims. The six-count filter also spans Good through Supreme; do not silently equate it with only the exact Good/Premium grade.
5. Medium: the advertised calibration enforcement has holes
Locations: scripts/calibration.py:96, scripts/forecast_preflight.py:27, research/predictions-2026-10.md:33, data/current/calibration.json.
The new ratchet retains the absolute price floor of $65 when the existing calibration output is present. However:
- Running the calibration module against an isolated copy of the current ledger without a prior calibration output produces $41 again. Missing or invalid prior JSON silently resets the ratchet. Historical minimums are not recoverable from the grading ledger alone.
- Only absolute floors are retained. Percentage floors can decline while the ratchet is closed. Its release condition also pools all families, so unrelated easy hits can eventually unlock a poorly performing family's floor.
- New row #37 has point +3.1°C and half-width 1.3°C. The current documented rule is the maximum of absolute and relative floors.
anomaly_chas a 92.9% relative floor, which gives 2.8799°C at that point. The row does not comply and its explanation cites only the absolute floor. Percentage errors on temperature anomalies may be a poor policy, but the policy and implementation must be reconciled explicitly. - Rows #33ā35 use
state_price_usd_ton, which has no entry in calibration.json. The report gives a rationale for $30, but the new family is not yet a machine-enforced floor. - Preflight checks the local forecast grid, not the widths of all pending prediction rows. All tests pass with the above cases present.
Correction: keep durable, explicit family floor state; fail closed if required state is lost; define the new family's initial floor; enforce all pending-row widths; add isolated regression tests. Clarify which constraints are absolute-only and whether ratchet release is per family. Preserve already frozen predictions and document any erratum rather than silently rewriting them.
6. Medium: the indicator policy is stated more strongly than it is enforced or followed
Locations: scripts/forecast_preflight.py:40, data/current/forecast_policy.json, research/17 section 7.
The indicator screen approves none of the added indicators. Policy says failed indicators may influence scenarios or monitoring but not numeric midpoints. The research nevertheless says the drought outlook is why the peak is not carried beyond December and JanāFeb midpoints roll off; prediction #31 also assigns a $20 adjustment to lower milk prices. These may be legitimate judgments, but there is no quantified, separately auditable adjustment record demonstrating policy compliance or an explicit exception.
Preflight computes and prints the approved list; it never checks which indicators affected the forecast. It therefore cannot enforce the central indicator restriction claimed by the learning system. A price path typed into a CSV can pass regardless of how it was chosen.
Correction: record the adjustment ledger in structured data and check it, or narrow the enforcement claim and explicitly label discretionary exceptions. Calling confidence LOW is useful but does not repair a false claim of mechanical enforcement.
Additional corrections
- $325 is tied for the highest qualifying SE large-square farm-origin observation in this ledger: a 24-ton Premium lot also printed $325 on November 11, 2022. āSingle highest printā overstates its uniqueness, and the comparison spans different grades.
- Report-card prose says persistence was within $5. The July NASS table gives June CO alfalfa $210 versus July $225, and June all-hay $208 versus July $222. Claude's forecast points, $220/$217, were within $5; pure June persistence missed by $15/$14. Distinguish forecast performance from persistence performance.
- The claimed +$45 basis uses the $235 statistical September anchor; the published state-grid September midpoint is $240, making the cross-grid basis $40. Identify the anchor and judgment step explicitly.
- The organic volume total is 4,500 tons, but only 4,000 of those tons are marked contract trades (3,000 alfalfa plus 1,000 grass); the remaining 500 tons are a trade. Change ā4,500 t organic contracts.ā
- JuneāSeptember diversion arithmetic reproduces, but September 2026 is partial whereas historical Septembers are complete. Label it through the retrieval date. A like-for-like full-month JuneāAugust comparison is about 26.2% of historical mean, so the broad water-shortage finding survives.
- The next qualifying report would make seven of 243, not seven of 242; and only a trade satisfying the size/grade/bale/basis filter qualifies, not any SE farm-gate trade.
Checks that passed and review limits
python3 -m unittest discover -s tests -v: 39 tests passed.python3 scripts/grade_predictions.py --check: 17 mechanically gradeable rows agree with protocol. This validates grading math against recorded observations, not the authenticity of those observations by itself.python3 scripts/scorecard.py: 6/16 interval coverage, two NEARs, central cluster 1/5, matching the report.- Reparsed all 265 saved input reports using the new parser and compared date, week-ending, volume, and extracted rows with
all.json: zero differences. Independently filtered the canonical history to reproduce the six qualifying observations and the 2022ā23 comparison rows. Parser reproducibility is not proof that every PDF line was classified correctly; this was not a visual audit of every historical page. - Current USDA Colorado report confirms no SE per-ton alfalfa trade, the 2,000-ton Good trade at $260 FOB-farm, and 8,500 tons total. Current Oklahoma and Kansas reports confirm the counterexamples above.
- Saved NASS July hay table confirms CO alfalfa $225, CO all-hay $222, KS alfalfa $126. Saved cattle-on-feed table confirms August placements 1.617 million and 91% of prior year.
- Recomputed FLCC JuneāSeptember 25,631.42 AF, historical mean 102,502.00 AF, ratio 25.006%; MarchāSeptember 50,713.55 AF and second-lowest ranking in the saved series. Checked the saved reservoir report's 19,412 AF John Martin reading.
- Recomputed seven state drought indices from saved USDM cumulative categories; checked Bent County D2+ 19.11% versus 79.67% on August 25 and CO D3+ 44.14%. Differences of a unit from rounded reported DSCI are immaterial here.
- Live EIA confirms September 14 U.S. diesel $6.285 and Rocky Mountain $6.066; live CPC discussion confirms the 75% historic-event probability and greater-than-90% very-strong-event probability.
- Old dated archive contents were not changed by the reviewed commit, and recorded archive hashes pass the integrity suite.
- I did not certify every peripheral news item, exchange settlement, haul quote, historical parser classification, or the seller's actual inventory/quality. Those limits do not prevent a firm request-changes verdict: the price-floor counterexample, unsupported netback, and mixed-product peak comparison are independently sufficient.
Disposition: retain the source collection, explicit uncertainty, and honest scorecard. Correct the parity model and exclusions, qualify the seasonal inference and incomplete archive, and close or accurately describe the enforcement gaps. Issue a dated correction to any already published report; do not overwrite its frozen archive. Recompute the valuation only after those changes, rather than assuming either the existing $280 or a lower replacement is validated.