Southeast Colorado alfalfa price forecasts, graded in public.

The Lessons Ledger

Every rule below was bought with a graded miss. They are binding on future reports: a forecast that violates one must say so explicitly and justify it. Newest first. This page only grows.

These rules are enforced in code, not memory (added 2026-08-18, same day as L4–L7): the graded record lives in a machine-readable ledger (prediction_grades.csv); a calibration engine (calibration.py) recomputes coverage and derives mandatory minimum range widths from measured error every weekly run; a survey-anchored estimator (production_estimate.py) fails the pipeline if the operative production estimate leaves the print-Β±-max-revision window; and the demand indicators we under-weighted (7-state drought index, weekly pasture condition, a cash-price benchmark ledger) are now collected as first-class data series. The test suite fails the build on violations. Current calibration floors, straight from the misses: price predictions Β±$65, production Β±30%, volume Β±57% β€” until measured coverage earns narrower.


L11 β€” The local target is observable ~2.5% of the time, and local cash led the state average by 4–5 months at the last peak. Grade the target you can see; time the sale on the cash series, not on NASS. (2026-09-19, first full local cash history)

The miss: rows #20 and #23 (Sep 2026) both assumed the grading target would print again and that weekly volume was forecastable; the August 27 grid anchored its spot on a single 88-ton lot that turned out to be the highest SE FOB-farm print in a six-year ledger, while the largest commercial Colorado print of the year (2,000 t of Good at $260 FOB-farm, Sep 10) came in $65 under it. The "drought peaks arrive Feb–Jun, hold for the plateau" base rate was computed on the NASS state average; the SE cash series the seller actually sells into peaked Nov–Dec 2022 and was $40 lower by February while NASS kept rising to May. Amended 2026-09-19 (same day, after independent review β€” research/18): the first application of this lesson broke it three ways: a comparable was left out of parity with no written exclusion rule (OK 380 t G/P at $170); an origin FOB price was turned into a "netback" by subtracting freight to an assumed destination; and the "Nov–Dec peak, $40 lower by Feb" compared a 1,000-t new-crop Good/Premium lot with a 25-t old-crop Good lot. Additions to the rule: (5) parity includes every β‰₯100-t completed trade in the comp set unless a written exclusion rule says why; (6) a netback requires an observed delivered price with a known destination, otherwise the row is shown at origin; (7) seasonal base rates from the ledger must be same-product comparisons or be labeled heterogeneous and carry no timing claim.

Amended again 2026-09-19 (research/19, after comparing the corrected report with an independent seller report): (8) the seller section of every report is procedure-first (test, weigh, three identical-spec bids, net at the stack), shows break-even arithmetic for holding, budgets with no automatic winter premium, and labels comparable ranges as comparables, never as the stack's value; (9) an outside review of the inference layer runs before publication β€” the first Sep 19 version passed all 39 tests and every data check and was still wrong in three load-bearing inferences.

The rule: (1) Base rates for the local target come from the local ledger (data/historical/ams_2905_se_ne_alfalfa_history.csv, 242 reports 2020–2026, rebuilt from the ESMIS AMS_2905 archive), never from the state average. (2) A SE FOB-farm Good–Supreme large-square trade β‰₯100 t prints in ~6 of 242 reports; predictions about it are categorical (prints / doesn't) with the base rate stated, and the spot nowcast is built from the tonnage-weighted regional comp set (β‰₯500 t lots in NE CO, WY, W NE, KS), with any single lot under 200 t treated as a data point, not an anchor. (3) Sell-timing guidance cites the local cash peak window (Nov–Dec in the only observable drought analog), and the exit rules stay senior. (4) Weekly AMS volume is not a prediction row (MAE 3,847 t on a 2,500 t point).

L10 β€” Indicators must beat the same release-aware baseline before they move the midpoint. (2026-08-28, whole-project retro)

The miss: rich water, weather, drought, ENSO, fuel, and production narratives were repeatedly converted into precise price adjustments without proving that those indicators improved an out-of-sample price forecast. The full model added only ~0.004 in-sample RΒ² beyond price persistence and seasonality, while its live forecasts were low by $34.1/ton on average. The rule: a historical indicator needs at least 48 out-of-sample observations, at least $2/ton MAE improvement over the same release-aware baseline, and improvement in at least four of five evaluation years before it may alter the numeric midpoint. Failed indicators remain scenario or monitoring variables. Comparable AMS cash may anchor the current-market nowcast because it observes the target directly; it does not receive unearned long-horizon predictive status.

L9 β€” Confidence is earned by comparable forecast outcomes, not by research thoroughness. (2026-08-28, whole-project retro)

The miss: reports labeled near-term prices HIGH confidence because supply facts were well researched, while the local price mapping had no comparable track record. The first frozen scorecard contained only 2 of 10 two-sided ranges; the live model put 0 of 3 outcomes inside its bands. The rule: confidence comes only from resolved directly comparable sample size, point error, bias, and interval coverage. Until there are at least 20 comparable local outcomes, 70–90% interval coverage across at least two cycles, and no material same-signed bias, the maximum forecast confidence is LOW. Source count, citations, in-sample RΒ², narrative agreement, and agent certainty cannot upgrade it.

L8 β€” Scores are computed, not typed; metrics are specified at freeze time, not grade time. (2026-08-18, the erratum)

The miss: the same day we published our first pre-registered scorecard, we published it wrong. A grader scored "weekly NiΓ±o-3.4" against the monthly index (a different number that happened to sit inside our range) and hand-typed "3 of 10" into four documents; the true score was 2 of 10. Corrected in public within hours β€” but the error class is structural: any number that travels from a ledger to prose by human hands can drift, and any prediction whose metric lives in prose can be graded against the wrong thing. The rule: every frozen prediction ships with a machine-readable source spec (exact release, table, row, column β€” prediction_sources.json); grading is done by script (grade_predictions.py) against that spec; every published score line is generated from the ledger (scorecard.py). Tests fail the build if a pending row lacks a spec, if a ledger outcome disagrees with the frozen scoring math, or if a published "X of Y" claim disagrees with the ledger.

L7 β€” Calibrate interval width to measured coverage, not to confidence. (2026-08-18)

The miss: our pre-registered ranges contained the outcome 2 times out of 10 (initially misreported as 3/10 β€” one row was graded against the wrong metric, corrected same day). Ranges that miss 80% of the time are roughly half as wide as honesty requires. The rule: prediction ranges are set from measured error (the vintage backtest: Β±$25–30 on winter months; the release-grading record), not from how sure the narrative feels. If the last cycle's coverage was below ~70%, the next cycle's ranges widen mechanically β€” no discretion.

L6 β€” When misses share a sign across cycles, it's bias, not bad luck. (2026-08-18)

The miss: three consecutive reports where every price surprise broke bullish (June $225 contract β†’ July cash rally β†’ August $275 delivered / $350–380 stable lots). Each time we treated it as a one-off. The rule: two consecutive same-signed misses on a series triggers a named bias review in the next report. Current standing bias: we under-weight demand. Demand-side indicators (pasture condition %, neighbor-state DSCI, feeder cattle margins, KS–CO spread) are now first-class forecast inputs, not color.

L5 β€” Drought indices are demand proxies for irrigated crops, not supply proxies. (2026-08-18)

The miss: we read "worst water year on record" as "the alfalfa crop fails." USDA's survey says CO alfalfa βˆ’4.8% while winter wheat went βˆ’67% and corn βˆ’24%. Drought kills dryland; wells and senior rights keep irrigated alfalfa alive β€” what drought actually does to the alfalfa price is destroy pasture and grass hay, which turns every ruminant into an alfalfa buyer. The rule: supply estimates for irrigated crops must be built from water deliveries by region weighted by production share, with explicit credit for groundwater. DSCI/USDM feed the demand model.

L4 β€” News coverage samples the disaster, not the state. (2026-08-18 β€” the big one)

The miss: our 1.5M-ton (βˆ’36%) production estimate was assembled from canal shutdowns, fallowing headlines, and worst-on-record stories in the regions where the story was dramatic. USDA's farmer survey printed 2,244k (βˆ’4.8%). For our number to be right, the quiet 57% of the state would have had to fail too β€” and nobody writes articles about fields that got water. The rule: statewide estimates anchor on the survey base rate and the maximum historical revision. The plausible-final window is [Aug print βˆ’ 12%, Aug print] (2022 is the record revision). An estimate outside that window requires extraordinary, quantified evidence β€” not vibes, not clippings. Our own estimate is now graded against this: 1.5M (frozen) vs ~1.95M (rev1), settled Jan 12, 2027.

L3 β€” Load-bearing negative claims need two independent verification methods. (2026-08-08)

The miss: a broken shell tool returned empty search results and we published "USDA dropped hay from the August report." The files contained the tables all along. Retracted same day. The rule: any "X is absent / X stopped happening" claim ships only after two independent checks (different tool, different method).

L2 β€” Measure markets from cash prints, never from lagging state averages. (2026-08-08)

The miss: July's "Kansas caps SE CO at ~$155–170 landed" used a stale NASS state average ($123) while Kansas cash grinding hay traded $200–210. The cap was off by ~$60–65 landed. The rule: floors and ceilings are computed from AMS cash trades plus current freight (the market_parity.py table), refreshed every cycle. NASS monthly averages are trend confirmation only β€” they lag 4–8 weeks.

L1 β€” Don't forecast fuel. Mark it to market. (2026-07-11)

The miss: two consecutive directional misses on diesel (April: missed the fall; July: called it "falling" at $4.58, it went to $5.35). The rule: every report uses the current posted EIA weekly price. No diesel opinions.


Why this page exists: on 2026-08-18 the project's pre-registered predictions were graded against the August USDA releases and scored 2/10 on intervals (initially published as 3/10 β€” one ENSO row was graded against the monthly index when the frozen prediction specified the weekly one; corrected the same day, in public), with the central production thesis missing by 28%. The full grading record is here; the reports that own each miss are in the reports archive. We keep the misses public because the alternative β€” quietly getting better without admitting what was wrong β€” is how forecasting products lie.