Cellar Index
Methods and results

How the model works

Wine list prices rarely change, so the model predicts whether each wine's list price will rise or fall over the next four quarters, and by how much. One model is trained across all 31,256 wines and scored on the most recent year, which it never saw during training. See the full system architecture →

Model version
hgbq-20260928-160312
Data through 2026-07-01
How often list prices change
2.8%
of wines get a new list price in a typical quarter. Guessing "no change" is right about 97% of the time, so the model's job is to spot the few that will move.
How reliable the price ranges are
97.4%
of actual prices landed inside the forecast's 80% range. The target is 80%, so the ranges are wider (more cautious) than they need to be.

Accuracy on the holdout year

Ranking accuracy is measured as the area under the receiver operating characteristic curve (ROC AUC): the chance that the model scores a wine whose price actually changed above one whose price didn't. 0.5 is a coin flip and 1.0 is perfect. Average precision shows how often the wines the model flags really moved; compare it with the share of all wines that moved, in brackets.

HorizonRise: ROC AUCRise: avg precision (share that rose)Rise: predicted vs actual rateCut: ROC AUC
1 quarter ahead0.8910.068 (0.011)0.51% vs 1.11%0.848
2 quarters ahead0.8770.093 (0.018)1.37% vs 1.77%0.856
3 quarters ahead0.8460.076 (0.022)2.7% vs 2.18%0.848
4 quarters ahead0.840.087 (0.026)4.53% vs 2.58%0.84

What drives the prediction of a price rise

Permutation importance: how much ranking skill (ROC AUC) is lost when a feature's values are shuffled. Longer bars matter more.

Years until the drinking window opens
0.100
Size of the current discount
0.014
Price level
0.013
Years past the drinking window
0.010
Quarters on the price list
0.008
Years since the vintage
0.003
Currently on clearance
0.003
Price tier
0.003
Quarters since the last price change
0.002
Quarter of the year
0.002
Currently on sale
0.002
Price change over 2 years
0.002
Style (red, white, …)
0.001
Price change over 1 year
0.001
Inside its drinking window
0.000

Next-vintage model

Predicts how a wine's next vintage will be priced against the current one, from 4,871 past vintage changes. Feature sets are compared over 4 rolling one-year test windows, each trained only on earlier data. Error is the mean absolute log price change (lower is better); the "same price" guess scores 0.093–0.128 across windows.

FeaturesError (mean)vs currentWindows betterRise ROC AUC
current (base + supply)0.0986——0.715
+both0.1007+0.00211/40.694
+brand0.0991+0.00042/40.723
+attributes0.1013+0.00260/40.697

Adopted: supply. California supply lowered error in every window in an earlier round; weather's effect flipped sign between windows, and brand momentum and wine attributes didn't help (attributes overfit the ~4,900 transitions).

Fair-price model

Estimates what a wine should cost from what it is: region, grape, classification, style, bottle size, age of the vintage, words in its name, and the producer's price level from its other wines. Each wine is priced by a model that never saw it or any of its vintages (5-fold cross-validation grouped by wine), so the gap between shelf price and fair price is an honest value signal. Scores are on all 31,256 wines; R² is the share of variation in (log) price the model explains.

ModelTypical errorR²
Full model (used)16.2%0.80
no name text18.2%0.72
+vintage weather16.4%0.80
no producer level19.1%0.77
baseline: region + grape median36.7%0.27
baseline: producer's other wines23.7%0.56

Typical error is the median absolute percentage error. Both the name text and the producer's price level add real information; vintage weather adds none, consistent with the other experiments: retail prices follow producer and appellation more than the growing season. The 80% fair range covers 80% of real prices after calibration; a wine is flagged as a deal or a premium only when its price falls outside that range.

Serving

The training job writes batch forecasts for every wine. The API also loads the same model artifact from the model registry and scores wines on demand: that powers the what-if panel on each wine page. Both paths share one feature pipeline, and an automated parity test checks that online predictions match the batch forecasts.