How the model works
Wine list prices rarely change, so the model predicts whether each wine's list price will rise or fall over the next four quarters, and by how much. One model is trained across all 31,256 wines and scored on the most recent year, which it never saw during training. See the full system architecture →
Accuracy on the holdout year
Ranking accuracy is measured as the area under the receiver operating characteristic curve (ROC AUC): the chance that the model scores a wine whose price actually changed above one whose price didn't. 0.5 is a coin flip and 1.0 is perfect. Average precision shows how often the wines the model flags really moved; compare it with the share of all wines that moved, in brackets.
| Horizon | Rise: ROC AUC | Rise: avg precision (share that rose) | Rise: predicted vs actual rate | Cut: ROC AUC |
|---|---|---|---|---|
| 1 quarter ahead | 0.891 | 0.068 (0.011) | 0.51% vs 1.11% | 0.848 |
| 2 quarters ahead | 0.877 | 0.093 (0.018) | 1.37% vs 1.77% | 0.856 |
| 3 quarters ahead | 0.846 | 0.076 (0.022) | 2.7% vs 2.18% | 0.848 |
| 4 quarters ahead | 0.84 | 0.087 (0.026) | 4.53% vs 2.58% | 0.84 |
What drives the prediction of a price rise
Permutation importance: how much ranking skill (ROC AUC) is lost when a feature's values are shuffled. Longer bars matter more.
Next-vintage model
Predicts how a wine's next vintage will be priced against the current one, from 4,871 past vintage changes. Feature sets are compared over 4 rolling one-year test windows, each trained only on earlier data. Error is the mean absolute log price change (lower is better); the "same price" guess scores 0.093–0.128 across windows.
| Features | Error (mean) | vs current | Windows better | Rise ROC AUC |
|---|---|---|---|---|
| current (base + supply) | 0.0986 | — | — | 0.715 |
| +both | 0.1007 | +0.0021 | 1/4 | 0.694 |
| +brand | 0.0991 | +0.0004 | 2/4 | 0.723 |
| +attributes | 0.1013 | +0.0026 | 0/4 | 0.697 |
Adopted: supply. California supply lowered error in every window in an earlier round; weather's effect flipped sign between windows, and brand momentum and wine attributes didn't help (attributes overfit the ~4,900 transitions).
Fair-price model
Estimates what a wine should cost from what it is: region, grape, classification, style, bottle size, age of the vintage, words in its name, and the producer's price level from its other wines. Each wine is priced by a model that never saw it or any of its vintages (5-fold cross-validation grouped by wine), so the gap between shelf price and fair price is an honest value signal. Scores are on all 31,256 wines; R² is the share of variation in (log) price the model explains.
| Model | Typical error | R² |
|---|---|---|
| Full model (used) | 16.2% | 0.80 |
| no name text | 18.2% | 0.72 |
| +vintage weather | 16.4% | 0.80 |
| no producer level | 19.1% | 0.77 |
| baseline: region + grape median | 36.7% | 0.27 |
| baseline: producer's other wines | 23.7% | 0.56 |
Typical error is the median absolute percentage error. Both the name text and the producer's price level add real information; vintage weather adds none, consistent with the other experiments: retail prices follow producer and appellation more than the growing season. The 80% fair range covers 80% of real prices after calibration; a wine is flagged as a deal or a premium only when its price falls outside that range.
Serving
The training job writes batch forecasts for every wine. The API also loads the same model artifact from the model registry and scores wines on demand: that powers the what-if panel on each wine page. Both paths share one feature pipeline, and an automated parity test checks that online predictions match the batch forecasts.