PrintRun

September's Best Print Run Models

Segmented Log won a close September race, but category leaders showed why PrintRun tests a library of models instead of expecting one approach to win everywhere.

#1Segmented Log finished first across 249 finalized September cards

Segmented Log finished first in September, but it did not run away from the field. Bounded Calibrated landed just 0.4 average-error points behind, and both models placed 79.1% of their estimates within 50% of the official result.

The more useful story appeared beneath the overall ranking. Different approaches led MLB, Formula 1, WWE, and standard cards. September strengthened the case for promoting repeat winners by category while retiring models that stop adding value.

What September told us

The overall winner matters, but the shape of the race matters more. September produced a narrow lead at the top, a strong second tier, and several category specialists that outperformed the overall leaders in their best territory.

The winnerSegmented Log35.2% average error across all 249 scored cards
The closest challengerBounded CalibratedOnly 0.4 percentage points behind
The shared consistency mark79.1%Both top models landed within 50% on nearly four of every five cards

September top 10

Models are ranked by lowest average percentage error across the complete finalized cohort. Median error breaks a tie.

#1 Segmented Log35.2% average error45.0% within 25% · 79.1% within 50%
#2 Bounded Calibrated35.6% average error44.6% within 25% · 79.1% within 50%
#3 Ensemble37.0% average error36.5% within 25% · 75.1% within 50%
#4 Market Log37.2% average error38.6% within 25% · 76.3% within 50%
#5 Position Baseline Consensus39.0% average error42.2% within 25% · 73.5% within 50%
#6 Surface Fusion39.5% average error34.1% within 25% · 71.1% within 50%
#7 Log Regression41.7% average error36.1% within 25% · 69.5% within 50%
#8 Power Law42.1% average error38.6% within 25% · 66.3% within 50%
#9 Popularity Consensus43.3% average error44.2% within 25% · 73.5% within 50%
#10 Isotonic43.8% average error38.2% within 25% · 70.7% within 50%

The overall winner succeeded on the broadest card pool

Standard cards made up 207 of the 249 finalized releases, so they were the main test of broad reliability. Segmented Log led that large group with 33.5% average error and placed 81.6% of its estimates within 50% of the final print run.

That broad-card performance carried it to the overall win. Bounded Calibrated stayed close by producing nearly identical full-cohort consistency rather than relying on one small specialty.

Standard cards scored207The largest card-type cohort
Segmented Log average error33.5%26.4% median error
Within 50% of official PR81.6%169 of 207 standard-card estimates

Different models won different sports

No model owned every category. Monotonic Boosted led the 114-card MLB cohort, Market Log led Formula 1, and Ensemble led WWE. Those category wins are the strongest argument for keeping the research library competitive rather than forcing one universal model onto every release.

MLBMonotonic Boosted24.6% average error across 114 cards · 92.1% within 50%
Formula 1Market Log24.6% average error across 9 cards · all 9 within 50%
WWEEnsemble22.3% average error across 8 cards · all 8 within 50%
SoccerBounded Calibrated39.7% average error across 87 cards · 77.0% within 50%

What happens next

PrintRun will keep tracking whether September's leaders repeat across larger cohorts and whether category winners hold their edge when the card mix changes.

Models that continue to win can earn a larger role. Some research models may be decommissioned as the evidence accumulates and their output stops improving the field.

This report compares recorded pre-close research outputs. It does not claim that one model was selected in advance as the official forecast for each September card.

The ranking covers 6,143 outputs from 25 models across 249 finalized cards. Full metrics and methodology are available below.

Frequently asked questions.

Did PrintRun select one official model before each September card closed?

No. This report compares the research model library rather than claiming one official September forecast.

What does median APE mean?

Median absolute percentage error is the middle percentage error after a model’s card-level errors are sorted. It is less sensitive to a handful of extreme misses than MAPE.

Why are some model sample sizes smaller?

Some surfaces were unavailable for a subset of cards, so those models did not produce a final eligible output for every release.

Will the September rankings change?

They can. Forty-four September releases were still unresolved in the frozen snapshot and are excluded until Topps publishes official print runs.

Sources

September model metricsOverall, sport, and card-type metrics with sample sizes, error bands, p90 error, and bias.Download ↓September report methodologyFrozen cohort definition, source timestamp, database hash, and model-evaluation boundary.Download ↓

Compare the full historical scorecards.

The live model page tracks accuracy bands across the broader completed-card archive.

Open model scorecards ↗

Get next month’s report.

Monthly Topps NOW benchmarks, standout releases, and recorded prediction results.

By joining, you agree to receive PrintRun results and product updates. Unsubscribe anytime.