September's Best Print Run Models
Segmented Log won a close September race, but category leaders showed why PrintRun tests a library of models instead of expecting one approach to win everywhere.
Segmented Log finished first in September, but it did not run away from the field. Bounded Calibrated landed just 0.4 average-error points behind, and both models placed 79.1% of their estimates within 50% of the official result.
The more useful story appeared beneath the overall ranking. Different approaches led MLB, Formula 1, WWE, and standard cards. September strengthened the case for promoting repeat winners by category while retiring models that stop adding value.
What September told us
The overall winner matters, but the shape of the race matters more. September produced a narrow lead at the top, a strong second tier, and several category specialists that outperformed the overall leaders in their best territory.
September top 10
Models are ranked by lowest average percentage error across the complete finalized cohort. Median error breaks a tie.
The overall winner succeeded on the broadest card pool
Standard cards made up 207 of the 249 finalized releases, so they were the main test of broad reliability. Segmented Log led that large group with 33.5% average error and placed 81.6% of its estimates within 50% of the final print run.
That broad-card performance carried it to the overall win. Bounded Calibrated stayed close by producing nearly identical full-cohort consistency rather than relying on one small specialty.
Different models won different sports
No model owned every category. Monotonic Boosted led the 114-card MLB cohort, Market Log led Formula 1, and Ensemble led WWE. Those category wins are the strongest argument for keeping the research library competitive rather than forcing one universal model onto every release.
What happens next
PrintRun will keep tracking whether September's leaders repeat across larger cohorts and whether category winners hold their edge when the card mix changes.
Models that continue to win can earn a larger role. Some research models may be decommissioned as the evidence accumulates and their output stops improving the field.
This report compares recorded pre-close research outputs. It does not claim that one model was selected in advance as the official forecast for each September card.
Frequently asked questions.
Did PrintRun select one official model before each September card closed?
No. This report compares the research model library rather than claiming one official September forecast.
What does median APE mean?
Median absolute percentage error is the middle percentage error after a model’s card-level errors are sorted. It is less sensitive to a handful of extreme misses than MAPE.
Why are some model sample sizes smaller?
Some surfaces were unavailable for a subset of cards, so those models did not produce a final eligible output for every release.
Will the September rankings change?
They can. Forty-four September releases were still unresolved in the frozen snapshot and are excluded until Topps publishes official print runs.
Sources
Compare the full historical scorecards.
The live model page tracks accuracy bands across the broader completed-card archive.
Open model scorecards ↗Get next month’s report.
Monthly Topps NOW benchmarks, standout releases, and recorded prediction results.
