"A full-season backtest of the FanRanked StreamMachine: 2,004 actual 2026 starts, tiered blind using frozen preseason inputs. Tier-by-tier results, a same-pitcher matchup test, and a hard look at whether the model calls the coin-flip arms correctly."
When we launched the StreamMachine, we made a bet: that a projection engine with real knowledge of the ballpark, the opponent, and the platoon matchup can tell you something useful about a single start, the noisiest event in fantasy baseball. A bet like that should be checked. So we ran the engine backwards over the entire 2026 season to date. That means 2,004 actual starts from April 1 through July 3, each one scored and tiered exactly as the live board would have scored it that morning.
A backtest is only as honest as its inputs, so we handicapped ourselves harder than the live board does:
Then each start got a Stream Score (50 is a league-average start) and a tier, and we looked up what actually happened.
| Tier | Starts | ERA | WHIP | K/9 | QS% |
|:---:|:---:|:---:|:---:|:---:|:---:|
| S | 216 | 3.44 | 1.13 | 10.1 | 43% |
| A | 560 | 3.83 | 1.20 | 9.3 | 39% |
| B | 521 | 4.14 | 1.28 | 8.4 | 35% |
| C | 389 | 4.58 | 1.38 | 7.5 | 32% |
| D | 318 | 4.80 | 1.37 | 6.8 | 31% |
Every stat improves in order at every rung. The gap from S to D is 1.36 runs of ERA, 3.3 strikeouts per nine, and 12 points of quality start rate. If you had started every S and A start and skipped every D start all season, that spread went straight into your ratios.
Tier tables can flatter a model. Good pitchers land in good tiers, and good pitchers pitch well. The honest question is whether the matchup call itself adds anything once you hold the pitcher fixed. For every arm with at least six starts, we compared his most favorable matchups (the top third by the model's own matchup delta) against his toughest (the bottom third), 580 starts per side:
| | ERA | K/9 |
|---|:---:|:---:|
| Favorable matchups | 3.95 | 8.8 |
| Tough matchups | 4.10 | 8.2 |
Same pitchers, different spots: the matchup layer alone is worth about 0.15 runs of ERA and 0.6 K/9. That's modest on ERA, since single starts are wildly noisy, and unambiguous on strikeouts.
"Always start the ace" is free advice. The whole point of a streamer model is the middling arm whose answer genuinely changes week to week. So we isolated exactly those pitchers: 8 or more starts, with at least 3 on each side of the start/sit line (score 50). Because the baseline here is frozen at preseason, it can never flip a verdict midseason. Every start-to-sit flip in this group is purely the matchup layer's call.
28 pitchers qualified, 408 starts:
| Verdict | Starts | ERA | WHIP | K/9 | QS% |
|---|:---:|:---:|:---:|:---:|:---:|
| START calls | 231 | 4.10 | 1.28 | 8.0 | 39% |
| SIT calls | 177 | 4.22 | 1.32 | 7.2 | 37% |
Same population of coin-flip arms, and the START-flagged outings win on every stat. Per pitcher, 17 of the 28 posted a better ERA in their START calls than their SIT calls.
The poster child is Aaron Civale: a 3.88 ERA across his 10 START calls and 10.03 across his 3 SIT calls. The model kept greenlighting him in soft spots and told you to bench him almost exactly the outings he got shelled. Jameson Taillon (4.17 vs 8.18), Griffin Canning (4.98 vs 8.85), and Jeffrey Springs (5.14 vs 7.92) follow the same pattern.
The misses are instructive too. The worst was Chris Bassitt, with a 6.67 ERA in his START calls against 3.07 in his SIT calls. Blame tiny 3-to-5 start samples where one gem flips the math, plus pitchers whose true talent moved sharply from March. That second failure mode is exactly what the live board doesn't have, since it re-baselines every pitcher daily.
The sharpest way to use a board like this is one pickup a day. So we asked a simpler question: what if you had taken the single highest-scored start under 50% rostered, every day since April 1? That pick now runs live on the board as the Stream of the Day, and the backtest gives it a season line.
| Picks | Hits | Blowups | ERA | WHIP | K/9 |
|:---:|:---:|:---:|:---:|:---:|:---:|
| 95 | 71% | 16% | 3.90 | 1.24 | 9.0 |
We grade each pick by the start itself, not the pitcher's decision, since nearly half of all streamer starts end in a no-decision. A hit is a start that helped you: three or fewer earned runs over five-plus innings, or one or fewer in a shorter outing. A blowup actively hurt: five or more earned runs, or four without finishing the fifth. Everything in between counts against neither.
A 3.90 ERA with a 1.24 WHIP and nine strikeouts per nine, from arms sitting on most waiver wires, is the whole pitch for streaming in one row. For deep leagues we also track a version restricted to arms under 20% rostered. It hits 64% of the time with a 5.12 ERA, which is worth knowing before you stream out of desperation: the wire gets thin fast below 20%, and we would rather publish that than pretend otherwise. Both picks and both season records update on the live board every day.
And you don't have to take a backtest's word for it. Since July 4 the live board freezes every prediction before first pitch and scores it against the real box score the next morning. The running scorecard is on the StreamMachine board.