Methodology
How KSplit builds a strikeout projection, and what the number actually means.
Most strikeout models hand you one number and let you assume it means something. KSplit builds a full probability distribution for every start, then shows you its shape. The projection is the center of that distribution.
Everything below is the order of operations: how a pitcher's raw strikeout rate becomes a matchup-adjusted rate, how his arsenal and location shift it, how nine individual hitters become one lineup, how batters faced turns a rate into a count, and how the count becomes a distribution you can price against a line.
For common questions about access, usage, and scope, see the FAQ. For a column-by-column walkthrough of the live board, read our How to Guide for Today's Dashboard. For formal definitions of every metric, open the Glossary.
The Plate Appearance Model
Every projection starts at the plate appearance, not the inning.
I model strikeouts as K per plate appearance, never K per inning. A pitcher does not face innings, he faces hitters, and each hitter is an independent opportunity for a strikeout.
The base rate for any matchup comes from two sides, the pitcher brings his own strikeout rate against the batter side he is facing, split L and R, because handedness splits are real and often large. Both sides are season specific and credibility weighted toward league baselines when the sample is thin, so a pitcher with forty plate appearances against lefties does not get treated as if he has four hundred. From there, all the rates get blended together to produce a probability per plate appearance that the batter is going to strikeout... There obviously is more to it, but I can't give away everything behind the curtain...
Pitch Arsenal and Matchup Quality
A blended K% tells you what two sides do on average, it does not tell you what happens when this specific arm throws this specific pitch to this specific bat.
That is the arsenal layer, and it is where kAdj (Strikeout Adjustment) comes from.
For every starter I carry his full pitch mix, split by batter handedness, because pitchers do not throw the same arsenal to lefties that they throw to righties. Usage is blended across a recent window and the current season, with prior year data as a Bayesian prior when the current sample is short. A rookie with three starts gets shrunk toward what he has actually shown. A veteran with a hundred innings does not.
Every pitch in the arsenal carries its own performance profile against each side:
- CSW%, called strikes plus whiffs per pitch
- Whiff%, swings and misses per swing
- Chase%, swings at pitches outside the zone
- PutAway%, how often a two-strike pitch finishes the at-bat
- Contact quality, ISO and xwOBA allowed
Each of these is measured against what the league does with that exact pitch to that exact side. A 30% CSW slider is not impressive in isolation, it is impressive relative to what the league gets from sliders against righties. Every signal in the model is a residual against its own baseline, never a raw number floating on its own.
Raw rates get shrunk toward league before they are trusted, pitch-mix shrinkage runs at 100 plate appearances of pseudo-count, which means a pitch with a small sample pulls hard toward the league rate for that pitch type and only earns its own identity as the sample grows. The same treatment runs on the hitter side.
The arsenal signals are then usage weighted into each hitter's blended K%. This is the part that matters and the part most models skip. A hitter's vulnerability to breaking balls only counts in proportion to how many breaking balls he is actually going to see. If a starter throws 8% sliders, a hitter's slider problem is worth 8%, the weighting is by real usage against that hitter's side, so the arsenal edge is always scaled to the arsenal he will actually face.
Location and Zone Mapping
Where a pitch goes is a separate question from what the pitch is, and it carries its own signal.
I map every starter's arsenal to a thirteen-cell grid built on the Statcast zone system: a three by three heart in the strike zone, wrapped by four shadow regions where chase and expansion live. For each pitch, against each batter side, I carry how often he goes to each cell, what his swinging strike rate is there, and what his put-away rate is there when he gets to two strikes.
Above the zone grid sits a regional view: Heart, Edge, Chase, and Waste. Regional distribution answers a coarser question than the grid does. It tells you whether a pitcher attacks or nibbles, whether he expands with two strikes, and whether the shape of his location profile supports a strikeout or a walk.
The hitter side mirrors it. For every hitter I carry his CSW, swinging strike, and put-away rates by zone and by region, split against left and right handed pitching, grouped by pitch family. Families are three: fastballs (four-seam, sinker, cutter), breaking balls (slider, sweeper, curveball, knuckle-curve, slurve), and offspeed (changeup, splitter).
The location fit signal comes out of the overlap. A pitcher who lives in the upper third with his fastball against a hitter who cannot touch anything up produces a different projection than the same pitcher against a hitter who feasts there. The grid is where that shows up, and the league location baselines are what make the comparison meaningful.
From Rate to Distribution
A rate is not a projection, the projection is the distribution.
Once every hitter in the lineup has a matchup-adjusted, arsenal-weighted, location-fit K% against the starter, I have nine individual probabilities, not one lineup average. That distinction is the engine.
Each plate appearance is its own Bernoulli trial with its own probability, and the sum across a variable number of trials produces the strikeout distribution for the start. Because the individual probabilities differ, the resulting distribution is not a clean Poisson. It carries more spread than a naive model would produce, and that spread is the entire point.
Batters Faced
A strikeout rate is worthless without knowing how many chances the pitcher gets.
Batters faced is modeled from workload rather than assumed. I carry two leash profiles, A normal leash centers around 23 batters faced. A short leash centers around 17, which is the difference between a starter who gets the third time through the order and one who gets pulled the moment he wobbles.
Batters faced is a distribution too, not a point estimate. A start that could plausibly last 18 batters or 26 has a wider strikeout distribution than one that will almost certainly last 24, even if the two share an identical mean projection. Workload uncertainty compounds into outcome uncertainty, and the model carries both.
Ceiling Profiles
Two starts can share a projection of 6.2 strikeouts and be completely different bets.
Every projection gets classified by the shape of its distribution, not the height of its center:
The profile is the reason the distribution matters. A 6.2 that is tail-driven and a 6.2 that is centered price identically on a projection sheet and price nothing alike in reality.
Ladder multipliers run on top of the profiles, damping the probability of the +1 and +2 rungs differently depending on the shape. Calibration showed the high-ceiling profile was over-projecting +1 hit rates, so the correction lives where the error lives.
Tail Environment
Beyond the individual start, I classify the tail environment of the slate itself.
Some days the distribution of outcomes across all starters is compressed, and the right tails are quiet league-wide. Other days the tails are live. Tail environment is a slate-level read that provides context for how much to trust an individual projection's upside, and it is why a play that looks identical on two different days is not actually identical.
Calibration Over Hit Rate
The model gets judged on calibration, never on last week's record.
A model that goes 7-3 is not a good model. A model where things projected at 60% happen 60% of the time is a good model, whatever this week's ledger says. Hit rate is an outcome, calibration is the property that produces outcomes over time.
Every projection lands in my Permanent Result Log, which is the authoritative record of everything the model has ever produced. Not the highlights, not the good weeks, everything. The diagnostics run against that log continuously:
- Edge bands. Do plays projected at a given edge actually return at that edge?
- Profile calibration. Does each ceiling profile hit its projected rates, or is one of them systematically lying?
- Ladder rungs. Do the +1 and +2 probabilities hold up, or are the tails over-projected?
- Side splits. Are overs and unders both honest, or is the engine leaning?
The Research Board
Everything above runs whether you look at it or not, the research board is where you can look at it.
Search any starter on the slate and you get the matchup the way the model reads it:
Lineup
Every hitter he is facing, in batting order, with their strikeout-relevant rates and contact quality against his hand, weighted by how often he actually throws each pitch.
Arsenal
Every pitch he throws, split against lefties and righties, with usage, CSW, whiff, put-away, chase, and contact quality. The supply side of the matchup before it meets a single hitter.
Hitter by pitch
The full grid. Every arsenal pitch against every hitter, filterable to a single weapon across the whole lineup.
Lineup read
For each hitter, the single pitch the model grades as the best strikeout option, that hitter's raw K% against this pitcher's hand, and his overall K-edge versus league baseline.
Zone maps
Where the starter lives with each pitch, and how each hitter performs by zone against the arsenal he is about to see. Toggle the metric, filter by pitch family, and the whole lineup re-reads itself.
Every number on the board is colored against league baseline for that exact pitch and that exact side. Green is a pitcher edge. Red is a hitter edge. There are no raw numbers floating without context, because a raw number without a baseline is not information.
The board updates itself as lineups firm up. Projected becomes confirmed, and the read changes with it.
What This Is, and What It Is Not
What it is. A distribution engine. The output is a shape, and the shape tells you where the risk actually sits. The projection is one property of that shape, and often not the most useful one.
What it is not. A pick service. I do not sell plays. I do not tell you what to bet. The tracked strategies exist to demonstrate that the model is calibrated, not to be followed blindly. If the only thing you take from KSplit is a list of sides, you have taken the least valuable thing here.
The model is wrong constantly, it is also right constantly. Every projection is wrong by some amount, and that is what a distribution is for. It is an honest statement about how wrong, how right, in which direction, and with what probability. A model that pretends otherwise is selling certainty, and certainty is the one thing this sport does not have.
Every number on this site comes out of the system described above. The Permanent Result Log holds the complete record, including the parts that did not work.
Ready to see the model live? View Today's Dashboard or open the How to Guide for Today's Dashboard.
Want full access to live dashboards? Create an account or sign in.