Archery is saturated with numbers. Draw weight, arrow speed, group size, spine rating, FOC percentage, ES on the chronograph, straightness tolerance, tip weight. Archers collect these numbers, argue about them, and make buying decisions based on them. But almost nobody asks the most important question first: what kind of number is this, and what does it actually tell me?
Almost every number in archery is a statistical value — a description of a distribution, an average from a population, or a measurement with inherent variance. Understanding that is not an academic exercise. It changes how you tune a bow, evaluate an equipment change, read a recommendation, and practice deliberately instead of just repeatedly.
The basic toolkit
You do not need a statistics course to use these concepts. You need to understand what each term means and be able to recognize it when it shows up.
Mean — the center of a distribution
The mean is the arithmetic average: add all the values, divide by how many you have. When a chart says "most archers at 70 lbs use a 340 spine," the word "most" is hiding a mean — the 340 is where the center of that population landed. It does not mean 340 is right for you. It means 340 is right for the average of the people they measured.
The mean is also the number that gets distorted by outliers. Ten archers who need 400 spine and one who needs 200 spine will average to a mean that doesn't describe any of them well. This is why archery recommendations that span a wide range of setups should be read with some skepticism — the population producing them may not be homogeneous.
Standard deviation — the width of the spread
Standard deviation (σ, "sigma") measures how spread out a set of values is around the mean. A small σ means values cluster tightly near the mean. A large σ means they spread widely. This is the most useful single concept in archery statistics, because almost everything you care about in archery is a measure of spread.
Your group size is σ. The variation in arrow velocity across shots is σ. The shot-to-shot consistency of your form is σ. The spread of arrow weights in a batch is σ. When you are trying to improve any of these things, you are trying to reduce σ — you are trying to make the distribution narrower.
Normal distribution — the bell curve
Many things in archery follow a normal distribution: arrow impact points, shot velocities, shaft spine values within a batch, draw weights across an archer population. The normal distribution is bell-shaped — most values cluster near the mean and fewer values appear as you move away from it in either direction.
The useful property is predictable: in a normal distribution, roughly 68% of all values fall within ±1σ of the mean, 95% fall within ±2σ, and 99.7% fall within ±3σ. This is called the 68-95-99.7 rule and it applies to almost everything you will measure in archery.
Your arrow impacts on a target face are normally distributed. If your σ in the horizontal direction is 1.5 inches at 40 yards, then 68% of your arrows land within 1.5 inches of center horizontally, 95% within 3 inches, and 99.7% within 4.5 inches. Those numbers fully describe your horizontal accuracy at that distance, without any measurement of a specific group.
Sample size — how much data you actually need
Any measurement you make is an estimate of the true underlying value. The more data points you have, the better your estimate. The fewer you have, the more your measurement is influenced by chance.
A 3-arrow group is a sample of size 3 from your shot distribution. A 10-arrow group is a sample of size 10. The 10-arrow group gives you a far more reliable estimate of your true σ — not because your shooting changed, but because statistics work better with more data. The standard error of your estimate decreases as sample size increases, following the square root rule: doubling your sample size makes your estimate √2 ≈ 1.4 times more reliable.
This has direct practical implications. If you shoot 3 arrows to test whether a tuning change improved your groups, you do not have enough data to know. If you shoot 3 arrows and they group tightly, you might conclude the change worked — but a 3-arrow cluster can happen by chance with any setup. You need at least 10, and preferably 20, to have a measurement that is more likely to reflect reality than luck.
Extreme spread — what it is and what it isn't
Extreme spread (ES) is the distance between the best and worst value in a sample — the highest minus the lowest. It is the standard reporting format for archery groups and chronograph sessions. ES is intuitive and easy to measure, and it is also the most statistically unreliable summary you can use.
ES is a range statistic. It depends entirely on the two most extreme values in your sample. Adding one more data point can only keep it the same or make it larger — it can never make it smaller. This means ES grows with sample size even if your underlying σ is not changing. A 3-arrow ES of 2" and a 10-arrow ES of 4" from the same archer shooting the same setup are completely consistent — the 10-arrow group didn't get worse, the sample just got big enough to include more of the distribution's tails.
The relationship between ES and σ for a normally distributed sample: at n=10, ES ≈ σ × 3.1. At n=20, ES ≈ σ × 3.7. At n=3, ES ≈ σ × 1.7 — but with enormous variance. That n=3 estimate can range from near zero to nearly as wide as the n=20 estimate by pure chance. This is why 3-arrow groups are nearly useless for comparison purposes.
Your group size is a distribution, not a number
When you shoot a group and measure it, you are getting one sample from an underlying distribution. The group you see on the target is not "your accuracy" — it is one realization of a random process. Shoot again and you will get a different number. The average of many groups, over many sessions, converges toward your true underlying σ. Any single group is just an estimate of that.
This matters most when you are evaluating changes. You shoot 10 arrows with your old rest and get a 4" group. You install a new rest, shoot 10 arrows, and get a 2.5" group. Did the rest improve your accuracy? Maybe. But your shot-to-shot σ produces natural variation in group size — two samples from the same distribution can easily differ by 1.5 inches by chance. To know whether the new rest actually changed anything, you need multiple groups from each setup, not one group from each.
The standard tool for this kind of comparison is the concept of statistical significance — whether the difference you observed is large enough, relative to the natural variation, to be unlikely to have happened by chance. Archery testing rarely uses this formally. But the intuition transfers: if your shot-to-shot variation is large and the improvement is small, you cannot see the improvement through the noise.
A practical rule: if the improvement in group size from a change is smaller than your natural session-to-session variation, you do not have evidence that the change worked. You have evidence that you had a better day.
Spine charts are population averages
Manufacturer spine charts are built from testing many archers across a range of draw weights and arrow lengths. The recommendation you see — "70 lbs, 28" arrow: use 340 spine" — is the spine that produced the best results for the average archer in that draw weight and length bucket. It is a mean, with a distribution around it.
The "range" you see on charts — 300 to 340, or 340 to 400 — is approximately the ±1σ window of that population. Roughly 68% of archers in that draw weight and length bucket will land somewhere in that range. The remaining 32% will land outside it. If your cam is unusually aggressive, your point weight unusually heavy, or your draw length unusually long relative to your weight, you may be one of the 32%.
This is why the ALR method exists. Rather than placing you in a population bucket by draw weight alone, it accounts for the variables that push you toward the tails of the distribution — cam aggressiveness, point weight, exact arrow length. It narrows the estimate from a population range to a single number calibrated for your specific setup. See the ALR spine selection article for the details.
The same logic applies to every chart
Point weight recommendations ("most hunters run 100 gr"), FOC ranges ("10–15% for target, 15–20% for hunting"), draw weight recommendations ("most adult men do well at 60–70 lbs"), stabilizer weight guidelines — every one of these is a statement about the center of a population distribution. They are not universal truths. They are the place where the most people in the relevant population got good results.
Whether a population average applies to you depends on where you sit on the relevant distribution. An archer with an unusually aggressive cam at a given draw weight needs a stiffer spine than the chart average. An archer with unusually strong shoulders for their weight can sustain a higher draw weight than the population recommendation. An archer with high-FOC arrows in a windy outdoor environment benefits more from a heavier front-end stabilizer than the average recommendation covers.
Arrow tolerances are distributions
When an arrow shaft is rated ±0.003" straightness, that is a maximum tolerance, not a description of every shaft in the batch. The actual distribution of straightness values within a quality batch is approximately normal and tighter than the spec — most shafts land closer to zero runout than the maximum allows.
The same applies to weight consistency. A batch of shafts rated at ±2 grains per shaft will have most shafts clustered near the batch mean, with fewer shafts at the ±2 grain extremes. When you sort arrows by weight and spine, you are pulling the center of that distribution — the shafts that landed closest to the mean on both dimensions simultaneously. The remaining shafts are not defective; they are at the edges of the distribution.
What premium arrows actually deliver, statistically: a tighter distribution, not a categorically different product. A batch of ±0.001" shafts has a σ roughly three times smaller than ±0.003" shafts. Every shaft in the batch is more likely to be near zero runout. The benefit compounds across a group of twelve — the worst shaft in a premium batch is typically better than the worst shaft in a standard batch.
Spine sorting within a batch
When archers spin-test or deflection-sort arrows within a batch of the same SKU, they are doing something statistically precise: identifying the shafts whose spine values cluster most tightly around the batch mean, and shooting only those. The rejected shafts are not bad arrows — they are at the tails of the batch's spine distribution. The selected arrows form a tighter sub-distribution, which produces a tighter group.
The improvement in group size from sorting is predictable from statistics: it depends on how wide the original batch distribution was (the batch-to-batch manufacturing σ) and how tightly you sort. Sorting to ±0.5 lb deflection in a batch with 2 lb deflection range cuts the spine variance by roughly 75% in the selected sample. That directly reduces the arrow-to-arrow contribution to group spread.
Community advice is mode data from a biased sample
When a forum thread reaches consensus — "everyone agrees the 350 is the right spine for that setup" — it has described the mode of a non-random sample. The mode is the most common answer. The sample is non-random because it consists of people who are engaged enough with archery to post on that forum, which systematically overrepresents experienced archers, competitive archers, and people whose results were good enough to share.
Two biases compound here. Self-selection bias: the sample is not representative of all archers — it is skewed toward people with strong opinions and relevant experience. Survivorship bias: archers whose equipment choices worked out are more likely to recommend those choices. The archer who tried a 340 and found it porpoised badly is less likely to be in the thread than the archer who tried a 340 and loved it.
None of this makes community advice useless. The mode of a large experienced sample is genuinely informative. But it tells you the most common answer, not the correct answer for your specific situation. Treat forum consensus as a prior — a reasonable starting point — not a conclusion.
"Everyone uses X" is not evidence that X is optimal
This specific claim deserves its own note. If everyone at a club runs a certain release aid, that is evidence that the release aid is functional and popular among that population. It is not evidence that it is optimal for you, or that alternatives are worse. The most common choice in a population reflects the mode of that population's preferences and constraints — budget, availability, coaching advice received, what the coach used, what won a tournament they watched. Many of these factors are unrelated to technical optimality.
Chronograph ES — speed variation is vertical spread
When you run arrows over a chronograph, you are measuring arrow velocity for each shot. Extreme spread (ES) on the chrono is the range of those velocities — fastest minus slowest. It is the speed equivalent of group size, with the same statistical properties: it is a range statistic that grows with sample size and is dominated by the two extreme values.
High ES means high shot-to-shot variation in the energy delivered to the arrow. This variation comes from several sources: inconsistent release timing, inconsistent draw length, inconsistent peep position, and inconsistent string contact — all of which affect how much energy the string transfers to the arrow. The ES you measure is the combined σ of all those sources, expressed as velocity variation.
The practical consequence is vertical spread at distance. An arrow traveling 5 fps faster follows a flatter trajectory; an arrow traveling 5 fps slower drops more. At 40 yards, 1 fps of velocity variation produces roughly 0.05–0.1" of vertical spread, depending on arrow weight and speed. A 10 fps ES — common in uncontrolled conditions — contributes half an inch or more to vertical group spread at that distance, entirely from energy inconsistency.
A low ES (under 5 fps in a 10-shot sample) indicates consistent energy delivery. It does not guarantee tight groups — that requires low σ in form and tuning as well — but it removes one source of variance from the error budget. The ES article goes into the specifics of what the chronograph actually measures and what values indicate a tuning problem versus a form problem. See Arrow ES: using the chronograph as a final tuning check.
Regression to the mean — your best session was partly luck
Regression to the mean is one of the most important and most ignored concepts in performance sports. The principle: after an extreme result — unusually good or unusually bad — the next measurement tends to be closer to the true underlying average. Not because anything changed, but because extreme results contain a luck component, and luck does not persist reliably.
If you shoot your best practice session ever — your tightest groups, your cleanest form, everything clicking — and then shoot the next day with nothing changed, you are statistically likely to perform closer to your average. This is not a failure. It is regression to the mean. The exceptional session was partly skill and partly favorable random variation. The favorable variation is unlikely to repeat at the same magnitude.
The danger in archery: making a change immediately after an exceptional session and then crediting the change for the maintained performance — or worse, making a change and then observing a regression, and blaming the change for the regression. Both conclusions are likely wrong. The regression would have happened regardless.
Regression to the mean also explains a common coaching observation: archers who receive a correction often improve on the very next shot. The coach praises the correction. But a bad shot, by definition, is likely to be followed by a shot closer to the mean — whether or not the correction was given. Coaching research accounts for this. Archery testing rarely does.
Confirmation bias — choosing when to stop collecting data
Confirmation bias in archery testing looks like this: you install a new rest, shoot three groups, stop when one of them looks good, and conclude the rest improved your accuracy. You have confirmed what you hoped to find — but you have not done a valid test. You collected data until you got the result you wanted and then stopped.
A valid test requires pre-specifying how many arrows you will shoot before looking at the results, shooting all of them regardless of what the interim results look like, and measuring the full result — not the best group of three attempts. This is the methodology that makes sports science research valid and informal archery testing largely unreliable.
The informal version — stop when the result looks good — is not obviously dishonest. It feels like prudent testing. The problem is purely statistical: when you stop collecting data based on the interim result, you are not measuring the equipment. You are measuring your patience. The probability that any two consecutive groups will look like an improvement, purely by chance, is substantial when your natural group-to-group variation is large relative to the effect you are looking for.
How much data is enough to see a real change?
A rough guide: if the improvement you are looking for is smaller than one standard deviation of your natural group-to-group variation, you cannot see it with a reasonable number of arrows. You need the improvement to be large relative to the noise.
If your groups vary naturally between 2" and 5" across sessions (a σ of roughly 0.75"), you cannot detect a real improvement of 0.5" without many sessions of data. The improvement is smaller than the noise. If a change produces a genuine 2" improvement — from 4" groups to 2" groups — you can see that clearly in two or three 10-arrow groups, because it is large relative to the noise.
This is why the most common advice — "buy better arrows and shoot tighter groups" — is so easy to believe and so hard to verify. The arrow improvement is real but modest. The natural variation in group size from session to session is large. The modest improvement sits inside the noise, and confirmation bias fills in the rest.
Correlation is not causation
Two things varying together does not mean one causes the other. This is the most commonly violated principle in archery equipment discussions.
Heavy stabilizer systems and tight groups appear together in elite setups. This is a genuine correlation. The incorrect conclusion: buy a heavier stabilizer and your groups will tighten. The correct analysis: elite archers also shoot more arrows per week, have better form, shoot better-matched equipment, have better coaches, and have been shooting longer. The stabilizer weight is correlated with tight groups because both are correlated with elite-level investment and practice. Whether the stabilizer itself is causing the improvement — as opposed to all the other things that come with that level of commitment — is a separate question that requires controlled testing to answer.
The confounding variable problem is everywhere in archery. Arrow speed is correlated with flat trajectory; it is also correlated with lighter arrows. Lighter arrows have lower kinetic energy and lower FOC. The speed benefit and the FOC cost come together. Archers who chase speed often get the trajectory improvement they wanted and the accuracy cost they didn't account for — and attribute the accuracy cost to something else, because the speed was supposed to help, not hurt.
Independent errors combine as root-sum-of-squares
When a group is opened by multiple independent sources of error — form variation, arrow inconsistency, tuning residue, wind — those errors do not simply add together. They combine as the square root of the sum of their individual variances, known as root-sum-of-squares (RSS).
The formula: σ_total = √(σ₁² + σ₂² + σ₃² + … + σₙ²)
The practical implication is counterintuitive and important: the largest error source dominates the total. If your form variation contributes a σ of 2" to your group and your arrow inconsistency contributes 0.5", the total is √(4 + 0.25) = √4.25 ≈ 2.06". The arrow inconsistency barely moved the needle. Eliminating it entirely would take you from 2.06" to 2.0" — a 3% improvement.
This means fixing the second-largest error source when the largest is still dominant produces almost no measurable improvement. The common archery mistake: buying premium arrows to tighten groups when form is still the dominant variance source. The math says the improvement will be invisible. The correct order is always to reduce the largest σ first.
The group size projection calculator uses this RSS approach — each factor (platform, arrows, tuning, FOC, fletching, wind) contributes an independent σ that combines into the total predicted group. The breakdown chart shows which factor is currently the biggest contributor, so you know where to focus effort.
What the manufacturer recommendations actually are
With the statistical framework in place, it is useful to go through the major recommendation types in archery and describe what each is actually measuring.
Spine chart recommendations — the mean spine value that worked for the population of archers at that draw weight and arrow length combination, across a moderate cam profile. Your setup's deviation from that average population shifts where you should land.
Point weight recommendations — typically "100 grains for most setups" or "125 for hunting" — these are modal recommendations from a population of archer+equipment combinations. Heavy-FOC setups run much heavier. Ultralight target setups run lighter. The recommendation is the center of the distribution, not its boundary.
Draw weight recommendations — "shoot as much as you can pull 30 times without form breakdown" is a practical threshold derived from fatigue physiology. The 60–70 lb range for adult men is where the population median lands for a draw weight that is sustainable through a practice session without form degradation. It is not a minimum or a target; it is a description of where most archers of that size and fitness level end up.
Physical bow weight and stabilizer weights — recommendations are based on what most archers in that shooting category find controllable through the hold and shot. Lighter = easier to hold; heavier = more inertia, less arc of movement during aiming. The recommended range is where the population balances those two trade-offs on average. Stronger archers benefit from more; weaker or newer archers from less.
FOC ranges — "10–15% for target, 15–20%+ for hunting" — these are the ranges where most archers in those categories observe good balance between stability and speed. Your specific arrow speed, stiffness, fletching size, and form all affect where your optimal FOC falls. The range is the ±1σ window of the population average, not a hard boundary.
Release aid recommendations — the most common releases in competitive archery are not necessarily optimal; they are the tools that the most competitive archers reached for, given historical availability, coaching tradition, and tournament success. The distribution of what works shifts as new designs appear. The recommendation describes the current modal choice, not the result of controlled comparison testing.
Tracking your own statistics
Most archers have no baseline. They know roughly what they shoot, but they do not know their σ, their session-to-session variance, or which factor in their error budget is largest. Without a baseline, every change is uninterpretable — you cannot tell if something helped, hurt, or did nothing, because you don't know what "nothing" looks like for your setup.
The minimum useful tracking: ten arrows per session at a fixed distance, measuring and recording the group center-to-center. Do this for ten sessions without changing anything. You now have a baseline distribution of your group sizes — you know your mean and your session-to-session σ. Any subsequent change can be compared against that baseline.
More useful tracking: record group size and group center separately. A group that is tight but consistently left tells you something different than a group that is tight and centered. Consistent offset points to a systematic error — equipment or form — that can be identified and removed. Random scatter around center points to random error — σ — which requires a different kind of work.
The most useful tracking: record session date, environmental conditions, equipment state (any changes since last session), and notes on what felt different. Over time, patterns emerge. You begin to see which variables correlate with tighter groups and which don't. You accumulate the kind of personalized data that no chart, no forum thread, and no recommendation can provide — because it is about your system, not a population average.
Applying this — the practical decisions
Before buying new arrows: determine whether form or equipment is currently your dominant error source. If your form σ is larger than your arrow σ, premium arrows will not produce a visible improvement. The RSS math says so.
Before concluding that a change worked: shoot at least ten arrows per condition, in multiple sessions, without selecting which sessions to include. Compare means and distributions, not individual groups.
Before accepting a recommendation: ask what population it was derived from and whether your setup and skill level place you near the center of that population or toward its edges.
Before acting on your best day of shooting: consider whether the result is a genuine improvement or regression to the mean playing out in the favorable direction. A baseline makes this answerable. Without one, it is not.
When reading a spine chart: understand that you are looking at a population mean and that your specific cam, point weight, and exact draw length may push you meaningfully toward one edge of the recommended range or outside it entirely.
When testing equipment: pre-specify how many shots you will take before looking at the result. Do not stop early because the interim data looked good. Do not keep going because it didn't. The sample size determines whether the measurement means anything.