Part of the complete guide: How to improve your archery score

It happens more or less the same way every time. An archer changes something in the shot, a slightly different bow grip, a corrected anchor, a few weeks of work on the back, and at the next competition scores four points more than at the one before. By Monday the conclusion is already drawn: it worked. Had the four points gone the other way the opposite conclusion would have been drawn with equal confidence, and in both cases the same instrument would have decided it, namely the total from one competition.

That total, though, holds two things together that are worth separating: the archer's ability and the luck of that particular afternoon. The share belonging to luck is a great deal larger than anyone is willing to believe before seeing it calculated. Knowing exactly how large changes several practical things: how a ranking should be read, when it makes sense to say one athlete is better than another, and above all how long it takes before a score can confirm a technical change you have decided to make.

The calculation can be done, and the answer is a number. What follows comes entirely from work I published in the Lab, where it is set out at length and the code to redo it can be downloaded.

Seventy-two arrows are a finite sample

A qualification round is made of 72 arrows, and 72 arrows are a finite sample in exactly the way 72 measurements of anything else would be. If the same person, with precisely the same ability, shot that round ten times over, they would post ten different totals, and the difference between the highest and the lowest would not be a negligible detail.

This happens because every arrow comes out of a distribution of impacts around the aiming point, and where any individual arrow lands inside that distribution is not decided by ability. It is decided by chance. Ability decides how tight the distribution is, meaning how close the arrows sit to one another over the long run, while the total from one specific competition also depends on how those particular 72 draws happened to fall. A very strong archer can have an afternoon where the arrows settle unluckily on the ring boundaries, and a weaker archer can have one where they all land on the good side of the line.

To see how much that effect weighs, you need a way of turning points into something physical, because as long as the argument stays in points the number remains abstract and everyone projects onto it whatever importance they prefer.

One point is seven tenths of a millimetre

The translation exists and is surprisingly stable. Across the whole band that matters in a serious competition, say from 630 to 700 points out of 720, the relationship between expected score and group width is almost perfectly linear, and a single qualification point corresponds to about seven tenths of a millimetre of group size, which is to say less than a millimetre.

The consequences of that equivalence show up immediately. Two archers separated by three points on a ranking, who on the grid look as though they belong to different worlds, differ by barely two millimetres of group size. Eight points, which by eye is a chasm, come to five and a half millimetres. To give the widths involved a concrete reference, an archer worth an average of 690 points shoots inside a group about four and a half centimetres across, one worth 650 inside a group of a good seven centimetres, and between those two worlds, which on a bracket are a very long way apart, run a little over two and a half centimetres of group size.

Asking 72 arrows to tell two millimetres of difference apart reliably is asking an instrument for more than it can give. Not because the arrows are few in absolute terms, but because the natural variability of the movement, arrow after arrow, produces swings in the total of the same order of magnitude as the difference you are trying to measure. The signal you want to read, in short, is smaller than the noise that comes with it, and the next step is to measure what that noise is worth over 72 arrows.

Two identical archers finish five points apart

The cleanest way to measure that noise is an experiment you cannot run on a field but can run by calculation. Take two top-level archers who really are exactly identical, meaning the same group, the same ability, no difference of any kind, and have them each shoot 72 arrows as in a real qualification round. Then look at the gap that appears between them on the ranking, knowing that gap cannot have been produced by anything other than chance, because by construction there is nothing to tell them apart.

The result is that the mean gap between two identical archers comes to five points, and that one time in three it exceeds six and a half. Between two people who on the target are the same thing exactly, you would instinctively expect a point or two of difference, and instead five appear, with tails that regularly reach further.

Anyone who has looked at a qualification ranking recognises the order of magnitude: five points is the kind of gap that in a tight grid separates positions several places apart, and it is the gap usually read as a difference in ability. The calculation says it can be produced entirely by chance, between two athletes whose ability is identical.

How much confidence a gap gives you

From the mean gap you reach the question that actually matters, which is how much confidence to place in an observed gap. When one archer sits five points above another on a ranking, the probability that they really are the more precise of the two, rather than simply the one whose afternoon went better, can be calculated for any gap.

Gap in qualification Confidence in the true order How to read it
1 point56%Barely better than a coin: the order is undetermined
3 points68%About one time in three the true order is reversed
5 points79%An indication, not yet evidence
8 points90%The signal starts to lift clear of the noise
12 points97%A difference in precision hard to put down to chance
16 points99%Practically certain

The conventional threshold at which any field says something has been established is 95%, and reaching it takes a gap of a good ten points or more. Below five points you are largely reading noise, and every time two consecutive archers on a ranking are separated by one, two, or three points, which is the ordinary situation in any tight grid, their order relative to each other is essentially undetermined.

FIG · 01 Below five points a ranking says almost nothing Probability that the archer ahead in qualification really is the more precise of the two 60% 70% 80% 90% 50% 100% 95% threshold 56%68%79%90%97%99% 13581216 GAP OBSERVED IN QUALIFICATION (POINTS OUT OF 720) Axes on a linear scale. Calculated values, not measured on the field: the curve is an optimistic estimate, because it counts only arrow-by-arrow chance.
Fig. 01 Confidence grows slowly with the gap and passes 95% only beyond ten points. The gaps seen most often in a tight grid all fall in the left-hand part of the chart.

Why these numbers are optimistic

Two things need saying about that table, and neither makes it more forgiving. The first concerns what the calculation assumes it does not know. The confidence is derived as though nothing were known about the two archers beforehand, meaning as though, before looking at the scores, every possible difference in ability between them were equally plausible. On a competition grid, though, it is perfectly well known that archers close together on a ranking are almost always close in ability too, and allowing for that would make an observed gap even less probative than the table claims.

The second is that the calculation measures only arrow-by-arrow chance, holding the archer's form perfectly constant from the first shot to the last. It is the noise that would remain if the only source of variation were the dispersion of the group, whereas in reality everything the model does not look at is added on top: the previous night's sleep, the condition of that afternoon, the head, and nearly all of those widen the range rather than narrowing it.

One exception runs the other way and is worth naming: a disturbance that strikes both archers alike, such as a wind that turns for everybody or light failing across the field, partly cancels in the difference between the two scores, and makes that gap more informative rather than less. The net sign of everything the model leaves out therefore remains an empirical question, but the sources that strike the two archers separately are the more numerous, which is why the five points of mean gap, and the ten that follow from it as a threshold, should be read as an optimistic estimate of the true ambiguity of a ranking.

Key point

Below five points, a ranking does not establish who is better.

Two exactly identical archers end up separated by five points on average over 72 arrows, purely by chance. It takes a good ten or more before you can say, with the conventional 95% confidence, which of the two is genuinely the more precise. Every smaller gap, which is the norm in a tight grid, leaves the order essentially undetermined.

What to do with a ranking

For anyone coaching, that optimistic estimate turns into a rule to apply every time a results sheet comes up. If an athlete finished fourth rather than second, and three points separate them from second place, that gap on its own establishes no stable difference in precision between the two, and building a technical analysis on top of it means explaining something that may not have happened. If instead they came twelfth rather than fourth, with fifteen points between them, that is a robust signal and it is worth finding out where it came from.

The same logic applies to comparing two competitions by the same archer, which is the commonest case of all. Last Sunday's total and next Sunday's will differ by a few points almost certainly, and that difference, taken on its own, says nothing about what happened in between. Which is why an archer who wants to know whether they are improving needs a strategy other than comparing two consecutive numbers.

How many competitions it takes to demonstrate an improvement

There is a way out and it does not require shooting better: it requires measuring more. The chance in a single competition cannot be removed, but it can be diluted by averaging across competitions, and the dilution is quick, because the variability of a mean falls with the square root of the number of rounds. Four competitions halve the uncertainty against a single one, and nine reduce it to a third.

From there you reach the practical question, which is how many competitions are needed before and after a technical change for an improvement of a given size to become demonstrable. For an archer shooting around 650 points, the calculation gives these figures.

True improvement Competitions needed per period What that means in practice
3 points29About two seasons of competitions, before and after
5 points11A full season per period
8 points5Within reach of a normal calendar
10 points3Visible within a handful of competitions
15 points2A jump the score recognises straight away

Three points gained over a season is excellent technical work, and demonstrating it through score would take about thirty competitions, which in practice is two seasons. Ten points, by contrast, show up in three competitions per period. Most real technical changes fall between those two extremes, and that is why score on its own rarely settles the question within the season in which the change was made.

Performance level then shifts things again, and in the least convenient direction. The same three-point improvement, for an archer scoring 690, takes thirteen competitions instead of the twenty-nine one at 650 needs, because the total of an archer who groups tighter fluctuates less. The archer still in the middle of changing things, who has most need of quick feedback, is therefore also the one who has to wait longest to get it.

Score is a complete measure and a slow one

All of this risks sounding like a demolition of the score, and that would be the wrong conclusion. The total from a competition remains the most complete measure available, because everything contributes to a final total at once: technique as much as staying power over the distance, the wind as much as what happens in your head when the points matter. No partial indicator, however precise, holds as many things in a single number.

What it lacks is speed. A measure that holds everything also collects all the variability of everything, and for that reason it needs time to settle. Score is therefore an excellent instrument over a season and a poor one over a Sunday, while being the same instrument: behind one Sunday's number sit 72 arrows, behind a season's average sit a few thousand.

A different way of keeping your own results follows from that. Instead of comparing the last competition with the one before, which is the least informative comparison available, keep a moving average of the last four competitions and watch how it moves across the months. That average carries exactly half the uncertainty of a single competition, and above all it moves very little, which makes it dull to check every week and reliable when it does move.

Comparison against other archers should be read in the same register. A qualification ranking sorts distant groups well and neighbouring ones badly, so it says something solid about an archer at 650 sitting below one at 690, and almost nothing about the order between two neighbours on the list. Anyone measuring themselves against the club mate a couple of points above or below is chasing a difference the ruler used to establish it cannot see.

There is one last consequence, and it bears on training decisions more than on reading results. If score takes months to register three points of improvement, then a season is not enough to try more than one or two different routes, and the archer who changes setup every six weeks because "there are no results" wipes out each time the accumulation of competitions the measurement would have rested on, and starts again from zero. For score to say anything about a technical choice, that choice has to stay in place for the number of competitions in the table, which for three points of improvement is about thirty.

What to judge a change on, then

An operating rule follows from all this, and it is worth stating in full because it is the most immediate practical consequence: a technical change is not judged on score in the short run, because in the short run score does not have the resolution to see it. The limit is in the instrument before it is in the patience of the person waiting, since a measurement is being asked to distinguish a difference smaller than its own margin of error.

Which does not mean going two seasons without feedback. It means changing the feedback. Repeatability of the movement shows up long before score does, because it is measured on the dispersion of the arrows rather than on the total, and dispersion responds to a change without having to wait for chance to average out. The shape of the group says as much just as quickly: a group stretching vertically or horizontally shows where the problem is, and changes shape when the cause is removed, as I describe in the article on how the group opens up with distance. And video shows directly whether the movement you meant to change has changed, which is a different question and a much easier one than "am I shooting better".

Score is asked for confirmation afterwards, once the competitions have accumulated. It is the exact reverse of how it usually goes, which is score first and sensations later, and in practice it changes one thing above all: it removes the temptation to start rebuilding your technique on the Monday after a bad competition, when the only data available is a single total, and a single total cannot tell a real decline from an afternoon in which the arrows fell on the wrong side of the line.

The same measurement, inside a match

One final consequence concerns what happens after qualification. If 72 arrows struggle to separate two archers who are close, an elimination match, which has at most fifteen arrows, separates far fewer of them, and that explains why upsets in the later rounds are frequent and are not anomalies at all.

This is where the argument joins a larger discussion, the one about how points are counted inside a match. In the essay on the set system I calculated what it costs, in precision, to decide a match by sets instead of by adding up the arrows, and the interesting result is not so much the cost itself as the proportion: the argument is about an arithmetic that moves outcomes by a couple of percentage points, while the instrument that decides who is the favourite, namely the qualification round, leaves the order of almost every neighbouring pair undetermined. An order of magnitude separates the two imprecisions, and the larger one sits upstream, in the qualification round that decides who enters the bracket as the favourite.

The essay carries the full calculation, the evidence that the result does not depend on the shape of the group, and the code package with which anyone can redo the sums and check whether the published numbers are the ones the calculation really produces.

Questions I get asked most

How many points of difference actually count on a ranking?
It takes a good ten points or more to say with 95% confidence which of two archers is the more precise. At eight points confidence is 90%, at five it is 79%, at three it is 68%. Below five points you are largely reading noise, and between archers separated by one or two points the order is undetermined.

Why don't two identical archers post the same score?
Because 72 arrows are a finite sample. Where any individual arrow lands inside the dispersion of the group is decided by chance rather than by ability, and over 72 draws the differences accumulate: between two exactly identical top-level archers the mean gap produced by chance alone is five points, and one time in three it exceeds six and a half.

I changed something in my technique and my score has not gone up. Is that normal?
Yes, and it is what you should expect in the short run. Demonstrating a three-point improvement through score takes about 29 competitions per period for an archer at 650, and 13 for one at 690. In the meantime the change is judged on repeatability, on the shape of the group, and on video, all of which respond far sooner than the total does.

How should I track my own scores, then?
Keep a moving average of your last four competitions rather than comparing the most recent with the one before. That average carries half the uncertainty of a single result, and because it moves slowly it is reliable when it does move. Judge technical changes on what you can observe directly, and leave the score to confirm them months later.

Where these numbers come from. Every figure in this article comes from the essay The set system in archery: the mathematics of who wins, published in full in the Lab, where the calculation is worked through step by step and the reproduction package holds the code to redo it. The practical readings for archers and coaches are mine: the essay deals only with the calculation.

Go to the bibliography
The score is slow. The group is not.

Dispersion answers months before the score does.

A technical change shows up in the shape of the group long before it shows up in a total. Send me a video of your shot and I will send back the biomechanical reading with the data and the priorities to work on.